Two posts in a row, I lined up decision models and Laya came last.
Jev got 92% of my test cases right. GPT-6 Luna and Clef were right behind it. Laya, the open-source one running on my desk, got 53%. Each time I wrote the same disclaimer underneath: it's untrained, its authors say fine-tuning is where it jumps.
At some point a disclaimer becomes an excuse. So I trained it.

That's three versions of Laya playing the same game. On the left, the model you get from pip install. On the right, the same model after about an hour of training on my RTX 3050. More on how it learned Pac-Man in a bit.
The short version
| untrained Laya | fine-tuned Laya | the reference | |
|---|---|---|---|
| Pigeonhole tests passed | 31 / 50 | 44 / 50 | Jev: 47 / 50 |
| Pac-Man games won | 0 / 200 | 87 / 200 | its teacher: 78 / 200 |
| Time per decision | ~100 ms | ~100 ms | Jev: ~500 ms |
| What it cost | $0.27 of API calls, all in |
On my own classifiers, fine-tuned Laya went from last place to level with Clef, three cases behind Jev, while still answering five times faster. At Pac-Man, it went from playing like a coin flip to matching the code that taught it.
Everything is open source, including a step-by-step guide for your own decisions: bnap00/laya-fine-tuning. The short version of that guide is at the end of this post.
Fine-tuning a decision model is copying a teacher
Laya doesn't learn from a pile of right answers the way you'd train a classifier in a course. It learns from probabilities. For every training example you tell it: here's the input, here's the question, and here's how likely each option is.
That sounds like more work. It's actually the cheat code. Something that's better than Laya at your job, but too slow, too expensive or too far away to use everywhere, can write those probabilities for you. Laya learns to copy it, then runs in 100 ms on your own machine.
So every fine-tune is the same four steps:
- Get inputs like the ones the model will see in production.
- Ask a teacher for its probabilities on every one of them.
- Train Laya to match those probabilities.
- Test it on something it has never seen.
Step 4 is where most of the honesty lives, and I'll come back to it.
6 GB is not enough. Here's what fits
The first thing I found out: a full fine-tune doesn't fit on my GPU.
Laya's English model has 421 million parameters. Training all of them needs the weights, a gradient for each one and two more numbers per weight for the optimizer. That's about 7 GB before it has even looked at a single input. My RTX 3050 has 6 GB, and two other containers already use one of them.
Laya's trainer has a --freeze-encoder flag that trains only the small decision head on top. That fits easily, but its own docs say it "gains much less than a full fine-tune".
So I went for the middle: freeze the bottom 22 of the encoder's 28 layers, train the top 6 and the decision head. That's 100 million trainable parameters, and it peaks at about 3.8 GB.
The nice part is that it doesn't need a fork. Laya's trainer already skips anything that isn't set to train, so freezing layers right after the model loads is enough:
for p in model.encoder.embeddings.parameters():
p.requires_grad_(False)
for layer in model.encoder.layers[:-top_layers]:
for p in layer.parameters():
p.requires_grad_(False)
Everything else is Laya's own training loop, including the part people skip: fitting the confidence calibration on a slice of data the model didn't train on.
python finetune.py data/train.jsonl checkpoints/mine --top-layers 6 --epochs 3
Round one: my own classifiers
Pigeonhole ships five template pipelines (email routing, GitHub issue labelling, lead qualification, moderation and support triage), each with ten test cases. Those 50 cases are the scoreboard.
Which means I can't train on them. Train on the test and you'll get a great score that tells you nothing.
The inputs: a cheap model writes them
I gave Claude Haiku 5.5 each pipeline's description and its questions, never its tests, and asked for realistic inputs in ten different styles: blunt one-liners, long rambling emails, typos everywhere, sarcasm, non-native English, cases that sit right on a boundary between two options.
we never received invoice 60012 for the october consulting block could you email a copy to our ap inbox
Dear Madam, I trust you are well. Our company offers commercial cleaning services across the
Greater Manchester area. Enclosed please find our brochure.
1,971 inputs for about ten cents. A filter dropped anything that looked too much like a test case. It caught exactly one.
The labels: Jev is the teacher
Then I asked Jev every question about every input and kept its full answer: not just "billing", but something like "billing 0.91, returns 0.06, shipping 0.02, other 0.01". Seven cents.
One detail mattered more than I expected. Pigeonhole doesn't send Laya the options the way you write them in the spec. It squashes each one into a single line first. So the training data uses that exact same squashed text. A model trained on a differently worded question is learning a different question.
The result
Training took 28 minutes. Then the same 50 test cases, on every decision model from the last post:
| Issues | Leads | Moderation | Support | Total | Typical response | ||
|---|---|---|---|---|---|---|---|
| Jev | 9 | 10 | 9 | 9 | 10 | 47 / 50 | 500 ms |
| GPT-6 Luna | 8 | 10 | 8 | 10 | 10 | 46 / 50 | 330 ms |
| Clef | 8 | 10 | 9 | 7 | 10 | 44 / 50 | 600 ms |
| Clef-flash | 8 | 10 | 8 | 7 | 10 | 43 / 50 | 680 ms |
| Laya, untrained | 6 | 5 | 4 | 10 | 6 | 31 / 50 | 100 ms |
| Laya, fine-tuned | 8 | 9 | 7 | 10 | 10 | 44 / 50 | 100 ms |
Thirteen more cases right. Support triage went from 6 to 10. Issue labelling went from 5 to 9, and it finally stopped calling everything a bug. Fine-tuned Laya now ties Clef, beats Clef-flash, and is still the only one on this list that's free, private and five times faster than Jev.
It's still three cases behind its own teacher, and that's expected. It learned from Jev's answers, so Jev is the ceiling. What's left is mostly lead qualification: it calls "TIL about window functions. Neat." a data-pipeline lead, and it misses "Airbyte keeps falling over on our largest table." It also still sends "Buy cheap followers" to the bug tracker.
Moderation is the one pipeline where untrained Laya was already perfect, and training didn't break it. I was half expecting it would.
Round two: Pac-Man
Classification is what Laya is built for. I wanted to see if it could do something that doesn't look like classification at all.
A game is a decision asked over and over again. Every tick, Pac-Man has two to four legal moves. That's a choice question:
{"move": {
"type": "choice",
"instructions": "You are Pac-Man. Pick the move that keeps eating pellets without getting caught. ...",
"criteria": {
"left": "eats a pellet, closest ghost 9 steps away",
"right": "closest pellet 3 steps away, closest ghost 2 steps away",
"down": "closest pellet 4 steps away, closest ghost 11 steps away, dead end",
}}}
The state is a status line and the maze in ASCII. Laya answers, the game moves, and we ask again. At 43 ms a move, it's a perfectly playable frame rate for Pac-Man.
I wrote a small version: one maze, 111 pellets, two ghosts. One ghost chases you 60% of the time, the other 30%.
Untrained Laya is a random number generator
I played 200 games with each player, on seeds no training game ever used.
| player | games won | pellets eaten (of 111) | picks the teacher's move |
|---|---|---|---|
| random moves | 0 / 200 | 10 | 42% |
| Laya, untrained | 0 / 200 | 12 | 38% |
Untrained Laya plays exactly like a coin flip. It's actually less likely than random to pick the sensible move. Telling it "the ghost is 2 steps away" in plain English means nothing to it yet.
The teacher costs nothing
For Pigeonhole, the teacher was Jev, and I paid per question. For Pac-Man I wrote one. It scores each move with distance arithmetic: dying is -100, eating a pellet is +6, each step to the nearest pellet is -1, being near a ghost costs more the closer it is, and a dead end near a ghost is a bad idea.
It's twenty lines and it's not great. It wins 39% of games. But it's free, so I can ask it about as many states as I like, and it gives a preference over every move, not just its favourite.
I played 300 games, with one move in five chosen at random so the data includes some bad positions, and saved 12,595 states with the teacher's probabilities. Training took 41 minutes.
| player | games won | pellets eaten | picks the teacher's move |
|---|---|---|---|
| Laya, untrained | 0 / 200 | 12 | 38% |
| Laya, fine-tuned | 63 / 200 | 81 | 91% |
| teacher | 78 / 200 | 82 | 100% |
From zero wins to 63. It eats as many pellets as the teacher does. But it wins less often, and I wanted to know why.
It had never seen its own mistakes
Every training state came from the teacher's games. The teacher rarely walks into a corner with a ghost behind it, so the student never saw what to do there. When the student made a small mistake, it ended up in a position it had never seen, and then made a bigger one.
There's a classic fix for this called DAgger. Let the student play, and have the teacher label every state the student gets itself into. The student learns what to do in its own messes.
I let the fine-tuned model play 150 games, collected 11,925 states, had the teacher label them, and trained for one more epoch, mixed with some of the original data. 28 minutes.
| player | games won | pellets eaten | picks the teacher's move |
|---|---|---|---|
| Laya, fine-tuned | 63 / 200 | 81 | 91% |
| Laya, + one DAgger round | 87 / 200 | 85 | 92% |
| teacher | 78 / 200 | 82 | 100% |
It wins more games than the teacher it learned from. I'd love to tell you the student surpassed the master. But over 200 games, 87 against 78 is within noise. What I can say is that it plays as well as the code it copied, and it learned that from text describing the board.
By the numbers
Every run, on the same desktop: an Intel Core i5-8500, 32 GB of RAM and the RTX 3050. Each run trains the top 6 of 28 encoder layers plus the decision head.
| run | training data | epochs | optimizer updates | time |
|---|---|---|---|---|
| Pigeonhole | 1,971 inputs, 6,358 questions labelled by Jev | 3 | 597 | 28 min |
| Pac-Man v1 | 12,595 states from 300 teacher games | 2 | 764 | 41 min |
| Pac-Man v2 (DAgger) | 11,925 states from v1's own games + 6,000 old ones | 1, from v1 | 548 | 28 min |
Laya's trainer holds a slice of the data back and reports before and after on it. For Pac-Man I also gave it 1,694 states from 40 separate games:
| Pac-Man v1, held-out states | before | after |
|---|---|---|
| Picks the teacher's top move | 62% | 90% |
| Calibration error (ECE) | 0.134 | 0.041 |
| Average confidence | 0.50 | 0.87 |
That last pair is the one I care about. After training, it isn't just right more often; when it says 87%, it means roughly 87%. That's what makes a rule like "below 0.4, ask a human" work.
Where the money went:
| what | cost | |
|---|---|---|
| Inputs | 1,971 synthetic inputs from Claude Haiku 5.5, plus a run that crashed halfway | $0.12 |
| Labels | Jev, every question on every input | $0.07 |
| Redo | relabelling after the template fix (see below) | $0.07 |
| Test runs | Jev, GPT-6 Luna, Clef and Clef-flash on the 50 cases | $0.01 |
| Total | $0.27 | |
| Pac-Man | everything | $0 |
What I got wrong along the way
My GIFs were lying. To record a game I copy its state every frame, and my copy function was pulling a random number from the game's own generator. Recording a game changed how the ghosts moved. The scoreboards were fine, because they don't record. But the first side-by-side GIF showed the DAgger model losing a game it actually wins. One line fixed it.
One lucky draw. After the first 50 games, the fine-tuned model had won 20 and the teacher only 16. I nearly wrote "the student beat its teacher" right there. Over 200 games, the teacher won 39% and the student 31.5%. Run more games before you write the headline.
The PyTorch install. pip install laya pulled a PyTorch built for CUDA 13. My driver supports CUDA 12.7. Nothing crashed. It just quietly ran everything on the CPU. If torch.cuda.is_available() says False on a machine with an NVIDIA card, install a matching build:
uv pip install torch --torch-backend=cu126
The templates I trained on were the broken ones. In the last post I found that some Pigeonhole templates had option descriptions cut off at the first comma, and fixed them. My copy of the templates was from before the fix. So my first fine-tune learned, among other things, that "damaged" means "The item arrived broken". It still scored 41 / 50. I pulled the fix, relabelled everything with Jev for another seven cents, retrained, and got 44. Garbage in, slightly worse out.
Is it fair?
Fairer than last time, not perfect.
- The tests stayed unseen. The generator never saw them, and nothing too close to one made it into training. Pac-Man was scored on seeds no training game used.
- The teacher has home advantage again. Pigeonhole's test cases were written and checked against Jev. Training Laya to copy Jev copies that advantage too, so on these tests Jev is the ceiling.
- These are five public pipelines, not all six. The Reddit self-promotion detector from the last two posts isn't in the Pigeonhole repo, so it's not here. That's 50 cases, not 74.
- One run each. A different seed would move a case or two. Treat one-case gaps as ties.
So should you train it?
If you were going to use Laya, yes. It isn't optional. The model out of the box is a demo. The model after half an hour on a gaming GPU and twenty cents of API calls is level with Clef on my own pipelines.
The pattern works anywhere you have a teacher:
- A hosted model you trust but can't afford at volume. Use it to label a few thousand inputs once, then run Laya for free.
- Code that already makes the decision but is slow or awkward to call. Rules, a search, an expert system. That's what the Pac-Man teacher is.
- People. A week of human-reviewed tickets is a training set.
And it runs where the data is. Nothing in either round of training left my desk except the inputs I sent to the teacher.
Shut up, how do I fine-tune it on my own decisions?
The repo's README has the long version. Here's the short one.
0. Set up, and check the GPU is actually used
git clone https://github.com/bnap00/laya-fine-tuning && cd laya-fine-tuning
uv venv --python 3.12 .venv
uv pip install --python .venv/bin/python -r requirements.txt
.venv/bin/python -c "import torch; print(torch.cuda.is_available())"
If that prints False on an NVIDIA machine, install a PyTorch that matches your driver (--torch-backend=cu126 for CUDA 12.x).
1. Write the decision as Laya questions
A choice, a score or a noul. Keep option descriptions to a line each: Laya fits all the options into about 192 tokens, and anything longer gets cut. Then freeze the wording, because the fine-tuned model learns this exact text.
2. Get a few thousand inputs
Production logs if you have them. If not, have a cheap model write them from your description in lots of styles, the way generate.py does. Never show it your test cases.
3. Have a teacher label them, with probabilities
A hosted decision model, code that already makes the decision, or people. Write one JSON line per input:
{"state": {"body": "I was charged twice for order 4412."},
"questions": {"department": {"type": "choice", "instructions": "Which team should handle this ticket?",
"criteria": {"billing": "charges, invoices, refunds", "shipping": "delays, tracking", "returns": "exchanges, damaged items"}}},
"gold": {"department": {"probabilities": {"billing": 0.91, "returns": 0.06, "shipping": 0.03}}}}
Only have hard labels? Use "expected": {"department": "billing"} instead of gold, or a plain CSV with a text column and a label column.
4. Train
.venv/bin/python finetune.py data/train.jsonl checkpoints/mine --epochs 3
# or from a CSV
.venv/bin/python finetune.py tickets.csv checkpoints/tickets --text-column body --label-column team \
--question-id team --instructions "Which team should handle this ticket?"
It prints how many optimizer updates the run will make before it starts. Aim for a few hundred; with a small dataset, add epochs. Stop anything else that's using the GPU while it trains.
5. Test it on data it has never seen
Not the calibration slice. Your own test cases, or fresh game seeds. Run enough of them that one lucky draw doesn't fool you.
6. Use it
import laya
agent = laya.load("checkpoints/mine")
print(agent.predict(state, questions)["answers"])
Or serve it over HTTP, Jev-style, with LAYA_EXTRA_MODELS='{"mine": "/path/to/checkpoints/mine"}' laya-serve.
Or rerun mine
cd pacman
../.venv/bin/python make_dataset.py --games 300 --out data/train.jsonl
../.venv/bin/python ../finetune.py data/train.jsonl ../checkpoints/pacman --epochs 2
../.venv/bin/python play.py --games 50 --players random teacher english ../checkpoints/pacman
The Pigeonhole steps, the DAgger round and every number above are in the README and RESULTS.md.
I spent two posts ranking models on how good they are out of the box. That turns out to be the wrong question for a model like Laya. A 421M model that answers in 100 ms isn't supposed to know your business. It's supposed to learn it cheaply from something that does, and then run anywhere, privately, for free.
The out-of-the-box score tells you where it starts. The question that matters is how much it costs to get it to where you need it. For me, the answer was an evening, a six-year-old desktop and twenty-seven cents.
Feel free to connect or reach out if you fine-tune it on your own decisions. I'd especially like to hear what you used as the teacher.
