One thing up front, before any numbers.
Having Claude Code or Codex write your training data is a great way to learn how fine-tuning works and to get a first model running. It is not how you should train the model you ship. For a real deployment you will be much more accurate with a golden dataset: real inputs from your own system, with labels a person has checked. Everything below is about the first part.
Where we left off
In the last post I fine-tuned Laya, the open-source decision model, on my five Pigeonhole classifiers. The teacher was Jev, a hosted decision model, through OpenRouter. A cheap model wrote 1,971 inputs, Jev labelled them, Laya learned to copy Jev, and it went from 31 to 44 of 50 test cases. The API bill was $0.27.
Twenty-seven cents is nothing. But it still means an account, an API key, a balance and a script that calls someone else's endpoint. Meanwhile, a lot of people reading this already pay every month for Claude Code or Codex, and mostly use it to write code.
So: can the coding agent you already pay for be the whole teacher? Writing the inputs, labelling them, and handing them to the trainer, without an API key anywhere?
What the agent is asked to do
One prompt per batch. It gets a template's description and its questions, exactly as Laya will see them, and never its test cases. Then:
- Write 25 realistic inputs in a given style (blunt one-liners, long rambling emails, typos, sarcasm, non-native English, cases right on a boundary).
- For each input, answer every question with honest probabilities. Spread them when an input is genuinely ambiguous, be confident only when it's clear.
Here's one row Claude Code wrote for the support-triage classifier, with its own labels:
this is the THIRD time i get the wrong shoes im done with this store refund me now or i dispute with my bank
| question | Opus 5.5's answer |
|---|---|
| department | returns 0.60, billing 0.30, other 0.06, shipping 0.04 |
| return reason | wrong item 0.75, wrong size 0.15, other 0.07 |
| angry | 0.98 |
| urgency | needs attention now 0.70, today 0.28 |
That's the shape Laya learns from best: not just "returns", but "returns, though there's a real chance billing should own the refund".
80 of those calls per agent gives 2,000 inputs, 400 per classifier. That's the same size as the Jev run, so the comparison is fair.
The repo does the rest
Three small scripts sit between the agent and the trainer:
teach.pyruns the agent headless (claude -porcodex exec), four calls at a time. It checks every reply against the template, fills in options the agent left at zero, drops anything malformed, and throws away any input that looks too much like one of the 50 tests, even though the agent never saw them.split.pyholds back 10% of each classifier's rows before training, so the trainer has data it never trained on to report against.finetune.pyand the scoring harness are the same ones from the last post.
python teach.py --agent claude --model claude-opus-5-5
python split.py data/claude/labelled.jsonl
python ../finetune.py data/claude/train.jsonl ../checkpoints/pigeonhole-claude --epochs 3 \
--eval-data data/claude/heldout.jsonl
python harness.py ../checkpoints/pigeonhole-claude -v
Or don't type any of that. Open your agent in the repo and say:
Read pigeonhole/README.md. Teach Laya the five templates using yourself as the teacher:
run teach.py with your own CLI, split the data, fine-tune, and score it on the tests.
One detail matters for cost. By default, every claude -p call sends Claude Code's whole system prompt: tool definitions, environment, skills. That's about 30,000 tokens before your prompt even starts. teach.py replaces it with a one-line prompt and turns tools off, which brings each call down to about 1,600 input tokens. Codex doesn't let you replace its built-in prompt, so every Codex call carries about 20,000 tokens, most of them cached.
The results
Same 50 Pigeonhole test cases as both previous posts, all trained and scored on my Mac mini (M4 Pro):
| Issues | Leads | Moderation | Support | Total | ||
|---|---|---|---|---|---|---|
| Laya, untrained | 6 | 5 | 4 | 10 | 6 | 31 / 50 |
| Laya, taught by Jev | 7 | 9 | 6 | 10 | 10 | 42 / 50 |
| Laya, taught by Claude Code (Opus 5.5) | 5 | 9 | 6 | 10 | 10 | 40 / 50 |
| Laya, taught by Codex (Sol 6.1) | 5 | 9 | 5 | 10 | 9 | 38 / 50 |
(The Jev-taught model scored 44 on my RTX 3050 and 42 on the Mac. Same data, same settings, slightly different arithmetic. That two-case swing is the noise floor here.)
With no API key, Laya picked up nine more test cases from Opus 5.5 and seven from Sol 6.1. The Opus-taught model is within that noise floor of the Jev-taught one. The Sol-taught one is a little further back.
Most of the gap is email routing (all of it for Opus, half for Sol), and both agent-taught models fail the same five emails there. Three of those also trip the Jev-taught model: a declined-card notice filed as junk, a recruiter's contractor pitch filed as sales, and an outage the tests expect not to be flagged urgent, which no model I've tried passes. The two only the agent-taught models miss: "What does the team plan cost for 40 seats?" goes to junk, and "How do I rotate an API key?" goes to other.
Both agents wrote noticeably more junk mail than Jev labelled: 20% and 17% of their email inputs, against Jev's 13%. My first guess was that they'd written scam emails that look like real billing mail and labelled them junk. I checked, and they hadn't. So I can tell you where it went wrong, not why.
What it cost
| Claude Code, Opus 5.5 | Codex, Sol 6.1 | |
|---|---|---|
| Inputs written and labelled | 1,997 | 1,998 |
| Wall time | 20 min | 26 min |
| Input tokens | 124,085 | 1,579,913 (83% cached) |
| Output tokens | 472,084 | 336,803 |
| The same work on the API | $10.47 | ~$4.04 |
| What I paid | nothing extra | nothing extra |
Opus 5.5 wrote more per input than Sol 6.1, and output is the expensive part, so it would have cost about two and a half times as much on the API. On a subscription, the number that matters isn't dollars, it's how much of your allowance it eats.
Codex on the $20 plan didn't make it in one go. After 64 of its 80 calls, it hit the plan's five-hour limit, and every support-triage batch failed. ChatGPT gives out the occasional free limit reset, so I used one, which put both the five-hour and weekly meters back to zero, and reran the 16 missing batches. Those 16 batches alone used 28% of the five-hour window and 4% of the week. Scaled up, one 2,000-input run needs about one and a half five-hour windows and a fifth of a week on Plus. Doable, but not "between two other tasks".
Claude Code, on my $100 Claude plan, finished in one go. That isn't a like-for-like comparison: it's five times the price of the Codex plan.
On the API, the Jev route cost $0.17 for the same job ($0.10 for a cheap model to write the inputs, $0.07 for Jev to label them). A hosted decision model is simply cheaper per label than a frontier coding model writing paragraphs of JSON. The subscription route doesn't win on cost per token. It wins on having no API key, no account and no new bill.
How the agents label
Going in, I expected a teacher that writes its own inputs to be overconfident, because it knows what it meant each one to be. Here's how often each teacher was near-certain:
| teacher | average top probability | labels at 0.99 or above | labels under 0.7 |
|---|---|---|---|
| Jev | 0.86 | 26% | 18% |
| Opus 5.5 | 0.84 | 13% | 21% |
| Sol 6.1 | 0.89 | 32% | 14% |
The worry was wrong for Opus 5.5: it hedged more than Jev, and half as often said "certain". Sol 6.1 was the most decisive of the three. On these tests the more cautious teacher produced the better student, but with one run each and a two-point gap, I wouldn't hang a rule on it.
Is it fair?
- The tests stayed unseen. Neither agent saw them, and anything too close to one was filtered out. Three inputs were.
- The held-out slice isn't a score. It was labelled by the same agent, so it measures how well Laya copied its teacher, not whether the teacher was right. Laya's agreement with Opus went from 0.61 to 0.70 on it; with Sol, 0.61 to 0.72. Only the 50 tests are the score.
- One run each. Two-case gaps are ties.
- The agents wrote the data and the answer key together. That's the point of the shortcut and also its weakness: there's no second opinion anywhere in the loop. Which brings me back to the top of this post.
So when would I use this?
To learn. It's the fastest way I know to get from "I have a classifier idea" to "I have a fine-tuned model and a number", using only tools you already have open. You see the whole pipeline (inputs, teacher, training, held-out data, a real test set) in an afternoon, and you see where it breaks.
Then I'd throw the training data away. The model you put in front of real users should learn from real inputs: production logs, support tickets, the emails you actually get, labelled or at least checked by someone who knows the job. That's your golden dataset, and it beats anything a model can imagine about your users. The agent is still useful there, for drafting labels a person then reviews, or for writing the edge cases your logs don't have yet.
The code is in bnap00/laya-fine-tuning: pigeonhole/teach.py, the generated data from both agents with every call's token usage, and all of these numbers in RESULTS.md.
Feel free to connect or reach out if you try this with your own decisions, and especially if you build the golden dataset afterwards. I'd like to know how far the real data moves the score.
