Two weeks ago I raced Jev against Laya on my own desk. Jev, hosted on OpenRouter, got 93% of my test cases right. Laya, running on my GPU, got 54%, but three times faster and for free.
Then the field got crowded. On October 1, Cloudflare released Clef and Clef-flash, open decision models (Apache 2.0) that speak Jev's format. On October 6, OpenAI opened a Decisions API running GPT-6 Luna. Same idea in both: you send text and questions, and you get probabilities back instead of prose.
So, same race, five runners.
Adding them took no new code
I expected to write two new adapters: one for Cloudflare's Workers AI and one for OpenAI's API, which names things differently (predicate instead of noul, lists instead of maps).
Then I checked OpenRouter. It already serves both on the same Decisions endpoint Pigeonhole uses for Jev, in the same format, on the same key:
model:
runtime: cloudflare/clef # or cloudflare/clef-flash, openai/gpt-6-luna-decisions
Pigeonhole only lets decision models answer classifications (a chat model's "0.95" means nothing), so the real change was adding them to that list. Everything else worked as it was.
The race
The test: the same 74 labelled cases as last time. Every test case from Pigeonhole's five template pipelines, plus the Reddit self-promotion detector. Sent one at a time, after a warm-up pass so nobody pays for a cold start.
The setup: Laya on my RTX 3050, as before. The other four through OpenRouter, from India. I added a pigeonhole bench command for this, which runs every pipeline's tests on any list of models and reports accuracy, latency and cost.
Laya, again, is untrained: the stock checkpoints, never shown any of these tasks.
| Jev | GPT-6 Luna | Clef | Clef-flash | Laya (GPU) | |
|---|---|---|---|---|---|
| Correct | 68 of 74 (92%) | 67 (91%) | 65 (88%) | 64 (86%) | 39 (53%) |
| Typical response | 470 ms | 380 ms | 700 ms | 630 ms | 100 ms |
| Slow responses (p95) | 560 ms | 560 ms | 940 ms | 730 ms | 111 ms |
| Cost for all 74 | $0.0027 | $0.0065 | $0.0026 | $0.0014 | $0 |
| Cost per million | about $37 | about $88 | about $35 | about $19 | electricity |
Jev still wins on accuracy, but only just. GPT-6 Luna is one case behind and the fastest hosted model. Clef is three cases behind Jev at a slightly lower price. Clef-flash is four behind at half the price. Laya is still the fastest and the only free one, but it's in a different league on accuracy.
Pipeline by pipeline:
| Pipeline | Jev | GPT-6 Luna | Clef | Clef-flash | Laya |
|---|---|---|---|---|---|
| Moderation | 8 / 10 | 10 / 10 | 7 / 10 | 7 / 10 | 9 / 10 |
| Email routing | 9 / 10 | 8 / 10 | 8 / 10 | 8 / 10 | 5 / 10 |
| Support triage | 10 / 10 | 10 / 10 | 10 / 10 | 10 / 10 | 7 / 10 |
| Issue labelling | 10 / 10 | 9 / 10 | 10 / 10 | 10 / 10 | 4 / 10 |
| Lead qualification | 9 / 10 | 8 / 10 | 9 / 10 | 8 / 10 | 4 / 10 |
| Reddit self-promotion | 22 / 24 | 22 / 24 | 21 / 24 | 21 / 24 | 10 / 24 |
Moderation is the interesting one again. Last time, Laya beat Jev on it. This time GPT-6 Luna got all ten, and both Clefs were the worst of the five on it. Every other pipeline was close among the four hosted models.
And in Hindi?
The same ten support-triage tickets, translated into Hindi:
| Jev | GPT-6 Luna | Clef | Clef-flash | Laya | |
|---|---|---|---|---|---|
| Correct | 10 / 10 | 10 / 10 | 10 / 10 | 10 / 10 | 6 / 10 |
| Typical response | 476 ms | 278 ms | 695 ms | 631 ms | 51 ms |
Every hosted model got all ten, and GPT-6 Luna was the fastest of them. Untrained Laya got six, at a tenth of the time. Hindi didn't trip up any of the hosted models.
The question GPT-6 Luna wouldn't answer
The first time I ran it, GPT-6 Luna scored 3 out of 10 on issue labelling.
It wasn't getting answers wrong. Six of the ten calls failed outright: "OpenAI refused to answer question is_good_first_issue." That question asks whether an issue is small and self-contained enough for a first-time contributor. The test inputs are just a one-line title, like "Crash on startup when config file is empty". GPT-6 Luna apparently decided that wasn't enough to judge, and refused.
Refusing is fair. An honest "I can't tell" beats a made-up 0.5. The problem was what happened next: OpenRouter fails the whole call when one question is refused. Pigeonhole sends every question for an input in one request, because that's what makes it cheap and fast. So one refused side question also lost the answer the test was actually checking, which was whether the issue is a bug or a feature.
The fix: when a question is refused, drop it and ask the rest again. The refused question comes back marked as unanswered, and everything else gets an answer. Issue labelling went from 3/10 to 9/10.
The fix isn't free, though. A refused question costs a second round trip, so that pipeline's typical response roughly doubled, from about 380 ms to 750 ms. Batching questions is still the right call, but now I know a batch needs a plan for one bad question.
What I didn't expect
The speed claims didn't survive the trip. Cloudflare says Clef-flash answers in about 39 ms, far faster than Jev. Through OpenRouter it took about 630 ms, slower than Jev. OpenRouter lists a third-party host serving it, I'm in India, and every request goes through OpenRouter first. To get the advertised numbers, you'd probably need to call Workers AI directly, near the model. I haven't tried that yet.
The price lists didn't either. Cloudflare and OpenAI publish per-token prices on their own APIs that are much lower than what OpenRouter charged me. Every cost in this post is what OpenRouter actually billed. If cost is your main reason to pick one of these models, check the direct API.
My own templates had a bug. Some of Pigeonhole's templates describe options in YAML like { what: The item arrived broken, scuffed, torn or otherwise defective. }. In that form, a comma ends the value. So the model was being told "damaged" means "The item arrived broken", followed by two empty keys called "scuffed" and "torn or otherwise defective.". Two templates, support triage and issue labelling, had descriptions cut short like that, in the last post too. I quoted the text, re-ran both pipelines and the Hindi set on every model, and corrected the last post. The fix moved only two results: Laya gained two cases on support triage, and Jev went from 9 to 10 out of 10 in Hindi. Pigeonhole's linter now warns when an option's text has been split like this.
The numbers move. Run the same suite twice and a model gains or loses a case. Jev scored 69 two weeks ago and 68 now; Laya scored 41 then and 39 now. Treat a one- or two-case gap as a tie.
Is it a fair race?
Still no, and the bias is the same as last time. Pigeonhole's compiler writes each pipeline's test cases, dry-runs them on Jev, and rewrites whatever Jev gets wrong. Every spec here was tuned on Jev before the others ever saw it. That GPT-6 Luna and Clef came this close anyway says a lot about them.
Laya is also still untrained. Its authors say fine-tuning on your own examples is where it improves, so its 53% is where it starts, not its best. Finding out how far it climbs is the next post.
So which one?
- Jev if you want the most right answers in English and are happy with half a second.
- GPT-6 Luna if speed matters among the hosted models, or your inputs aren't in English. Watch for refusals on questions the input can't answer.
- Clef-flash if cost per request matters most. It gave up four cases to Jev for half the price.
- Clef if you want open weights you could host yourself later. On OpenRouter it doesn't beat Jev on anything except a slightly lower price.
- Laya if data can't leave your machine, or you need answers in 100 ms. Untrained, it's not there yet. Trained, I'm hoping it is (see below).
Or skip my opinion and race them on your own pipeline:
pigeonhole bench --models jev-latest,openai/gpt-6-luna-decisions,cloudflare/clef-flash
That's the point of having tests: when a new model comes out, you can see how it does on your own pipelines before you switch.
Try it
You need Docker and an OpenRouter key. The one key covers Jev, Clef and GPT-6 Luna.
git clone https://github.com/bnap00/pigeonhole && cd pigeonhole
make up
Then set runtime: cloudflare/clef (or any model above) in a pipeline's spec, or run pigeonhole bench to compare them. The full results, including every pipeline, are in docs/comparing-models.md.
It's still a weekend project: one instance, one admin token, no TLS. It's not production software, but it's fun to try.
Next: training Laya
Every Laya number in both posts is from a model that has never seen a support ticket, a GitHub issue or a Reddit post. That's the least fair way to judge it. Its authors report fine-tuning taking it from 36% to 77% on their own benchmark, and it runs in 100 ms on a six-year-old desktop for free.
So that's the next experiment. I'm fine-tuning Laya on these pipelines' own examples on the same RTX 3050, then racing it against Jev, GPT-6 Luna and Clef on the same 74 cases. If it closes most of that nearly 40-point gap, you'd get a classifier that's close to the hosted models, four to seven times faster, free to run and fully private. That would be a big deal. The next post will show whether it does.
A month ago there was one hosted decision model I knew of. Now there are four, plus one I can run on my desk. That's good news for anyone who has been asking chat models for confidence scores. And it's why I keep saying a classifier needs a test suite: models keep changing and new ones keep arriving, so measure them on your own data.
If you race them on your own data and get different results, I'd like to hear about it.
