Agentic Research

A Reality Check on Decision Models: Why They Seem Miraculous Online but We Measured Only 54%: A Full Comparison of JEV / LAYA / KEV / CLM-8B and a Deployment Formula

2026/09/2779 min readBryan Chan閱讀中文原文
TopicsSystem OneDecision ModelsJEVLAYAKEV

The One-Sentence Version

The value of this class of models lies in "speed × cost × gateability", not in "being smarter". The officially advertised 9x latency advantage and near-zero marginal cost are real; the claim of "more accurate judgment" is the exaggerated part. After running our own hold-out set to completion, the conclusion is simple: strip away "labeled data + a gating threshold + a human fallback", and it is just a 54% raw classifier; apply this formula, and it can absorb 50-70% of the mechanical judgments in a system, at close to zero cost.


1. First, Understand What It Actually Does Through Six Examples

Before discussing architecture, let us look at what problems it is actually used to solve. The six scenarios below are all typical uses of this class of model. Note what they have in common: a fixed set of options, a very short answer, and high volume.

Example 1 | Which team should this support ticket go to?

Input (a customer message):

"The shoes arrived two weeks late, and the size was wrong; also, I noticed my credit card was charged twice."

Question: Which team should this ticket go to? Candidates: returns/exchanges / logistics / billing

Model output: returns/exchanges 0.47 | logistics 0.28 | billing 0.25

System decision: the highest probability is only 0.47, below the threshold → do not auto-assign, route to human review

👉 The most important thing about this example is not whether it got the answer right, but that it says on its own that it is not confident. Three issues are mixed into one passage (late arrival, wrong size, duplicate charge), and it did not force a single choice and call it done.

Example 2 | Three hundred emails a day, which project should each be filed under?

Question: Which project does this email belong to? Candidates: due diligence / property bidding / financial reconciliation / other

Model output: property bidding 0.91

System decision: above the threshold → auto-file and notify the owner

👉 This judgment used to be sent to a large language model, waiting half a second to two seconds per email and paying each time; now it is 20 milliseconds per email with a marginal cost of zero.

Example 3 | Are these two notes talking about the same thing?

Question: Are note A and note B duplicates? (yes / no)

Model output: yes 0.88 → added to the auto-merge candidates Model output: yes 0.55 → flagged for human confirmation, no auto-merge

👉 Deduplication is a classic "one mistake will not kill you, but many mistakes are annoying" task, which makes it a perfect fit for this kind of model.

Example 4 | Can this document be sent out directly?

Question: Does this text contain customer personal data (ID number, passport number, bank account number)?

Model output: yes 0.97 → blocked, requiring redaction before resending

👉 The guardrail scenario has one key difference: a false positive is better than a miss. So its threshold is lowered in the opposite direction (for example, block at 0.3), the exact opposite of the logic in Example 2.

Example 5 | How severe is this incident?

Input:

"The production database is completely down, every API endpoint returns 500, and no user can log in."

Question: Rate it on the scale trivial → low → medium → high → critical

Model output: critical 0.69 → triggers the on-call escalation process

👉 Note that this is a "scoring" type question, which is also the weakest area for this class of model (measured data follows in Section 5), so in practice it should be used only for relative ranking, not treated as an absolute score.

Example 6 | Pick the best of four candidate plans

Flow: first have a generative model write 4 plans → then use the decision model to pick the best one

👉 This is a "do not generate, only select" use case, suited to teams that already have a high-quality generator. It does not write code or articles; it only answers "which of these is best".

The Common Structure Across the Six Examples

ExampleQuestion TypeModel ConfidenceSystem Action
Ticket dispatchChoose one of three0.47 (low)Route to human
Email filingChoose one of four0.91 (high)Auto-file
Note deduplicationyes / no0.88 / 0.55Auto / pending confirmation
Safety guardrailyes / no0.97Blocked (threshold deliberately lowered)
Incident ratingOrdered score0.69On-call escalation
Plan re-rankingChoose one of many-Use the highest-scoring candidate

See the pattern? This class of model never "gives you an answer"; it gives you an answer with a confidence score, and the system then decides based on that confidence whether to "execute automatically, verify further, or route to a human". That is its true product form.


2. Why a New Category Emerged Within Two Weeks

The concept comes from Kahneman's "Thinking, Fast and Slow" and its "System 1": intuitive, fast, without deliberation. In implementation terms, it refers to models that do not generate token by token, but output a structured judgment directly in a single forward pass.

All models of this kind share three primitives:

PrimitiveSemanticsCorresponding Example Above
noulProbability that a proposition is true or falseExample 3 deduplication, Example 4 guardrail
choiceChoose one from a fixed candidate set, with a probability distributionExample 1 ticket, Example 2 email, Example 6 re-ranking
scoreScore against an ordered scaleExample 5 incident rating

The timeline is quite tight:

DateEvent
2026-09-15JEV released (TypeSafe AI, closed-source API, $40M seed round announced the same day)
2026-09-17KEV open-sourced (based on Qwen3.5/3.8 plus LoRA and a decision head)
2026-09-18LAYA open-sourced (ModernBERT / mmBERT, supporting 100+ languages)
2026-09-23CLM-8B open-sourced (contrastive learning route, encoding state and action separately)

The pain point they address is real: what is truly high-frequency in today's automation systems is not "write a report", but the six kinds of small judgments above. These judgments have fixed options, short answers, and high volume, yet they are all wrapped into prompts and thrown at a large language model, making them slow, expensive, and unauditable.


3. The Four at a Glance

DimensionJEVLAYAKEVCLM-8B
FormClosed-source cloud APIOpen source Apache-2.0Open source Apache-2.0Open source Apache-2.0
BackboneUndisclosedModernBERT / mmBERT (0.32B)Qwen3.5 / 3.8 + LoRA + decision headFrozen Qwen3-8B + 20M projection head
SizeUndisclosed0.32B0.8B / 4B / 9B / 27B8B
Training methodRLCDRLCD (proper scoring rules)RLCD-style + in-house decision datasetContrastive learning (bidirectional InfoNCE)
Latency70-500ms (including network)~33msTens of msFastest (officially claimed up to 9x)
Local deployment✗✓✓✗ (requires NVIDIA + vLLM)
Fine-tuning✗✓✓✓ (head is only 75MB)
MultilingualMostly English100+Mostly EnglishMostly English

An underrated engineering legacy: the four API contracts are isomorphic. Change a URL and you change the model, with no change to business code, which turns "start with the cheap one, and switch to the expensive one if it is not enough" into a one-line config change.


4. Between the Marketing and Our Measurements, Six Differences in Definition

The miraculous numbers you see online and the 54% we measured are actually not contradictory. The difference comes mainly from six issues of "definition":

#Definition DifferenceThe Reality
1Different comparison targetIt is compared against "using an LLM to do classification" for cost and latency, not against human experts for accuracy
2A self-selected test setSome leading numbers are only on the vendor's own development set; independent third-party evaluations on the same class of task have produced clearly lower results
3Only the "confident portion" is reportedThe post-gating 91.7% is often presented as overall accuracy, but it covers only about one third of the questions (see Section 5.4)
4Fine-tuned results are cited without mentioning the un-tuned baselineThe official documentation states it plainly: the un-tuned baseline is 0.362, and after fine-tuning it is 0.766. Most re-tellings cite only the latter
5The headline keeps only the single brightest numberHeadlines like "overtakes" or "sets a new SOTA" usually reflect a single metric
6Reading "the format will not be wrong" as "the judgment will not be wrong"What it guarantees is that the output format is constrained (it will always return only the options you gave), not that the judgment content is correct

Put more plainly: the marketing says "I only answer the questions I am sure about, and I am right 90% of the time"; your own test says "every question must be answered, and I am right 54% of the time". Both numbers are true; they just do not describe the same thing. Understanding that means you will not feel deceived, and you will have a better idea of how to use it.


5. Our Own Measurements: Method and Results

5.1 Method

ItemSetting
ModelsLAYA 0.3.20 (English and multilingual checkpoints), KEV-0.8B
Set oneGeneral 150 questions: 50 each of choice / score / noul
Set twoFinance 80 questions: Chinese-English translation proofreading (74 noul + 6 choice)
Set threeAbstention probe, 22 questions: 10 questions where "information is insufficient and the correct answer should be uncertain" + 12 clear questions, with options yes / no / uncertain
IsolationAll three sets are hold-out, never used for any training or tuning
ConditionsAll zero-shot (no fine-tuning), run in a local environment with no cloud GPU

5.2 Results

ConfigurationSetnAccuracyLatencyBrierECEOverconfident-wrong rate
LAYA EnglishGeneral 15015054.0%20ms0.4000.38930.0%
↳ by typenoul / choice / score50 each54.0% / 80.0% / 28.0%
KEV-0.8BGeneral 15015070.7%19ms0.1720.0872.0%
↳ by typenoul / choice / score50 each72.0% / 88.0% / 52.0%
LAYA EnglishFinance 808063.7%21ms0.2630.2415.4%
LAYA multilingualFinance 808067.5%15ms0.2390.20314.9%
KEV-0.8BFinance 808061.3%20ms0.2640.2289.5%
LAYA EnglishAbstention probe2250.0%17ms---
KEV-0.8BAbstention probe2236.4%17ms---

5.3 Six Findings

  1. The choice type is clearly stronger than the others: KEV 88%, LAYA 80%. Compared with the examples in Section 1, email filing and ticket dispatch are exactly its most mature use cases.
  2. Calibration is KEV's biggest verifiable advantage: Brier 0.172 versus LAYA's 0.400; ECE 0.087 versus 0.389; overconfident-wrong rate 2.0% versus 30.0%. This is consistent with the official claim that "every checkpoint carries a fitted temperature".
  3. On Chinese tasks, the multilingual checkpoint beats the English checkpoint (67.5% vs 63.7%), supporting routing by language.
  4. Abstention ability is the exact opposite: on ambiguous questions, LAYA's correct abstention rate is 50%, while KEV-0.8B is only 10%. The former tends to say "uncertain", while the latter tends to answer anyway. If you are building the kind of guardrail in Example 4, this difference matters more than accuracy.
  5. Both vendors are weak on the score type. When compared by argmax, KEV appears to lead (52% vs 28%), but when measured with continuous values the gap almost disappears (MAE 0.83 vs 0.90, with rounding accuracy both at 30%). This corresponds to Example 5: incident ratings can be a reference, but should not be executed directly as an absolute score.
  6. Latency is 15-21ms across the board, so local deployment is entirely viable with no cloud GPU required.

5.4 The Most Counterintuitive Table: After Adding Gating

ConfigurationFull CoverageThreshold 0.5 (coverage / accuracy)Threshold 0.7Threshold 0.9
LAYA EnglishGeneral 15054.0%79% / 63.6%61% / 71.4%
LAYA multilingualFinance 8067.5%98% / 67.9%84% / 73.1%
KEV-0.8BGeneral 15070.7%69% / 81.6%49% / 85.1%
KEV-0.8BFinance 8061.3%99% / 60.8%74% / 67.8%

How to read this table: taking KEV-0.8B on the general set as an example, when the threshold is set to 0.9, only 32% of the questions are executed automatically, but 91.7% of that 32% are correct; the remaining 68% are all routed to a human or a fallback process. This is how Example 1 works: better to do less than to do it wrong.


6. Why Accuracy Is Below Expectations: Four Causes

We ran two diagnostic experiments to locate the problem.

6.1 Diagnostic One: Re-testing with the Official Example Question

Taking the original example from the official documentation (the Example 1 in Section 1: shoes two weeks late, wrong size, credit card charged twice), we re-ran it locally:

ItemOfficial Example (4B class)Our Re-test (0.8B)
Dispatch resultreturns/exchangeslogistics
Needs human0.930.56
Output structureAll three primitives presentCompletely identical

Conclusion: the interface and flow are correct (structure, primitives, and probability distribution are all normal); the difference comes from model size: the 0.8B got the primary issue wrong on a "multi-issue" question. Notably, it gave itself only 0.45 confidence. In other words, simply adding gating would send this question into the "route to human" branch of Example 1, keeping the system level safe.

6.2 Diagnostic Two: Replacing Absolute Scores with Continuous Values

Section 5.3, point 5, already explained this: a substantial part of the score type's "52%" reflects argmax's preference for probability mass. After switching to continuous metrics, the two vendors are nearly even.

6.3 Four Causes, Ranked by Weight

#CauseHolds?Evidence
1Evaluation method is strict (full-coverage argmax)✅ The single largest factor+20pp after adding gating
2No domain fine-tuning✅Official same set: base 0.362 → fine-tuned 0.766
3Usage details (language, question type, instruction text)✅ PartiallyMultilingual +3.8pp; choice 80-88% vs score 28-52%
4Insufficient model size✅ Real variableOfficial same question: 0.8B got it wrong, 4B got it right
5Interface/flow misused❌ Does not holdThe official same-question re-test produced a completely identical structure

7. Four Real-World Uses and an Implementation Checklist

UseCorresponding ExampleUsable LevelNeeds Fine-Tuning?
Dispatch / routingExample 1, Example 2choice 80-88%No
Guardrail / reviewExample 4Depends on abstention ability; recall must be self-testedRecommended
Re-ranking (best-of-N)Example 6Requires dedicated fine-tuning and a GPURequired
Scoring / rankingExample 5Only suitable for relative ranking, not absolute scoresRequired

The shared prerequisites for all four uses: ① the same judgment repeats hundreds of times or more per day ② the options can be enumerated ③ labeled data exists (or can accumulate) ④ abstention to a human is possible.

The Threshold Is Not Guessed

The threshold must be computed from your own data, in just three steps:

  1. Collect 200-400 questions with real answers already labeled;
  2. Run the model on all of them, recording each question's "maximum probability" and "whether it was correct";
  3. Plot the "coverage vs accuracy" curve (this is the source of the table in Section 5.4) and pick a trade-off point you can accept.

The acceptance criterion is also just one sentence: for example, if you set "the automated portion's accuracy must not fall below 95%", then choose the corresponding threshold; if coverage at that threshold is only one tenth, then this scenario is not yet worth adopting.


8. The Value Formula: Not Smarter, but Better Value

Using Example 2 from Section 1, let us do a real calculation, assuming 2,000 emails per day need classification:

ApproachTime per emailTotal time per dayMarginal cost
Sent to a large language modelAbout 400-500msAbout 13-17 minutesBilled each time
Dedicated decision modelAbout 20msAbout 40 secondsNear zero

Add gating on top: about half the emails can be auto-filed (at roughly 85% accuracy), and the other half goes through the original process. That is the entire value: not judging better than a human, but making an adequate judgment a hundred times cheaper and twenty times faster.


9. A Side Note: Why Some People Use It to Play Games

This section is a bit of a digression, but I find it quite interesting, because the most widely circulated demonstration online is not support ticket dispatch, but "letting the model play a game".

First, Let Us See What Everyone Is Doing

The earliest wave of attention came from a Super Mario demonstration an engineer posted in a community. The picture is convincing: the model does not "read the screen"; it receives a set of numbers, the character's position, velocity, enemy distance, whether there is a pit ahead, and then picks one from a few fixed actions. That demonstration reported a response time of 51 milliseconds; it chose "jump right" with 73% confidence.

The official team also has a DOOM demonstration, with the same emphasis: "hundred-millisecond-level responses, enough to support real-time control".

Later, a Japanese engineer turned it into a laboratory you can actually play: reproducing World 1 of Mario himself while running four different "decision brains": a cloud model, a small local model, a LightGBM distilled from the model, and Laya. His conclusion was notably restrained: after adding search and safety assistance, all four could clear the stage; remove the assistance, and none could. He did not credit the model.

Another author went further: recording a "teacher model's" actions and training a decision tree (LightGBM), which then cleared all 100 runs in real-time mode, with each decision taking only 0.25 milliseconds. His observation was direct: if the input is not natural language in the first place (game state is inherently a pile of numbers), traditional machine learning may be handier.

Some in the community also did Snake: condensing the map size, the snake's head, its body, and food coordinates into a passage of text, feeding it to the model, and letting it choose among up, down, left, right. Others, demonstrating Laya, emphasize "local, free, 30 decisions per second", using lightweight games such as Breakout, Flappy Bird, and Tetris.

These demonstrations share one thing: the model is not the player; it is "the one holding the controller". Reading the screen, computing collisions, drawing animations, all of that is done by the game program itself; the model is only responsible for answering "which one next" at each instant. The middle layer that translates game state into text and then translates the choice into a button press is called a harness by the community. How well the harness is written often matters more to success than which model you pick.

We Also Ran a Small Experiment Ourselves

After seeing all this, I could not resist building the simplest possible version: a one-dimensional runner. There are two types of obstacles, low ones to jump and high ones to duck, with only three actions (run, jump, duck), asking the model once per tick.

First, a diagnostic of "distance × obstacle type", asking three times per cell:

Obstacle DistanceTypeShould ChooseModel Actually ChoseAverage Confidence
5 / 4 / 3 / 2 / 1lowjumpjump (all three correct)0.71-0.75
5 / 4 / 3 / 2 / 1highduckduck (all three correct)0.42-0.44
0lowjumpjump0.59
0highduckrun (wrong all three times)0.36

In other words, when the obstacle was not yet in its face, it actually answered correctly; its one clear blind spot was when the obstacle was already right in front and was the "duck" kind: it would choose "run", with only 0.36 confidence.

Then I tried another setting: letting its action take effect a few ticks late, to see how much response time mattered.

Action DelayAverage obstacles cleared out of 40All five levels clearedSingle-decision latencyAverage confidence
0 ticks (real time)2.00 / 523ms0.58
1 tick40.05 / 522ms0.56
2 ticks40.05 / 522ms0.56
3 ticks40.05 / 522ms0.56
4 ticks11.40 / 523ms0.60
5 ticks6.00 / 522ms0.60

This result looked counterintuitive at first: "slightly delayed" turned out to be far better than "zero delay". After thinking it over, it made sense, because actions can "queue up". The model had already answered correctly when the obstacle was still two or three cells away (jump when it should jump, duck when it should duck), and that action took effect one or two ticks later, just in time. "Zero delay", by contrast, forced it to decide at the exact moment the obstacle was in its face, which happens to be its one blind spot.

As for collapsing at four or five ticks of delay, that is easy to understand too: the action was taken too early and had already expired by the time the obstacle actually arrived.

Honestly, this little game's rules are too clean to prove any big conclusion. At most it shows two things: this route works; and in a game setting, what is genuinely hard is usually not "is the model accurate" but "when do you ask it, and when does its answer take effect". Ask too late and even an accurate answer arrives too late; ask too early and the action expires.

So, Is Game Control Actually Suited to This Class of Model?

SituationSuitable?
Turn-based, or a decision interval of tens of milliseconds or moreSuitable
Game state can be read directly from inside the program (no need to look at the screen)Suitable
Needs a judgment every frame, and state can only come from the screenNot very suitable; rules, tree models, or reinforcement learning are usually less trouble

Most of the more mature demonstrations actually do the same thing: rules or a strong model handle tactics and safety boundaries, and the decision model makes only small judgments at branch points. This is the same point made in the earlier sections: finding the right place matters more than finding the right model.


10. Three Iron Rules

  1. Uncertainty must be preserved. A judge that dare not say "uncertain" is more dangerous than one that makes mistakes. The first acceptance item for any adoption should be the ability to "route to a human on low confidence", as in Example 1.
  2. The amount of evidence sets the ceiling, not the recipe. Our experience is this: when a hold-out has only thirty-odd positive examples, a single question can cause several percentage points of swing. Build an evaluation set of 400+ real questions first, then talk about model selection and fine-tuning.
  3. Replaceability takes priority over performance. This class of model is still iterating on a weekly basis. The isomorphic API contract provides rare freedom to swap, and any design should preserve the ability to "switch vendors next week".

One Question to Ask Yourself Before Adopting

"If this question is answered wrong, what is the cost?"

  • The cost is bearable (wrong category → retry, as in Example 2, Example 3) → usable
  • The cost is unbearable (compliance, money, external documents, as in Example 4, Example 5) → use for first-pass screening only; a human final review cannot be removed

Appendix: Notes on Method and Sources

  • All measured numbers in this article are zero-shot measurements taken locally on 2026-09-27. The question sets and evaluation protocol differ from each project's official ones, and cannot be compared directly with official benchmarks.
  • "Gating" means using the model's maximum output probability as a threshold, adopting automatically only when it is above the threshold, and handing the rest to a human or a fallback process.
  • Official numbers are quoted from each project's public documentation; third-party evaluations are quoted from public independent analyses.
  • The game experiment in Section 9 is a self-built one-dimensional runner environment (fully deterministic rules), used to observe the effect of "when to ask" and "when the action takes effect" on outcomes. It is not any official benchmark, nor does it represent the model's performance in real games.
  • This article is a technical evaluation and does not constitute any investment or procurement advice; all numbers are subject to the latest public documentation of each project.