Agentic Research

The "Non-Text-Generating Model": Jev and the New System One Category, and How Agent Architecture Changes When AI Only Answers Multiple Choice

2026/09/1978 min readBryan Chan閱讀中文原文
TopicsJevTypeSafe AISystem OneRLCDAgent Architecture

The One-Sentence Version

Jev is a frontier model that generates no text at all. It only answers the multiple-choice questions you give it, returning a calibrated probability distribution, processing hundreds of questions in parallel in a single forward pass, with about 100ms latency and a price of $0.042 per million input tokens.

This is not an incremental improvement. It represents a new category: the System One model, purpose-built for fast, structured, calibratable decisions, cleanly separating "judgment" from "writing". For teams building agent architectures, this is a foundational component worth seriously evaluating.


Project Overview

ItemValue
MakerTypeSafe AI (San Francisco, founded 2024)
Release date2026-09-15 (exiting stealth after two years under cover)
Funding$40M seed, led by DCVC; valuation around $200M (per Forbes)
FounderDiogo Almeida (CEO, co-inventor of RLHF / InstructGPT, ex Google Brain / OpenAI)
Team backgroundOpenAI, Google Brain, Meta FAIR, Stripe, Airbnb, Plaid, Docker
First modelJev (version jev-1.13.0, aliases jev-latest / jev-preview)
Model categorySystem One Model (generates no text, only outputs typed decisions)
Question typeschoice (≤255 options), score (2-10 ordered levels), noul (yes-no probability)
Latency70-500ms, mostly ~100ms
Pricing$0.042 / million input tokens, output is free
Rate limits250,000 tokens/s, 1,200 requests/min
AccessTypeSafe SDK (early access via waitlist) or Vercel AI Gateway (typesafe-ai/jev)

Why This Matters: A Counterintuitive Model

For the past three years, the industry's intuition has been "a bigger model → stronger generation → a better agent". Every lab has been chasing the next GPT, the next Claude, competing on longer context windows, faster generation, and broader multimodal capability. In that narrative, "a better model" equals "a model that can generate more and better text".

Jev takes the exact opposite route:

Deliberately give up text generation in exchange for the speed, cost, and determinism of structured decisions.

It cannot write replies, generate code, summarize, or explain its reasoning. But on high-frequency, high-volume decisions that require judgment, such as "which department does this support ticket belong to?", "what is the toxicity level of this comment?", or "is this transaction suspicious?", it is two orders of magnitude faster and several times cheaper than a frontier LLM, and its structured output error rate is zero (by construction, because there is no generation process at all, so the answer necessarily falls within the schema you defined).

The implication for agent architecture is deep: not every intelligent behavior requires text generation. A large share of an agent's internal "judgment" steps, routing classification, entity selection, guard conditions, priority scoring, are really just classification, scoring, and boolean decisions. These are precisely the territory of System One models, not the sweet spot of LLMs. Using an LLM for these is like using a cannon to kill a mosquito: slow, expensive, and with probabilities that cannot be calibrated.


What Jev Is: The Three Question Types of a System One Model

Jev's name comes from Daniel Kahneman's Thinking, Fast and Slow. System 1 is the fast, intuitive, pattern-matching thinking system responsible for most everyday judgments; System 2 is the slow, deliberate, logical reasoning system. Jev deliberately covers only the System 1 dimension, leaving System 2 to the LLM.

The Three Question Types

Question TypeInputOutputTypical Use
choiceA state description + ≤255 optionsA probability distribution over the optionsClassification, routing, entity selection, intent recognition
scoreA state description + 2-10 ordered levelsA probability distribution over the levelsToxicity scoring, urgency, priority, satisfaction
noulA state description + a yes/no questionP(yes) probability valueGuard conditions, boolean judgments, trigger rules

A single API request can contain hundreds of questions, all answered in parallel in one forward pass. This is a fundamental contrast with the token-by-token autoregressive generation of an LLM: Jev has no "generation" process; what it does is encode the state → score in parallel → output a probability distribution. The whole process is a deterministic feedforward computation, with no sampling, no decoding, and no text assembly.

A concrete example: suppose you have a support ticket classification scenario where you need to determine which department the ticket belongs to (choice, 5 options), its urgency (score, levels 1-5), and whether it needs immediate escalation (noul). With an LLM, you need to construct a prompt → wait for autoregressive generation → parse the JSON → validate the schema → retry on failure. With Jev, you construct one request containing three questions, and about 100ms later you get three calibrated probability distributions, with no parsing layer and no retry logic.

Calibrated Probabilities

Jev's core selling point is not just "fast", but calibration: when it says 90%, it is right about 90% of the time. This is not accuracy; accuracy is whether a single judgment is right or wrong. Calibration is the long-run match between probabilities and actual outcomes.

An intuitive analogy: a well-calibrated weather forecast model, on the days it says "there is a 70% chance of rain tomorrow", will indeed see rain on roughly 70% of those days over the long run. This does not mean it predicts correctly every time, but those 30% of days without rain are the uncertainty it honestly expressed, not systematic overconfidence.

For agent architecture, calibration is the real basis for decisions: you can set thresholds based on probability, with high confidence (>95%) executing automatically, medium confidence (70-95%) routed to an LLM for a second check, and low confidence (<70%) escalated to a human. This is what LLM logprobs have long failed to deliver: in decision scenarios, LLM token probabilities are generally overconfident, saying 99% for things that are only 70% accurate. Using LLM logprobs for confidence gating is like using an uncalibrated thermometer to control a thermostat: roughly the right direction, but not precise enough.


Why 200x Faster and 5x Cheaper

TypeSafe officially claims Jev is 193.6x faster and 4.6x cheaper than frontier LLMs (a company claim, an order-of-magnitude estimate, not yet independently verified; SiliconANGLE reported another comparison showing 445x cheaper, with the difference depending on the baseline). But even discounted, the order-of-magnitude difference is real. The reason is at the architectural level:

First, no autoregressive generation. LLM latency comes mainly from token-by-token decoding: a 500-token reply requires hundreds of forward passes, each running the entire model. Jev completes all questions in one forward pass, with latency fixed at 70-500ms, nearly independent of the number of questions. When you need 100 answers on the same prompt, LLM cost grows linearly with the number of questions, while Jev's cost is nearly unchanged.

Second, a parallel sampler. TypeSafe designed a dedicated parallel sampler that expands hundreds of typed queries from a single prompt at once. This is closer to an extreme optimization of batch inference than traditional chat serving. Judging from the numbers of the open-source replica OpenJev (10 questions in 28ms on an H100), this parallelization has already been fully realized at the hardware level.

Third, free output. The pricing model is based entirely on input tokens ($0.042/MTok), with output at zero cost. Compared with GPT-5.6 Luna at $0.20-10/MTok input plus considerable output fees, Jev's cost advantage in high-throughput decision scenarios is structural. TypeSafe's website gives a comparison: $0.39 per 1,000 workflows versus $3.31 for OpenAI GPT-5.6 Luna, a gap of about 8.5x.

Fourth, a smaller effective model. Although TypeSafe has not disclosed Jev's parameter count, judging from the performance of its open-source replicas (0.4B-0.6B), Jev's underlying model is far smaller than a frontier LLM; it only needs to be large enough for typed decisions, not to carry world knowledge and generation capability. A smaller model → faster inference → lower hardware cost → cheaper pricing. This is a virtuous cycle.


RLCD: The Fundamental Difference from RLHF

Jev's training method is called RLCD (Reinforcement Learning for Calibrated Decisions). This is TypeSafe's core technical contribution, and it is worth understanding, because it represents a new answer to the fundamental question of "what should a model optimize for".

What Does RLHF Optimize?

RLHF (Reinforcement Learning from Human Feedback) optimizes human preference: give the model two replies, have a human pick the better one, train a reward model to predict human choices, then optimize the policy with algorithms such as PPO. The result is a model that learns to "write text humans like", but the relationship between that "liking" and "correctness" is indirect and fuzzy. Human preference is influenced by surface features such as wording, length, tone, and format, and the reward model easily learns those surface features rather than deep correctness. This is why an LLM can "confidently say something wrong": it optimizes for "looking right", not "being right".

What Does RLCD Optimize?

RLCD optimizes the match between probability and actual outcome: if the model gives an option a 90% probability, then across many similar scenarios, that option really should be correct in about 90% of cases. The loss function directly measures this calibration gap, with common metrics including the Brier score (the mean squared distance between the predicted probability distribution and the actual outcome) or ECE (Expected Calibration Error).

This means RLCD does not care what the model "says"; it only cares whether the probability distribution the model gives matches reality. An RLCD-trained model may "not know" that the capital of France is Paris (because it does not need to generate text), but it can accurately tell you that the probability this passage contains false information is 87%.

The Fundamental Difference

DimensionRLHFRLCD
Optimization targetHuman preference rankingProbability-outcome match
Output formA text token sequenceA typed probability distribution
CalibrationNo guarantee (generally overconfident)The core design goal
Hallucination riskHigh (can confidently say something wrong)Low (no text generation, only probabilities)
Applicable scenarioOpen-ended generation, dialogue, creationClosed-ended decisions, classification, scoring, guarding
Representative modelsGPT-5.6, Claude 4.5, Gemini 3Jev (and open-source replicas)

The fact that Diogo Almeida, as a co-inventor of RLHF, went on to design RLCD is itself noteworthy. He is not rejecting RLHF, which remains the optimal solution for open-ended generation, but rather finding a more suitable optimization target for a specific scenario (structured decisions). That attitude of "knowing where the method you invented falls short, and then designing a better one" says more than any technical detail.


Deliberately Abandoned Capabilities = Design Philosophy

Jev's most interesting design decision is not what it can do, but what it does not do:

  • ❌ Cannot write replies
  • ❌ Cannot generate code
  • ❌ Cannot summarize
  • ❌ Cannot explain its reasoning
  • ❌ Cannot hold open-ended conversation

These "cannots" are not technical limitations but design choices. TypeSafe's argument is this: when you remove "generation" capability from the model, what you get is:

  1. Deterministic output: the answer necessarily falls within the schema you defined, with a zero structured-output error rate. No JSON mode, no schema validation, no retry logic.
  2. Calibratable probabilities: the model does not try to "persuade" you, only gives a probability distribution. No wording embellishment, no tone manipulation, no "I am certain but actually not" hallucination.
  3. Extreme speed and low cost: no autoregressive decoding, no large parameter count carrying world knowledge. The model does one thing, and does it to the extreme.
  4. Composability: Jev's output is typed probability values that can be fed directly into deterministic code, with no parsing layer. This is the idea of "AI as software primitive": the model is not a chat partner, but a function call in your code.

The philosophy behind this is the separation of decision and generation:

Jev decides, the LLM writes.

In agent architecture, this means a clear division of labor:

  • Routing, classification, guarding, scoring → Jev (System One): fast, cheap, calibratable, deterministic output
  • Reply generation, code writing, reasoning explanation, open-ended dialogue → LLM (System Two): slow, expensive, flexible, creative

The two model layers each do their job, and the overall system becomes faster, cheaper, and more controllable. This is not a narrative of "Jev replaces the LLM", but one of "Jev and the LLM form a complementary architecture".


The Open-Source Replica Ecosystem

Jev's API is still in early access (waitlist), but the open-source community has already produced a batch of replicas in a very short time. As of 2026-09-19, the main projects are:

ProjectBase ModelFeaturesLink
OpenJev (kotobalabs)DeBERTa-v3-large (0.4B)512 token context; M1 Max CPU 1.8s/4 questions; H100 28ms/10 questions; three-seed in-domain 0.847±0.005, OOD 0.678±0.012; the most transparent methodologyHuggingFace
openjev (AlexWortega)Qwen3.5-4B / 35B-A3B (MoE)An NLI cross-encoder form; zero-shot entailment for rerank/grade/guard; can play Doom directly from pixelsHuggingFace
NanoJev (C-Tianyu)Qwen3-0.6B + a decision headFull weights open-sourced; multi-state parallel decisionsGitHub
LightJev-0.6B (rongxinzy)Qwen3-0.6BA synthetic-domain research release, with complete training logsGitHub
JevlikeQwen2.5-0.5BClaimed ~100x speedup (research at an early stage)GitHub
open-alternative-jev (so1)Any open-weights LLMA self-hosted GPU approach, not depending on the TypeSafe APIGitHub
daf-jev (Zenodo, MIT)A Python toolkitIncludes an MCP server + agent skill integrationZenodo
jev (PyPI v0.3.0)A Python 3.14+ packageA @jev.fn decorator compiles a function signature into a Jev query; passes pyright/mypy strictPyPI

Several Observations Worth Noting

OpenJev's numbers are the most solid. kotobalabs provided a complete three-seed experiment, an in-domain vs OOD split, a comparison with a LoRA-tuned LLaDA-MoE-7B-A1B (0.835 in-domain, 16x latency), and a head ablation (the marker-token head does not learn; the span head does). This is currently the most methodologically transparent of the open-source replicas. Its conclusion is honest: an in-domain 0.847 shows the Jev-style architecture works on familiar scenarios, while an OOD 0.678 shows that generalization remains an open problem.

openjev's NLI angle is interesting. AlexWortega reduces Jev-style decisions to natural language inference (NLI), a three-way classification of entailment / contradiction / neutral, then uses this primitive for rerank, grade, and guard, and even to play Doom directly from pixels. This demonstrates the reducibility of "Jev-style decisions": a sufficiently strong NLI cross-encoder can serve as a general decision engine. The 35B MoE version can play Doom zero-shot, which is an impressive demo.

The PyPI jev package takes API access to a Pythonic extreme: one @jev.fn decorator where the function signature is the decision spec, the docstring is the state, the return annotation is the schema, with no prompt strings, no JSON schema, and no parsing layer. A bool field automatically compiles to noul, a Literal[...] automatically compiles to choice, and an int with Field(ge=, le=) automatically compiles to score. This is a complete realization of the "function signature as decision spec" idea, and a rare piece of elegant design in the Python ecosystem where "the type system is the interface".


Implications for Agent Architecture

1. Structured Decisions Do Not Require an LLM

A large share of steps inside an agent are "judgments" rather than "generation": routing classification, entity selection, guard conditions, priority scoring, intent recognition, toxicity detection. Doing these with an LLM is like using a cannon to kill a mosquito: slow, expensive, and with probabilities that cannot be calibrated.

Jev-style models offer a new option: push System One decisions down to a dedicated model, and call the LLM only when open-ended generation is needed. This can significantly reduce an agent's overall latency and cost while improving the controllability of decisions.

2. The Viability of Small Local Models

OpenJev's numbers, 28ms/10 questions on an H100 and 1.8s/4 questions on an M1 Max CPU, show that Jev-style decisions can run on small local models, without depending on a cloud API. For latency-sensitive or privacy-sensitive scenarios (financial transaction guarding, medical classification, legal document classification), this is an important option.

A model size of 0.4B-0.6B means it can run on consumer-grade GPUs or even CPUs, echoing our measured result of WeMM-Embedding 2B at 47ms on a Mac Studio: the next generation of agent infrastructure components is migrating toward consumer hardware.

3. Calibration Is the Correct Basis for Agent Decisions

LLM logprobs have long been used for confidence gating, but in decision scenarios LLM token probabilities are generally overconfident, saying 99% for things that are only 70% accurate. Jev's RLCD training makes calibration the direct optimization target, which is the correct basis for agent autonomy (auto-routing, confidence-gated escalation).

Imagine an agent's triage flow: Jev gives a decision probability for every input, with >95% handled automatically, 70-95% routed to an LLM for a second check, and <70% escalated to a human. The reliability of this flow depends directly on the quality of calibration: if Jev says 95% for things that are only 80% accurate, then the auto-handling threshold needs to be raised, more traffic leaks to the LLM and the human, and the cost advantage shrinks.

4. "Separating Decision from Generation" Is an Architectural Pattern

Jev's philosophy is not merely "use another model for classification". What it proposes is an architectural pattern:

Agent Loop:
├─ System One (Jev / OpenJev): routing, classification, gatekeeping, scoring → fast, cheap, calibratable
  └─ System Two (LLM): generation, reasoning, explanation, dialogue → slow, expensive, open-ended

This pattern maps directly onto the dual-process theory of human cognition (Kahneman's System 1 / System 2), and it also matches the layered needs of "fast thinking / slow thinking" in agent architecture. It is not about replacing the LLM, but about freeing the LLM from the "high-frequency, low-latency decisions" it is not good at, so it can focus on "open-ended generation and reasoning".


A Cold-Eyed Look

Company-Claimed Numbers Need Verification

TypeSafe claims Jev is 193.6x faster and 4.6x cheaper than frontier LLMs (or 445x, depending on the baseline). These numbers are not yet independently verified, and depend heavily on workload, baseline, and network environment. The SiliconANGLE report states plainly that "Those figures have not been independently verified and will vary by workload, network location and comparison method." Until you run the benchmark yourself, these numbers should be treated as order-of-magnitude estimates, not engineering facts.

Calibration ≠ Accuracy

Calibration is the match between probability and outcome, not the accuracy of a single judgment. A model can be perfectly calibrated while having only 60% accuracy (saying 60% for things that are indeed 60% correct), which is acceptable on low-difficulty tasks but may not be enough for high-stakes decisions (medical, financial, legal). Developers need to verify Jev's calibration and accuracy on their own data, rather than blindly trusting official numbers.

Track Competition

Jev is not the only player aiming at the "structured decisions" track. Conformal Prediction, the Guidance framework, Outlines, and LLM-based structured output (JSON mode, tool use) all address similar problems. Jev's differentiation lies in an end-to-end dedicated model + RLCD calibration, but whether this moat is deep enough is a question worth watching: the rapid emergence of open-source replicas (7 projects within a week) is a notable signal, indicating that the architectural barrier of "Jev-style decisions" is not high, and the core barrier may lie in data and calibration quality rather than the architecture itself.

The Performance Gap of Open-Source Replicas

OpenJev's in-domain 0.847 / OOD 0.678 numbers indicate that open-source replicas still suffer a significant performance drop on OOD scenarios. TypeSafe's Jev claims frontier LLM-level intelligence, but that claim is based on private tests and cannot be compared directly with the open-source numbers. Conclusion: open-source replicas have already proven the "Jev-style decision" architecture viable, but the performance gap still needs the closed-source model and the open-source community to push forward together.


Closing

Jev represents not just a new model, but the emergence of a new category: the System One model, purpose-built for fast, structured, calibratable decisions, complementing the open-ended generation capabilities of LLMs.

For agent architects, it proposes a pattern worth serious consideration: separate decision from generation, using a System One model for high-frequency, high-volume judgment steps and an LLM for steps that require open-ended generation. This can significantly reduce latency, cost, and uncontrollability.

The open-source community's rapid follow-up (7 replica projects emerging within a week) proves the category's appeal and architectural viability. But the company-claimed performance numbers still need independent verification, the distinction between calibration and accuracy must still be kept in mind, and track competition is still at an early stage.

Our next step: run the Jev API and the local OpenJev model on a real agent workflow (support ticket triage), measure real latency, cost, and calibration, and compare against the existing LLM-based solution. Let the data speak.


References: TypeSafe AI official blog · Flavio Copes deep dive · SiliconANGLE report · The Register report · OpenJev (kotobalabs) · openjev (AlexWortega) · jev PyPI v0.3.0 · OpenJev browser demo