TypeSafeSystem One ModelsJevAgent InfrastructureLLMLangChain

TypeSafe's Jev: System One Models — Frontier-Intelligence Decisions in 70ms, Without Hallucinating

TypeSafe AI — founded by an InstructGPT co-author after two years in stealth — is shipping a new class of frontier model that gives up text generation entirely. In exchange it produces type-safe, structured decisions with calibrated confidence, in 70–500ms at a fraction of a cent, claiming 40x–200x faster and ~445x cheaper than frontier LLMs on production workflow tasks. This report covers the model, the evidence, the company's own caveats, and how LangChain's Sydney Runkle wired it into agent loops.

September 18, 2026Michel Laclé12 min read
🎧 Audio Version — Listen to This Report

1. The Question: Where Is All the Automation?

TypeSafe AI launched on September 14, 2026 with a question its founder has been circling for four years: "Models have been superhuman at chat for years, so where is all the automation?" The framing is worth pausing on, because it's the thesis of the whole company [Link].

The founder, Diogo Almeida, was an OpenAI researcher on the instruction-following work behind InstructGPT — the RLHF paper that became the research foundation of ChatGPT [Link] [Link]. His conclusion after watching chat models dominate: "I thought maybe chat models would lead to AGI, but despite the hype it became obvious to me that there was something really big missing." After two years in stealth, TypeSafe's answer is a new stack — a new model architecture, a parallel sampler, and a training method they call Reinforcement Learning for Calibrated Decisions (RLCD) — built entirely for automation rather than conversation. The first model out of that stack is Jev, now in early access.

2. What a "System One Model" Is

The name borrows from Kahneman's Thinking, Fast and Slow: System 1 is fast, intuitive thinking; System 2 is slow, deliberate reasoning. LLMs have spent a decade becoming System 2 workhorses — fluent, general, powerful, and slow. TypeSafe's bet is that most decisions software needs to make are System 1 shaped: classify, route, score, extract, branch. And that a model class built for exactly those decisions can be dramatically faster and cheaper without matching an LLM on open-ended generation [Link].

The contract, in the company's own words: a System One model is "a class of AI models built to make fast, structured decisions that software can use directly. A System One model evaluates a state and returns typed answers and probabilities." Concretely, you send it a state (unstructured context: a ticket, a page, a log, a profile) and a set of questions; it returns type-safe structured values — the possible outputs and their structure are defined in advance in a schema [Link] — each answer carrying a calibrated probability and confidence score.

Existing LLMsSystem One + Jev
Trained withRLHF (human preference) / RLVR (verifiable rewards)RLCD — reinforcement learning for calibrated decisions
Optimizes forHuman-preferred text; programmatically verifiable outputsEpistemically honest probabilities on decision tasks
InputUnstructured text, sequential messagesUnstructured data as structured program state
OutputStrings — flexible but require parse + validate; can hallucinateType-safe values; the model cannot make a type error; every answer carries calibrated confidence
SamplingSequential, one token at a timeParallel — all outputs in a single query
CostInput $0.20–$10/MTok; output ~5x inputInput $0.042/MTok ($42 per billion tokens); output tokens free ("too cheap to meter")
Latency~3–329s end-to-end for frontier models70–500ms — 40x–200x faster at comparable task intelligence
ConfidencePrompted estimates are overconfident and inconsistentAlways communicated; calibrated — higher confidence means higher accuracy
Best atHuman-in-the-loop work (chat, copilots, agents), verifiable problems, demosAI-powered workflows ("smart if-statements"), map-reduce over big data, real-time apps, verifying/guardrailling other LLMs

The sharpest line in the launch post is the one-sentence definition: "Think of Jev as a frontier-intelligence function call: unstructured state in, typed probabilistic decisions out." Three question types are supported — Choice (pick from options, probability per option plus overall confidence), Score (rate against ordered levels, returning a continuous score, the underlying distribution, and confidence), and Yes/No (probability that a statement is true) — and all questions in a request are evaluated in parallel, so adding questions barely changes response time and costs only the cheap input tokens for the extra questions [Link].

3. Jev: The First Model, and Its Numbers

Model: Jev (early access, Sept 14, 2026)  ·  Stack: new architecture + parallel sampler + RLCD  ·  Positioning: frontier-intelligence decisions, not generation

Jev's headline numbers, straight from the launch post: similar levels of intelligence to existing frontier LLMs on System One tasks, at two orders of magnitude lower cost and 70–500ms end-to-end. The pricing is the most unusual part of any model launch I've seen: input at $0.042/MTok — i.e. $42 per billion input tokens — with output tokens listed as free, "too cheap to meter" [Link]. For comparison, that's roughly 1/5 to 1/240 of typical frontier input pricing, before you even count that there's no output bill at all. The company concedes it can't yet prove the pricing is sustainable long-term and expects it to go down, not up — a notably honest line for a launch post.

Two design consequences fall out of giving up strings. First, hallucination becomes structurally impossible at the output layer: a hallucinated tool call in a chatbot is inconvenient, but in a system with latency guarantees — or buried several layers deep in a dependency chain — it's a deal-breaker. Jev's outputs are schema-constrained, so "type safety is table stakes for automation," in the company's words. Second, parallel sampling means the model isn't paying the autoregressive tax on every decision — which is why a single call can answer many questions at once and still land under half a second.

Two naming stories carry the thesis. The class name comes from Kahneman, with an explicit inversion: "System 1 thinking" has historically implied error-prone — TypeSafe's claim is that System One models can be made more reliable than their alternatives. And "Jev" is named after William Stanley Jevons, after the Jevons paradox: when the efficiency of steam engines rose, coal consumption increased, because cheaper efficiency unlocked more demand [Link]. TypeSafe's expectation is the same curve for machine intelligence: "Every order of magnitude drop in the cost of intelligence unlocks orders of magnitude more use cases." That's not just a marketing name — it's the economic argument for why this model class matters more than its raw benchmarks.

4. The Evidence — and the Company's Own Caveats

The launch post is unusually explicit about what can and can't be verified, and it's worth mirroring that structure rather than just repeating the marketing.

Easy to verify

  • Speed per call: the published evals are run from laptops on the US West Coast (where the service currently lives), so latency claims are reproducible in kind [Link].
  • Cost per call: pricing is public. The company's own caveat: it can't prove the pricing isn't subsidized; long-term sustainability is an open question.
  • No type errors: "This would be an easy thing to falsify with just a single counter-example, but it is mathematically impossible." Schema matching is guaranteed by construction, so the 0% type-error number is a design property, not an empirical measurement.

The workflow evals

The central evidence is a new evaluation type: instead of a ground-truth classification (which would let teams overfit via harness engineering), every model gets the same workflow — a correct compute graph written in code — and predictions are scored against the average of the largest, most expensive external models (GPT-6 Astra and Fable 5.1 in this case) as the reference probabilities [Link]. Jev "ow[ns] the Pareto frontier for almost 2 orders of magnitude" on the accuracy-vs-cost plot, with the launch numbers of 193.6x faster and 444.6x cheaper. The LLMs are run through TypeSafe's own System One LLM adapter (open-sourced), which constrains chat models to output structured decisions compatible with the API — the most accurate way to get decisions from LLMs, the team found, but slower and more expensive than decisions without probabilities [Link].

The caveats the company publishes alongside these results are the part most launch posts skip:

  • The 193.6x/444.6x figures are "on the higher end of real-world gains," by the company's own say.
  • The four workflows were not in the training distribution and not constructed to flatter the model — but they were built by the company's own model-capabilities team, so some bias is possible.
  • The reference answers average two of the largest models (OpenAI + Anthropic), which biases results against DeepSeek and other models — TypeSafe says it probably underestimates their relative performance.
  • The side-by-side demo's input is "highly simplified" and the short state "paints our model in an advantageous light"; the only recorded disagreement with GPT-5.6 Terra (used as the most comparable baseline) was a "churn likelihood" question the team calls genuinely ambiguous.
  • The LLM hallucination/type-error comparison numbers come from OpenRouter, which may route complex queries to better models — so the gap may be overstated.

⚠️ The skeptical read

This is a two-year-stealth launch from a credible former OpenAI alignment researcher, with an unusually candid self-audit — but it is still a launch post. The evals are self-hosted, the reference set is self-chosen, and the workflows are in-house. The claims that hold up independent of any of that are the design-property ones: outputs can't be type errors, confidence is always emitted, sampling is parallel, and the price list is public. The benchmark magnitudes — 193x, 444x — should be treated as "direction and order of magnitude, pending third-party replication." The company's own FAQ acknowledges the obvious question — "these results are kinda crazy, how is it possible?" — without yet answering it in the public post.

The demos

The two showcase demos are the clearest window into the new use-case surface. Doom: a bot that plays Doom from structured game state, making 10 queries per second at a cost of roughly $7/hour — the kind of latency-and-price profile where an LLM simply can't enter the conversation (a non-AI Doom bot plays better, the team admits; the point was reactive, instruction-following intelligence per second). Wikiracing: navigate from one Wikipedia page to another following only in-page links, where each step is a choice among hundreds or thousands of options — a direct demonstration of "intelligence per second" and the compounding benefit of not hallucinating on high-cardinality choices. Jev supports cardinality up to 255 in a single call; beyond that it falls back to a two-stage score-then-choose pass, which shows up as the occasional slowdown [Link].

5. The Harness: Jev Inside Agent Loops (LangChain)

The launch is a model, but the ecosystem response arrived within days: LangChain's Sydney Runkle published "Building a Harness with Jev", the first serious treatment of where a System One model actually sits in an agent architecture [Link]. The argument is the most important part of this whole story, because it reframes Jev from "a fast classifier" to the missing decision layer of the agent loop.

The setup: agents run in a loop — an LLM decides, a tool executes, a model evaluates, repeat. Tool calling and structured outputs made LLMs integrable, but every decision in that loop still costs a full LLM call. Jev slots into exactly those decision points: send it the state (text, structured data, or LangChain messages), ask the questions, and use the typed probabilistic answer to guide what the agent does next — without a chat LLM call per decision. The LangChain integration ships as langchain-typesafe, exposing Jev through a TypeSafeClassifier: you pass state and questions to .invoke() and get classification results rather than a chat response [Link].

Pattern 1 — Model routing

Routing middleware lets Jev assess each incoming request and pick the model that fits: cheap and fast for trivial tasks, frontier-class for hard ones. The router selects from the latest user message and uses that model for the whole run, with the probabilities and confidence left in agent state for downstream logic. This is the Jevons-paradox play at the architecture level: the expensive model stops being the default, and the bill reflects the difficulty distribution of actual traffic [Link].

Pattern 2 — Auto Mode: safety gating that's finally affordable

This is the claim that generalizes furthest. Coding harnesses (Claude Code, Codex, Cursor) have all quietly shipped some version of "classify the dangerous action before it runs" — but that classifier has been locked inside the closed-source part of each harness. Runkle's point: now that a cheap, fast, calibrated classifier model exists as a commodity API, every agent harness can adopt the same trust pattern. AutoModeMiddleware uses Jev to check tool calls for risky decisions and block them before execution — turning "agents are untrustworthy, including under prompt injection" from a philosophical problem into a priced, measurable middleware step [Link].

The early community examples Runkle cites show the cost curve doing what the Jevons name predicts: Kyle Jeong (Browserbase) running browser-use agents for fractions of a cent, Jarrod Watts building a live trading agent, and Ryan Vogel doing email triage at scale [Link] [Link] [Link]. Note what these three have in common: high decision volume, low per-decision stakes, real-time or near-real-time constraints — the exact System 1 profile the model class is built for.

6. Why This Matters for the Agent Stack

Three implications, in order of proximity:

The agent loop gets a cheap decision layer — and the cost model of agents changes shape

Today, an agent that makes 200 LLM calls a day on 30% of which it only needs to classify/route/score is paying frontier prices for System 1 work. Jev's price point (public input pricing around $0.042/MTok, free outputs, sub-500ms) moves those calls off the LLM bill entirely. The LangChain routing pattern is the template: frontier models for the 20% that needs them, System One models for the rest. This is the most direct cost lever on agent economics since structured outputs.

Calibrated confidence is the unlock for "smart if-statements"

LLM prompt-hacking for confidence is well documented to be unreliable; Jev's confidence is a trained, calibrated property of the output. That's what makes the workflow pattern real: branching logic that says "if the model is 92%+ confident the ticket is urgent, route to tier-1; otherwise to human review" is only safe when the 92% means something. The docs' support-ticket example returning a 99.9% urgency probability is the unit of this new category of code [Link].

Safety gating becomes a commodity primitive

The Auto Mode pattern inverts the history of agent safety: the risk classifier moves from a proprietary harness feature to a callable API. If cheap, fast, calibrated classifiers become commodity infrastructure, then "gating every tool call" stops being a luxury feature of closed-source IDEs and becomes a default line in every agent middleware stack. That's a real security-infrastructure shift, and it's worth tracking separately from the model itself.

⚠️ What this is not

Jev is not a drop-in LLM replacement, and Runkle is explicit about it: no text generation, no open-ended reasoning, no chat. It's a complement — "use an LLM for open-ended reasoning and generation, and Jev for fast, structured decisions along the way." It's also not yet battle-tested: early access just opened, the training-data and generalization questions are still FAQ stubs in the launch post, and the Doom demo's structured state is text, not images, "yet." The category thesis (cheap decisions, Jevons economics, commodity safety gating) is where the real argument lives; the specific model's third-party validation is still pending.

7. What to Watch

  • Third-party replication of the workflow evals. The open-sourced System One adapter and the public evals site make this unusually easy — the 193x/444x figures will either survive replication or get revised down. The evals were built in-house; the adapter being open is the strongest counterweight to that bias.
  • Pricing sustainability. "Output tokens: free" and ~$42/B input tokens need a cost structure that hasn't been published. If the price holds as models improve, the Jevons-paradox use-case explosion is real. If it's subsidized, watch for repricing.
  • Image and multimodal state. The demos are text-structured-state only, explicitly "yet." Multimodal System One decisions (frames, screenshots, sensor data) would move the category into robotics and perception — where the latency budget is even tighter.
  • The safety-gating pattern spreading. Auto Mode in LangChain is the reference implementation. The interesting question is whether Claude Code, Codex, and Cursor replace their in-house classifiers with an external calibrated one — or double down on proprietary. Either outcome is informative about where agent trust lives.
  • What "System One" displaces. Beyond LLM calls: rule engines, scoring models, and hand-written decision logic. The "smart if-statement" framing is an explicit attack on decades of brittle hand-coded branching — and on the ML-ops stack around small supervised models.

The one-paragraph summary: a former InstructGPT researcher spent two years in stealth on the observation that chat superhumanity never became automation, and shipped a model class that treats the answer as a decision, not a string — parallel, typed, calibrated, and priced per billion tokens. The LangChain harness response within days is the proof that the ecosystem already knows which agent-loop slot it fills. Watch the replication; the design properties are already real.

References

  1. TypeSafe AI — Introducing System One Models and Jev (Diogo Almeida, Sept 14, 2026) — launch post: architecture, RLCD, pricing, evals, demos, caveats, FAQ
  2. Sydney Runkle (LangChain) — "Building a Harness with Jev" (X Article, Sept 18, 2026) — Jev question types, LangChain integration, model routing, Auto Mode middleware, community examples
  3. LangChain docs — TypeSafe provider — langchain-typesafe, TypeSafeClassifier, model-routing and tool-risk-gating middleware
  4. TypeSafe documentation — state concepts, question schemas, quickstart request format
  5. TypeSafe Quickstart — support-ticket state/questions example (99.9% urgency probability)
  6. system-one-adapter-python (GitHub) — open-source adapter that constrains LLMs to structured decisions compatible with the API
  7. TypeSafe Workflow Evals — the four public workflows, examples, disagreements, full queries
  8. OpenAI — Aligning language models to follow instructions — InstructGPT paper (Diogo Almeida co-author)
  9. Diogo Almeida — Google Scholar
  10. Jevons paradox — Wikipedia — the efficiency→demand-rebound effect the model is named for
  11. LLM benchmarks (Diego Romero) — the frontier-latency reference (3–329s end-to-end) cited in the launch post
  12. Kyle Jeong (Browserbase) — browser-use agents at fractions of a cent
  13. Jarrod Watts — live trading agent
  14. Ryan Vogel — email triage at scale
  15. Daniel Kahneman — Thinking, Fast and Slow — the System 1 / System 2 distinction behind the class name