1. The 30-Second Version
If you've seen our launch coverage of Jev, this report is the practical follow-up: how to actually get your hands on it, and what to do if you'd rather run the open equivalents on hardware you own. In one line: Jev is the first System One model — a frontier model that gives up text generation entirely and instead returns typed, probabilistic decisions that your code can branch on directly, in 70–500ms [TypeSafe announcement].
The three facts that make it unlike anything else in the model zoo:
- No text out. You send a state (text or JSON) plus a set of multiple-choice questions —
choice,noul(yes/no), orscore— and get back a probability distribution over the options for each question, all computed in one parallel pass [TypeSafe docs] [Wikipedia]. - Calibrated by construction. It's trained with a new method, Reinforcement Learning for Calibrated Decisions (RLCD), on synthetic data — so the confidence it reports is meant to be epistemically honest: high confidence should actually mean high accuracy.
- Cheaper than the meter can see. Input tokens cost $0.042 per million; output tokens are free — "too cheap to meter," in the company's own words — versus input pricing of $0.20–$10/MTok and output tokens costing ~5× input on typical frontier LLMs.
This report walks both paths to a working Jev-class API: the cloud path (TypeSafe's early access — playground, raw API, Python SDK) and the local path (Unsloth serving the open Laya decision model through a TypeSafe-compatible endpoint, plus fine-tuning your own decision model from Qwen or Gemma weights in under an hour).
2. Why Jev Is Different
2.1 A discriminative model, not a chatbot
The cleanest way to see what Jev is: it is a discriminative model — a model for classification and regression — not a generative one. It does not produce natural-language text at all; it returns typed values with probability estimates and confidence scores, and that output is designed to be consumed by software, not read by a human [Wikipedia]. There is no parsing step, no schema validation, no refusal, no hallucinated field name. The output space is defined in advance, so type errors are not unlikely — they are structurally impossible.
The second difference is in the sampler. An LLM generates one token at a time, each conditioned on the last; that sequential decode is where most of the latency in a "fast" LLM actually lives. Jev evaluates every question against the state in parallel, in a single query — a hardware-aware parallel sampler rather than an autoregressive loop. Adding questions to a request barely changes response time, because there's no context-rot and no token budget to burn [TypeSafe docs].
What that buys, side by side with an LLM:
| Dimension | Frontier LLM | Jev (System One) |
|---|---|---|
| Output | Strings — chat, code, or structure; must be parsed and validated | Type-safe structured values + calibrated probabilities |
| Sampling | Sequential, token by token | Parallel — all questions in one pass |
| Training objective | RLHF (human preference) / RLVR (verifiable rewards) | RLCD — calibrated, honest probabilities |
| Confidence | Overconfident and inconsistent when prompted | Reported on every answer; calibrated by training |
| Type errors | Possible, even at 99% accuracy | Mathematically impossible |
| Latency | 3–329s end-to-end on frontier models | 70–500ms end-to-end |
One nuance worth keeping straight: unlike classical discriminative models (the linear classifiers of old), Jev does not require data-specific training for general use — the RLCD training on synthetic data is meant to give it broad decision competence out of the box, and you can still fine-tune it (or its open siblings) on your own labeled decisions [Wikipedia].
2.2 Calibration is the product
Most of the "why" of Jev is actually about confidence. An LLM you prompt for a confidence score tends to be overconfident and inconsistent — it'll tell you 95% on a Tuesday and 60% on a Wednesday for the same input. TypeSafe's framing is blunt: if a model can do a task 95% of the time but doesn't reliably tell you when it's in the 5%, it cannot automate that task. You can't build a production workflow on a guess about when the guess is good. RLCD is the answer to that: the optimization target is calibrated decisions — probability values that are honest about themselves — which makes it possible to write if p > 0.85: file_ticket() and mean it [TypeSafe announcement].
2.3 The constraints are part of the design
Jev's type system is deliberately narrow, and the limits are documented:
- All three question types are multiple-choice in spirit: pick from options (
choice), yes/no with a probability (noul), or rate on a scale (score). - No dependencies between questions. Everything is evaluated in parallel and in isolation against the same state — so if question B depends on question A's answer, you make two sequential requests. That's a real workflow constraint, not a bug [Wikipedia].
- Up to 64 questions per request, up to 255 options per choice, 10 levels per score. Choices above ~255 cardinality fall back to a two-stage score-then-pick, which is why high-cardinality demos occasionally slow down [Unsloth docs].
2.4 The name is the thesis
Jev is named after economist William Stanley Jevons — of Jevons paradox, where coal efficiency gains made coal cheaper and total coal consumption went up. The bet: every order-of-magnitude drop in the cost of intelligence unlocks more use cases than it saves. A decision call that costs a few hundredths of a cent changes the math on everything from per-frame game AI to per-paragraph document triage [TypeSafe announcement].
3. Why It Matters
Here's the argument in four moves.
3.1 It changes the unit economics of a decision
Today, if your software needs a judgment — route this ticket, flag this document, score this lead — the standard answer is "call the LLM." That means paying input and output tokens (with output running ~5× input on most frontier pricing), parsing the returned string, validating it against your schema, and handling the occasional hallucinated field or refusal. Jev collapses that to $0.042 per million input tokens and zero output cost — the company's published claim is up to 193.6× faster and 444.6× cheaper than the reference LLM workflow on its own evals, and it flags that as the high end [TypeSafe announcement]. Even at a quarter of the claimed multiple, the economics cross a line: decisions become cheap enough to make per event, not per batch.
3.2 Latency opens real-time categories
70–500ms end-to-end is a different world from 3–329 seconds. It's the difference between "AI in the workflow" and "AI in the loop" — i.e., per-frame, per-keystroke, per-packet. The company's own demos make the point: a Doom-playing agent making 10 structured decisions per second (~$7/hour at their pricing), and a Wikiracing bot picking from thousands of links at interactive speed, where not hallucinating on high-cardinality choices compounds over every step [TypeSafe announcement]. For a business, the same property applies anywhere UX or SLA is the constraint: real-time fraud signals, live sentiment on a call, in-flight routing of support traffic.
3.3 It automates what LLMs can only assist
The calibration point is the deep one. An overconfident model can't be automated even when it's mostly right, because your code has no honest signal for when to defer to a human. Jev's contract is that confidence is meaningful: higher reported confidence correlates with higher accuracy, and similar inputs get similar answers. That turns "AI review" into a smart if-statement — a fuzzy decision rule that ordinary software can branch on, with a defensible threshold and a clean escalation path for the low-confidence tail [TypeSafe confidence docs]. That's what "type-safe automation" actually means in practice: the surrounding code constrains the model's freedom, so the whole thing composes into a reliable system.
3.4 It completes the agent stack, not replaces it
Jev is not a drop-in replacement for the LLM. It doesn't write prose, code, or plans. The architecture that's emerging — and that LangChain's team has already written up — is a harness: the LLM stays the brain for open-ended reasoning and generation, while a System One model handles the fast, structured decisions along the way — routing, classification, guardrail checks, branch selection — without a full chat call for each one [LangChain]. An agent loop that would have made 20 LLM calls now makes a handful, plus a stream of cent-free, 100ms decision calls. That's a direct hit on agent cost and latency, the two numbers every operator is trying to bring down.
The one-sentence version
Jev is a frontier-intelligence function call: unstructured state in, typed probabilistic decisions out — 40–200× faster than an LLM for the same decision, and it can't hallucinate a field that isn't in the schema.
4. Path One: Jev via TypeSafe's Cloud
4.1 Getting in
Jev is in early access. The flow: sign up at typesafe.ai (developers are being pulled off the waitlist), then work from the console at console.typesafe.ai. The fastest orientation is the Playground: paste any text as the state — a support ticket, an email thread, a paragraph of code — add one or more questions (Noul / Choice / Score), and watch every question come back in parallel with its probabilities [Quickstart].
4.2 The raw API
One endpoint, one shape. Grab an API key from Settings → Keys and POST:
curl -X POST https://api.typesafe.ai/v1/systemone \
-H "Authorization: Bearer $TYPESAFE_API_KEY" \
-H "Content-Type: application/json" \
-d @- <<'EOF'
{
"state": "Hi, I've been trying to connect my Stripe account for 3 days and it keeps failing. I'm losing sales. Please help ASAP.",
"model": "jev-latest",
"questions": {
"department": {
"type": "choice",
"instructions": "Which team should handle this",
"criteria": {
"billing": "Payment or subscription issues",
"technical": "Bugs or integration problems",
"sales": "Pricing or account questions"
}
},
"frustration": {
"type": "score",
"instructions": "How frustrated the customer appears",
"criteria": [
"Calm, just stating facts",
"Frustrated but civil",
"Very angry, strong language"
]
},
"is_urgent": {
"type": "noul",
"instructions": "The message conveys urgency or time-sensitivity"
}
}
}
EOFThe response is what your code actually branches on — typed, no parsing required:
{
"model": "jev-latest",
"answers": {
"department": {
"type": "choice",
"choice": "technical",
"probabilities": { "billing": 0.159, "technical": 0.84, "sales": 0.001 },
"confidence": 0.596
},
"frustration": {
"type": "score",
"score": 1.035,
"legend": { "0": "Calm, just stating facts", "1": "Frustrated but civil", "2": "Very angry, strong language" },
"confidence": 0.842
},
"is_urgent": { "type": "noul", "noul": 0.999 }
},
"usage": { "input_tokens": 312, "output_tokens": 48 }
}How to read the two numbers
probabilities tells you which option won; confidence tells you how strongly the distribution is peaked (close to 1 = one answer stands out, close to 0 = evenly spread). Act on the per-option probability (e.g., file a ticket only when technical > 0.8), and use confidence to decide whether to auto-act or escalate. Noul returns a single 0–1 probability of "yes."
4.3 The Python SDK
For anything real, use the typed SDK (Python ≥ 3.10):
pip install typesafe-sdk
from typesafe_sdk import Choice, Noul, Score, TypeSafeClient
client = TypeSafeClient() # reads TYPESAFE_API_KEY from the environment
response = client.system_one(
state=ticket_text,
questions={
"department": Choice(
instructions="Which team should handle this",
criteria={
"billing": "Payment or subscription issues",
"technical": "Bugs or integration problems",
"sales": "Pricing or account questions",
},
),
"frustration": Score(
instructions="How frustrated the customer appears",
criteria=["Calm, just stating facts", "Frustrated but civil", "Very angry, strong language"],
),
"is_urgent": Noul(instructions="The message conveys urgency or time-sensitivity"),
),
)
print(response.answers["department"].choice) # "technical"
print(response.answers["frustration"].score) # 1.035
print(response.answers["is_urgent"].noul) # 0.9994.4 The agent skill
TypeSafe also ships an agent skill so your coding agent knows the API by heart — install it in Claude Code with claude plugin marketplace add typesafe-ai/skills + claude plugin install typesafe@typesafe-ai, or in any other agent with npx skills add typesafe-ai/skills --skill typesafe-ai, then just tell your agent to build with it [Agent skill docs].
4.5 Cost and limits at a glance
| Item | Value |
|---|---|
| Input tokens | $0.042 / MTok |
| Output tokens | Free ("too cheap to meter") |
| Latency | 70–500ms end-to-end |
| Questions per request | Up to 64 (all in parallel) |
| Options per choice | Up to 255 (two-stage above that) |
| Score levels | Up to 10 |
| Model alias | jev-latest |
5. Path Two: Running Decision Models Locally
Here's the part most of the coverage skipped. Jev itself is a closed cloud model, but the architecture is now being replicated in open models, and Unsloth ships a complete local stack: serve the open decision models Laya and Clef through a TypeSafe-compatible Jev API, or fine-tune your own decision model from Qwen or Gemma weights [Unsloth decision models]. "Laya" is the reference open-weight model (678MB — smaller than a large photo library file); "Clef" is Cloudflare's open decision model, which shares the scoring-head design.
5.1 Why local works at all
The unsupervised reason this is practical: a decision model never spends compute writing text. Unsloth's implementation puts your state and every question into one prompt, lets the base LLM read it once, then a small scoring head — the same design as Cloudflare's Clef — looks at the LLM's output over each question and option and scores them all together. Because it never does autoregressive text sampling, a sub-billion-parameter model on a plain CPU answers in well under a second [Unsloth training guide].
5.2 Serve Laya locally (10-minute setup)
Step one, get Unsloth running — either the desktop app (macOS / Windows / Linux) or the one-liner install:
# macOS, Linux, or WSL
curl -fsSL https://unsloth.ai/install.sh | sh
# Windows PowerShell
irm https://unsloth.ai/install.ps1 | iexStep two, in the app: Settings → API → Decision API → turn on Serve requests. Confirm the download of the default Laya model (678MB). Your API key lives on the same page under Access tokens — or enable Keyless API access → Chat and inference to skip the key on localhost [Unsloth docs].
Step three, send a decision request. This classifies a support ticket, detects a refund request, and scores urgency — three question types in one call, same payload shape as the cloud:
curl http://localhost:8888/v1/systemone \
-H "Authorization: Bearer sk-unsloth-YOUR_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "laya",
"state": "Hi, I was charged twice for my March invoice (#4411). Please refund the duplicate today.",
"questions": {
"team": {
"type": "choice",
"instructions": "Which team should handle this?",
"criteria": {
"billing": "invoices, payments, refunds",
"technical": "bugs, outages, errors",
"sales": "pricing, new plans",
"other": "everything else"
}
},
"refund": { "type": "noul", "instructions": "Does the customer ask for a refund?" },
"urgency": { "type": "score", "instructions": "How urgent is this?", "criteria": ["not urgent", "soon", "today"] }
}
}'The first request loads the model (10–20 seconds); after that, requests take well under a second on most CPUs — no GPU required for inference. The model field accepts laya, default, or jev-latest (whatever you selected in settings), plus the specific variants laya-multilingual, laya-english, and laya-typed-decisions.
5.3 Point existing Jev code at your machine
This is the elegant part: the local server speaks the TypeSafe protocol. If you already wrote cloud Jev code, migrating it to your own hardware is two environment variables:
pip install typesafe-sdk
export TYPESAFE_BASE_URL=http://localhost:8888
export TYPESAFE_API_KEY=sk-unsloth-YOUR_KEYThe SDK's default model, jev-latest, maps to whichever model you picked in Unsloth's settings, and answers come back in the same format as the cloud. That means a privacy-first deployment is a config change, not a rewrite: support tickets, HR documents, or client data stop leaving the building, and you've got a TypeSafe-compatible decision API running on a laptop.
Don't port your confidence thresholds blindly
Laya and Jev calculate confidence differently. If your cloud code had a Jev-tuned threshold (e.g., "auto-act when confidence > 0.9"), re-tune it against local Laya, or — better — act on the per-option probabilities instead, which behave more consistently across the two [Unsloth docs].
5.4 Train your own decision model (40 minutes, 4GB VRAM)
Unsloth can fine-tune regular LLMs into decision models — same dataset format as Laya and Clef, and you can even further fine-tune Laya or Clef themselves. Current reference numbers from the docs:
| Base model | Test accuracy | VRAM | Training time |
|---|---|---|---|
| Qwen3.5-0.8B | 78% | 4GB | 42 min |
| Qwen3.5-2B | 81% | 8GB | 40 min |
| Llama 3.2 3B | 79% | 4.1GB | 30 min |
| Gemma 4 E4B | 77% | 14.4GB | 49 min |
| Laya (fine-tuned) | 77% | 2.5GB | 10 min |
The settings that reproduce those numbers: 2 epochs, LoRA rank 16, learning rate 2e-4, dataset subset all. Unsloth reports accuracy on held-out decisions before and after training, and calibrates the model's confidence as part of the run. Dataset rows are the same triple you've been seeing: a state, the questions to decide, and the gold answers [Unsloth training guide].
To serve a model you trained from code, start Unsloth pointing at the merged weights and turn on the Decision API:
UNSLOTH_SYSTEMONE_MODEL=/path/to/qwen-decisions-merged unsloth studio -H 0.0.0.0 -p 8888The model needs a GPU and loads on the first request (a few minutes); while a training run has the GPU, the Decision API waits and resumes when the run ends. Note what this means for the stack: the cloud path gets you the frontier Jev; the local path gets you your own decision model — domain-specific, private, and free to run — through the identical API surface.
6. The Pattern: A Harness, Not a Replacement
LangChain's write-up is the clearest articulation of how these pieces fit together, and it's worth internalizing because it decides how you should architect around Jev-class models [LangChain]:
- The LLM is the System 2 layer. Open-ended reasoning, drafting, planning, anything that needs to say new things in new words. Keep it for what it's actually good at.
- The System One model is the System 1 layer. Every structured judgment the agent makes along the way: which tool, which branch, does this output pass the guardrail, is this a refund request, what's the urgency. These used to be cheap LLM calls; they're now 100ms decision calls with honest probabilities.
- Code composes them. Atomic questions in, combined by your own logic out. The docs' own guidance: if a judgment needs extended reasoning or weighs multiple independent factors, decompose it into separate questions and combine the results with a formula in your code — "when priorities shift, change a coefficient in your code rather than rewriting a prompt" [TypeSafe docs].
- Confidence is the control plane. High-confidence answers auto-act; low-confidence ones escalate to the LLM or to a human. Because calibration is the training objective, that control plane is something you can actually trust — which is the entire reason the model class exists.
KodeKloud's walkthrough of Jev lands the same architecture from the practitioner side: use it for routing and classification inside LLM workflows, turn the probabilities directly into software decisions, and deploy it as guardrails ahead of — or behind — the models you already run [KodeKloud].
7. What to Watch
- The early-access queue. Jev is still being released to developers in waves. If you can get in, the highest-value experiment is cost-per-workflow: take one real automation you currently run on LLM calls (ticket triage, document scoring, guardrail checks) and measure the actual latency and cost multiple — not TypeSafe's 193×/444× headline, your own.
- Local parity. Laya and Clef are open today; whether their accuracy keeps up as Jev's frontier moves — and whether Unsloth's fine-tuning pipeline closes the gap on your own data — is the question that decides whether "local Jev" becomes a real category. The 77–81% reference accuracies on a benchmark set are a promising start [Unsloth].
- Multi-model routing. The System One LLM wrapper TypeSafe published (which constrains regular LLMs to the same structured decision output) means you can A/B Jev against your existing LLM on identical workflows [System One adapter]. Expect vendor-neutral harnesses — LangChain included — to standardize this soon.
- The no-dependency limit. Sequential dependencies require two requests. As workflows get more deeply chained, watch whether the API grows a lightweight dependency mechanism or whether decomposition stays the sanctioned pattern — it affects how deep you can push real production graphs.
- Sustainability of the pricing. TypeSafe is candid that it can't yet prove the $0.042/MTok input price isn't subsidized. If it holds, the Jevons logic kicks in hard: decision counts go up orders of magnitude. If it doesn't, the local path stops being a privacy option and becomes the cost option.
References
- TypeSafe AI — Introducing System One Models & Jev (launch post, Sept 15, 2026: architecture, pricing, workflow evals, caveats)
- Wikipedia — Jev (AI model) (discriminative model classification, I/O spec, RLCD training on synthetic data)
- TypeSafe docs — Introduction · Quick start (Playground, API, Python SDK, agent skill) · Confidence
- LangChain — Building a Harness with Jev (System One as the fast layer in agent harnesses)
- Unsloth — Run and Serve Decision Models Locally: Laya + Jev API (local serving, TypeSafe-compatible endpoint, migration) · Train your own Decision Model with Unsloth (fine-tuning Qwen/Gemma/Llama, reference accuracy table)
- KodeKloud — What Is Jev? The AI Model That Doesn't Generate Text (calibrated decisions, probabilities → software decisions, guardrails, LLM combination)
- TypeSafe — System One LLM adapter (GitHub) · Workflow evals site
Companion report: TypeSafe's Jev: System One Models — Frontier-Intelligence Decisions in 70ms, Without Hallucinating covers the launch itself. This report covers access, architecture, and the local stack.