JevSystem One ModelsTypeSafeUnslothLocal AIDecision ModelsAgent Infrastructure

Jev in Practice: The System One Model That Returns Probabilities, Not Text — Cloud Access and Local Laya/Jev API

TypeSafe's Jev gives up text generation entirely and returns typed, calibrated probabilities in 70-500ms. This report is the hands-on companion: how to get cloud access (playground, API, Python SDK), why the architecture is structurally different from any LLM, and how to run the open Laya/Clef decision models locally through a TypeSafe-compatible Jev API - or fine-tune your own in 40 minutes on 4GB of VRAM.

October 8, 2026Michel Laclé14 min read
🎧 Audio Version — Listen to This Report (11:36)
📺 Video Version — Watch This Report (11:36)

1. The 30-Second Version

If you've seen our launch coverage of Jev, this report is the practical follow-up: how to actually get your hands on it, and what to do if you'd rather run the open equivalents on hardware you own. In one line: Jev is the first System One model — a frontier model that gives up text generation entirely and instead returns typed, probabilistic decisions that your code can branch on directly, in 70–500ms [TypeSafe announcement].

The three facts that make it unlike anything else in the model zoo:

  • No text out. You send a state (text or JSON) plus a set of multiple-choice questions — choice, noul (yes/no), or score — and get back a probability distribution over the options for each question, all computed in one parallel pass [TypeSafe docs] [Wikipedia].
  • Calibrated by construction. It's trained with a new method, Reinforcement Learning for Calibrated Decisions (RLCD), on synthetic data — so the confidence it reports is meant to be epistemically honest: high confidence should actually mean high accuracy.
  • Cheaper than the meter can see. Input tokens cost $0.042 per million; output tokens are free — "too cheap to meter," in the company's own words — versus input pricing of $0.20–$10/MTok and output tokens costing ~5× input on typical frontier LLMs.

This report walks both paths to a working Jev-class API: the cloud path (TypeSafe's early access — playground, raw API, Python SDK) and the local path (Unsloth serving the open Laya decision model through a TypeSafe-compatible endpoint, plus fine-tuning your own decision model from Qwen or Gemma weights in under an hour).

2. Why Jev Is Different

2.1 A discriminative model, not a chatbot

The cleanest way to see what Jev is: it is a discriminative model — a model for classification and regression — not a generative one. It does not produce natural-language text at all; it returns typed values with probability estimates and confidence scores, and that output is designed to be consumed by software, not read by a human [Wikipedia]. There is no parsing step, no schema validation, no refusal, no hallucinated field name. The output space is defined in advance, so type errors are not unlikely — they are structurally impossible.

The second difference is in the sampler. An LLM generates one token at a time, each conditioned on the last; that sequential decode is where most of the latency in a "fast" LLM actually lives. Jev evaluates every question against the state in parallel, in a single query — a hardware-aware parallel sampler rather than an autoregressive loop. Adding questions to a request barely changes response time, because there's no context-rot and no token budget to burn [TypeSafe docs].

What that buys, side by side with an LLM:

DimensionFrontier LLMJev (System One)
OutputStrings — chat, code, or structure; must be parsed and validatedType-safe structured values + calibrated probabilities
SamplingSequential, token by tokenParallel — all questions in one pass
Training objectiveRLHF (human preference) / RLVR (verifiable rewards)RLCD — calibrated, honest probabilities
ConfidenceOverconfident and inconsistent when promptedReported on every answer; calibrated by training
Type errorsPossible, even at 99% accuracyMathematically impossible
Latency3–329s end-to-end on frontier models70–500ms end-to-end

One nuance worth keeping straight: unlike classical discriminative models (the linear classifiers of old), Jev does not require data-specific training for general use — the RLCD training on synthetic data is meant to give it broad decision competence out of the box, and you can still fine-tune it (or its open siblings) on your own labeled decisions [Wikipedia].

2.2 Calibration is the product

Most of the "why" of Jev is actually about confidence. An LLM you prompt for a confidence score tends to be overconfident and inconsistent — it'll tell you 95% on a Tuesday and 60% on a Wednesday for the same input. TypeSafe's framing is blunt: if a model can do a task 95% of the time but doesn't reliably tell you when it's in the 5%, it cannot automate that task. You can't build a production workflow on a guess about when the guess is good. RLCD is the answer to that: the optimization target is calibrated decisions — probability values that are honest about themselves — which makes it possible to write if p > 0.85: file_ticket() and mean it [TypeSafe announcement].

2.3 The constraints are part of the design

Jev's type system is deliberately narrow, and the limits are documented:

  • All three question types are multiple-choice in spirit: pick from options (choice), yes/no with a probability (noul), or rate on a scale (score).
  • No dependencies between questions. Everything is evaluated in parallel and in isolation against the same state — so if question B depends on question A's answer, you make two sequential requests. That's a real workflow constraint, not a bug [Wikipedia].
  • Up to 64 questions per request, up to 255 options per choice, 10 levels per score. Choices above ~255 cardinality fall back to a two-stage score-then-pick, which is why high-cardinality demos occasionally slow down [Unsloth docs].

2.4 The name is the thesis

Jev is named after economist William Stanley Jevons — of Jevons paradox, where coal efficiency gains made coal cheaper and total coal consumption went up. The bet: every order-of-magnitude drop in the cost of intelligence unlocks more use cases than it saves. A decision call that costs a few hundredths of a cent changes the math on everything from per-frame game AI to per-paragraph document triage [TypeSafe announcement].

3. Why It Matters

Here's the argument in four moves.

3.1 It changes the unit economics of a decision

Today, if your software needs a judgment — route this ticket, flag this document, score this lead — the standard answer is "call the LLM." That means paying input and output tokens (with output running ~5× input on most frontier pricing), parsing the returned string, validating it against your schema, and handling the occasional hallucinated field or refusal. Jev collapses that to $0.042 per million input tokens and zero output cost — the company's published claim is up to 193.6× faster and 444.6× cheaper than the reference LLM workflow on its own evals, and it flags that as the high end [TypeSafe announcement]. Even at a quarter of the claimed multiple, the economics cross a line: decisions become cheap enough to make per event, not per batch.

3.2 Latency opens real-time categories

70–500ms end-to-end is a different world from 3–329 seconds. It's the difference between "AI in the workflow" and "AI in the loop" — i.e., per-frame, per-keystroke, per-packet. The company's own demos make the point: a Doom-playing agent making 10 structured decisions per second (~$7/hour at their pricing), and a Wikiracing bot picking from thousands of links at interactive speed, where not hallucinating on high-cardinality choices compounds over every step [TypeSafe announcement]. For a business, the same property applies anywhere UX or SLA is the constraint: real-time fraud signals, live sentiment on a call, in-flight routing of support traffic.

3.3 It automates what LLMs can only assist

The calibration point is the deep one. An overconfident model can't be automated even when it's mostly right, because your code has no honest signal for when to defer to a human. Jev's contract is that confidence is meaningful: higher reported confidence correlates with higher accuracy, and similar inputs get similar answers. That turns "AI review" into a smart if-statement — a fuzzy decision rule that ordinary software can branch on, with a defensible threshold and a clean escalation path for the low-confidence tail [TypeSafe confidence docs]. That's what "type-safe automation" actually means in practice: the surrounding code constrains the model's freedom, so the whole thing composes into a reliable system.

3.4 It completes the agent stack, not replaces it

Jev is not a drop-in replacement for the LLM. It doesn't write prose, code, or plans. The architecture that's emerging — and that LangChain's team has already written up — is a harness: the LLM stays the brain for open-ended reasoning and generation, while a System One model handles the fast, structured decisions along the way — routing, classification, guardrail checks, branch selection — without a full chat call for each one [LangChain]. An agent loop that would have made 20 LLM calls now makes a handful, plus a stream of cent-free, 100ms decision calls. That's a direct hit on agent cost and latency, the two numbers every operator is trying to bring down.

The one-sentence version

Jev is a frontier-intelligence function call: unstructured state in, typed probabilistic decisions out — 40–200× faster than an LLM for the same decision, and it can't hallucinate a field that isn't in the schema.

4. Path One: Jev via TypeSafe's Cloud

4.1 Getting in

Jev is in early access. The flow: sign up at typesafe.ai (developers are being pulled off the waitlist), then work from the console at console.typesafe.ai. The fastest orientation is the Playground: paste any text as the state — a support ticket, an email thread, a paragraph of code — add one or more questions (Noul / Choice / Score), and watch every question come back in parallel with its probabilities [Quickstart].

4.2 The raw API

One endpoint, one shape. Grab an API key from Settings → Keys and POST:

curl -X POST https://api.typesafe.ai/v1/systemone \
  -H "Authorization: Bearer $TYPESAFE_API_KEY" \
  -H "Content-Type: application/json" \
  -d @- <<'EOF'
  {
    "state": "Hi, I've been trying to connect my Stripe account for 3 days and it keeps failing. I'm losing sales. Please help ASAP.",
    "model": "jev-latest",
    "questions": {
      "department": {
        "type": "choice",
        "instructions": "Which team should handle this",
        "criteria": {
          "billing": "Payment or subscription issues",
          "technical": "Bugs or integration problems",
          "sales": "Pricing or account questions"
        }
      },
      "frustration": {
        "type": "score",
        "instructions": "How frustrated the customer appears",
        "criteria": [
          "Calm, just stating facts",
          "Frustrated but civil",
          "Very angry, strong language"
        ]
      },
      "is_urgent": {
        "type": "noul",
        "instructions": "The message conveys urgency or time-sensitivity"
      }
    }
  }
EOF

The response is what your code actually branches on — typed, no parsing required:

{
  "model": "jev-latest",
  "answers": {
    "department": {
      "type": "choice",
      "choice": "technical",
      "probabilities": { "billing": 0.159, "technical": 0.84, "sales": 0.001 },
      "confidence": 0.596
    },
    "frustration": {
      "type": "score",
      "score": 1.035,
      "legend": { "0": "Calm, just stating facts", "1": "Frustrated but civil", "2": "Very angry, strong language" },
      "confidence": 0.842
    },
    "is_urgent": { "type": "noul", "noul": 0.999 }
  },
  "usage": { "input_tokens": 312, "output_tokens": 48 }
}

How to read the two numbers

probabilities tells you which option won; confidence tells you how strongly the distribution is peaked (close to 1 = one answer stands out, close to 0 = evenly spread). Act on the per-option probability (e.g., file a ticket only when technical > 0.8), and use confidence to decide whether to auto-act or escalate. Noul returns a single 0–1 probability of "yes."

4.3 The Python SDK

For anything real, use the typed SDK (Python ≥ 3.10):

pip install typesafe-sdk

from typesafe_sdk import Choice, Noul, Score, TypeSafeClient

client = TypeSafeClient()  # reads TYPESAFE_API_KEY from the environment

response = client.system_one(
    state=ticket_text,
    questions={
        "department": Choice(
            instructions="Which team should handle this",
            criteria={
                "billing": "Payment or subscription issues",
                "technical": "Bugs or integration problems",
                "sales": "Pricing or account questions",
            },
        ),
        "frustration": Score(
            instructions="How frustrated the customer appears",
            criteria=["Calm, just stating facts", "Frustrated but civil", "Very angry, strong language"],
        ),
        "is_urgent": Noul(instructions="The message conveys urgency or time-sensitivity"),
    ),
)

print(response.answers["department"].choice)  # "technical"
print(response.answers["frustration"].score)  # 1.035
print(response.answers["is_urgent"].noul)     # 0.999

4.4 The agent skill

TypeSafe also ships an agent skill so your coding agent knows the API by heart — install it in Claude Code with claude plugin marketplace add typesafe-ai/skills + claude plugin install typesafe@typesafe-ai, or in any other agent with npx skills add typesafe-ai/skills --skill typesafe-ai, then just tell your agent to build with it [Agent skill docs].

4.5 Cost and limits at a glance

ItemValue
Input tokens$0.042 / MTok
Output tokensFree ("too cheap to meter")
Latency70–500ms end-to-end
Questions per requestUp to 64 (all in parallel)
Options per choiceUp to 255 (two-stage above that)
Score levelsUp to 10
Model aliasjev-latest

5. Path Two: Running Decision Models Locally

Here's the part most of the coverage skipped. Jev itself is a closed cloud model, but the architecture is now being replicated in open models, and Unsloth ships a complete local stack: serve the open decision models Laya and Clef through a TypeSafe-compatible Jev API, or fine-tune your own decision model from Qwen or Gemma weights [Unsloth decision models]. "Laya" is the reference open-weight model (678MB — smaller than a large photo library file); "Clef" is Cloudflare's open decision model, which shares the scoring-head design.

5.1 Why local works at all

The unsupervised reason this is practical: a decision model never spends compute writing text. Unsloth's implementation puts your state and every question into one prompt, lets the base LLM read it once, then a small scoring head — the same design as Cloudflare's Clef — looks at the LLM's output over each question and option and scores them all together. Because it never does autoregressive text sampling, a sub-billion-parameter model on a plain CPU answers in well under a second [Unsloth training guide].

5.2 Serve Laya locally (10-minute setup)

Step one, get Unsloth running — either the desktop app (macOS / Windows / Linux) or the one-liner install:

# macOS, Linux, or WSL
curl -fsSL https://unsloth.ai/install.sh | sh

# Windows PowerShell
irm https://unsloth.ai/install.ps1 | iex

Step two, in the app: Settings → API → Decision API → turn on Serve requests. Confirm the download of the default Laya model (678MB). Your API key lives on the same page under Access tokens — or enable Keyless API access → Chat and inference to skip the key on localhost [Unsloth docs].

Step three, send a decision request. This classifies a support ticket, detects a refund request, and scores urgency — three question types in one call, same payload shape as the cloud:

curl http://localhost:8888/v1/systemone \
  -H "Authorization: Bearer sk-unsloth-YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "laya",
    "state": "Hi, I was charged twice for my March invoice (#4411). Please refund the duplicate today.",
    "questions": {
      "team": {
        "type": "choice",
        "instructions": "Which team should handle this?",
        "criteria": {
          "billing": "invoices, payments, refunds",
          "technical": "bugs, outages, errors",
          "sales": "pricing, new plans",
          "other": "everything else"
        }
      },
      "refund": { "type": "noul", "instructions": "Does the customer ask for a refund?" },
      "urgency": { "type": "score", "instructions": "How urgent is this?", "criteria": ["not urgent", "soon", "today"] }
    }
  }'

The first request loads the model (10–20 seconds); after that, requests take well under a second on most CPUs — no GPU required for inference. The model field accepts laya, default, or jev-latest (whatever you selected in settings), plus the specific variants laya-multilingual, laya-english, and laya-typed-decisions.

5.3 Point existing Jev code at your machine

This is the elegant part: the local server speaks the TypeSafe protocol. If you already wrote cloud Jev code, migrating it to your own hardware is two environment variables:

pip install typesafe-sdk
export TYPESAFE_BASE_URL=http://localhost:8888
export TYPESAFE_API_KEY=sk-unsloth-YOUR_KEY

The SDK's default model, jev-latest, maps to whichever model you picked in Unsloth's settings, and answers come back in the same format as the cloud. That means a privacy-first deployment is a config change, not a rewrite: support tickets, HR documents, or client data stop leaving the building, and you've got a TypeSafe-compatible decision API running on a laptop.

Don't port your confidence thresholds blindly

Laya and Jev calculate confidence differently. If your cloud code had a Jev-tuned threshold (e.g., "auto-act when confidence > 0.9"), re-tune it against local Laya, or — better — act on the per-option probabilities instead, which behave more consistently across the two [Unsloth docs].

5.4 Train your own decision model (40 minutes, 4GB VRAM)

Unsloth can fine-tune regular LLMs into decision models — same dataset format as Laya and Clef, and you can even further fine-tune Laya or Clef themselves. Current reference numbers from the docs:

Base modelTest accuracyVRAMTraining time
Qwen3.5-0.8B78%4GB42 min
Qwen3.5-2B81%8GB40 min
Llama 3.2 3B79%4.1GB30 min
Gemma 4 E4B77%14.4GB49 min
Laya (fine-tuned)77%2.5GB10 min

The settings that reproduce those numbers: 2 epochs, LoRA rank 16, learning rate 2e-4, dataset subset all. Unsloth reports accuracy on held-out decisions before and after training, and calibrates the model's confidence as part of the run. Dataset rows are the same triple you've been seeing: a state, the questions to decide, and the gold answers [Unsloth training guide].

To serve a model you trained from code, start Unsloth pointing at the merged weights and turn on the Decision API:

UNSLOTH_SYSTEMONE_MODEL=/path/to/qwen-decisions-merged unsloth studio -H 0.0.0.0 -p 8888

The model needs a GPU and loads on the first request (a few minutes); while a training run has the GPU, the Decision API waits and resumes when the run ends. Note what this means for the stack: the cloud path gets you the frontier Jev; the local path gets you your own decision model — domain-specific, private, and free to run — through the identical API surface.

6. The Pattern: A Harness, Not a Replacement

LangChain's write-up is the clearest articulation of how these pieces fit together, and it's worth internalizing because it decides how you should architect around Jev-class models [LangChain]:

  1. The LLM is the System 2 layer. Open-ended reasoning, drafting, planning, anything that needs to say new things in new words. Keep it for what it's actually good at.
  2. The System One model is the System 1 layer. Every structured judgment the agent makes along the way: which tool, which branch, does this output pass the guardrail, is this a refund request, what's the urgency. These used to be cheap LLM calls; they're now 100ms decision calls with honest probabilities.
  3. Code composes them. Atomic questions in, combined by your own logic out. The docs' own guidance: if a judgment needs extended reasoning or weighs multiple independent factors, decompose it into separate questions and combine the results with a formula in your code — "when priorities shift, change a coefficient in your code rather than rewriting a prompt" [TypeSafe docs].
  4. Confidence is the control plane. High-confidence answers auto-act; low-confidence ones escalate to the LLM or to a human. Because calibration is the training objective, that control plane is something you can actually trust — which is the entire reason the model class exists.

KodeKloud's walkthrough of Jev lands the same architecture from the practitioner side: use it for routing and classification inside LLM workflows, turn the probabilities directly into software decisions, and deploy it as guardrails ahead of — or behind — the models you already run [KodeKloud].

7. What to Watch

  • The early-access queue. Jev is still being released to developers in waves. If you can get in, the highest-value experiment is cost-per-workflow: take one real automation you currently run on LLM calls (ticket triage, document scoring, guardrail checks) and measure the actual latency and cost multiple — not TypeSafe's 193×/444× headline, your own.
  • Local parity. Laya and Clef are open today; whether their accuracy keeps up as Jev's frontier moves — and whether Unsloth's fine-tuning pipeline closes the gap on your own data — is the question that decides whether "local Jev" becomes a real category. The 77–81% reference accuracies on a benchmark set are a promising start [Unsloth].
  • Multi-model routing. The System One LLM wrapper TypeSafe published (which constrains regular LLMs to the same structured decision output) means you can A/B Jev against your existing LLM on identical workflows [System One adapter]. Expect vendor-neutral harnesses — LangChain included — to standardize this soon.
  • The no-dependency limit. Sequential dependencies require two requests. As workflows get more deeply chained, watch whether the API grows a lightweight dependency mechanism or whether decomposition stays the sanctioned pattern — it affects how deep you can push real production graphs.
  • Sustainability of the pricing. TypeSafe is candid that it can't yet prove the $0.042/MTok input price isn't subsidized. If it holds, the Jevons logic kicks in hard: decision counts go up orders of magnitude. If it doesn't, the local path stops being a privacy option and becomes the cost option.

References

Companion report: TypeSafe's Jev: System One Models — Frontier-Intelligence Decisions in 70ms, Without Hallucinating covers the launch itself. This report covers access, architecture, and the local stack.