> ## Content Index
> Fetch the complete content index at: https://debugly.dev/llms.txt
> Use this file to discover other available public pages before exploring further.

# The AI Model That Cannot Hallucinate and Can Still Be Wrong
- URL: https://debugly.dev/jev-decision-model-schema-guarantee/
- Published: 2026-10-07T13:30:00.000Z
- Updated: 2026-10-07T13:30:00.000Z
- Author: Rohit Bhadani
- Tags: AI, Debugging, Architecture

The model cannot hallucinate, and the cannot was the pitch, and the pitch was the true, and the true was the narrow, and the narrow was the schema, and the schema was the declared options, and the declared options were the only outputs, and the only outputs were the guarantee, and the guarantee was the not the correctness, and the not correctness was the wrong answer, and the wrong answer was the valid, and the valid was the incident.

Jev is worth understanding this month, and not because of the benchmarks. It is the first model to ship a genuinely different contract to application code, and the contract is being described in a way that will cause a specific and predictable category of production bug.

TypeSafe AI released Jev on 15 September 2026, the same day it came out of stealth with a $40M seed led by DCVC. The company was founded by Diogo Almeida, who co-invented RLHF and InstructGPT at OpenAI. The name is after William Stanley Jevons, of the Jevons paradox, and the bet is exactly that paradox: when machine decisions get cheap enough, you stop rationing them.

## What Jev actually is

Jev is not an LLM and it does not generate text. You pass it state, meaning text or JSON, plus one or more typed questions you declare in code, and it returns answers to those questions and stops. There is no prose, no rationale, and no code.

There are three primitives, and every call is built from them. A **Choice** picks from a list of options you declared. A **Score** rates against levels numbered from zero, and the returned value is the probability-weighted mean, so it can land between levels, which means 1.4 on a three-level scale is a distribution and not a verdict. A **Noul** is a yes-or-no question returning a probability between zero and one.

The architecture is non-autoregressive. A parallel sampler computes all outputs in a single pass against a single shared read of the state, which is where the speed comes from: TypeSafe reports 70 to 500 milliseconds end to end, against seconds-to-minutes for a frontier chat model. Adding a tenth question costs input tokens but almost no extra time. Pricing is $0.042 per million input tokens with output unmetered and free. The context window is 64,000 tokens with a 32,000-token budget for the state. Training uses a method TypeSafe calls RLCD, Reinforcement Learning for Calibrated Decisions.

On TypeSafe's own four-workflow benchmark Jev scores 67.8 per cent, level with GPT-5.6 Terra at 67.9 and behind GPT-5.6 Sol at 74.1 and Claude Opus 5 at 73.1, at roughly a two-hundredth of the cost. That is an honest framing and it matters: Jev wins on latency and price, not on intelligence.

## The guarantee people are mishearing

"Cannot hallucinate" is doing a lot of work in the marketing, and it is technically defensible if you read it narrowly. What TypeSafe means is that Jev has no free-form output surface, so it cannot return a malformed value or an option you did not declare. Structured-output errors are zero by construction, which is a real and useful property, because the parse-and-validate layer that wraps every LLM call in production simply disappears.

What it does not mean is that the answer is right. Jev can return a perfectly valid schema containing a confidently wrong choice. The type system guarantees the shape of the answer and says nothing about its truth. Anyone who reads "cannot hallucinate" as "cannot be wrong" will remove the verification step, and the removal is the bug.

This is the same trap as [temperature zero not being determinism](https://debugly.dev/temperature-zero-is-not-deterministic/), where a real technical property was read as a much stronger promise than it makes, and the gap between the two became the production incident.

There is a second misreading that is subtler. For Choice and Score, confidence measures how concentrated the probability distribution is, not the chance the judgement is correct. A model can be sharply confident and wrong, so a threshold that auto-acts above 0.9 confidence is a threshold on sharpness, and sharpness is not accuracy until you have measured it on your own data.

## Where the bugs actually move

Three properties of the interface shift failure into your code rather than removing it.

**Questions are independent within a call.** One answer cannot feed another question in the same request. That is fine architecturally, but it means dependencies, control flow, calculations and side effects all stay in your application, and that is exactly the layer with the least test coverage in most pipelines.

**Constrained output is not injection-proof.** TypeSafe's own model jaggedness notes for Jev 1.13 state that the model does not treat state as hostile by default, so injected instructions or deliberately misleading framing can move the answer. A classifier that cannot emit a malformed value can still be talked into picking the wrong valid one.

**The state is your job.** Jev does not fetch context. Whatever you pass is all it sees, so retrieval quality becomes the ceiling on decision quality, and a decision model will give you a crisp typed answer about a stale or incomplete state without any indication that the input was thin.

There is also an honest data point in the early demos. Someone built a real-time autonomous trading bot on Jev in an evening and reported that it had, so far, lost them $31,680\. That is the lesson in one sentence: fast, cheap, well-typed decisions are still decisions, and a decision engine will execute a bad strategy at a frequency no human could match.

## What to do before you ship it

**Calibrate on your own data before trusting the confidence.** Run a few thousand labelled examples through your real questions and plot confidence against accuracy. If the curve is not monotonic, the confidence score is not a gate you can act on.

**Keep a verification path for the consequential cases.** Route the high-stakes or low-confidence decisions to a frontier model or a human. Jev's economics make it the right tool for the ninety-five per cent that is routine, which is an argument for a tiered design rather than a replacement.

**Treat the state as untrusted input.** If any part of it originated from a user, sanitise it, because the schema constraint protects the output shape and nothing else.

**Version the questions.** The questions are declared in code, which means they can be reviewed, diffed and rolled back. Use that. A silent change to a Choice's option list is a silent change in behaviour, and it will not show up as a type error if you only added an option.

## The rule

A decision model that returns typed values guarantees the shape of the answer and nothing about its truth, and its confidence measures distribution sharpness rather than correctness. Calibrate against your own labelled data, keep a verification path for consequential calls, treat the state as untrusted, and version the questions like code.

Jev is the most interesting model release of the quarter precisely because it gives up generation, and the trade is real: zero malformed outputs, sub-second latency, and a cost low enough to call ten times a second. But "cannot hallucinate" describes the schema and not the judgement, and the teams that will get hurt are the ones that heard the first and deleted the second, because a wrong answer that arrives in 80 milliseconds and type-checks perfectly is much harder to notice than a wrong answer that arrives as a paragraph.