Jev AI Can’t Hallucinate, But It Can Still Be Wrong
Jev AI cannot hallucinate in the sense of inventing values, because every answer is a probability over options the caller defined in advance. It cannot return an unlisted number, category or tool call. It can still select the wrong option, so the failure becomes a bounded classification error rather than fluent invention.
A model that can only answer from a menu you wrote cannot invent a number, but it can still pick the wrong item on the menu with great conviction, and the difference between those two failures is the whole story.
On 14 September I published AI Hallucination Is Not Lying. That’s Why It’s Dangerous, which walked through a frontier assistant inventing my LinkedIn follower count, inventing a citation to support it, folding under pushback, and then attaching a fabricated provenance to a figure that happened to be true. The argument ended in three places: that a hallucinating model is a measurement instrument that does not reliably expose its error bars, that our evaluation regimes reward a confident guess over an honest abstention, and that the real enterprise risk is provenance collapse, where retrieved evidence, user assertions and outright invention arrive in one paragraph with nothing to mark the seams. The following day, TypeSafe AI came out of stealth with a model called Jev and a launch post whose headline claim is that the model can’t hallucinate. The timing was a coincidence, but the claim lands squarely on the argument I had just made, so it deserves a careful look at what it actually means, which of my three problems it addresses, and which it quietly hands back to the people building on top of it.
1. What Jev actually is
Jev is not a chatbot, not a coding model, and not a smaller language model in the usual sense, and it helps to be concrete about the interface before evaluating any claim made about it. You send it a block of state, which is text or a JSON object of text, together with a set of typed questions, and it returns one structured answer per question rather than a string. There are three question types: a Noul, which returns a probability that a statement about the state is true; a Choice, which selects one option from a list you define and returns the probability of every option alongside a confidence figure; and a Score, which places the state on an ordered scale whose levels you describe, again with probabilities and confidence. TypeSafe describes this class as System One models, after Kahneman’s distinction between fast intuitive judgment and slow deliberate reasoning, and says the model was trained with a method it calls Reinforcement Learning for Calibrated Decisions, which on the company’s account optimises the returned probabilities against outcomes rather than against the preferences of human raters (TypeSafe, 2026).
The company was founded in 2024 by Diogo Almeida, Erik Gafni and Sasha Sheng, raised a $40 million seed round led by DCVC, and Almeida previously worked at OpenAI on reinforcement learning from human feedback, InstructGPT and ChatGPT (TechStock², 2026). That background matters to this discussion, because the founder’s stated motivation is that the preference training he helped pioneer produced systems that are extremely good at pleasing people and much less good at making reliable decisions, which is a close cousin of the sycophancy argument in section 9 of the original post. The published pricing is $0.042 per million input tokens with output unmetered, and the company quotes end to end response times of 70 to 500 milliseconds. TypeSafe has not published a technical paper, the weights, or the architecture in any detail, and it has chosen not to report results on public benchmarks, so almost everything we know about how the model behaves comes either from the company or from early users.
2. What “cannot hallucinate” actually guarantees
The guarantee is real, it is narrow, and TypeSafe is more honest about its narrowness than most of the coverage has been. Because every answer is a probability distribution over options that the caller defined in advance, the model cannot return a value outside the supplied schema, which means it cannot produce a malformed output, cannot invent a category that does not exist, and cannot emit a tool call to a function that is not there. In the launch post the company plots its type error rate at zero next to empirical rates for language models, and then states plainly in the accompanying notes: “Our number is not empirical.” The zero is structural, a property of the interface rather than an observed outcome, and that is precisely why it holds. It would be falsified by a single counterexample and there cannot be one.
Now replay my follower count exchange against that interface. To ask Jev how many followers I have, you would have to supply the candidate answers yourself, perhaps as a Choice over bands such as under 10,000, 10,000 to 50,000, 50,000 to 150,000 and over 150,000, alongside whatever state you had assembled about me. The model could not return 13,210, because 13,210 is not on the menu, and it could not attach a fabricated hyperlink, because it does not emit links, and it could not narrate “verified platform data”, because it does not narrate anything. What it could do is select the wrong band. If the state you gave it contained nothing relevant, it would still distribute its probability across your four options, because a Choice must choose, and the answer that came back would be one of your own options with a probability attached.
So the precise claim is that Jev converts an unbounded failure into a bounded one. A language model that does not know an answer can fail in an effectively infinite number of ways, most of them fluent and many of them plausible, whereas a Choice over four options can only fail by selecting one of the other three. That is not the same as being right, and the independent coverage has been careful to say so: a fixed schema makes a malformed output unlikely to break the surrounding program, but it does nothing to make the underlying decision correct (TechStock², 2026). What the conversion does buy you is something quite valuable, which is that a wrong selection from a closed set is a classification error, and classification error is a failure mode the industry has known how to measure, threshold and monitor for decades. Hallucination has been renamed into misclassification, and misclassification is governable in a way that open ended invention never was.
3. Error bars, finally, but only if the probabilities are honest
The original post described a hallucinating system as a measurement instrument that does not reliably expose its error bars or its failure indicator, and noted that making uncertainty legible was an active research area rather than an abandoned one. Jev is a commercial attempt at exactly that, and its documentation states the principle more bluntly than most vendors would: a system that cannot express honest uncertainty cannot be trusted (TypeSafe docs). Every Choice and Score answer returns the full probability distribution across the options, and every Noul is itself a probability, so the uncertainty is not something you have to prompt for and then disbelieve. It is the output.
The part worth reading slowly is how the confidence figure is produced. The documentation explains that confidence is a statistic derived from the shape of the returned distribution, high when the probability is concentrated on one option and low when it is spread out, and that the caller is free to compute a different statistic from the raw probabilities if that suits the domain better. This is a sensible design, but it has a consequence that is easy to miss, which is that the confidence figure contains no information that is not already in the probabilities. If the model’s probabilities are well calibrated, meaning that answers given 90% probability turn out correct about 90% of the time on the population you care about, then the confidence figure is a genuinely useful gate. If the probabilities are overconfident, the confidence figure is overconfident in exactly the same way, and you have a confident looking number derived from a confident looking distribution with nothing independent underneath either of them.
That is why the most interesting thing about Jev is not the schema at all, but the training objective. Section 8 of the original post drew on two findings: that the GPT-4 technical report observed calibration being reduced by post training, and Kalai and colleagues’ argument that accuracy based grading makes a confident guess strictly better than an admission of uncertainty. RLCD, as described, is an attempt to change the incentive at its source by rewarding probabilities that match outcomes rather than answers that please a rater or score on a leaderboard. If it works as described, it addresses the root cause I identified rather than a symptom of it. The honest qualification is that TypeSafe has published no calibration metric, no reliability diagram and no paper describing RLCD, so the claim that the probabilities are calibrated is, at the time of writing, a claim. Calibration is also a property of a population rather than of a model in the abstract, so even a model that is well calibrated on the distribution it was trained on can drift when it meets your support tickets, your transaction narratives or your contract clauses. TypeSafe’s own guidance is to start with conservative thresholds and tune them against your own data, which is the right advice and also an acknowledgement that the calibration you need is the calibration you measure yourself.
4. There is no apology to extract
Section 9 of the original post argued that the assistant’s apology was the same failure one level up, and leaned on research showing that challenging a model with “are you sure?” moves its answers regardless of whether they were right. Jev changes the shape of that problem in a way I did not expect to find so clean. There is no conversation to push back in, no reassurance to demand, and no commitment to extract, because the model returns probabilities and stops. The documentation also states that each question in a request is evaluated independently and in parallel against the same state, so one answer never becomes context for another, which removes the mechanism by which an earlier concession contaminates a later answer within the same call.
That does not make Jev immune to steering, and it would be a mistake to read it that way. Text inside the state can still be written to move the answer, and TypeSafe’s own published list of known weaknesses for the current model version acknowledges that adversarial content in the state can shift results and says it expects to improve on this (Copes, 2026). The sycophancy vector has moved from the person asking the question to whoever controls the content being evaluated, which in an enterprise pipeline is frequently a customer, a counterparty or an attacker. That is a more familiar threat model, and a more tractable one, but it is still a threat model, and anything that sits in front of an irreversible action has to be tested with hostile inputs before it goes anywhere near production.
5. Provenance does not collapse, it becomes your problem
This is where the claim of eliminating hallucination most needs a correction, because the central argument of the original post was never really about hallucination as a discrete event. It was about provenance collapse, the flattening of supplied, retrieved, recalled and invented material into one undifferentiated surface. Jev does not solve that, but it does something I think is more useful than solving it, which is that it makes the problem impossible to ignore.
A language model answering a question draws on its weights, its retrieval layer, the conversation history and its own inferences, and nothing in the output tells you which. Jev evaluates questions against the state you supplied and returns nothing but judgments about that state, which means the evidence boundary is a JSON object that your code assembled, can log, and can reproduce. It cannot pretend to have browsed a profile it could not reach, because it does not browse, and it cannot silently upgrade a user’s screenshot into “observed on the live public profile”, because it does not report sources at all. The provenance of every Jev answer is, by construction, whatever you put in the request.
The catch is that everything in section 6 of the original post, the unrendered pages, the crawler controls, the login walls and the stale indexes, now sits upstream of the model in whichever component builds the state. If that component feeds Jev a weak or incomplete source, the model will return a carefully distributed judgment about weak or incomplete material, and the confidence figure will describe how clear the answer looks given the state, not whether the state was any good. A low confidence answer is a useful warning that the state did not contain enough to go on, and the documentation says as much, but a high confidence answer about the wrong document is still a high confidence answer. The retrieval problem has not been solved. It has been moved to a place where it is visible, testable and owned by you, which is the correct place for it, but only if someone actually takes ownership.
6. The benchmarks measure agreement, not truth
TypeSafe’s headline figures of 193.6 times faster and 444.6 times cheaper come from its own workflow evaluations, and the company’s notes on those evaluations are unusually candid. The workflows were built by members of its own model capabilities team, which it acknowledges could introduce bias; the company describes the gains as likely to sit at the high end of real world results; and its latency figures are measured from the US West Coast, where the service currently runs (TypeSafe, 2026). For anyone calling it from Johannesburg, the network round trip alone will be a material fraction of the advertised response time, so the speed claim needs to be measured from where you actually sit.
The more important detail for this discussion is what the evaluations use as the reference answer. There is no ground truth answer key; instead, each model’s outputs are compared against the averaged probabilities of two large frontier models, GPT-6 Astra and Fable 5.1, run through a wrapper that forces them to produce Jev compatible structured decisions. That is a defensible way to measure whether a cheaper model reproduces the judgments of expensive ones inside a fixed workflow, and TypeSafe notes the bias it introduces towards those two vendors. What it does not measure is correctness, because the reference is itself the output of the class of systems whose reliability was the subject of my original post. The fair reading is narrower than the headline: on TypeSafe’s own workflows, Jev produced judgments close to a frontier consensus at a fraction of the cost and latency. That is a meaningful result, but it tells you nothing about how often the consensus was right, and the only way to learn that is against labelled data from your own domain.
The published weaknesses are also worth taking seriously precisely because the company published them. The current version does not count reliably, cannot do arithmetic, reads dates as text rather than as dates, loses accuracy through indirection and double negatives, and degrades as the state fills with irrelevant material (Copes, 2026). None of those is surprising for a model of this kind, and all of them point to the same design rule, which is that anything exact belongs in code and only the fuzzy judgment belongs in the model.
7. Where the accountability goes
Section 10 of the original post separated the chain into three layers, the model, the organisation that designs and ships it, and the organisation that consumes its output, and argued that the concept of recklessness applies perfectly well to the two layers made of people. It also argued that the design layer mostly chooses not to surface uncertainty, because a system that frequently says it could not check something demos worse than one that always has an answer. Jev is a clear counterexample on that specific point, and it deserves credit for it. The vendor has chosen to surface uncertainty on every answer, to publish a list of what its model does badly, and to annotate its own benchmark claims with the reasons to discount them.
The effect is to move a great deal of responsibility onto the consuming organisation, and to make that responsibility legible in a way it has never been before. The worked example in TypeSafe’s confidence documentation is, conveniently for a banker, a user asking to approve a pending withdrawal: below a confidence of 0.5 the request goes to a human, and the transfer only proceeds without explicit user confirmation above 0.9. Those are the documentation’s illustrative values rather than recommendations, but the structure is the point. The risk tolerance is no longer buried in the tone of a generated paragraph. It is a number, written in code, reviewable in a pull request, and attributable to whoever approved it. If an organisation sets a threshold of 0.5 in front of an irreversible payment, that is not an unlucky hallucination. It is a documented decision about how much error to accept, made by people, and the Derry v Peek analogy from the original post applies to it without any strain at all, because the person who chose the number knew exactly what the number meant.
I think that is the healthiest thing about the design. With a chat model, an organisation can always claim, at least rhetorically, that the machine surprised it. With a decision model that reports its uncertainty on every call, the only way to be surprised is to have chosen not to look.
8. Jevons, applied to errors
The name is not incidental. TypeSafe named the model after William Stanley Jevons, whose observation about coal was that more efficient steam engines increased total consumption rather than reducing it, and the company says it expects the same of machine intelligence: every order of magnitude reduction in the cost of a decision unlocks far more places to make one. I suspect they are right, and I think the same logic applies to the errors, which is the part nobody puts on a launch slide.
If a judgment that used to cost a meaningful fraction of a cent and several seconds now costs a small fraction of that and a tenth of a second, organisations will not make the same number of decisions more cheaply. They will make vastly more decisions, in places where no model call was ever economic before, on every transaction narrative, every log line, every document paragraph and every inbound message. Even a model that is well calibrated and correct 98% of the time produces two wrong decisions in every hundred, and when the volume rises by two orders of magnitude, the absolute number of wrong decisions rises with it, most of them made in pipelines where no person ever looks at an individual answer. Calibration is what lets you route the uncertain cases to a human, but it only works if the thresholds are set against measured accuracy, if the review capacity actually exists at the new volume, and if someone is monitoring whether the confident answers are still right as the input distribution drifts. Cheap decisions are an operational risk problem wearing an efficiency argument, and the governance has to scale with the call volume rather than with the budget line.
9. How I would evaluate it
None of this is an argument against using Jev, and several of its properties map directly onto the controls the original post recommended. That post suggested decomposing an answer into atomic claims, verifying each one separately and surfacing the uncertainty through the surrounding system, because the fluent paragraph will never carry it on its own, and a decision model is almost purpose built for the verification half of that pattern. A language model writes the summary; Jev asks, one Noul per claim, whether the retrieved source actually supports each sentence; code decides which low probability claims go back for review. That is a far more honest architecture than trusting the generator’s own citation chips, and it keeps the three provenance paths, supplied, retrieved and generated, separate at the point where a person or a process can still see the seams.
If I were putting it anywhere near a production decision, I would run it in shadow mode beside the existing process and change nothing for several weeks, logging the full state, the full probability distribution and the exact model version for every call, since the response reports which versioned model answered and the thresholds you tune are only valid for that version. I would build a labelled set from our own data and plot confidence against observed accuracy before trusting any threshold, because the only calibration that matters is calibration on our population. I would keep every question and every threshold in one reviewed file, so the risk tolerance lives somewhere an auditor can read it. I would test the state with hostile inputs before it evaluates anything a customer or counterparty can write. And I would reserve automatic action for decisions that are both high confidence and cheaply reversible, keeping anything irreversible behind a human or a stronger process however good the numbers look, at least until we have months of measured accuracy rather than weeks.
10. What Jev actually fixes
The original post ended with the observation that hallucination is not lying, and that this should make us more concerned rather than less, because a liar has a model of the truth you can extract and a hallucinating system has no brake that scrutiny can engage. Jev does not give the machine a model of the truth either, and it does not make its answers correct. What it does is narrower and, I think, more important than its marketing line. It removes the open ended failure by closing the answer space, it attaches an explicit and inspectable uncertainty to every answer instead of leaving it buried in tone, it removes the conversational loop in which a model can be talked into agreement, and it forces the provenance boundary into a data structure that the calling organisation owns and can log.
Put more directly, it turns hallucination back into ordinary error, and ordinary error is something organisations already know how to govern if they choose to. The questions that remain are the ones that were always there underneath the language: whether the probabilities are genuinely calibrated on your data, whether the state you assembled deserved to be trusted, and whether the people setting the thresholds understood what they were accepting. Those were never questions a model could answer for you, and the most useful thing Jev does is refuse to pretend otherwise.
11. A closing note on this post
As with the original, I wrote this with AI assistance, and every source linked above was opened and read before it went in. Jev launched on 15 September and has been publicly available for about two weeks, TypeSafe has published no paper or independent evaluation, and several of the figures quoted here are the company’s own, which I have tried to label as such in the sentence rather than in a disclaimer. If any of this changes as independent measurements appear, which I expect it will, the claims above should be read as conditional on what was published at the time of writing.