Which Jev Claims Are Verified?

Every headline figure from the launch, sorted into what is checkable today, what is company-reported, and what TypeSafe itself declined to publish.

Jev benchmark verification: none published No third party has reproduced the Jev benchmark figures. Last checked September 19, 2026.

Jev claims, one by one

Every Jev claim sorted into what you can check yourself, what rests on the company's own testing, and what was simply not published. This is not a verdict on whether the Jev claims are true — several may well be. It is a record of which Jev claims currently have evidence behind them.

70–500ms end-to-end latency Company-reported

Published in the launch blog. Consistent with the Doom demo's 10 queries per second, though that is also a company demo.

40x–200x faster than frontier models Company-reported

A range rather than a measurement. TypeSafe notes results "are likely to sit at the high end of real-world results".

193.6x faster in workflow evals Self-tested

From the company's own eval suite. No third-party has reproduced it as of September 2026.

444.6x cheaper in workflow evals Self-tested

Same suite, same caveat. The comparison wraps external models to produce compatible output, which raises their latency and cost.

$0.042 per million input tokens Verifiable

A published rate card, not a benchmark. You can check it against your own invoice once you have access.

Output tokens free Verifiable

Also a rate-card fact. Follows naturally from responses being typed values rather than long sequences.

0% type errors Asserted

TypeSafe explicitly states this figure "is not empirical" — it follows from schema design rather than from measurement.

Cannot hallucinate Partly true

Cannot return a value outside your schema. Can still return the wrong valid value, which is a different failure and not covered by the claim.

$40M seed led by DCVC Verified

Confirmed by the company announcement and independent reporting from multiple outlets on September 15–16, 2026.

Revenue, customers, valuation Not disclosed

Absent from the launch materials. Named customers and error rates on real data are what would settle the economics question.

What is wrong with the Jev benchmark

The headline Jev benchmark figures — 193.6 times faster, 444.6 times cheaper — come from a workflow evaluation the company built and ran. That is normal at launch. What is worth understanding is the specific Jev benchmark design choice that makes the ratio hard to interpret.

The Jev benchmark has no answer key

The workflow evals do not score against ground truth. They measure each model against the average probabilities returned by two large external models — agreement with a committee, not correctness.

Jev benchmark baselines were wrapped

To produce compatible structured output, the external models were wrapped by TypeSafe. Wrapping adds latency and cost to the baseline, which flatters the ratio being reported.

The company says so itself

TypeSafe acknowledges possible bias and states the figures sit at the high end of real-world gains. That candour is worth more than the numbers, and it is easy to miss in coverage that quotes only the multiple.

What would settle the Jev claims

Independent evaluation on third-party data, published error rates from named deployments, and pricing that holds after early access ends. None of those exist yet.

Credit where the Jev claims deserve it

It would be easy to write this Jev claims audit as a takedown, and it would be unfair. The company states in its own materials that the results carry possible bias and are likely to sit at the high end of real-world gains. It says the zero-hallucination Jev claim is not empirical. It lists, by name, the Jev questions its launch post does not answer.

That is more disclosure than most launches manage, and it inverts the usual problem. Normally the caveats have to be reconstructed by outsiders reading between the lines. Here the Jev caveats are in the source material, and the distortion happens downstream — in coverage that quotes the Jev benchmark multiple and drops the sentence beside it.

So the honest summary is not "the Jev claims are inflated". It is "the Jev claims are unaudited, and the company says so". Those are different situations, and only the second is compatible with the Jev benchmark numbers turning out to be broadly right.

What would settle the Jev claims

Three kinds of Jev evidence, in rough order of how much each would move the picture. First, independent Jev evaluation on data the company did not select — accuracy and, more importantly, calibration quality measured by someone with no stake in the result.

Second, named Jev deployments with published error rates. A customer saying "we route this volume at this accuracy and here is what it costs us" is worth more than any Jev benchmark, because it includes all the Jev integration friction that evaluations leave out.

Third, Jev pricing that survives the end of early access. A Jev rate card offered to a waitlist is a hypothesis about unit economics. The test is what it looks like a year after general availability, under real load, with margin expectations attached.

Until then the reasonable posture on the Jev claims is neither dismissal nor adoption on faith. Jev is cheap enough to test that you can generate your own evidence for the price of an afternoon — and given that everything public traces back to one source, your own numbers are worth more than anyone's summary of theirs.

Status: This Jev claims audit is maintained. When independent evaluations appear the table above is updated and the reasoning kept. Status as of September 19, 2026: no third-party Jev benchmark reproduction published.

How to test the Jev claims yourself

The unusual thing about this launch is that the Jev claims are cheap to check. You do not need a research team or a benchmark suite. You need a few hundred cases you already have labels for, a schema, and an afternoon once your invitation lands.

Run the sample, record the answer and the confidence on each case, then bucket by confidence and compute observed accuracy per bucket. That single curve tests the two Jev claims that matter — whether it is right often enough, and whether the confidence figure is honest — on your data rather than the company's.

Time the calls while you are at it and total the spend. Those two numbers test the remaining Jev claims directly, and unlike the published multiples they are measured against your current pipeline rather than against a wrapped baseline somebody else chose. Your own figures are worth more than anyone's summary of the Jev benchmark, and they are the only evidence your own decision should rest on.

Jev claims questions

Has anyone independently benchmarked Jev?
Not as of September 2026. Every performance figure in circulation traces back to TypeSafe's own eval suite.
What is wrong with the benchmark design?
There is no objective answer key. Models are scored against the average probabilities of two large external models, which TypeSafe wrapped to produce compatible output — a step that raises the baseline's latency and cost.
Is the 0% hallucination figure real?
TypeSafe states directly that it is not empirical. It follows from schema constraints rather than measurement, and it covers fabrication only — not wrong answers inside your schema.
What would change the picture?
Third-party evaluation on independent data, published error rates from named customers, and pricing that holds once early access ends.

Keep reading