$40M Bet on Calibrated Confidence, Yet TypeSafe Publishes No Calibration Data
prakersh · reddit · 2026-09-21
TypeSafe AI emerged from stealth on September 15 with $40M led by DCVC at a reported $200M valuation, founded by ex-OpenAI researcher Diogo Almeida (InstructGPT, ChatGPT, GPT-4). Its model Jev doesn't generate text: it takes program state plus typed questions and returns typed answers with a probability on each, trained via what it calls Reinforcement Learning for Calibrated Decisions. The pitch is that those probabilities are calibrated — 70% should be right about 70% of the time.
The author's core objections:
- Calibration is the whole product, with no public measurement: no expected calibration error, no reliability diagrams, no standard public benchmark, no architecture paper. The only published eval is a company-designed one over four workflow tasks where the correct answer is defined as the average of GPT-6 Astra and Claude Fable 5.1 at high thinking — that measures agreement with two competitors, not correctness. Competing models run at default reasoning settings, so the 193.6x speed and 444.6x cost multipliers are measured against reasoning-off configs while the accuracy target comes from reasoning-on ones.
- "Cannot hallucinate" is true only narrowly: the output space is fixed before decoding, so it can't emit an off-list option, eliminating fabricated output but not guaranteeing the chosen option is correct; constrained decoding isn't new either, OpenAI has guaranteed JSON Schema conformance since August 2024.
Where skepticism should stop: Vercel reported Jev as the fastest-adopted model in AI Gateway history — a tenth of paid teams within 18 hours, nearly 13% by hour 24, roughly six times Fable 5.1's first-day share. And TypeSafe's own docs are unusually honest, publishing a jaggedness page admitting Jev doesn't count reliably, underperforms on hex and RGB values, can't judge whether two values are near each other, reads dates as text, and that there's no guaranteed mathematical relationship between semantically related outputs — so a statement and its negation need not sum to 1. Verdict: real product, real adoption, and a headline property nobody outside the company has measured.
Related event: TypeSafe Raises $40M but Calibration Claims Face Scrutiny(2 posts)→
More from Venture
- AI-assisted reread of Yandex leak: behavioral signals outnumber content signals 3 to 1 — tracyingram · 2026-09-21
- Smaller LLMs Aren't Always Cheaper: Review and Rework Costs Blow Up the Bill — zeuslac · 2026-09-21
- Indie maker claims Outrank grew SuperX impressions from 2k to 23k per day — tibo_maker · 2026-09-21
- China's top AI labs raised ~$35B in five months; combined haul could hit $60B by 2027 — FinanceYF5 · 2026-09-21
- AI is making startups harder, not easier — picking the right idea is everything — 0xsachi · 2026-09-21
- 'Spray-and-pray marketing reads as desperate': dev critique of 1M-view demo that converted no users — willcb · 2026-09-21