TypeSafe exits stealth with $40M at $200M valuation, but Jev's calibration claims ship with zero public evidence
prakersh · reddit · 2026-09-21
TypeSafe AI exited stealth on September 15 with $40M led by DCVC at a reported $200M valuation, launching Jev, founded by ex-OpenAI researcher Diogo Almeida (InstructGPT, ChatGPT, GPT-4). The post dissects its two headline claims:
Claim 1: no hallucinations — true but narrow. Jev doesn't generate free text; it takes program state plus typed questions and returns typed answers with probabilities, so it physically can't emit fabricated output. But well-formed ≠ correct. The technique isn't new either: constrained decoding has guaranteed 100% JSON Schema conformance in OpenAI Structured Outputs since August 2024, and fixed-option scoring is how MMLU has been evaluated since 2020. Jev's real contribution is doing this zero-shot across domains.
Claim 2: calibrated probabilities — unmeasured publicly. The actual product is a probability on every answer, trained via "Reinforcement Learning for Calibrated Decisions." Yet there's no ECE, no reliability diagrams, no public benchmarks, no architecture paper. The company's own eval measures agreement with GPT-6 Astra and Claude Fable 5.1 rather than correctness, and the 193.6x speed / 444.6x cost multipliers compare against reasoning-off competitors — while the funding press release says only up to 100x.
What's genuinely strong: third-party adoption data. Vercel reports Jev as the fastest-adopted model in AI Gateway history — a tenth of paid teams within 18 hours, 13% by hour 24, double the first-day share of the GPT-5.6 family. TypeSafe's own docs are candid: Jev can't count reliably, struggles with hex/RGB, reads dates as text, and probabilities of a statement and its negation needn't sum to 1 — awkward next to a calibration pitch. Until someone independent measures calibration, the strongest evidence is fast developer adoption, which proves usefulness, not calibration.
Related event: TypeSafe Raises $40M but Calibration Claims Face Scrutiny(2 posts)→
More from Models
- Confident Nonsense: why polished AI answers are the ones you should trust least — nikola_mr64990 · 2026-09-21
- Qwen3.8-27B in native 8-bit hits 37-55 tok/s on Apple Silicon, avoiding the 4-bit reasoning cliff — SnooPredictions515 · 2026-09-21
- djev-run Deploys DiffusionGemma on Cloud Run's RTX PRO 6000 Blackwell for ~$3/hr — bodonoghue85 · 2026-09-21
- Zero-shot classification is cool again, and vboykis is happy the market noticed — vboykis · 2026-09-21
- Will ChatGPT merge chat and work modes like Claude — and what happens to limits? — koltregaskes · 2026-09-21
- Agents feel less reliable since September model updates, devs log failure rates to find out — john_snow_21 · 2026-09-21