TypeSafe hits $100M ARR and $7.5B valuation as decision models become a category overnight
Latent Space · rss · 2026-10-10
Headline: TypeSafe/Jev
- TypeSafe's "Series AI" round comes with Sequoia "leaking" that it crossed $100M ARR in week one and a $7.5B valuation just 3 weeks after launch, despite astroturfing accusations.
Decision Models Become a Category
- OpenAI Decisions API: three request types (probability, pick-from-list, score), running on GPT-6 Luna at $0.10/M input tokens with no output charge, up to 10x faster.
- Microsoft Decision-1 targets LLM judges and hypothesis screening; early evaluators flag consistency issues.
- Perplexity pplx-decider-v1.1-27b: 94.5% on Decision Bench (1,071 cases) at $0.017/1K decisions.
- Cloudflare clef-omni handles audio/video/image/text; clef-flash undercuts Jev, weights open-sourced. Liquid d1 ships on Vercel AI Gateway with vision.
- DIY: Unsloth's free notebook turns Qwen3.5-4B into a decision model on 8GB VRAM; Qwen3.5-0.8B accuracy rose 37%→65% in 60 steps (10 min on 4GB).
- Why it matters: many agent steps are yes/no calls; LangChain reports routing to the cheapest adequate model cut median Open SWE cost per task by 64%.
- Research: Apple/CMU's Selection-based Structured Reasoning scores six strategies in one batched pass sharing the KV cache — per-turn latency down 90%+ but end-to-end only 28–54%.
Agents & Coding Tools
- Claude Managed Agents (public beta): lead agent fans out to up to 1,000 agents per run; Anthropic warns about token burn. Claude Code Projects fully rolled out; Opus 5.5 fast mode bills separately.
- Vals AI: agent teams cost 1.8–5.1x more; only GPT-6 Sol at medium effort improved significantly (+7.3 pts). Opus 5.5 ran 1,140 subagent tool calls per app with no significant gain.
- Prime Agent self-rewrote in Rust: 2,000+ agents, 10K+ sandboxes, 200B+ GLM-5.3 tokens — 13x faster usable input, 83% less startup memory.
- Codex: new Windows sandbox (MXC), Pro-only Composer next-message predictions (users push back), daylong outages reported; DHH says GPT-6.1 Sol made Codex his primary tool over Claude.
- Devin: spawns trees of managed Devins; accepts personal ChatGPT plans.
Model Releases & Evals
- Qwen-Image-2.1-Turbo (open weights): 8-step 2K generation + NL editing.
- StepFun Step 5 Preview: 600B/27B-active sparse MoE, 1M context, Hermes Index 33.89 matching GPT-6 Luna; open weights due Oct 15.
- Upstage Solar Mini 4: 35B MoE/3B active, 524K context, 208 tok/s, AAII 24 — best at 3B active.
- Gemini 4 Argon: 77.9% on DeepSWE v1.1 vs Opus 5.5's 74.2%; unconfirmed "Carbon" checkpoint reportedly nears Opus on coding.
- Voice: HeyGen Voice tops TTS arena (Elo 1,201); Whistle is a 16.9MB on-device STT rivaling Whisper base.
- Multi-turn image editing: after 30 chained edits, Ideogram 4.5 and FLUX 3 retain 95%+ local fidelity.
More from coding & agent
- Dev Builds an "ML Team" of Agents: AutoML Pipeline From Feature Engineering to Deployment — kmeanskaran · 2026-10-10
- MCP for Blender hits 26.6k stars as author shares Opus 5.5 rendering workflow tricks — sidahuj · 2026-10-10
- Open-source self-hosted connector layer gives personal agents Claude Code's toolset — shensi · 2026-10-10
- Team ranks top 100 AI agent skills across 12 categories from 5,200+ reviewed — NathanWilbanks_ · 2026-10-10
- Review AI code in a fresh session — ideally a different model — to catch bugs the agent misses — DanielLockyer · 2026-10-10
- Amp now supports Claude Pro/Max subscriptions for free via Claude Agent SDK — iannuttall · 2026-10-10