Dev finds typesafeai too small for spec-quality assessment; Gemini Flash wins but costs
julianharris · x · 2026-09-19
julianharris spent two days deep in @typesafeai for content-quality assessment of software specs and found it miles short — clearly a small model lacking the needed intelligence. Gemini Flash 3.8 blows it away, but is too expensive for his taste. He's asking for alternative suggestions.
More from Models
- LLM fails coin-flip probability test, says 60/40 coin is 99/1 — JnBrymn · 2026-09-19
- Deep Dive: Open-weight models now within ~4 months of best closed frontier models — markjeffrey · 2026-09-19
- Veteran ML engineer: Jev may push agent tool-calling back to discriminative models — multiply_matrix · 2026-09-19
- First Jev benchmark announced as Scoble touts 4M-post AI report database — airesearch12 · 2026-09-19
- ChatGPT co-inventor launches Jev, a frontier model claiming 20-200x speed and 40-400x cost gains — JFPuget · 2026-09-19
- 'Still using Claude Code? So 2025' — devs joke about switching to GLM — TheZachMueller · 2026-09-19