I ran 16 models to vet one tool: one task is not a benchmark
AlexKim · x · 2026-09-19
To decide if TypeSafe's Jev belonged in his stack, the author benchmarked 16 models — Jev came 10th on accuracy. His first draft claimed "Haiku was more accurate," true on one task only; adding two more made it a split: Haiku wins business categories 83.2–79.9, Jev wins commit types 50.0–42.0, dead tie on prose at 66.0. Lesson: one task is not a benchmark.
More from coding & agent
- Jevable site catalogs 342 demos of Typeface's Jev, from $0.0039 flight search to intent-scored spreadsheets — gaganghotra_ · 2026-09-19
- GitHub runs beginner livestream series on Copilot SDK for building agentic apps without the agent loop — 0xkarasy · 2026-09-19
- Stripe opens Muse connector program for agentic payments, first connectors go live — jeff_weinstein · 2026-09-19
- AI agent plays Subway Surfers at superhuman speed, 50 games at once, for under a cent — ravithejads · 2026-09-19
- User proposes Codex thread usage panel to track per-thread impact on weekly quota — CtrlAltDwayne · 2026-09-19
- MSP IT Admins Want Agents That Can Patch, Heal and Report — Copilot Isn't It — PerfectReflection155 · 2026-09-19