Simonw on 'decision models': evals matter even more than for regular LLM projects
HamelHusain · x · 2026-09-22
Simon Willison published notes on Jev and a new "system one" category he calls decision models. Isaac Flath amplified the key takeaway, endorsed by Hamel Husain: in practice, evals and structured experiments matter even more for these systems than for regular LLM projects.
The core argument: when models make decisions directly rather than generate content, rigorous evaluation and experiment design become the decisive factor.
More from coding & agent
- Dev Swaps Opus for Mimo-v2.6 in Cline on Client Projects: 'It's a Beast' — MicahBerkley · 2026-09-22
- Dev ships Chrome extension driving WebMCP tools with small model Jev, splitting agent workloads — _AustinCalvert_ · 2026-09-22
- Kev refactored onto Qwen3.5: open-source decision models now at 0.8B, 4B and 9B — alexcovo_eth · 2026-09-22
- Community 1000x-ing official docs: jevify prompt makes coding agents scout Jev use cases — alexcovo_eth · 2026-09-22
- Fireworks: routing 18 models per task hits 97.6% solve rate at $1.88 vs best single model's 74.1% at $6.52 — sophiamyang · 2026-09-22
- LlamaIndex adds calibrated confidence scores to LlamaParse Extract, field by field — llama_index · 2026-09-22