Dev benchmark: Astra beats sol and fable on LOC, API cost and code quality in 10-15 feature test
robleclerc · x · 2026-09-19
Rob Leclerc shares his method for evaluating new models: have different models at the same thinking level implement the same 10-15 small features, then judge the results.
Astra clearly beat both sol and fable on lines of code, implied API cost, and quality — in several cases sol wrote 10x as much code as Astra.
Quoting theo: for real-world code work, the benefits of Fable and Astra massively outweigh the cost, though the benefit isn't better code per se — it's something subtler.
More from coding & agent
- Split coding and testing across two agents to save weekly quota — ___Patrice___ · 2026-09-19
- onPanda: open-source token-level LLM editor that can harness Claude Code and Codex — Fancy_Fanqi77 · 2026-09-19
- Muse use-case directory adds 90 new entries, now 678 total with ready-to-run prompts — armand_ruiz · 2026-09-19
- Skip the waitlist: running Jev's evaluate API via Vercel AI Gateway for $0.00074 — AlchainHust · 2026-09-19
- Simulated Worlds for Sonnet 3.6: Bureaucracy, Commutes and TONS OF KINDNESS — repligate · 2026-09-19
- Anthropic's Fable 5.1 lands in Kiro, built for long-running context-heavy agentic sessions — DigitalColmer · 2026-09-19