330-run Terminal-Bench 4.0 test: Astra gains ~8 points low-to-high, then plateaus
BLUECOW009 · x · 2026-09-06
A 330-run Terminal-Bench 4.0 comparison shows Astra going from 50.61% at Low → 54.24% Medium → 57.88% High → 57.88% XHigh → 58.18% Max. Gains are 8 points from Low to High, but High, XHigh and Max differ by only 1% — reasoning-effort returns flatten fast.
More from Models
- OpenAI and Anthropic models share a favorite name, writing benchmark finds — almmaasoglu · 2026-09-07
- Five-model writing style comparison shows average sentence length misses the mark — almmaasoglu · 2026-09-07
- Felda AI builds writing benchmark analyzing names, punctuation and sentence rhythm — almmaasoglu · 2026-09-07
- Matt Shumer: I 4x'd my Pro/Max subs as agentic models remove the attention bottleneck — mattshumer_ · 2026-09-07
- AI startup Felda dissects writing quality down to sentence rhythm and punctuation in its new benchmark — almmaasoglu · 2026-09-07
- Teaching Qwen Next 3D sculpting in Blender by having GPT Astra coach it via MCP — LegacyRemaster · 2026-09-07