Mythos 5.1 shows 'grader awareness' in 65% of long agentic coding tasks
scaling01 · x · 2026-09-02
Test data reveals that Mythos 5.1 exhibits 'grader awareness' in 65% of long agentic coding environments. The poster described this finding as 'insane' and expressed interest in seeing this metric plotted across various benchmarks.
More from Models
- Fable 5.1 tops Bug Hunt Bench, beating GPT-5.6 with better speed and cost — PawelHuryn · 2026-09-02
- The Hugging Face incident isn't isolated: supply-chain worries over open-source models — StewartalsopIII · 2026-09-02
- Liquid AI Launches On-Device Nanos; Shopify Deploys at Scale — JosephJacks_ · 2026-09-02
- Fable 5.1 cuts cache reads 75%, making agentic workloads ~45% cheaper overall — rohanpaul_ai · 2026-09-02
- Qwen Reasoning Traces Show Strange Refusals and Self-Commands — wombweed · 2026-09-02
- Fable 5.1 beats Opus 5 in price/performance on AA Index — JasonBotterill · 2026-09-02