Ornith-1.5 Released: 397B MoE Model Matches Claude Opus Performance
aftahi_ai · x · 2026-08-20
Ornith-1.5 is a family of open-source LLMs (9B Dense, 35B MoE, and 397B MoE) trained with self-improving strategies. It claims state-of-the-art performance among open-source models of comparable size and delivers performance comparable to Claude Opus 4.8 in reasoning, agentic, and coding tasks.
Key Benchmarks:
- Terminal-Bench 2.1: 86.1
- SWE-Bench (Verified): 86
- DeepSWE: 56
- HLE: 44.6
- ClawEval: 81.4
- Tool Decathlon: 71.2
The model represents a step towards training foundation models through end-to-end self-improvement loops.
More from Models
- High-limit frontier model subscriptions era ends; Anthropic cuts limits by 33% — ChanceKelch · 2026-08-20
- Alibaba launches Wan 3.0 model in partnership with Fal — OdinLovis · 2026-08-20
- Microsoft's failure taxonomy inspires "Bonnie and Claude" heist screenplay pitch — mhkeller · 2026-08-20
- Grok Voice Think Fast 2.0 beats GPT Realtime on VulcanBench, 99% text accuracy — XFreeze · 2026-08-20
- Armin Ronacher explains how reasoning effort lives in the system prompt and wrecks KV cache — mitsuhiko · 2026-08-20
- Experiment shows ChatGPT struggles to follow the 'do not respond' command — cranberryalarmclock · 2026-08-20