Open-source Ornith-1.5 achieves Claude Opus-level performance via self-improvement
aftahi_ai · x · 2026-08-19
Ornith-1.5, a family of open-source LLMs (9B Dense, 35B/397B MoE), introduces a complete end-to-end self-improvement loop. It achieves SOTA among comparable open-source models and rivals Claude Opus 4.8 in reasoning, agentic, and coding tasks.
Key benchmarks include:
- Terminal-Bench 2.1: 86.1
- SWE-Bench (Verified): 86
- DeepSWE: 56
- HLE: 44.6
The model extends the self-scaffolding strategies from Ornith-1.0.
More from Models
- Grok Voice Think Fast 2.0 beats GPT Realtime on VulcanBench, 99% text accuracy — XFreeze · 2026-08-20
- Armin Ronacher explains how reasoning effort lives in the system prompt and wrecks KV cache — mitsuhiko · 2026-08-20
- Experiment shows ChatGPT struggles to follow the 'do not respond' command — cranberryalarmclock · 2026-08-20
- In real C++ debugging, open-source Qwen 3.8 27B showed better engineering judgment than Gemini 3.7 Flash — GravyPoo · 2026-08-20
- Qwen 3.8 27B thinks until context runs out, never produces output locally — WyattTheSkid · 2026-08-20
- LiquidAI releases LFM 2.5 QAD, GGUF weights on Hugging Face — jacek2023 · 2026-08-20