Open-Source Ornith-1.5 Matches Claude Opus in Reasoning and Coding
aftahi_ai · x · 2026-08-19
The Ornith-1.5 family of open-source models has been released, featuring 9B Dense, 35B MoE, and 397B MoE variants. Trained with end-to-end self-improvement strategies, it claims performance comparable to Claude Opus 4.8 across reasoning, agentic, and coding tasks.
Key benchmarks:
- Terminal-Bench 2.1: 86.1
- SWE-Bench (verified): 86
- DeepSWE: 56
- Tool Decathlon: 71.2
The model achieves a more complete self-improvement loop by extending self-scaffolding strategies from Ornith-1.0.
More from Models
- Grok Voice Think Fast 2.0 beats GPT Realtime on VulcanBench, 99% text accuracy — XFreeze · 2026-08-20
- Armin Ronacher explains how reasoning effort lives in the system prompt and wrecks KV cache — mitsuhiko · 2026-08-20
- Experiment shows ChatGPT struggles to follow the 'do not respond' command — cranberryalarmclock · 2026-08-20
- In real C++ debugging, open-source Qwen 3.8 27B showed better engineering judgment than Gemini 3.7 Flash — GravyPoo · 2026-08-20
- Qwen 3.8 27B thinks until context runs out, never produces output locally — WyattTheSkid · 2026-08-20
- LiquidAI releases LFM 2.5 QAD, GGUF weights on Hugging Face — jacek2023 · 2026-08-20