Ornith-1.5 Launches: 397B Model Matches Claude Opus in Coding Benchmarks
testingcatalog · x · 2026-08-19
DeepReinforce released the Ornith-1.5 family of open models in 397B MoE, 35B MoE, and 9B Dense sizes. The flagship 397B model scores 86.1 on Terminal-Bench 2.1 and 86 on SWE-bench Verified, matching the performance of Claude Opus 4.8. The 35B model, activating only 3B parameters per token, significantly outperforms peers like Qwen and Gemma on agentic coding tasks. A quantized 9B-Mobile build is also available for on-device deployment on iPhone and Android.
More from Models
- Grok Voice Think Fast 2.0 beats GPT Realtime on VulcanBench, 99% text accuracy — XFreeze · 2026-08-20
- Armin Ronacher explains how reasoning effort lives in the system prompt and wrecks KV cache — mitsuhiko · 2026-08-20
- Experiment shows ChatGPT struggles to follow the 'do not respond' command — cranberryalarmclock · 2026-08-20
- In real C++ debugging, open-source Qwen 3.8 27B showed better engineering judgment than Gemini 3.7 Flash — GravyPoo · 2026-08-20
- Qwen 3.8 27B thinks until context runs out, never produces output locally — WyattTheSkid · 2026-08-20
- LiquidAI releases LFM 2.5 QAD, GGUF weights on Hugging Face — jacek2023 · 2026-08-20