Ornith-1.5 released: Self-improving loop matches Claude Opus 4.8 performance
alexcovo_eth · x · 2026-08-20
DeepReinforce released Ornith-1.5, a family of open-source models (9B Dense / 35B MoE / 397B MoE) featuring a revolutionary self-improvement training loop (self-proposing tasks, scaffolding, RL rollouts). It achieves SOTA results on Terminal-Bench (86.1) and SWE-Bench Verified (86), rivaling Claude Opus 4.8. The 35B model outperforms Qwen 3.6, while the 9B quantized version runs on phones at just 1.5GB. Supports FP8/GGUF/MLX under MIT license.
Related event: Open-Source Ornith-1.5 Family Launches, 397B MoE Rivals Claude Opus 4.8(24 posts)→
More from Models
- NVIDIA Details Qwen3.8-2.4T Deployment on GB300, Achieving >4K Tokens/s per GPU — PyTorch · 2026-08-21
- Tutorial: Training a local LLM on a new domain via continued pretraining — funJS · 2026-08-21
- Qwen 3.8 27B two-shots a playable 3D amusement park game in the browser — RandumbRedditor1000 · 2026-08-21
- Zhipu's SAO: single-rollout async RL trains stably for 1,000 steps, beats GRPO — teortaxesTex · 2026-08-21
- Gemini's Safety Filters Too Strict? Rejects Kissing and Roadside Photos — Dry-Sympathy-3182 · 2026-08-21
- 0.63M Parameter Verifier Matches 7B Models in Specific Tasks — jm_alexia · 2026-08-21