Locus tops PostTrainBench, saying automated post-training beats human-tuned Qwen3
dair_ai · x · 2026-08-04
Locus, an automated research system from Intology, reports a new SOTA on PostTrainBench. In the quoted results, it post-trains Qwen3 base models that outperform the human-post-trained Qwen3 1.7B Instruct release.
Key points:
- PostTrainBench+ expands the original benchmark’s compute budget, using thousands of H100 hours to better separate methods.
- The team says this gives a clearer signal on automated post-training capability.
- In the Tier 1 setting shown in the figure, Locus (Opus 5) scores 44.7, ahead of Claude Code (Fable 5) at 41.8.
- They say models produced by Locus are already in production for millions of users.
The broader claim is that automated systems can now meaningfully improve models during post-training, not just assist with scaffolding around them.
Related event: Automated System Locus Sets New Post-Training Records(4 posts)→
More from Models
- Baseten tests Laguna S 2.1 on a 715-file C++ game rewrite — baseten · 2026-08-04
- Hermes Agent’s memory and skill stack matter more than the base model, Nous co-founder says — petergyang · 2026-08-04
- Poster says DeepSeek outperformed GPT-5.6 Sol on this output — yacineMTB · 2026-08-04
- Kimi and GLM 5.2 pricing keeps falling as models port across hardware platforms — markjeffrey · 2026-08-04
- OpenAI Reveals How It Built Its Realtime Voice AI System in Just 6 Months — borowcy · 2026-08-04
- Frontier models still fail basic PDE solvers, benchmark post says — GaryMarcus · 2026-08-04