Locus says its automated post-training system beats human-tuned Qwen3 with thousands of H100 hours
_akhaliq · x · 2026-08-04
Locus scales automated post-training with more H100 compute
An automated AI research system called Locus claims state-of-the-art results on PostTrainBench and shows that post-training performance keeps improving as compute scales.
- On PostTrainBench+, the authors expand the benchmark’s compute budget to better distinguish automated post-training methods.
- They report that with thousands of H100 hours, methods trained by Locus on Qwen3 1.7B-Base can outperform the official human post-trained Qwen3 1.7B model.
- The system is also tested on live Kaggle competitions to examine generalization.
- According to the post, Locus-generated post-trained LLMs are already in production for millions of users.
More from Research
- Open-source CAD computer-use environments add 50 engineering tasks — DevvMandal · 2026-08-04
- AI writing can often be spotted by long sentences, nominalizations, and overused “and” — Afinetheorem · 2026-08-04
- Stanford researcher calibrates synthetic data using historical tasks — arena · 2026-08-04
- How Athena Crisis Built a Fast, Deterministic Game AI Without LLMs — cnakazawa · 2026-08-04
- Bittensor subnet 107 says OpenAI co-authored a field report on agentic scientific computing — markjeffrey · 2026-08-04
- Humanoid robot vaults and climbs unseen terrain in real time with a real→sim→real loop — chris_j_paxton · 2026-08-04