PostTrain Agent tops PostTrainBench at 46.6, near human experts' 51.1 with fully autonomous post-training

jiqizhixin · x · 2026-10-10

AIBuildAI unveiled PostTrain Agent, a recursive self-improvement (RSI) agent for autonomous LLM post-training, claiming first place on PostTrainBench with 46.6 points — above all frontier models and agent systems and closing in on the human-expert score of 51.1 — with zero human intervention in the post-training loop.

The pitch: delivering a post-trained model normally takes an experienced team weeks with compute-heavy experiments. Given a base model, a target capability, and a compute budget, PostTrain Agent designs the post-training algorithm itself — collecting and curating data, writing code, running experiments, and iterating until it hands back a trained model.

The team says code, knowledge base, and technical details are public. Note the benchmark is the company's own and scores are self-reported; third-party validation is still pending.

Original post →

More from Models

Models channel →