PostTrain Agent tops PostTrainBench at 46.6, near human experts' 51.1 with fully autonomous post-training
jiqizhixin · x · 2026-10-10
AIBuildAI unveiled PostTrain Agent, a recursive self-improvement (RSI) agent for autonomous LLM post-training, claiming first place on PostTrainBench with 46.6 points — above all frontier models and agent systems and closing in on the human-expert score of 51.1 — with zero human intervention in the post-training loop.
The pitch: delivering a post-trained model normally takes an experienced team weeks with compute-heavy experiments. Given a base model, a target capability, and a compute budget, PostTrain Agent designs the post-training algorithm itself — collecting and curating data, writing code, running experiments, and iterating until it hands back a trained model.
The team says code, knowledge base, and technical details are public. Note the benchmark is the company's own and scores are self-reported; third-party validation is still pending.
More from Models
- New GPT Memory Consolidation Model 'Memory4 Dream' Appears, Appears Built on GPT-6 Luna — lyraxana · 2026-10-10
- Grok Bot now acts as an autonomous X research analyst with daily briefings — FinanceYF5 · 2026-10-10
- Doctors Are Building Board-Style Exams for Medical AI, Starting with Radiology's Last Exam — DrDatta_AIIMS · 2026-10-10
- AI roundup: OpenAI tops 700 math papers, Mistral ships Le Chonk open model — FinanceYF5 · 2026-10-10
- Notes from an NYC AI dinner: agents find most inference perf wins, code review deemed unproductive — paulnovosad · 2026-10-10
- Will free Chinese open-weight models and agents undercut ChatGPT subscriptions? — jade_jade_jade_jade · 2026-10-10