Amazon open-sources Rufus-Air: full 8-stage post-training recipe on GLM-4.5-Air-Base
amazon · hf · 2026-09-25
Amazon released Rufus-Air, an open and reproducible post-training recipe built on GLM-4.5-Air-Base (106B-A12B), organized as a serial eight-stage pipeline: SFT, Reasoning RL, Coding RL, Instruction-Following RL, General Agent, Coding Agent, Search Agent, and RLHF.
- Full documentation of data, reward design, infrastructure, stage ordering, and stagewise results; built on open-source components and public data, with no new human annotation or in-house distillation teacher.
- Key findings: (i) diverse, high-quality SFT sets a strong capability floor; (ii) difficulty filtering keeps RL prompts in a productive range; (iii) reward reliability is the practical principle for ordering stages; (iv) infrastructure choices are part of the recipe, not a detail.
- Outperforms the official GLM-4.5-Air post-trained release and is competitive with similarly sized open models.
More from Research
- NVIDIA Open-Sources Nemotron-Terminal: Boosts Qwen3-32B from 3.4% to 27.4% on Terminal-Bench 2.0 — _weiping · 2026-09-25
- Cusp AI co-founder Max Welling on AI-powered materials design — NandoDF · 2026-09-25
- Sharpa's World Synesthesia Model for dexterous hands accepted at CoRL 2026 — jiqizhixin · 2026-09-25
- DeltaWAM cuts video model cost for bimanual manipulation, lifting RoboTwin success to 85.4% — Han Yan · 2026-09-25
- Researcher proposes agent-suggests-human-executes loop for real-world science experiments — suragnair · 2026-09-25
- MLPerf Training v6.1 adds first LLM post-training benchmark: agentic RL on a 397B model — TheKanter · 2026-09-25