AI Research Needs More Failed Experiments: Kibitzer Training Recap
BlackHC · x · 2026-07-31
The post calls for academia to document and share failed AI experiments, arguing that current papers often only present 'just-so' stories of success, making it hard for new researchers to learn the actual research process.
The cited blog details the training recap of the Kibitzer model. Highlights include achieving an excellent size-to-strength ratio purely through supervised training, data scaling, and search (without RL). It covers the architecture (including an SSM hypothesis), final training recipe, Elo evaluation, and transparently shares failed RL experiments and what went wrong.
More from Research
- DeepSeek V4's fully compressed layers might bottleneck long-context performance — bookwormengr · 2026-07-31
- New Review on Opportunities for Legged Robots by Jonas Frey et al. — ChongZzZhang · 2026-07-31
- AI Compresses Months of Data Analysis into 1.5 Hours, Shifting Audiences to Agents — nathanbenaich · 2026-07-31
- Designing and Post-Training Edge Agentic Models: Slides & Talk — maximelabonne · 2026-07-31
- TTT3R: Test-Time Training Enhances Length Generalization in 3D Reconstruction — rsasaki0109 · 2026-07-31
- NTU Introduces Σ-Mem: Online Reliability Memory for Multi-Agent Systems — NanyangTechnologicalUniversity · 2026-07-31