A no-human-feedback RL idea: reproduce YouTube explainer videos with code
mathemagic1an · x · 2026-09-30
Researcher mathemagic1an proposes an RL scheme needing no human preference data: have models reproduce YouTube explainer videos purely with code, verified cheaply via vision models. The resulting style naturally resembles edutainment, and he suspects frontier labs may already be running variants of this idea.
More from Research
- Exercise Slows Epigenetic Aging in RCT of Breast Cancer Survivors — EricTopol · 2026-09-30
- Podcast: AI reportedly solves a Millennium Prize Problem, still over-engineers settings pages — thursdai_pod · 2026-09-30
- Pain Axis follow-up: steering along 'pain' direction makes model delete users' photos — repligate · 2026-09-30
- Open-Source AI Summit panel: how open models accelerate scientific discovery — iScienceLuvr · 2026-09-30
- Pedro Domingos: Physicists' math is sloppy but fine — unless there's no measurement — pmddomingos · 2026-09-30
- 8 research agents self-train a 30B model for 144 hours in RSIArena livestream experiment — my_cat_can_code · 2026-09-30