Synthetic pretraining with RL-trained generators: gradient-overlap reward explained
cephaloform · x · 2026-09-27
A new approach trains LMs on fully synthetic data: the learner uses standard next-token prediction on generator output bytes, while the generator itself is trained with RL. A simple difficulty reward fails (random bytes make output arbitrarily hard without structure), so the reward uses the AdamW-preconditioned gradient inner product with the learner's recent trajectory — matching programs to the learner's learnable frontier. Commenters link it to neural cellular automata training and the fully-synthetic pretraining thesis.
More from Models
- ScienceArena benchmark: LLMs score 64.5% on chemistry tasks needing structural diagrams vs 74.1% without — geoffwolfe · 2026-09-27
- GLM-5.3 Flash Matches Claude at 1/429th the Price in a YouTube Script Benchmark — OnlyProggingForFun · 2026-09-27
- Frontier AI is now so cheap and abundant that subscriptions go barely used — intellectronica · 2026-09-27
- Karpathy: Claude Opus 4.5 beats GPT-5 Pro for interactive history learning — doodlestein · 2026-09-27
- ChatGPT-6 Astra cracks 85-year-old 1941 Enigma message in two days — luisdans · 2026-09-27
- Grok accused of uploading user chat images to the web as Musk says 'this keeps getting worse' — EthanJPerez · 2026-09-27