LLM judge quality hinges on pretraining scale; full online RL alone won't crack continual learning
herbiebradley · x · 2026-08-29
Herbie Bradley responds to two technical questions. On evaluating longer tasks with LLM judges, the difficulty isn't length but whether the judge is good enough to capture subtle qualitative judgements — which depends heavily on pretraining scale and grinding out data coverage. On continual learning, it depends on the in-weight method, but challenges like RL sample inefficiency make him skeptical that full online RL alone would solve it.
Related event: Researchers Debate Whether AI Can Keep Stacking S-Curves(16 posts)→
More from Research
- Learn Positional Encodings derivation from first principles — zainhas · 2026-08-30
- COLM Paper Traces Capability Provenance in LLMs via Gradient Attribution — ziv_ravid · 2026-08-30
- Toby Ord paper argues recursive self-improvement has physical limits — Exponential View (Azeem Azhar) · 2026-08-30
- AI Formalization Tools Fable and Sol Spot First Repairable Error in Published Literature — Sauers_ · 2026-08-30
- Mark Schmidt Posts ICML Tutorial Video: Is Numerical Optimization Theory Irrelevant to ML Practice in 2026? — MarkSchmidtUBC · 2026-08-30
- SDF Donut in 46 Lines of Python — voooooogel · 2026-08-30