LLM judge quality hinges on pretraining scale; full online RL alone won't crack continual learning

herbiebradley · x · 2026-08-29

Herbie Bradley responds to two technical questions. On evaluating longer tasks with LLM judges, the difficulty isn't length but whether the judge is good enough to capture subtle qualitative judgements — which depends heavily on pretraining scale and grinding out data coverage. On continual learning, it depends on the in-weight method, but challenges like RL sample inefficiency make him skeptical that full online RL alone would solve it.

Related event: Researchers Debate Whether AI Can Keep Stacking S-Curves(16 posts)→

Original post →

More from Research

Research channel →