Hobbyist says nightly RL on just 400-600 rollouts makes model gains noticeably deployable

cephaloform · x · 2026-09-04

@cephaloform shared a hands-on impression of small-scale reinforcement learning: running only 400-600 rollouts per night, yet claiming to feel the model improve with every deployed checkpoint. In a follow-up he mused he "maybe doesn't need a 35B" model — hinting small models plus light RL fine-tuning may suffice. A relatable data point for hobbyist fine-tuners.

Related event: Small-Scale RL: A Few Hundred Rollouts Per Night May Be Enough(2 posts)→

Original post →

More from Research

Research channel →