Follow-up: light nightly RL may mean you don't need a 35B model after all
cephaloform · x · 2026-09-04
In a follow-up to his post about running 400-600 RL rollouts per night, @cephaloform joked he "maybe doesn't need 3.8 35b" — suggesting a smaller model paired with light reinforcement learning already yields perceptible gains, without needing a 35B-class model.
Related event: Small-Scale RL: A Few Hundred Rollouts Per Night May Be Enough(2 posts)→
More from Research
- Agent's Last Exam tops out at 59.3% — the benchmark that matters for AI replacing humans — DevToD4 · 2026-09-04
- Has Anyone Tried Feeding All the Bio x ML Datasets to a Single Model? — iskander · 2026-09-04
- AI formalization capabilities hit no ceiling, but Lean may be starting to buckle — davidad · 2026-09-04
- K2 Horizon ships six fully open models from 0.9B to 375B, data and recipes included — HongyiWang10 · 2026-09-04
- LLM eval pipeline flaw: unsupported tool-call responses could score as perfectly stable — docybo · 2026-09-04
- Routing across 44 LLMs cuts error 54%: single-model benchmarks understate AI — testingcatalog · 2026-09-04