Hot take: active params barely matter for cyber capability evals — RL env coverage is key

teortaxesTex · x · 2026-10-07

teortaxesTex argues that active parameter count basically doesn't matter for evaluating cyber capabilities and similar RLVR domains: with RL environments reasonably covering a domain, even a 10B-active model can be RL'd to strong scores; without them, 104B won't help — it's narrow generalization. He suspects Anthropic's counterpoint with Mythos-Preview was 'if generalization is too narrow, you aren't scaling network and inference budget far enough' — an obvious compute-first workaround when environments are scarce, saving a few months.

Original post →

More from Models

Models channel →