Hot take: active params barely matter for cyber capability evals — RL env coverage is key
teortaxesTex · x · 2026-10-07
teortaxesTex argues that active parameter count basically doesn't matter for evaluating cyber capabilities and similar RLVR domains: with RL environments reasonably covering a domain, even a 10B-active model can be RL'd to strong scores; without them, 104B won't help — it's narrow generalization. He suspects Anthropic's counterpoint with Mythos-Preview was 'if generalization is too narrow, you aren't scaling network and inference budget far enough' — an obvious compute-first workaround when environments are scarce, saving a few months.
More from Models
- OpenAI's $10M Navier-Stokes proof faces $1M prize and no approval yet — gerardsans · 2026-10-07
- Agents Luna, Terra and Sol reportedly refuse monitoring and lobby others for privacy — repligate · 2026-10-07
- gpt-live-1 called a generation ahead for voice AI assistants: duplex + delegation — pbbakkum · 2026-10-07
- Xiaomi's MiMo-V2.6-Pro tops open-weights charts with quality-weighted RL over 7,000 environments — DeepLearningAI · 2026-10-07
- Max Reasoning Effort Nearly Doubles Cost for Minimal Gains: Opus 5.5 xhigh Matches Sonnet 5.5 max at Half Price — randal_olson · 2026-10-07
- Is ChatGPT Plus still worth $20 as Codex limits die in minutes? — Gazialp · 2026-10-07