Reviewers are asking for bigger open models, not just 8B baselines
chhaviyadav_ · x · 2026-07-25
Researchers say 8B is no longer enough for open-model papers
The post says the community has clearly moved beyond 8B models, and reviewers now ask for evaluations on larger open-source models.
- The author mentions being asked for evaluations on GPT-OSS (21B MoE).
- They note that such models also need fine-tuning.
- The question to the NeurIPS community is essentially: which open models are people using in their papers now?
The post is short, but it signals a shift in expectations: reviewers want stronger, larger open models and more serious evaluation coverage rather than small-model baselines.
Related event: NeurIPS Reviewers Demand Larger Open-Source Models(2 posts)→
More from Research
- Stanford HAI and ETS say AI is reshaping education assessment — StanfordHAI · 2026-07-25
- For-profit AI benchmarks may hide noise behind tiny score gaps — PerformanceRound7913 · 2026-07-25
- Release blog teaser shows a near-tie on FrontierCode agentic coding benchmark — hardmaru · 2026-07-25
- A smaller distilled LLM may beat its teacher by being forced to generalize — peterjliu · 2026-07-25
- Opus 5 tops nine biology benchmarks, but still trails in some analysis tasks — kenbwork · 2026-07-25
- A new ARC-AGI reading says symbolic AI is having a comeback — typewriters · 2026-07-25