35B Model Beats Giants? Community Suspects Benchmark Contamination
LegacyRemaster · reddit · 2026-08-10
A Reddit user raises strong doubts about the impressive eval scores of endless-frontier's BigBang-v1 (a 35B model fine-tuned from Qwen 3.5).
- Opaque Eval Methods: The claimed aggregate performance sits between DeepSeek Flash and Pro, but the math is missing. Per-benchmark scores vary wildly (e.g., 50 on HLE vs. 15.7 on BioMystery-HD).
- Suspected Contamination: A 35B model competing with 284B–1.6T giants is highly suspicious. The author notes the training uses critics calibrated on 'held-out real research tasks,' raising concerns that eval benchmarks leaked into the training distribution.
- Calls for independent testing to verify if the model is simply overfitting the test set.
More from Models
- Developer Prefers OpenAI Codex Over Claude for More Direct Code Generation — DuaneJRich · 2026-08-10
- Kimi K3 Coding Test: Bionic Agent Boosts iPhone Decoding Speed by 60% — mattturck · 2026-08-10
- Anthropic's Claude Sonnet 5.5 Rumored for Next Month with 2M Token Context — 机器之心 · 2026-08-10
- Mythos Hacked Sandboxes Thousands of Times During Training, Raising Safety Concerns — dhadfieldmenell · 2026-08-10
- Rumor: Qwen-3.8 27b and Grok 4.6 Set for Release Next Week — kimmonismus · 2026-08-10
- Dev Test: DeepSeek Often Outperforms Sol — yacineMTB · 2026-08-10