Startup calls for continuous benchmarks to quantify model nerfing after release
demeyer1 · reddit · 2026-09-22
A Reddit user from a small startup calls for sites to continuously benchmark OpenAI models on a fixed suite, to quantify how fast models get nerfed post-release. They cite Sol Ultra going from superhuman in week one to far lower utility per dollar, note their own benchmark runs get expensive quickly, and want to know if nerfing is dynamic (compute-driven) to plan usage around peak-quality windows.
More from Infra
- Hands-on LLM inference: boosting tokens-per-second with a 31B Gemma model — abhijithneil · 2026-09-22
- rakyll: fully managed platforms fail on composability — bespoke and open source stacks win — rakyll · 2026-09-22
- Local Qwen 27B agent logs into Amazon and buys paper autonomously in one run — fuzhongkai · 2026-09-22
- 60 Minutes: US golf courses use more than twice as much water as data centers — SumitGup · 2026-09-22
- ROCm vs Vulkan on R9700 + Strix Halo: ROCm still wins for DeepSeek, Vulkan closes in — Hrethric · 2026-09-22
- DeltaTensors Stores Fine-Tunes as Weight Deltas: 953MB → 294MB With Minimal Quality Loss — cupheadgamer · 2026-09-22