Startup calls for continuous benchmarks to quantify model nerfing after release

demeyer1 · reddit · 2026-09-22

A Reddit user from a small startup calls for sites to continuously benchmark OpenAI models on a fixed suite, to quantify how fast models get nerfed post-release. They cite Sol Ultra going from superhuman in week one to far lower utility per dollar, note their own benchmark runs get expensive quickly, and want to know if nerfing is dynamic (compute-driven) to plan usage around peak-quality windows.

Original post →

More from Infra

Infra channel →