Dev: if your eval costs $20 for 100 runs, the problem set isn't hard enough
pvncher · x · 2026-09-23
Developer pvncher argues that if an eval suite only costs $20 for 100 runs, it's probably not a hard enough problem set — genuinely difficult tasks should cost far more to evaluate. The remark came amid discussion of new pricing doing away with Ultra/Xhigh/High reasoning tiers.
More from coding & agent
- Can Jev Write Graphical Models? LLM Meets Factor Graphs and Bayesian Nets — blaizedsouza · 2026-09-23
- Jev explained: the semantic decision layer between rules and LLMs — blaizedsouza · 2026-09-23
- A complete AI workflow for turning ideas into product launch videos, prompts included — Aiden_Tech_Ai · 2026-09-23
- A 19-year-old built a Claude Code trading bot in 2 days, claims $750K profit — Aiden_Tech_Ai · 2026-09-23
- Dev sends coding tasks from 35,000 feet with Codex Remote: "this will never cease to amaze me" — Dimillian · 2026-09-23
- PM vibes coding at a big tech firm: 1-2 months of work now done in a week — vista8 · 2026-09-23