Model evals need much better tooling defaults, says Aiden Bai: costly rollouts, rampant reward hacking
aidenybai · x · 2026-10-10
Aiden Bai argues that model evals need much better tooling defaults to become practical at scale:
- Rollouts are extremely expensive
- False positives/negatives and reward hacking are very common and hard to catch
- Bad instructions and overly strict tests abound
- Inspecting data manually is labor-intensive
His conclusion: for every company doing posttraining, good evals must come first. He notes this is an industry-wide problem, not specific to any single framework like Harbor.
Related event: Aiden Bai: AI eval tooling falls short industry-wide(2 posts)→
More from coding & agent
- Vibe-coded slop vs human slop: quality comes from taste, intention and time — cameronstow · 2026-10-10
- Engineer Argues Abstractions Hurt AI Agents: 95% of Cases Better Raw — zack_overflow · 2026-10-10
- Testing proactive AI agents on my email and calendar: dot caught a date mixup, Muse is noisy — hazelcough · 2026-10-10
- Personal AI agents are stickier than you think: one now learns work habits and builds its own CRM — thisiskp_ · 2026-10-10
- Creator finds hand-tweaking generative models faster than prompts, sees room beyond text UIs — keenanisalive · 2026-10-10
- "Do better!" prompting stalls fast; even top VLMs understand images unevenly — keenanisalive · 2026-10-10