NVIDIA Team on Benchmarks and ARC-AGI
ziv_ravid · x · 2026-07-14
An episode of The Information Bottleneck featuring Jean-Francois Puget from NVIDIA: head of the Kaggle Grandmasters team and ranked 3rd historically on Kaggle.
Key discussion points:
- Why many LLM benchmarks reward overfitting
- How they discovered O3 was "solving problems without reading the code" on SWE-bench tasks
- How the team cracked ARC-AGI using a 4B model at roughly 20 cents per problem
- Extended topics on agent skills
The value here is that it moves beyond mere model gossip, using concrete cases to re-examine classic questions like "are evaluations trustworthy?" and "do models actually understand the task?".
Related event: NVIDIA Distinguished Engineer Discusses Coding Agents and Benchmarks(3 posts)→
More from AGI Musings
- The Evolution of LLM Business Models: Selling Outcomes Over Tokens — yacineMTB · 2026-07-22
- Bindu Reddy says GPT-6 is coming soon, with Alibaba, DeepSeek and Kimi close behind — bindureddy · 2026-07-22
- Bindu Reddy says the industry still lacks a way to train 20T models and scale post-training RL — bindureddy · 2026-07-22
- Advanced AI Models Are Becoming Impossible to Plug and Play — emollick · 2026-07-22
- AI suggested a better composition, and that made one user uneasy — Sydde · 2026-07-22
- The Thimble and the Waterfall: AI's Data Bottleneck and Feedback Loops — dyamins · 2026-07-22