NVIDIA Engineer Talks Benchmarks and Coding Agents
ziv_ravid · x · 2026-07-14
A repost of an episode of The Information Bottleneck podcast, featuring Jean-Francois Puget from NVIDIA—head of Kaggle Grandmasters team, #3 in Kaggle rankings.
Topics discussed:
- Why many LLM benchmarks reward overfitting
- How the team discovered that O3 might not actually read code when solving SWE-bench
- How they won ARC-AGI with a 4B small model at $0.20 per problem
- Agent skills and how coding agents are 'killing' AutoML
- Direct opinions on frontier labs' marketing narratives
The full episode is available on the website, YouTube, and podcast platforms.
Related event: NVIDIA Distinguished Engineer Discusses Coding Agents and Benchmarks(3 posts)→
More from coding & agent
- Grok 4.5 is now free inside Cursor, the popular AI coding IDE — mark_k · 2026-07-21
- Travel MCP Server adds flight, hotel, weather and budget tools for agents — modelcontextprotocol · 2026-07-21
- OpenAI hackathon project stalls as Codex struggles on voice, while Claude spots the issue — ColleenMBrady · 2026-07-21
- Gemini CLI adds skill_name telemetry to tool-call metrics — dimpavloff · 2026-07-21
- Douyin Video Analysis MCP turns share links into structured video summaries — modelcontextprotocol · 2026-07-21
- An indie builder open-sourced 50+ AI apps and says users should only pay for tokens — matchaman11 · 2026-07-21