NVIDIA Engineer on Coding Agents and AI Evaluation
ziv_ravid · x · 2026-07-14
这期《The Information Bottleneck》采访了 NVIDIA 的 Jean-Francois Puget——他是 Distinguished Engineer,也带领 Kaggle Grandmasters 团队。
访谈重点包括:
- 为什么很多 LLM benchmark 会奖励过拟合
- 他们如何发现 O3 在没读代码的情况下解 SWE-bench
- 用一个 4B 模型、每题约 20 美分拿下 ARC-AGI 的过程
- agent skills 的作用,以及为什么 coding agents 让 AutoML 失去意义
- 他对 frontier lab 公关叙事的直率看法
这期内容同时覆盖了评测、编码智能体和行业方法论,适合关注 agent 与 benchmark 的人。
Related event: NVIDIA Distinguished Engineer Discusses Coding Agents and Benchmarks(3 posts)→
More from coding & agent
- Building a Secure AI Agent Gateway: Self-Hosting OAuth for Multiple SaaS Apps — Defiant_Cod_2654 · 2026-07-22
- Rowboat launches as an open-source, local-first AI coworker with memory — ycombinator · 2026-07-22
- Reddit user chains Ideogram 4 and Krea2 to mimic bbox-based image positioning — v3lh0t05c0 · 2026-07-22
- Apollo Cuts AI Assistant Skill Dev Time by 85% with Deep Agents — LangChain · 2026-07-22
- Scoble says AI “loops” really means long-running multi-agent workspaces — Scobleizer · 2026-07-22
- Kimi Code opens a waitlist as Moonshot rolls out its coding product — Fabulous_Bonus_8981 · 2026-07-22