Jay Alammar at PyData: ten questions that explain benchmark score gaps

JayAlammar · x · 2026-09-15

Jay Alammar spoke at PyData Amsterdam alongside co-author Maarten Grootendorst (agents fundamentals vs. agent evaluation). His core message: technical users must read benchmark scores critically — ten questions often explain as much of the gap between two reported scores as the models themselves. Lessons drawn from building Cohere's North Mini Code and its code agents, where grading code exposed assumptions about harnesses, partial credit, and retries.

Original post →

More from coding & agent

coding & agent channel →