Scale AI NYC Meetup: Advancements in Benchmarking Coding Agents
jyangballin · x · 2026-07-30
Scale AI announced its third Research Meetup in NYC, diving deep into benchmarking for coding agents. The event features talks from researchers at Meta, Columbia University, and Scale AI, presenting updates on various benchmarks including ProgramBench, BranchBench, Terminal Bench, and SWE Atlas.
More from coding & agent
- Developer Codes with AI for Two Days Using an Xbox Controller — noahsolomon · 2026-07-30
- PaperPush: Open-Source Tool Automates Academic Paper Submission with LLMs — lpachter · 2026-07-30
- Agents still struggle with mathematical work: Codex spirals into 'proof certificates' and inventories — doodlestein · 2026-07-30
- Contour: Open-Source Tool Turns 2D Maps into 3D Terrain with Gemini Voice Guide — tom_doerr · 2026-07-30
- The AI Coding Dictionary: Clarifying Tokens, Inference, and Agent Boundaries — luisdans · 2026-07-30
- Keep Coding with Kimi K3 in Claude Code During Anthropic Outages — HowDevelop · 2026-07-30