A research talk compares SWE-bench, CodeClash, and ProgramBench for agentic coding
OfirPress · x · 2026-07-29
Krzysztof? not named in the post, but the talk focuses on agentic coding benchmarks.
It starts with SWE-bench and then covers CodeClash and ProgramBench, both of which were done with collaborators. The post points to a full research spotlight talk hosted by AGI House.
More from coding & agent
- Verdent says Kimi K3 is tuned for real-world agentic coding workflows — socialwithaayan · 2026-07-29
- Hermes agent now connects to Buzz for workflow automations — intellectronica · 2026-07-29
- How to wire custom MCP servers into ChatGPT and Claude chat interfaces — jeremyjordan · 2026-07-29
- HANDBOOK.md benchmarks whether agents obey long SOPs over extended tool use — dair_ai · 2026-07-29
- AI coding agents often solve problems by adding code instead of rewriting it — bendee983 · 2026-07-29
- A Blender plugin built with ChatGPT automates LEGO-style assembly animations — azed_ai · 2026-07-29