Scale AI launches SWE-Bench Pro V2, a harder agentic coding benchmark
bigblueboo · x · 2026-09-23
Scale AI announced SWE-Bench Pro V2 is live, with a thread outlining what's new. SWE-Bench Pro is a harder expansion of the mainstream SWE-Bench, and V2 aims to better probe real agentic coding capability.
More from Research
- 6 serving-side techniques that make LLM inference faster - from prefix caching to PD disaggregation — techNmak · 2026-09-23
- ReFigBench paper: same model scores swing on identical tasks across Claude Code and Codex harnesses — omarsar0 · 2026-09-23
- NVIDIA's Skill2Env turns 3.4k Agent Skills into 8k RL environments, boosting Qwen-27B by 4.7 points — burny_tech · 2026-09-23
- Szegedy on OpenAI proof controversy: journals might become irrelevant — ChrSzegedy · 2026-09-23
- Judea Pearl, AI's causal reasoning pioneer, turns 90 as UCLA hosts symposium — yudapearl · 2026-09-23
- On OpenAI's proof prize, Szegedy argues journals may become irrelevant — JFPuget · 2026-09-23