SWE-Rebench: Benchmarking 13 LLMs and 4 Agents on SWE Tasks

ibragim_bad · hn · 2026-07-31

This project benchmarks 13 large language models and 4 agents on Software Engineering (SWE) tasks across multiple programming languages including Go, Java, Python, Rust, and TypeScript, aiming to evaluate their performance in real-world coding and bug-fixing scenarios.

Original post →

More from coding & agent

coding & agent channel →