Real-SWE benchmarks AI models on private, real-world enterprise codebases
theanonymousone · hn · 2026-09-13
Real-SWE is a new benchmark from withspecific.com that evaluates AI models on private, real-world enterprise codebases, aiming to avoid the training-data contamination that plagues public-repo benchmarks and to better reflect how models perform on messy, real engineering work.
More from Research
- FlyGPT lands on Hugging Face: trainable LM wired as a fruit-fly connectome — QuixiAI · 2026-09-14
- FlyGPT trains a language model on a real fruit-fly brain connectome topology — QuixiAI · 2026-09-14
- 4 sink tokens + 64-token window matches distilled linear attention, no training needed — burny_tech · 2026-09-14
- Oak Ridge NL builds fully automated AI system that assembles molecules atom-by-atom for 25+ hours — jiqizhixin · 2026-09-14
- CMU launches Benchmark Radar, a living search engine for AI benchmarks — CarnegieMellonU · 2026-09-14
- Two weeks, two courses: writing a flow matching post by timing what to hide — ariG23498 · 2026-09-14