1360 runs: Edinburgh researcher benchmarks local LLMs against Aider, Claude Code, OpenCode and more

PMinervini · x · 2026-09-04

Dr. Pasquale Minervini (University of Edinburgh, CTO of Miniml.AI) published a WIP benchmark, harness-bench, pairing local LLMs (served via llama.cpp's llama-server) with five agent harnesses — Aider, Claude Code, OpenCode, Pi, Qwen CLI — on 16 software-engineering tasks across Python, PyTorch, JAX, C, C++, Rust, and SQL. Current sweep: 17 model-quants × 5 harnesses × 16 tasks = 1360 runs on a single M3 Max / 128 GB laptop.

Key points:

A rare systematic reference for developers choosing local-model + agent-harness combos on a budget.

Original post →

More from coding & agent

coding & agent channel →