HarnessTester finds 100+ real bugs in LLM agent harnesses like OpenClaw
LingmingZhang · x · 2026-10-07
New work from Yiling Lou's team introduces HarnessTester, a harness-oriented test generator for LLM agents. It builds contract-faithful tests to exercise LLM-dependent harness code, targeting bugs at the interaction boundary between the LLM and the harness.
- It has already surfaced over one hundred real-world harness bugs across widely used agent systems (e.g., OpenClaw).
- The framing: are we testing the agent harness itself enough, rather than just the model?
- Paper link available.
More from coding & agent
- One prompt, one 3D scene: Codex + three.js + fal workflow shared — OdinLovis · 2026-10-08
- A fine-tuned 9B beats a 31B model: 600 labels, $0.12, 91% accuracy — julsimon · 2026-10-08
- Matt Pocock: Agents Are Good at Strategic Programming, They Just Aren't Trained to Care — mattpocockuk · 2026-10-08
- LlamaIndex Launches OpenDocRouter: Swap Between 10 Document Parsing Models in One Line — llama_index · 2026-10-08
- Microsoft open-sources Agent Lightning v1.0: 3,500-line RL framework boosting SWE-bench +14.6 pts — Microsoft Research · 2026-10-08
- OverclaimBench: Opus 5 skipped assigned files in 61% of runs, then claimed full review — hugo_larochelle · 2026-10-07