Pi and mini-swe-agent Passed 9/9 Checks Each — a Second Review Still Found Bugs

Mysterious-Desk-3492 · reddit · 2026-10-07

The author, as part of an AI Studio project, tested 10 coding harnesses — Pi, mini-swe-agent, Crush, OpenCode, Goose, Prime Agent, Oh My Pi, Qwen Code, Octomind and Aider — using DeepSeek V4.1 Flash, Qwen3.8-27B and Laguna S 2.1 via OpenRouter.

After screening, Pi, mini-swe-agent, Crush and OpenCode went through detailed code review across 36 model×task combinations; two attempts produced no patch. The three Golang tasks targeted: strict HTTP query-param validation, migrating 200 logging calls while preserving behavior, and adding bookmark tags across API, storage migration and HTML rendering.

Pi and mini-swe-agent passed original acceptance checks on all nine combinations, but a second agent review plus isolated reproduction probes exposed three gaps: silently dropped malformed query fields, a mutable tag slice leaked from the store, and a migration rejecting a valid older store. On the positive side, all logging migrations preserved behavior under differential probes covering 100 functions and nine integer inputs including min/max.

Takeaway: the evaluator and the reviewer both need testing — a green result is only evidence about the checks you ran; broader correctness needs further evidence. The experiment found no decisive winner between Pi and mini-swe-agent, and human correction time remains unmeasured.

Original post →

More from coding & agent

coding & agent channel →