Four open models fix the same three real bugs, with a 27B model 14× faster

MaziyarPanahi · x · 2026-07-25

Four open models were tested on the same three real bugs, and all of them passed.

The author says Bonsai-27B, Gemma-4, and Inkling all fixed the bugs locally, while Kimi K3 looked the sharpest but could not be run until Monday.

A notable detail: the 27B model produced tokens 14× faster than the 950B model. The post asks which one people would trust in CI, making this a practical coding-workflow comparison rather than a pure benchmark post.

Related event: Open-Source Models Fix Real-World Bugs in ReAct Framework Test(2 posts)→

Original post →

More from coding & agent

coding & agent channel →