Local 2-bit model as coding-agent judge: 66% agreement, double the built-in judges

Cultural_Self8980 · reddit · 2026-10-06

A developer built a local "System One" judge on Ternary-Bonsai-4B (1.1GB packed 2-bit weights on Apple Silicon via MLX): one forward pass answers multiple-choice, rating, or yes/no questions about a record. A tree attention mask shares the record prefix across questions while keeping question branches separate; weights are frozen, no fine-tuning.

One application is a coding-agent judge: an omp plugin uses the local model for auto thinking-effort selection, with an optional model router. On 80 real coding requests, Bonsai matched Claude Opus reference labels 66% of the time, versus 29% for omp's built-in LFM2-1.2B judge and 25% for LFM2.5-230M; median latency on M2 was 0.8s / 3.0s / 0.3s respectively.

Caveats: reference labels came from a single Opus model, only 80 examples, differing label spaces (four effort levels for Bonsai vs. three for built-ins), Bonsai tends to rate one level low, and the router's accuracy is untested. Code and benchmarks are open-sourced on GitHub.

Original post →

More from coding & agent

coding & agent channel →