Local 2-bit model as coding-agent judge: 66% agreement, double the built-in judges
Cultural_Self8980 · reddit · 2026-10-06
A developer built a local "System One" judge on Ternary-Bonsai-4B (1.1GB packed 2-bit weights on Apple Silicon via MLX): one forward pass answers multiple-choice, rating, or yes/no questions about a record. A tree attention mask shares the record prefix across questions while keeping question branches separate; weights are frozen, no fine-tuning.
One application is a coding-agent judge: an omp plugin uses the local model for auto thinking-effort selection, with an optional model router. On 80 real coding requests, Bonsai matched Claude Opus reference labels 66% of the time, versus 29% for omp's built-in LFM2-1.2B judge and 25% for LFM2.5-230M; median latency on M2 was 0.8s / 3.0s / 0.3s respectively.
Caveats: reference labels came from a single Opus model, only 80 examples, differing label spaces (four effort levels for Bonsai vs. three for built-ins), Bonsai tends to rate one level low, and the router's accuracy is untested. Code and benchmarks are open-sourced on GitHub.
More from coding & agent
- Claude Opus 5.5 took 90,000 screenshots of its own game over two days to polish the visuals — prasenx · 2026-10-06
- classif: shell scripts branch on meaning via one token's logprobs on a local 12B — piotr1215 · 2026-10-06
- Agent 3D-renders its own game assets to build Star Control Hyper Melee unsupervised — draginol · 2026-10-06
- Markov: a pi-style, transparent LLM harness in a single Bash script with sub-agents — biller23 · 2026-10-06
- Meta paper: dual coding agents cross-review lift correct patches from 45.8% to 62.5% — rohanpaul_ai · 2026-10-06
- Rewriting from scratch beats depending on a single maintainer with feelings — l4rz · 2026-10-06