riderless: A Zero-Token Decision API on Gemma 4 Hits 93.1% Accuracy on One RTX 5090

Pale-Soil-2524 · reddit · 2026-09-22

The author open-sourced riderless (Apache-2.0), testing the hypothesis that a stock instruction-tuned model can serve as a decision engine if you skip generation and read label logits directly.

How it works

Benchmarks on one RTX 5090 (Unsloth UD-Q4KXL GGUF)

Comparisons: Kev-9B scored 0.898, Laya's typed-decisions checkpoint 0.618, and so1 landed at 0.921 (separate) / 0.901 (packed) on the same suites—prompt/readout differences were worth only a point or two. Caveats: one run each, synthetic suites, no confidence intervals.

Original post →

More from coding & agent

coding & agent channel →