Testing 'code verifies, model generates' on 4B models

retarded_770 · reddit · 2026-08-21

A month-long solo experiment tested if a verification harness could substitute for scale at 4B. The pattern splits requests: model extracts facts -> code verifies them (numbers verbatim, ≥60% word coverage) -> code does arithmetic -> next reasoning stage receives only verified facts.

Results:

Residual Issue: Cross-stage incoherence persists—the model names a constraint correctly in stage 2 but violates it in stage 4. Reasoning-trained weights only partially mitigate this.

Original post →

More from Research

Research channel →