SanSi: A Looped Typed Decision Model for System 1.5 Thinking
Shuyu Gan, Young-Jun Lee, Dongyeop Kang
cs.CL, cs.AI, cs.LG
2026-10-06
SanSi turns a looped 1.4B LM into a typed decision model, re-running its layers up to 8 times before one readout: 72.0% accuracy, 13.5 points over a same-shape single-pass model.
A growing share of LLM calls want a decision, not text: is this message spam, does this answer follow from the document. Typed decision models serve these calls directly. The caller declares the options and the model returns one probability per option in a single forward pass, with no generated text to parse. The weakness is that one pass spends a fixed amount of compute on every decision. The paper's example: on a chain of statements ("Omar is honest, Bert says Omar tells the truth, Ben says Bert lies"), the commercial Jev API scores 100% at chain length one, 62.5% at three, and chance from six on. Going deeper normally means a bigger model (more parameters and memory) or generated reasoning (more tokens and latency, and the typed contract is gone). SanSi takes a third route: run the same stack of layers several times, then read one answer. More compute, no new parameters, no tokens. The authors call this regime System 1.5.
The base is Ouro-1.4B, a language model pre-trained to loop: one stack of 24 transformer layers counts as one loop, and each loop's normalized output is the next loop's input. That pre-training is the entry ticket; the ablations show that bolting loops onto an ordinary model yields nothing.
Training has 12,800 decisions across five types (classification, multi-step reasoning, uncertain evidence, long documents, sentence pairs). The test set has 10,027 decisions from 59 sources, 30 of them never seen in training (FOLIO, BBH, MMLU and others). 382 unanswerable items, whose key evidence was deleted, target the uniform distribution: they probe whether the model knows it has no evidence.
| Model | Params | Loops | Relative cost | Accuracy |
| SmolLM2-1.7B (same-shape control) | 1.7B | 1 | 1.1x | 58.4% |
| Qwen3.5-2B | 1.9B | 1 | 1.2x | 66.7% |
| Qwen3.5-4B | 4.2B | 1 | 2.4x | 73.8% |
| SanSi | 1.4B | 8 | 7.7x | 72.0% |
| SanSi-2.6B | 2.7B | 8 | 14.8x | 75.8% |
The comparison that matters is rows one and four: same shape, data, recipe and seeds, differing only in loops. All 13.5 points come from looping. Trained and run with a single loop, Ouro-1.4B reaches 58.6%, indistinguishable from SmolLM2, which rules out "the backbone is just better". The gain is uneven: classification adds only 3.1 points, multi-step reasoning +15.1, long documents +16.7, knowledge questions +17.6. The bill is explicit: eight loops cost 7.7x a single pass at inference and 9.5x at training, and reading at loop 3 already banks 88% of the gain.
The depth-controlled experiments cut sharpest. Liar chains and object swaps were trained only to depth 8 and tested to 16. On liar-chain depths 9-16, SanSi scores 72.6% while the 3x-larger Qwen3.5-4B scores exactly 50.0%, chance; on object swaps the gap on unseen depths is 29.5 points. Parameters do not buy this kind of length extrapolation.
As an RL judge: GRPO trains a SmolLM2 generator on 2WikiMultiHopQA with SanSi's option probabilities as the only reward, gold answers unused. Read at loop 8, the generator's F1 rises from 39.5 to 47.3. Read at loop 1 it damages the generator (29.1): the single-pass reward separates good from bad answers at AUROC 0.78, versus 0.91 and 0.90 at loops 4 and 8.
The probabilities are a mixed story. Detecting missing evidence improves a lot (evidence AUROC 0.835 to 0.935, close to the 3x model), but confidence rises on wrong answers as well as right ones, self-knowledge barely moves (right/wrong AUROC 0.760 to 0.795), and ECE bottoms at loop 3 then climbs back. Averaging the eight loops' probabilities, with no retraining, halves ECE from 0.093 to 0.044.
Trading test-time compute for parameters is proven for chat models with long chains of thought. SanSi moves that trade into typed decisions and shows the depth-reasoning gains arrive without emitting a single token. For moderation, routing and confidence-tiered pipelines, one 1.4B model that scales from 1 to 8 loops on demand, with nothing to parse, is an attractive deployment shape. With restraint: at equal compute (three loops) the 4B single-pass model still leads by 3.4 points.
The authors' own list: everything rests on the Ouro family of looped pre-trained backbones, with most analysis on the 1.4B model; a fixed loop count is read, with no per-item early exit; the depth claims rest on two program-generated tasks; the verifier study covers one dataset and one generator; the suite is English-only, and pre-training contamination of public datasets cannot be excluded.
Reading closely adds more. Most of the gain lands by loop 3; loops 5-8 together add 0.4 points, and a four-loop model lands only 0.8 points lower at loop 4, so training eight loops is marginal. Running past the trained count hurts (70.1% at loop 16), so the compute knob has a hard cap. The commercial Jev API reaches 78.9% overall and 83.5% on sources none of the models trained on, well ahead of SanSi-2.6B's 71.9%, though its size and data are private. The sharpest constraint: retrofitting loops onto SmolLM2, in either of two ways, yields nothing (33.5% and 58.3%). This line of work needs looped pre-trained backbones, and that shelf is nearly empty today.