Swarm added just +0.05: splitting context into blocks did the real work for tiny local models

deadatreides1 · reddit · 2026-10-05

A pre-registered experiment on a GTX 1660 SUPER 6GB ran six 1-2B local models as a 'Soviet design bureau' swarm on a long-context extraction task (18 records, pick matching ones), with inputs split into 3 blocks by script, per-block voting, and code-based id checks to kill hallucinations.

Key findings:

Advice: don't dump document piles into small models; cut, run pieces, and verify with code, not a judge model (author says model-as-judge is worse than you think). Caveats: only 1-2B models, one task family, 40 tasks.

Original post →

More from coding & agent

coding & agent channel →