Swarm added just +0.05: splitting context into blocks did the real work for tiny local models
deadatreides1 · reddit · 2026-10-05
A pre-registered experiment on a GTX 1660 SUPER 6GB ran six 1-2B local models as a 'Soviet design bureau' swarm on a long-context extraction task (18 records, pick matching ones), with inputs split into 3 blocks by script, per-block voting, and code-based id checks to kill hallucinations.
Key findings:
- Reading full context, recall falls monotonically from 0.609 at position 0 to 0.214 at position 17 — models increasingly answer 'no' the deeper they read.
- Splitting into blocks shrank the head-tail gap 10x, from 0.204 to 0.021; position stopped mattering.
- Splitting beats swarming: the six-model swarm only beats the best single model through the same pipeline by +0.05 (0.900 vs 0.850); strict emergence was zero.
- Weaker models gain more: qwen2.5-coder-1.5b jumped 0.175 → 0.700; gemma-3-1b from 0 to 0.050; smollm2-1.7b stayed at 0 — org charts can't fix missing capability.
Advice: don't dump document piles into small models; cut, run pieces, and verify with code, not a judge model (author says model-as-judge is worse than you think). Caveats: only 1-2B models, one task family, 40 tasks.
More from coding & agent
- Same prompt showdown: Claude vs ChatGPT vs Figma Make — Tegadesigns · 2026-10-05
- Dev ports Muse voice assistant to a 15-year-old PSP using a coding agent — alexandr_wang · 2026-10-05
- Coding agents burn more tokens finding code than changing it — the interface is to blame — Wise_Reflection_8340 · 2026-10-05
- MicroFactory ships an AI robotic integrator that designs tooling and writes vision code — ihorbeaver · 2026-10-05
- Linear CEO shows AI agent triaging his bug and opening a fix PR autonomously — jeff_weinstein · 2026-10-05
- TunnelGPT connects a local project folder to ChatGPT over MCP — carlosrodera · 2026-10-05