FlashLoop exploits cross-loop redundancy to speed up Looped Transformers by 1.65x
KyeGomezB · x · 2026-09-26
- Looped Transformers enable deeper computation without more parameters, but every extra loop adds compute and KV cache cost.
- The authors find only a small fraction of features actually change across loops, identifying three types of computational redundancy that can be exploited for faster inference.
- Results: 1.65x lower latency and 6x less memory.
- Open sourced, installable via pip install flashloop.
More from Infra
- SpaceX's Memphis supercomputer: millions of GPUs and over two gigawatts of compute — CurieuxExplorer · 2026-09-26
- The handiest GPU this dev ever bought is a ~$300 Intel Arc A310, not NVIDIA or AMD — TheZachMueller · 2026-09-26
- Running image generation in the browser on local hardware: 10s pixel art on an RTX 3060 — Bartholomheow · 2026-09-26
- Alibaba's T-Head unveils Zhenwu V900 chip: 216GB per card, sales in Q1 2027 — shashib · 2026-09-26
- llama.cpp fork dedups repeated prompts losslessly, cutting 108k to 71k tokens in agent loops — Odd_Cauliflower_8004 · 2026-09-26
- Why rent servers when agents can run your terminal? — StewartalsopIII · 2026-09-26