KAUST's training-free Periscope lets a 27B model read 4.5M-token contexts on one 80GB GPU
KAUST · hf · 2026-10-06
KAUST introduces Periscope, a training-free inference method that factorizes long-text reading for frozen LLMs:
- Method: arranges N chunks on a K×K grid, probes a frozen model with the same question over K local and K strided spans, reading answer log-odds at one token; best local/strided scores yield a free evidence map whose peak localizes the supporting chunk
- Complexity: a window of W tokens reaches W²/c tokens at s^1.5 cost; only one probe cached per call
- Results: on LongBench v2, reading just the top-ranked 9k tokens matches the same model's best window read across 32k–1M windows; leads by 5 points on InfiniteBench (150k median contexts); best NDCG@10 of six methods on BRIGHT
- Practical: a 27B model reads 4.5M-token contexts on one 80GB GPU, where a single pass would need 296GB of cache
Takeaway: long reads need a GPU that holds the model, not one that holds the text.
Related event: Periscope Extends Frozen LMs to Million-Token Contexts Without Training(2 posts)→
More from Research
- Decision grader replaces LLM judge: 32x cheaper, 8x faster, 94% agreement on evals — rhythmrg · 2026-10-06
- Raghunathan lab to present pretraining safety and adaptation papers at COLM 2026 — AdtRaghunathan · 2026-10-06
- Retinal Imaging AI Predicts Preeclampsia Before Onset in New Nature Biotechnology Study — EricTopol · 2026-10-06
- New COLM paper: faithful LLMs decide and report with the same layers — dhadfieldmenell · 2026-10-06
- Uni-LaDiR unifies image, text and 3D reasoning via latent diffusion thoughts — Lianhuiq · 2026-10-06
- Critic Calls Out 'Research Taste Doubles Every 3 Months' Eval as Merely Metric Optimization — dhadfieldmenell · 2026-10-06