Research: Inference-time Recurrence Cuts LLM Perplexity by 23% Without Weights Update
heghbalz · x · 2026-08-20
Introduces an inference-time recurrence mechanism allowing off-the-shelf, frozen LLMs to act as dynamical systems. By leaking deep representations back to shallow layers, it achieves significant performance gains without touching model weights: 23% reduction in perplexity and a 21% relative accuracy increase on GSM8k.
More from Research
- Mini Kimi-K3 Replicated Under $250 Beats GPT-2 Benchmark — OtherRaisin3426 · 2026-08-20
- llama.cpp PR Uses AVX2 to Speed Up Large Batch IQ Quantization — pmttyji · 2026-08-20
- Qwen 2.5 72B Aces ACT Exam with Perfect Reading Score — on_line187 · 2026-08-20
- GitHub repo curates 400+ free AI/ML books and resources in PDF — mdancho84 · 2026-08-20
- NVIDIA Integrates TriAttention: Trigonometric KV Compression for Long Context — 青稞AI · 2026-08-20
- NVIDIA Taiwan Hiring Research Interns in Embodied AI & 4D Vision — CMHungSteven · 2026-08-20