Leaking deep residual vectors into early layers may fix state tracking in frozen LLMs, zero retraining
burny_tech · x · 2026-09-18
A research thread argues Transformers fail at state tracking because circuit depth is bounded. The proposed fix treats a frozen LLM as a dynamical system: at the next step, leak deep residual vectors back into early layers — with zero retraining. Full method and results are in the linked thread.
More from Models
- ChatGPT co-inventor launches Jev, claiming 200x faster, 400x cheaper frontier model — multiply_matrix · 2026-09-18
- Tencent's Hy4 Preview ranks #4 among open-weight models, cheapest in top ten — mariofilhoml · 2026-09-18
- Simple letter-counting test exposes huge gap: GPT-6-Astra hits 93%, Fable 5.1 flounders — scaling01 · 2026-09-18
- Jev, a 'System One' model by Typesafe, launches on OpenRouter with typed decisions instead of text — majidmanzarpour · 2026-09-18
- Zhipu claims AI autonomously discovered a WeWorm-exploitable vulnerability — teortaxesTex · 2026-09-18
- User comparison: Gemini nailed an insurance-law question that ChatGPT defended with circular reasoning — Hatrct · 2026-09-18