Redditor claims new LLM architecture improves loss, speed, and memory simultaneously
AlternativeSure2891 · reddit · 2026-09-06
A Redditor reports early results from an unnamed new LLM architecture showing lower validation loss alongside fewer parameters, fewer FLOPs, less memory, a smaller KV cache, and faster training and inference. Trained on 1B+ Python tokens; the author withholds architecture details, declines to call it a breakthrough, and says official benchmarks are being scheduled. Unverified claim.
More from Models
- Andrew Carr: now is the day for huge, well-documented proprietary datasets — andrew_n_carr · 2026-09-06
- GPT-6 Astra autonomously beats Portal, echoing OpenAI's 2016 game-solving goal — scaling01 · 2026-09-06
- Astra vs Fable 5.1 on real ML tasks: rigorous vs readable, with a mojibake twist — returnity · 2026-09-06
- Astra vs Fable 5.1 on real ML tasks: Astra grinds agentically, Fable writes better — returnity · 2026-09-06
- Insider: GPT-6 Astra wowed internals with speed and scale before release — kagigz · 2026-09-06
- GPT-6 Astra appears to leak neuralese internal thinking traces openly — ivan_bezdomny · 2026-09-06