Loop Transformer's Real Winner Is SRAM-Based Ultrafast Inference, Argues Bing Xu
bingxu_ · x · 2026-09-03
Responding to The Information's report that OpenAI's upcoming Astra model uses a technique making its reasoning harder for researchers to inspect — improving coding performance and cutting costs while complicating dangerous-behavior detection — Berkeley researcher Bing Xu argues the real story of Loop Transformer is SRAM-based inference à la Cerebras/Groq. His take: the only thing to worry about is affordability of ultrafast inference, or you'll run 10× slower.
More from Infra
- FastVideo-FastH3 appears on MLX listing ahead of actual model release — Structure-These · 2026-09-03
- Leaked-style chart claims to show OpenAI's pretraining compute capacity, past and future — midgaze · 2026-09-03
- South Korean exports jump 69% YoY in August on AI hardware demand — VraserX · 2026-09-03
- RTX 5090 writes nightly stock briefs with a numbers gate so the LLM can't invent figures — JakeChj · 2026-09-03
- CXMT reaches 10% global DRAM market share in Q2, Counterpoint Research says — zephyr_z9 · 2026-09-03
- Open-Source RL Framework Miles Debuts for Enterprise LLM and VLM Post-Training — AravSrinivas · 2026-09-03