Trillion-Parameter Models Abandon Transformers as Architectures Diverge
imjustnewatai · x · 2026-08-26
Recent breakthroughs highlight a shift away from traditional Transformer architectures:
- Trillion-Parameter Neural Operator Model: Completely removes the Transformer architecture. It uses neural operators to learn continuous physical fields across space and time, capable of ingesting over 5 trillion input elements and generating an entire 4D trajectory in one shot.
- Pathway's Dragon Hatchling: A tiny 1.5M parameter model using recurrent latent reasoning and Hebbian memory. It updates internal state from demonstrations and scored 29.5% pass@2 on public ARC-AGI-1 at a cost of $0.00070 per task.
- Core Automation: Led by the teams behind o1/o3 and Gemini pretraining, this initiative aims to meta-learn long-horizon learning directly into the architecture, allowing models to keep adapting after deployment.
Related event: Ex-NVIDIA Scientists' Startup Launches Neural Operator Physical AI Model(14 posts)→
More from Models
- Critique of Claude style: Verbose but information-dense — TheZvi · 2026-08-27
- GLM 5.3 Flash benchmarks hit 881 tok/s untuned — HankYeomans · 2026-08-27
- GLM-5.3 Flash review: Fixes 13 bugs with high cost-performance — PawelHuryn · 2026-08-27
- Opus 5 writing regression confirmed by benchmarks — gleech · 2026-08-27
- Gemini 3.7 Flash beats benchmarks at 96% lower cost — DynamicWebPaige · 2026-08-27
- User Touts DeepSeek Reasoning Model Performance on Complex Logic Test — vista8 · 2026-08-27