Speculation: 2T Training Run Points to V4.1 Pro, as Critics Lament Scaling Is Still Alchemy
teortaxesTex · x · 2026-09-21
teortaxesTex speculates (unconfirmed) that a reported 2T-parameter training run corresponds to V4.1 Pro, with roughly 540GB of engram memory if it follows V4.1's design — questioning whether engram memory saturates or turns fragile at scale. He expects roughly Fable 5-level performance and laments that scaling remains alchemical: Meta ran 405B active params on 16K H100s two years ago, yet rushing this scale today can still waste months of compute.
More from Infra
- AI intern cuts job boot time ~60%, saving ~5 CPU-days per day — DanielLockyer · 2026-09-21
- Tim Dettmers' lab announces open-source week: 2 frameworks, 4 papers for frontier AI on local hardware — Tim_Dettmers · 2026-09-21
- I was wrong: you can run LLMs on this device and even do development — gnukeith · 2026-09-21
- Backblaze CEO on CoreWeave deal: 'a land grab in the AI ecosystem' — TiernanRayTech · 2026-09-21
- Training engineer tip: save checkpoints to K8s PVC — and don't alert on that PVC — cto_junior · 2026-09-21
- Baseten CEO: agents burning tokens will make inference 'core IP' for every company — baseten · 2026-09-21