Trillion-parameter model trained on own physics lab data hits memory-efficiency SOTA
vwxyzjn · x · 2026-09-16
zijiey shared that his team trained a trillion-parameter model on data from their own physics labs, noting "there's a lot of training infra hidden in that sentence." Long scientific traces stress memory and parallelism; the team hit SOTA in memory efficiency and long-context training performance, letting them try more ideas faster — some of which went into Neon. Bahdanau called him "the GOAT of frontier-level training" and said the team is hiring a new colleague, with a job posting coming soon.
More from Infra
- Running Qwen 27B and DeepSeek v4 Flash together on one heterogeneous machine — samsja19 · 2026-09-16
- JPMorgan sees 25M+ GPU/ASIC shipments by 2028, ASICs dominate — a 'narrative violation' — bookwormengr · 2026-09-16
- iamtrask: The Endgame Is a Trust Web of Personal LLM Servers, Not One AGI — iamtrask · 2026-09-16
- Program-as-Weights: 0.6B interpreter matches Qwen3-32B prompting with 1/50 the memory — yuntiandeng · 2026-09-16
- Intel bets on CPU-side KV Cache offloading and QAT hardware compression for agent inference — 量子位 · 2026-09-16
- Six efficiency breakthroughs labs didn't see coming upend semiconductor demand assumptions — bookwormengr · 2026-09-16