Mach-1 Additive: 35B Model Runs at 120 t/s on Laptops Using 1.7-bit Weights
pbaylies · x · 2026-08-04
Mach-1 Additive is a 35 billion parameter model that infers without traditional matrix multiplication, using an extreme 1.7 bits per weight approach.
- Performance & Size: Recovers 95% of the original full-precision model (Qwen 35b) performance across 12 agentic and reasoning benchmarks while being 10x smaller.
- Local Deployment: At 7GB, it comfortably fits on consumer laptops, achieving speeds up to 120 tokens per second.
- Training Cost: Unlike algorithms like BitNet, this approach requires minimal retraining (under 15 GPU hours), making it highly scalable to massive LLMs (up to 3 trillion parameters).
Users can currently try it via web browser or download the desktop app.
More from Infra
- Tested: MiniMax H3 Runs Locally on 64GB MacBook Pro via Phosphene — cocktailpeanut · 2026-08-05
- NVIDIA Joins NSF Regional AI Hubs to Expand Computing Access Nationwide — nordicinst · 2026-08-05
- Agentic RL Bottlenecked by Inference: SkyPilot Halves Training Time — skypilot_org · 2026-08-05
- Hardware Architecture Debate: Why Vertical Power Delivery Over Vertical Optical IO? — jwt0625 · 2026-08-04
- CoreWeave Announces Fully Connected 2026: Fei-Fei Li & NVIDIA to Keynote — wandb · 2026-08-04
- Agentic AI Triggers a Storage Shock: Enterprise Data Becomes the New Bottleneck — BenBajarin · 2026-08-04