Running 2.4T Params on Mac Ultra: 1-bit Quantization Real-world Test
Ok_Technology_5962 · reddit · 2026-08-14
A developer successfully ran a massive 2.4-trillion-parameter model on a 512GB Mac Ultra using Unsloth's 1-bit quantization (iQ1S), consuming around 508GB of memory.
In medium reasoning mode, the quantized model achieved 50 tokens/s generation speed, dropping to 5 t/s at 50k tokens. The author conducted several generation tests:
- Game Code: Successfully generated a flight simulator with gravity and physics, requiring up to 96k tokens to fully output the code.
- SVG Generation: Showed impressive results in creating complex SVGs (e.g., a panda picnic, capybara in an onsen). The model autonomously used Python for noise and texture generation, maintaining high structural coherence.
Related event: Developers Test Run 2.4T Parameter Qwen Model on Consumer Hardware(3 posts)→
More from Infra
- Explainer: How KV Cache Eliminates Redundant Attention Math for Fast LLM Inference — blaizedsouza · 2026-08-17
- Podcast: AU Govt's View on Compute and AI Economy Strategy — joecole · 2026-08-17
- Tsinghua Startup Boosts Domestic Chip Adaptation by 30x with Self-Evolving AI Infra — 量子位 · 2026-08-17
- Omarchy: DHH-endorsed, agent-first Linux operating system — vista8 · 2026-08-17
- Llama-3.0-Flash Leaked to Run End-to-End on a Single DGX Spark — Affectionate-File-26 · 2026-08-17
- AWS Trainium 4 Projected to Deploy 5M Units by 2H27, 12M by 2028 — zephyr_z9 · 2026-08-17