Running 276B Inkling Small on a Single Mac Studio via nvfp4 Quantization
pcuenq · x · 2026-07-31
A developer demonstrated how to efficiently deploy the newly released Inkling Small (276B parameters) within the Apple ecosystem.
- Massively Reduced Requirements: Using the official nvfp4 checkpoint, the model requires only 150 GB of memory to run on a single Mac Studio, or it can be split across two 128GB laptops.
- More Aggressive Quantization: The author noted that with more aggressive quantization, the model should fit entirely on a single laptop soon.
- Tooling: A PR has been submitted on GitHub to natively support nvfp4 and mxfp4 models, enabling direct inference via mlx-vlm.
Related event: Thinking Machines Releases Inkling-Small Open-Source Model(23 posts)→
More from Infra
- AWS Operating Income Margin Jumps to 39%, Killing AI ROI Myth — RihardJarc · 2026-07-31
- Martin Shkreli on AI Infra Trade Unwinding: 4x Leverage and Weak Hands Panic — ivan_bezdomny · 2026-07-31
- AWS Revenue Surges 37% YoY, Crushing Market Estimates — RihardJarc · 2026-07-31
- Race for Space Datacenters Rockets Forward, Faces Laws of Physics — Grady_Booch · 2026-07-31
- Google Reportedly Plans to Backstop and Supply Chips to Anthropic — Wonderful_Buffalo_32 · 2026-07-31
- AI Lab Economics: Frontier Labs Pursue Vertical Integration, Open Labs Leverage Interoperability — kevinsxu · 2026-07-31