Local 276B Multimodal Model Inference: 93GB VRAM at 32k Context
MaziyarPanahi · x · 2026-08-01
MaziyarPanahi shared real-world performance metrics for running the Inkling 276BA12B multimodal model locally (3-bit quantization).
At the default 1M context, the model requires 127GiB of resident memory; capping it at 32k reduces this to 93GiB. The weights take up 91GB, audio encoding takes 575ms, and generating an answer takes 103s. He recommends capping the context length for local runs.
Related event: Local Testing of Inkling 276B: Accurate Heart Failure Diagnosis(3 posts)→
More from Infra
- antirez looks into deploying LLMs on DGX Spark — antirez · 2026-08-01
- Modal Releases Comprehensive GPU Glossary Covering Hardware to Software Stack — charles_irl · 2026-08-01
- ARM Introduces FEAT_CSSC: Native Popcount for General-Purpose Registers — lemire · 2026-08-01
- Train Your Own Model When Inference Exceeds $750/Day: Pallet's Playbook — marcbhargava · 2026-08-01
- Local Inference of 91GB Audio Model: 127GB RAM Needed for 1M Context — andimarafioti · 2026-08-01
- Llama 3.1 405B Hits 5.6k t/s on Cerebras for Select Customers — kimmonismus · 2026-08-01