Local 276B Multimodal Model Inference: 93GB VRAM at 32k Context

MaziyarPanahi · x · 2026-08-01

MaziyarPanahi shared real-world performance metrics for running the Inkling 276BA12B multimodal model locally (3-bit quantization).

At the default 1M context, the model requires 127GiB of resident memory; capping it at 32k reduces this to 93GiB. The weights take up 91GB, audio encoding takes 575ms, and generating an answer takes 103s. He recommends capping the context length for local runs.

Related event: Local Testing of Inkling 276B: Accurate Heart Failure Diagnosis(3 posts)→

Original post →

More from Infra

Infra channel →