Local Inference of 91GB Audio Model: 127GB RAM Needed for 1M Context

andimarafioti · x · 2026-08-01

A developer shared real-world hardware metrics for running a massive audio model locally. The model uses 91GB of weights and requires 127GiB of resident memory at its default 1M context length. Capping the context at 32k reduces memory usage to 93GiB.

In terms of performance, audio encoding takes 575ms, while generating an answer takes 103 seconds. The test was conducted using llama.cpp (Metal backend) with Unsloth quants.

Related event: Local Inference Tested on Inkling 276B Multimodal Model(4 posts)→

Original post →

More from Infra

Infra channel →