Inkling-276B Multimodal Model Local Test: 3-bit Quantization Hits 37 t/s Decode

MaziyarPanahi · x · 2026-08-02

Developer Maziyar Panahi shared local benchmark results for the multimodal model Inkling-Small (276B A12B architecture) running on Metal hardware.

Using a 3-bit quantization, the model occupies 91.19 GiB of memory. It achieved a prefill speed (pp512) of roughly 510 tokens/s and a decode speed (tg64) of 37 tokens/s. In a practical test, the model successfully processed simultaneous image and audio inputs, accurately analyzing a 2-minute medical audio to deduce that a patient's heart failure worsened due to stopping their diuretic.

Related event: Inkling 276B Multimodal Model Local Inference: 3-bit Quantization Runs Medical Diagnosis(5 posts)→

Original post →

More from Models

Models channel →