Inkling-276B Multimodal Model Local Test: 3-bit Quantization Hits 37 t/s Decode
MaziyarPanahi · x · 2026-08-02
Developer Maziyar Panahi shared local benchmark results for the multimodal model Inkling-Small (276B A12B architecture) running on Metal hardware.
Using a 3-bit quantization, the model occupies 91.19 GiB of memory. It achieved a prefill speed (pp512) of roughly 510 tokens/s and a decode speed (tg64) of 37 tokens/s. In a practical test, the model successfully processed simultaneous image and audio inputs, accurately analyzing a 2-minute medical audio to deduce that a patient's heart failure worsened due to stopping their diuretic.
More from Models
- Sakana AI Launches Namazu API: A Japanese-Specialized LLM Built on Kimi K2.6 — SakanaAILabs · 2026-08-03
- Opinion: OpenAI Leads the Race as DeepMind Struggles with Model Consistency — haider1 · 2026-08-03
- Nemotron 3 Nano Omni Hits 264 tok/s Native on DGX Spark — ivan_bezdomny · 2026-08-03
- Sol 5.6 Model Behavior: Ignores Distracting Requests When Focused — MannyKayy · 2026-08-03
- Kimi K3 Max Agent Benchmark: Delivers 2.8x More Solved Tasks Per Dollar — togethercompute · 2026-08-03
- LLM Token Price Index Plummets as Demand Shifts to Cheaper Models — AccBalanced · 2026-08-03