DeepSeek V4 MLX 8-bit Gets Local Inference Speedup

No_Run8812 · reddit · 2026-07-06

A user optimized the DeepSeek V4 Flash 8-bit MLX model on oMLX using Codex, achieving around 1.6x prefill and 3x decode speedups. They utilized the 302GB 8-bit affine MLX version from mlx-community to avoid the precision loss associated with 4-bit. The author, noting they are not a Metal kernel expert, shared the modification details for community review, confirming that tool-calling currently works normally with reportedly no loss in precision.

Original post →

More from Infra

Infra channel →