vMLX: OpenAI-Compatible Inference Server for Apple Silicon with JANG Quantization

tom_doerr · x · 2026-08-15

vMLX is a self-hosted inference server for LLMs, VLMs, and image generation on Apple Silicon. It features an OpenAI, Anthropic, and Ollama compatible HTTP API, requiring no third-party keys. The project introduces JANG 2-bit quantization, which achieves 74% on MMLU compared to MLX 4-bit's 26.5%, with a smaller model size. Additional features include native MTP artifact detection, hybrid SSM scheduler, and continuous batching.

Original post →

More from Infra

Infra channel →