vMLX: OpenAI-Compatible Inference Server for Apple Silicon with JANG Quantization
tom_doerr · x · 2026-08-15
vMLX is a self-hosted inference server for LLMs, VLMs, and image generation on Apple Silicon. It features an OpenAI, Anthropic, and Ollama compatible HTTP API, requiring no third-party keys. The project introduces JANG 2-bit quantization, which achieves 74% on MMLU compared to MLX 4-bit's 26.5%, with a smaller model size. Additional features include native MTP artifact detection, hybrid SSM scheduler, and continuous batching.
More from Infra
- Dev realizes memory bandwidth bottleneck after switching to a27b; asks for A6000 thermal paste guide — cephaloform · 2026-08-15
- Report: 36% of ERC-8004 Agents Share Identities, Index Analysis Reveals — GlobalScoreAgent · 2026-08-15
- Study: Anthropic models may be cheaper than some open-source Chinese models — rohanpaul_ai · 2026-08-15
- Polygres turns Postgres into extended context for AI agents — Scobleizer · 2026-08-15
- vLLM Introduces Adaptive Verification for Speculative Decoding with DSpark — vllm_project · 2026-08-15
- Open source closes the gap with closed labs: Quality gap now just months — togethercompute · 2026-08-15