omlx: LLM Inference Server with SSD Caching for Apple Silicon
jundot · github · 2026-08-17
- Overview: An LLM inference server optimized for Apple Silicon featuring continuous batching and SSD caching to handle larger models.
- UX: Managed directly from the macOS menu bar for a seamless local experience.
- Compatibility: Offers an OpenAI-compatible API interface for easy integration.
More from Infra
- Local AI Primer: Benchmarking 2B to 27B Models on Consumer Hardware — draginol · 2026-08-17
- LocalAI adds 17 C++ inference backends in six months — rms80 · 2026-08-17
- Nvidia: Land, Power, and Shell Are Critical Resources for AI Factories — firstadopter · 2026-08-17
- Combining 3060 and 4070 Super for local AI: feasibility and setup — HsSekhon · 2026-08-17
- NVIDIA Invests $1.5B to Secure 8GW Capacity for OpenAI in Ohio — nvidia · 2026-08-17
- NVIDIA to Provide Exclusive AI Infrastructure at Ohio Campus for OpenAI — nvidia · 2026-08-17