Macaron-V1-Venti ships day one on vLLM with open-source multi-LoRA routing
vllm_project · x · 2026-07-22
Macaron-V1-Venti launched with day-one support on vLLM and SGLang.
- The team’s open-source MoL Harness routes each request to the right LoRA specialist behind a single OpenAI-compatible endpoint.
- It relies on vLLM’s native Multi-LoRA serving.
- The post highlights an open implementation built for modular expert routing and easy deployment.
More from Infra
- OpenFPM CUDA-style kernels now run on Apple Silicon GPUs via Metal — Scobleizer · 2026-07-22
- SkewAdam cuts MoE optimizer memory by 97.4% and fits 6.78B on a 40GB GPU — Kooky-Ad-4124 · 2026-07-22
- Samsung is said to weigh a €1B Mistral investment at a €20B valuation — rohanpaul_ai · 2026-07-22
- NVIDIA and ETH Zürich cut small-message AllReduce latency by deleting barriers — thoefler · 2026-07-22
- Grok Build adds token usage, batching and diagnostics for developers — elonmusk · 2026-07-22
- Qwen3.8-Max-Preview ranks No. 1 on NVIDIA’s FlashInfer benchmark — Scobleizer · 2026-07-22