liuliu warns: claimed 4x-10x speedups over MLX or llama.cpp on Apple hardware are noise
teortaxesTex · x · 2026-09-23
Researcher liuliu followed up on his earlier take: be suspicious of any claims of 4x, 8x, or 10x speedups over MLX or llama.cpp from "custom / model-specific inference engines" on Apple hardware — such numbers are mostly noise.
More from Infra
- Flash-dLLM accelerates diffusion LLMs up to 11x with IO-aware KV caching — MBZUAI · 2026-09-23
- vLLM v0.30.0 ships with 762 commits: watermarking, HiSparse, Model Runner V2 — vllm_project · 2026-09-23
- Sea first ASEAN company to adopt NVIDIA Vera Rubin as Nemotron spreads across Southeast Asia — NVIDIA Blog · 2026-09-23
- DeepSeek DSec cluster BoM estimated at ≤$20M; $1B could buy 50 clusters and 19M concurrent sandboxes — teortaxesTex · 2026-09-23
- B200 Rental Prices Hit All-Time High at $7.88 per GPU-Hour — sudoraohacker · 2026-09-23
- ASML shipped zero new systems to European customers in Q2; Korea took 43% — FinanceYF5 · 2026-09-23