Ex-SWE Turned Inference Engineer: Ollama Is Never Optimal — A Bottleneck Guide
Postmodern_Plunger · reddit · 2026-10-02
A former software engineer who moved into inference engineering and took his side hustle full time shares common inference misconfigurations.
Runtime choice:
- Ollama is never optimal, just simplest — fine for quick setups with spare RAM, but not fastest
- Use llama.cpp when CPU offload is needed; VLLM is usually best for pure GPU
- SGLang may be worth it for Langgraph workloads; otherwise the two above cover 90% of cases
- Stick to mainline branches; community forks target niche models/configs and are often under-optimized and poorly maintained
Runtime optimization:
- Compile with hardware-specific flags; architecture-specific optimizations are the most commonly missed
- Recompile after system updates, kernel/driver updates, running a model newer than your last compile, or if you haven't recompiled in over a month
Quantization:
- IQ quants are generally best for the size; IQ4XS has a smaller VRAM footprint and higher complexity than Q4KM
- Nonlinear (NL) quants only win with CPU offload, and IQ often offers more context space
- Read the docs for unfamiliar quants; searching or asking ChatGPT "gives the wrong answer every single time" for custom quants
- Standard QKM quants have the best path optimizations when there are no hardware constraints — but then you could run a better model
Task fit: For coding via API the author picks Anthropic — Opus and Sonnet 5.5 perform well with low token usage per task, making them cheaper than prior iterations; DeepSeek is also mentioned (text truncated).
More from Infra
- Cloudflare launches SQL API to query Workers logs and traces, replacing GraphQL plans — irvinebroque · 2026-10-02
- Cloudflare launches Web Search API via AI Gateway with Exa, Linkup and Ceramic — michellechen · 2026-10-02
- Lightpanda 1.0 ships: a Zig-built browser for AI agents, out of beta — jedisct1 · 2026-10-02
- Qwen3.8-27B coder quant fits a 24GB GPU with 262k context at 40 t/s — W61k3r · 2026-10-02
- Suffix Cache Reuse: fixing KV cache for agents that edit context in place — RulinShao · 2026-10-02
- NVIDIA's New DGX Spark SKU Halves Memory to 64GB While Selling at Nearly the Same Price — firstadopter · 2026-10-02