Debugging latency in Ollama + OpenWebUI stacks: is the bottleneck the UI or the model?
daddyMaterialBolte · reddit · 2026-08-23
A user is optimizing an internal deployment using NVIDIA Nemotron, Ollama, and OpenWebUI/AnythingLLM, seeing 20 tokens/sec generation with 95% GPU utilization. They suspect the bottleneck lies in the orchestration layer rather than the model itself.
Key questions raised:
- Do interfaces like OpenWebUI add significant latency through prompt construction, RAG, or API overhead?
- How to accurately measure TTFT, prompt processing time, and generation time?
- What methods can systematically benchmark the model directly via Ollama vs. through the UI layer?
- How to verify if the UI is sending larger prompts than visible and measure the impact of context history?
More from Infra
- Mistral reportedly plans up to 1 GW of European compute capacity by 2030 — emmanuelvivier · 2026-08-23
- Nvidia hikes AI product prices by over 15% amid rising memory costs — emmanuelvivier · 2026-08-23
- Nvidia hikes some AI product prices by over 15% on surging memory chip costs — emmanuelvivier · 2026-08-23
- Contextual News Search APIs: A Deep Comparison for AI, RAG, and Research — ermanos12 · 2026-08-23
- Qwen3.8-27B MTP Grafted to Unsloth Saves RAM, Requires Thinking Mode — Nyghtbynger · 2026-08-23
- Running Kimi K3 on 8x B300: $190 per million tokens, full cost breakdown — OtherRaisin3426 · 2026-08-23