Debugging latency in Ollama + OpenWebUI stacks: is the bottleneck the UI or the model?

daddyMaterialBolte · reddit · 2026-08-23

A user is optimizing an internal deployment using NVIDIA Nemotron, Ollama, and OpenWebUI/AnythingLLM, seeing 20 tokens/sec generation with 95% GPU utilization. They suspect the bottleneck lies in the orchestration layer rather than the model itself.

Key questions raised:

Original post →

More from Infra

Infra channel →