Kimi K3 and DeepSeek V4.1 Keep Failing on NVIDIA NIM While GLM 5.3 Works Fine
Professional_Log1367 · reddit · 2026-10-09
A user reports persistent "took too long to respond" errors when calling Kimi K3 and DeepSeek V4.1 Flash via NVIDIA NIM in opencode, while GLM 5.3 and other models return fine (just slowly). Diagnostics show status:200 failures producing zero reasoning tokens and one or two content items — starting then stopping — and a DeepSeek failure that ran exactly 200s with empty content, well before their 30-minute timeout and not matching rate limits. All flagged models have reasoning: true, so reasoning mode isn't the differentiator. The post asks why these specific models fail on NIM.
More from Infra
- lithos-metal Megakernels Beat Ollama on M5 Pro, Tuned for M5 Max — JiaZhihao · 2026-10-09
- Evoke open-sourced: 30M-param model in Postgres matches 0.6B embedding model on recall — bdsqlsz · 2026-10-09
- Dev discovers vLLM accepts input embeds, speculates finetuning tokens are underused — cephaloform · 2026-10-09
- Kevin Kwok essay: How TSMC manufactured its own market via Abstract Warfare — kevinakwok · 2026-10-09
- Self-hosted SearXNG behind 67 proxy chains, accessible only via tailnet — haydendevs · 2026-10-09
- iPad + AMD R9700 eGPU Runs Qwen3.8-27B at 159 tok/s via Open-Source LSE Engine — TheOriginalG2 · 2026-10-09