Kimi K3 and DeepSeek V4.1 Keep Failing on NVIDIA NIM While GLM 5.3 Works Fine

Professional_Log1367 · reddit · 2026-10-09

A user reports persistent "took too long to respond" errors when calling Kimi K3 and DeepSeek V4.1 Flash via NVIDIA NIM in opencode, while GLM 5.3 and other models return fine (just slowly). Diagnostics show status:200 failures producing zero reasoning tokens and one or two content items — starting then stopping — and a DeepSeek failure that ran exactly 200s with empty content, well before their 30-minute timeout and not matching rate limits. All flagged models have reasoning: true, so reasoning mode isn't the differentiator. The post asks why these specific models fail on NIM.

Original post →

More from Infra

Infra channel →