DeepSeek V4.1 Flash Keeps Timing Out on 2-Hour Agentic Benchmarks, Author Shares Failure Logs
sebnadeau · reddit · 2026-09-16
A benchmark author reports DeepSeek V4.1 Flash via OpenRouter never finished a 2-hour agentic coding task while GLM and Qwen completed in 27-110 minutes. Detailed logs: two 502 providerunavailable errors from Together per hour, 32 empty replies (7 hitting the 32,768-token cap), 600s timeouts — despite impressive cost ($8.34 for 143M input tokens). Asks where to reliably run long agentic sessions with DeepSeek.
More from Infra
- Unlocking CMP 170HX to 80GB VRAM: most users hit a stable wall at 40GB — Feralzi · 2026-09-16
- Skip the $9K RTX 5090: fly to Taipei, buy at ~$4,093, vacation for two weeks — DegenDataGuy · 2026-09-16
- Ben Bajarin doubles down on Credo: photonics-plus-SerDes vertical integration drives growth — BenBajarin · 2026-09-16
- Ocean Protocol launches hourly dedicated-GPU inference, H200 from $2.16/hour — w1kke · 2026-09-16
- NVIDIA Vera CPU completes agentic task lifecycles 1.64x faster than x86, Signal65 finds — ryanshrout · 2026-09-16
- Run your AI agents from anywhere with Tailscale and a simple PWA — johnlindquist · 2026-09-16