DSv4 Flash Tool Call Stability Issues at Long Context
captaintobs · reddit · 2026-08-21
A user running DSv4 Flash via jasl/vllm on 4 Spark instances reported that tool calls start failing (e.g., missing characters like '>') after about 100k context tokens. While the model offers high throughput (50 tps single, >100 tps aggregate), it lacks stability. GLM 5.2 is found more efficient and correct but significantly slower (25 tps).
More from Models
- New Ox Alpha model leaked with 1M context and multimodal input — External_Mood4719 · 2026-08-21
- Mystery Model Ox Alpha Suspected to be Zhipu GLM — teortaxesTex · 2026-08-21
- OpenAI open-sources Codex evaluation harness — Armmani · 2026-08-21
- OpenRouter's stealth model Ox Alpha sparks guesses about Chinese AI vendors — lxfater · 2026-08-21
- Ox Alpha model test: Impressive performance in Minecraft — max_paperclips · 2026-08-21
- Reviewer notes 1T model shares similarities with GLM 5.3, adopts V4-Flash style reasoning — teortaxesTex · 2026-08-21