Ask HN: Best way to add vision support to Deepseek v4 Flash on DGX?
StartupTim · reddit · 2026-08-24
A user is running Deepseek v4 Flash 0731 on 2x DGX Sparks with 1M context and vLLM TP2 but needs vision support.
Current blockers:
- Available vision encoders either require disabling the thinking mode in vLLM or rely on sglang, which is difficult to deploy on their RDMA cluster.
- Workarounds like Image MCP and vision plugins underperform compared to native multimodal models.
The user seeks advice on the best engineering approach to enable image/vision capabilities for DSv4F without sacrificing performance.
More from coding & agent
- Agent Memory Should Track Outcomes, Not Just Steps: Adding Success Counters to Workflows — No_Advertising2536 · 2026-08-24
- AI Agents struggle to fail fast, often avoiding error responsibility — mattpocockuk · 2026-08-24
- MCP servers criticized as new grift wrapping scripts in paywalled APIs — jasonkneen · 2026-08-24
- Open-Source 14-Agent System Runs 24/7 Company on Laptop, No Vector DB — thisdudelikesAI · 2026-08-24
- Agno Builder Now Supports Visual Construction of Agents and Code Export — pritisinghhhh · 2026-08-24
- 48h Hackathon Winner: Draw infrastructure, AI audits it — aahiknsv · 2026-08-24