DeepSeek V4 Flash Vision impresses via API, but 305B needs 4x GB300 to run
No_Issue_8224 · reddit · 2026-09-04
DeepSeek's V4-Flash-Vision-Exp fills the model family's vision gap. Informal API tests via ZenMux showed strong image recognition, but the 305B checkpoint requires a node with four GB300 GPUs per DeepSeek's vLLM recipe — viable for high-volume, latency-sensitive teams, impractical for hobbyists. Quantized local performance remains unknown.
More from Infra
- Video generation isn't solved: fal h3 costs $22/hour vs $0.24 mobile gamer spend, 100x gap remains — IndraVahan · 2026-09-04
- Apple's 512GB M5 Ultra can run 93 of 97 open-weight models locally, up to ~50% faster inference — bigaiguy · 2026-09-04
- Extropic founder jokes Astra on Cerebras will feel like riding a Bugatti at 400kph — beffjezos · 2026-09-04
- Paddock Open-Sourced: Rust/C++ Inference Engine With Own CUDA Kernels Beats vLLM in All 13 Cells — saltexx · 2026-09-04
- Tim Sweeney says gaming faces worst crash since the 1980s as AI drives RAM costs — tekbog · 2026-09-04
- Transformers.js podcast: browser background removal, transcription and local agent workflows — nicodotdev · 2026-09-04