Looking for the fastest local vision model to run on an RTX 5090
StartupTim · reddit · 2026-09-15
A game developer wants to build an AI that plays their own game: feed it screenshots of a top-down view with simple prompts and RAG, and have it decide what to do next.
The hard constraint is raw speed on an RTX 5090 — general intelligence matters far less than latency. They're asking the community for recommendations on the fastest local vision models for this kind of real-time screen-reading workload.
More from Infra
- Orthrus Study: Lossless Speculative Decoding Holds Only at High Numerical Precision — Ilya Koziev · 2026-09-15
- DeepSeek V4.1 Flash on M3 Ultra nearly doubles decode to 31 t/s with first public DSpark Metal port — IngeniousIdiocy · 2026-09-15
- Trigger.dev's chat.agent turns AI chats into durable tasks that survive crashes and redeploys — CodeByPoonam · 2026-09-15
- Oracle Executes Pre-Dawn Mass Layoffs as AI Data Center Spending Balloons — 量子位 · 2026-09-15
- Same GB300, same workload: switching serving engine moved the benchmark result by 11x — Slight_Republic_4242 · 2026-09-15
- Solar panels now ~$0.12/watt, down from $5-6, seen as key to data center growth — kimmonismus · 2026-09-15