A company with 2 H200s asks which coding model and vLLM setup can serve 4–10 users
redblood252 · reddit · 2026-07-28
- The author says their company has access to 2 H200 GPUs and wants recommendations for an agentic coding model.
- They are targeting 256k context and 4–10 concurrent users via vLLM, usually around 4 and rarely above that.
- They ask which model and settings offer the best balance between coding quality and speed.
- They also ask whether it is a good idea to reboot the vLLM container daily.
More from coding & agent
- Ruff 0.16.0 Lets One Prompt Fix an Entire Codebase — xeophon · 2026-07-28
- OpenAI is redesigning engineering workflows for a world where agents write code — every · 2026-07-28
- OpenAI is redesigning code review as AI-generated code floods engineering teams — every · 2026-07-28
- OpenCode adds Kimi K3 to its coding-agent model lineup — zainhas · 2026-07-28
- A practical cloud-agent stack: Devin, Cursor, Open-Inspect, Infisical and more — vinvan · 2026-07-28
- Stripe upgrades Directory with instant listings and MCP search for agents — jeff_weinstein · 2026-07-28