Seeking advice for local dev setup on dual RTX 6000s
alexp702 · reddit · 2026-08-31
A user with a Threadripper (128GB RAM) and two RTX 6000 Pro Max-Q GPUs wants to set up a local coding harness for a small team. They are considering GLM, Deepseek R4, and Qwen next, preferring models with vision capabilities. Since quantization is needed for dual-card deployment, they are asking for advice on model choice and VLLM configurations.
More from coding & agent
- Codex struggled for 3 hours on Shortcuts; Worker solved it in 5 — banteg · 2026-08-31
- How to implement an agent's decision to gather more information? — Smart_Promise441 · 2026-08-31
- Agent Audit Log Security: Don't let agents control the logs — amu4biz · 2026-08-31
- Defining the Harness: Distinguishing model logic from agent scaffolding — JeremyCMorgan · 2026-08-31
- Templater debugging tip: Use debugger statement — dSebastien · 2026-08-31
- Developers miss the accelerated coding flow state from early AI models — Dimillian · 2026-08-31