Redditor builds dual AI workstation with 2x RTX 3090 and 4x Tesla P100, asks how to optimize
FearFactory2904 · reddit · 2026-09-14
Reddit user FearFactory2904 assembled a two-system local AI workstation on aging Xeon x99 boards and asked the community for optimization advice.
System 1 (running): Xeon E5-2630V4 with 128GB DDR4 and two RTX 3090s (no NVLink, PCIe 3.0 x16):
- One GPU runs Qwen 3.8 27B Q4 quant served via Open WebUI
- The other runs ComfyUI for image generation/editing, called through Open WebUI's API
- CPU handles speech-to-text and text-to-speech
- A separate Docker container runs SearXNG to add web search to Open WebUI
System 2 (being rebuilt): Same Xeon with 256GB DDR4 and four Tesla P100 16GB cards. Plans include a Q8 quant of Qwen 3.8 27B with larger context, or trying Qwen 3.8 Flash, plus experimenting with agentic workflows.
The goal is a local "Swiss army knife" of AI features; the author asks how others would reallocate the same hardware.
More from Infra
- TSMC paces the frontier by accident whenever it underestimates chip demand — dan_s_becker · 2026-09-14
- Dev warns: shady cheap-token inference providers send fake tool calls and resell your traces — Nils_Reimers · 2026-09-14
- SemiAnalysis: 4-hi HBM wins on $/bandwidth, cutting inference cost amid DRAM shortage — dylan522p · 2026-09-14
- Used RTX 5090 listed at £3,900 (~$5,200), more than double its MSRP — julianharris · 2026-09-14
- Why Amazon and Microsoft Are Taking Communities' Side Against Utilities — pstAsiatech · 2026-09-14
- agi-memory: SQLite-only persistent memory MCP server for coding assistants, 32MB RAM — Rude_Gate7599 · 2026-09-14