Running 256k-context open models on 2x RTX 3090 for months: a home server LLM retrospective
knighty1981 · reddit · 2026-10-03
The author details a local LLM setup: Supermicro GPU server (PCIe 3.0, 8-slot), dual Xeon Gold 6230R, 230GB RAM, 2x RTX 3090 24GB, running huihui-27b and qwen3.8-27b at 256k context via opencode remotely.
Real-world usage so far:
- Built a route-planner webserver on another machine
- Extensive Home Assistant automation work
- Designed a FreePBX + Whisper VoIP system (awaiting hardware)
- Extracted and summarized trends from thousands of Excel delivery sheets/invoices
- Competitor research; spent 3 days autonomously configuring prowlarr/radarr/sonarr/qbittorrent behind a VPN on a Synology NAS — a task the author had failed at manually
Lessons: 256k context needs repeated compression; the agent frequently uses wrong SSH commands and misconfigures settings requiring manual rollback; the author reflects that smaller, pre-planned task decomposition would have worked better, and believes >256k context increases hallucination and loops.
Upgrade questions posed: add 2 more 3090s to split the model faster, run different models per card ('council of AI'), whether larger models are worth it, and whether PCIe 3.0 bottlenecks multi-card tensor parallelism.
More from coding & agent
- Dev orchestrates ~150 coding agents from his car via Tailscale and a basement server — haydendevs · 2026-10-03
- Dev orchestrates 150 Opus sub-agents from his car via a home server and OpenAI Dot — haydendevs · 2026-10-03
- Gmail MCP Server brings secure email management to MCP clients — modelcontextprotocol · 2026-10-03
- ChapterPal dev: frontier vision models consistently fail to spot obvious webpage conversion artifacts — burkov · 2026-10-03
- Dev finds Argon enough for nearly all coding tasks, misses it after switching to Opus 5.5 — m2saxon · 2026-10-03
- Air-gap file transfer via animated QR codes flashing at 10-30 frames per second — Thionne_WTZ · 2026-10-03