Running 256k-context open models on 2x RTX 3090 for months: a home server LLM retrospective

knighty1981 · reddit · 2026-10-03

The author details a local LLM setup: Supermicro GPU server (PCIe 3.0, 8-slot), dual Xeon Gold 6230R, 230GB RAM, 2x RTX 3090 24GB, running huihui-27b and qwen3.8-27b at 256k context via opencode remotely.

Real-world usage so far:

Lessons: 256k context needs repeated compression; the agent frequently uses wrong SSH commands and misconfigures settings requiring manual rollback; the author reflects that smaller, pre-planned task decomposition would have worked better, and believes >256k context increases hallucination and loops.

Upgrade questions posed: add 2 more 3090s to split the model faster, run different models per card ('council of AI'), whether larger models are worth it, and whether PCIe 3.0 bottlenecks multi-card tensor parallelism.

Original post →

More from coding & agent

coding & agent channel →