Dual RTX 3090 owner asks which local LLM stack actually works today

ruffus_or · reddit · 2026-07-25

A user with two RTX 3090s asks how to choose local LLMs for a 48 GB VRAM setup, comparing vLLM, Ollama, and llama.cpp, plus quantization formats like GGUF, AWQ, GPTQ, and FP8. The thread also asks for concrete coding workflows: model choice, IDE integration, and tools such as Continue, Cline, Roo Code, Aider, and Open WebUI.

Original post →

More from coding & agent

coding & agent channel →