Dual RTX 3090 owner asks which local LLM stack actually works today
ruffus_or · reddit · 2026-07-25
A user with two RTX 3090s asks how to choose local LLMs for a 48 GB VRAM setup, comparing vLLM, Ollama, and llama.cpp, plus quantization formats like GGUF, AWQ, GPTQ, and FP8. The thread also asks for concrete coding workflows: model choice, IDE integration, and tools such as Continue, Cline, Roo Code, Aider, and Open WebUI.
More from coding & agent
- ECC is a harness layer for Claude Code, Codex and Cursor agents — affaan-m · 2026-07-25
- Andrew Ng’s aisuite offers one interface for multiple GenAI providers — andrewyng · 2026-07-25
- AFK Pilot links Grok Build in VS Code to your phone with no tunnel or config — PawelHuryn · 2026-07-25
- Kimi K3 Coding Review: Is the Subscription Better Than the Pricey API? — agentcubed · 2026-07-25
- Four open models fix the same three real bugs, with a 27B model 14× faster — MaziyarPanahi · 2026-07-25
- Baichuan livestream argues scenario-specific skills must be customized, not reused — aigclink · 2026-07-25