Best local uncensored vision+text models for an RTX 5070 Ti 16GB setup?
Mystvearn_ · reddit · 2026-09-12
A Reddit user wants a fully local, uncensored vision+text workflow on an RTX 5070 Ti (16GB VRAM) with 64GB DDR4, mimicking cloud models like Gemini and Claude for image description and prompt brainstorming. They ask which open-weight multimodal models, quantization levels, and software stacks (Ollama, LM Studio) best fit the 16GB constraint.
More from Infra
- Frontier model for planning, local Qwen for coding: a hybrid dev workflow experiment — kirisoraa · 2026-09-12
- Qwen3.8-Flash-Next only hits 15 tok/s on 4x RTX 5060 Ti 16GB setup — Ambitious_Fold_2874 · 2026-09-12
- Local-First AI Inference Cuts PDF Processing API Costs 75% Across 4,700 Documents — bibryam · 2026-09-12
- From Accuracy to Latency: Why Inference Engineering Is a Different Discipline — mdancho84 · 2026-09-12
- CheckCle: self-hosted open-source full-stack monitoring platform hits 2.9k GitHub stars — tom_doerr · 2026-09-12
- AI in Space Is Mostly Inference: '$10/Month 200-IQ Employees' Means Infinite Demand — JOBhakdi · 2026-09-12