A fully local AI girlfriend runs on 15GB VRAM with Whisper, llama.cpp, and Qwen3-TTS

max_paperclips · x · 2026-07-27

A fully local AI girlfriend runs end to end with no internet and no API keys, using Silero VAD v5, Whisper, llama.cpp, and Qwen3-TTS.

The project reportedly fits all four models into just 15GB of VRAM and supports hot-swapping models on the fly, showing how consumer hardware can now handle smooth real-time voice conversations that once required cloud infrastructure.

Related event: Fully Local AI Girlfriend Voice Demo Runs on 15GB VRAM(2 posts)→

Original post →

More from Infra

Infra channel →