A fully local AI girlfriend runs on 15GB VRAM with Whisper, llama.cpp, and Qwen3-TTS
max_paperclips · x · 2026-07-27
A fully local AI girlfriend runs end to end with no internet and no API keys, using Silero VAD v5, Whisper, llama.cpp, and Qwen3-TTS.
The project reportedly fits all four models into just 15GB of VRAM and supports hot-swapping models on the fly, showing how consumer hardware can now handle smooth real-time voice conversations that once required cloud infrastructure.
Related event: Fully Local AI Girlfriend Voice Demo Runs on 15GB VRAM(2 posts)→
More from Infra
- New Q8_CR GGUF format keeps Krea 2 diffusion speed near INT8 while shrinking VRAM pressure — molbal · 2026-07-27
- Nvidia signs $1.5B multi-year Amkor deal to expand US chip packaging capacity — Beth_Kindig · 2026-07-27
- PyTorch DDP misses a tiny-parameter NVIDIA GPU optimization out of the box — gordic_aleksa · 2026-07-27
- OpenAI reportedly plans to spend over $30B on a 3.2GW Georgia data center — Beth_Kindig · 2026-07-27
- Two RTX 5080s failed to beat one RTX 5090 in ComfyUI single-render tests — Geekdomo · 2026-07-27
- Open-source profiler tracks every STT, LLM, and TTS call in self-hosted voice agents — mahimairaja · 2026-07-27