Local voice assistant auto-picks its Ollama model by VRAM, shows status on Arduino

PrimeEmre · reddit · 2026-10-04

A developer shares a local JARVIS-style voice assistant: a Flask app answers questions via a local model through Ollama, grounds answers with DuckDuckGo search, reads replies with a local TTS server, and drives an Arduino Uno status display (one serial byte: B=researching, Y=debating, W=done, with blue/yellow LEDs and a sweeping servo).

The handiest part is VRAM-based model tiering: at startup it reads total VRAM from nvidia-smi and picks a tier — 24GB→qwen2.5:32b (16k context, 8 search results), 16GB→14b, 8GB→llama3.1:8b, floor→qwen2.5-coder:1.5b, up to 72b at 48GB; an OVERRIDEMODEL env var skips detection. Search-result counts scale with model size so smaller models aren't flooded with context.

Voice uses OmniVoice Studio's OpenAI-compatible local API; with under 5GB free VRAM, switch to CPU-only KittenTTS. Code and a full setup write-up are open-sourced; the author also debates VRAM tiering vs letting Ollama offload layers to RAM.

Original post →

More from coding & agent

coding & agent channel →