Inflect-Nano-v2 packs a full English TTS stack into just 3.96 million parameters
_akhaliq · x · 2026-07-28
- Inflect-Nano-v2 is a complete English text-to-speech model in 3.96M parameters.
- It targets ultra-compact local deployment, with punctuation-aware long-form speech, deterministic output, and CPU/CUDA support.
- The model card also highlights a public adaptation toolkit for preparing data, auditing splits, adapting a voice or language, resuming training, evaluating checkpoints, and exporting PyTorch or ONNX packages.
- The author says the project was built and funded independently and hopes a real audience will justify a broader v3.
Related event: Inflect-Nano-v2: A 3.96M Parameter Fine-Tunable TTS Model(2 posts)→
More from Multimodal
- Claude wrote an entire song purely in code, no Suno involved — ctjlewis · 2026-09-23
- Opus 5.5 turns a single image into a Three.js game menu in one simple prompt — majidmanzarpour · 2026-09-23
- PixVerse's R2 world model goes hands-on: endless exploration, but compute limits cap play time — Xianbao_QIAN · 2026-09-23
- Local image generation tests: Qwen base model combined with a Flux refiner — freshstart2027 · 2026-09-23
- Reddit users find Qwen 2.1 a major disappointment for text-to-image — -becausereasons- · 2026-09-23
- Given a 3-hour budget and a one-line prompt, Opus 5.5 produced a full Kowloon horror short itself — rainbird · 2026-09-23