parakeet_cuda: local Parakeet transcription on a 2GB VRAM GTX 750 Ti, ~20x realtime
mgostIH · x · 2026-10-02
A developer wrote a custom CUDA runtime (parakeetcuda, Apache-2.0) so NVIDIA's Parakeet TDT 0.6B v3 and Nemotron diarization models run on a GTX 750 Ti with just 2GB VRAM. Exact FP32 inference with cuBLAS: 280 MiB for ASR, 384 MiB with diarization, roughly 20x realtime — local audio transcription on nearly any hardware.
More from Infra
- Local AI roundup: 27B reasoning in 5.9GB, phone-class 35B, and dozens more — vramkickedin · 2026-10-02
- Liquid AI's Decision Model D1 Hits OpenRouter: Typed Answers With Probabilities at $0.04/M Input, $0 Output — maximelabonne · 2026-10-02
- Teenager tapes out a chip, rebuilds GPU interconnects with optics and poaches Nvidia veterans — ai · 2026-10-02
- huggingface_hub v2.1.0 ships: 17x faster downloads, job retries, rerun and port exposure — huggingface · 2026-10-02
- Infra engineers compare notes on when self-hosting LLMs beats paying for APIs — One_Mention_5385 · 2026-10-02
- Gemini 4 isn't even out yet — but Google's TPU advantage is being underappreciated — haider1 · 2026-10-02