parakeet_cuda: local Parakeet transcription on a 2GB VRAM GTX 750 Ti, ~20x realtime

mgostIH · x · 2026-10-02

A developer wrote a custom CUDA runtime (parakeetcuda, Apache-2.0) so NVIDIA's Parakeet TDT 0.6B v3 and Nemotron diarization models run on a GTX 750 Ti with just 2GB VRAM. Exact FP32 inference with cuBLAS: 280 MiB for ASR, 384 MiB with diarization, roughly 20x realtime — local audio transcription on nearly any hardware.

Original post →

More from Infra

Infra channel →