AI agent cuts Orpheus 3B TTS RTF from 1.03 to 0.87 via TensorRT tuning

TheMoonMidas · x · 2026-09-27

A developer had an AI agent named Astra test and optimize an old Orpheus 3B TTS deployment, cutting RTF from 1.03 to 0.87 at 16 concurrent generations. The same attempt with GPT last December failed. The accompanying open-source repo (orpheus-tts-api) serves speech over WebSocket using TensorRT-LLM plus SNAC decoding, with male/female voices, request cancellation, silence trimming, optional telemetry, and Docker deployment (Linux, NVIDIA GPU, CUDA 13.1).

Original post →

More from coding & agent

coding & agent channel →