Running Inkling Small on a Single Node: Native Voice Interaction Under 500ms

andimarafioti · x · 2026-07-31

A developer tested the newly released Inkling Small, highlighting that it fits on a single node (8x RTX Pro 6000 Blackwell), whereas the flagship requires 2TB of VRAM.

Plugged into HF's speech-to-speech pipeline, the model takes raw audio input directly—bypassing text transcription—and replies via faster-Qwen3TTS with an end-to-end latency under 500ms. Because the model processes native audio, it can pick up on the user's tone and emotions, showcasing strong potential for real-time multimodal interaction.

Original post →

More from Infra

Infra channel →