Open-source C-based LLM engine runs Gemma voice conversations on Jetson Orin, faster than llama.cpp

cortexist · reddit · 2026-09-08

A developer demoed a fully open-source two-device voice conversation system: Gemma 12B on an RTX PRO 4500 Blackwell and Gemma E2B on a Jetson Orin NX 16GB (similar performance expected on Orin Nano Super 8GB), both using a reSpeaker Flex 4-mic array and 3W speaker.

Inference is handled by Cortexist Little Gemma, a small LLM engine written in C for CUDA devices — claimed to be faster than llama.cpp on Jetson Orin with no degradation on long voice prompts. The pipeline supports lip sync, expressions, and gestures. Everything is open source.

Original post →

More from Embodied

Embodied channel →