Open-source C-based LLM engine runs Gemma voice conversations on Jetson Orin, faster than llama.cpp
cortexist · reddit · 2026-09-08
A developer demoed a fully open-source two-device voice conversation system: Gemma 12B on an RTX PRO 4500 Blackwell and Gemma E2B on a Jetson Orin NX 16GB (similar performance expected on Orin Nano Super 8GB), both using a reSpeaker Flex 4-mic array and 3W speaker.
Inference is handled by Cortexist Little Gemma, a small LLM engine written in C for CUDA devices — claimed to be faster than llama.cpp on Jetson Orin with no degradation on long voice prompts. The pipeline supports lip sync, expressions, and gestures. Everything is open source.
More from Embodied
- Robotics researcher: harness + tool calls beat raw VLA on control benchmarks, at 2.3x lower cost — rbhar90 · 2026-09-08
- Grok bot plays a Roland piano, composing romantic-style music from a text prompt — yunta_tsai · 2026-09-08
- 360CityArena at ECCV 2026: best AI agent scores 17.1% vs humans' 77.3% on city navigation — CSProfKGD · 2026-09-08
- Robotics veteran Jitendra Malik challenges LLM robotics hype with a dexterity benchmark — JitendraMalikCV · 2026-09-08
- Self-driving demo runs on nothing but logic gates, moves like an insect — burny_tech · 2026-09-08
- Xpeng announces first automated production line where robots mass-produce robots — Anxious-Yoghurt-9207 · 2026-09-08