Forking exllamav3 for Jetson Orin: 1136 tok/s prefill on Gemma4 E4B, beating llama.cpp

cortexist · reddit · 2026-10-12

A developer optimized Gemma E4B on Jetson Orin Nano Super 8GB: after little-gemma V1.0 beat llama.cpp on decode (26.36 vs 19.49 tok/s), they forked exllamav3 and ported key code to get EXL3 running — 1136 tok/s prefill on a 930-token prompt (vs 622 for llama.cpp with FA) and 878ms first token, with slightly lower decode but better overall voice-chat latency. Fork and quantized weights are public.

Original post →

More from Infra

Infra channel →