Forking exllamav3 for Jetson Orin: 1136 tok/s prefill on Gemma4 E4B, beating llama.cpp
cortexist · reddit · 2026-10-12
A developer optimized Gemma E4B on Jetson Orin Nano Super 8GB: after little-gemma V1.0 beat llama.cpp on decode (26.36 vs 19.49 tok/s), they forked exllamav3 and ported key code to get EXL3 running — 1136 tok/s prefill on a 930-token prompt (vs 622 for llama.cpp with FA) and 878ms first token, with slightly lower decode but better overall voice-chat latency. Fork and quantized weights are public.
More from Infra
- OpenRouter token traffic explodes from 2T to 379T/month, open-weight models at 75% — Beth_Kindig · 2026-10-12
- ASUS RTX 5090 listed at €11,079 in Germany, only 3 units left in stock — janusch_patas · 2026-10-12
- Modular Claims Mojo Kernels Beat FlashAttention 4, Launches MAX Inference Framework — AI Engineer · 2026-10-12
- DeepSeek-V4.1-Flash on 2x DGX Spark: TP2 Patches Cut First Token From 33.8s to 4.9s — rez0__ · 2026-10-12
- Ed Zitron: AI needs $425B annual revenue to justify hyperscaler capex, $260B short — SumitGup · 2026-10-12
- DiffusionBear: MLX-powered macOS app runs FLUX.2 & Krea 2 on 16GB Macs — Puzzleheaded_Note739 · 2026-10-12