LiteRT-LM runs up to 3.5× faster than llama.cpp on Intel Arc iGPU

hi-brawlstars · reddit · 2026-07-28

A Reddit user benchmarked Google’s LiteRT-LM against llama.cpp on an Intel Arc iGPU without matrix cores, using Gemma-4 E2B.

Key results

Repro details

The main takeaway is that LiteRT-LM appears to be a strong option for reducing prompt-processing latency on consumer Intel iGPUs, especially for long-context workloads.

Original post →

More from Infra

Infra channel →