OpenVINO-Powered Laya Hits 40ms per Question on CPU, 3.4x Faster Than PyTorch

simpleuserhere · reddit · 2026-09-25

Developer rupeshs added OpenVINO backend support to Laya, a lightweight LLM runtime, achieving 40ms per question on CPU — 3.4x faster than PyTorch. Code is open-sourced (laya-openvino), with a Flappy Bird demo showing CPU-only low-latency inference.

Original post →

More from Infra

Infra channel →