Liquid AI's 600M Decision Model Runs in the Browser via WebGPU at 45ms per Query

FinancialAd1961 · reddit · 2026-10-08

A developer ported Liquid AI's newly released d1-omni-600M decision model to runntime, the WebGPU inference library he is building. It is plain TypeScript on top of TypeGPU, with no WASM and virtually no export step: the model is written directly from core ops (matmul, attention, norms, a few elementwise bits) and weights load straight from HF safetensors.

The demo moderates a fake comment feed live, asking four questions per comment (toxic? spam? asking something? overall tone?) and removing toxic and spam ones. Performance is 180ms per comment for all four questions, 45ms per question, with the author expecting much better after tuning the engine for this model.

He also tried making it play Snake, which "did not go well." The d1 port is not in the npm release yet; the rest of runntime (detection, segmentation, speech-to-text, embeddings and more) is, with docs and live demos at docs.swmansion.com/runntime.

Original post →

More from Infra

Infra channel →