d1-omni-600M sorts your voice notes in ~160ms, fully in-browser on WebGPU

iamrobotbear · x · 2026-10-09

d1-omni-600M, LiquidAI's experimental model combining LFM2.5-Encoder-350M with vision and audio encoders, powers a demo that files spoken voice notes into reminders, lists, messages, travel, questions or music in 160ms, running entirely in-browser on WebGPU.

No transcript is produced and nothing is generated — just typed JSON — and your voice never leaves the tab. The model leads Liquid's text benchmark comparison in toxicity detection and paraphrase identification, and targets voice-command routing, on-device moderation, and intent classification. A live demo and ONNX builds (text + vision + audio) are available.

Original post →

More from Infra

Infra channel →