d1-omni-600M sorts your voice notes in ~160ms, fully in-browser on WebGPU
iamrobotbear · x · 2026-10-09
d1-omni-600M, LiquidAI's experimental model combining LFM2.5-Encoder-350M with vision and audio encoders, powers a demo that files spoken voice notes into reminders, lists, messages, travel, questions or music in 160ms, running entirely in-browser on WebGPU.
No transcript is produced and nothing is generated — just typed JSON — and your voice never leaves the tab. The model leads Liquid's text benchmark comparison in toxicity detection and paraphrase identification, and targets voice-command routing, on-device moderation, and intent classification. A live demo and ONNX builds (text + vision + audio) are available.
More from Infra
- Bain projects 38.6M GPU and custom silicon shipments by 2030, 10x 2023's 3.9M — Beth_Kindig · 2026-10-09
- AWS reference architecture: multi-team GPU cluster sharing on SageMaker HyperPod — AWS ML Blog · 2026-10-09
- Mistral slammed for training open models on datacenters powered ~70% by coal — wavefnx · 2026-10-09
- Why do we resend the whole conversation every turn? Server-side KV slots proposal sparks debate — Vasili_Sk · 2026-10-09
- NVIDIA's NeMo-DCR cuts 1T-model weight sync from 87.5 min to 150s, 12-40x faster checkpoint transfer — dair_ai · 2026-10-09
- Zyphra Speeds MoE Expert Routing Communication 2.63x on AMD MI300X GPUs — QuentinAnthon15 · 2026-10-09