Audio8 ASR Infinite: streaming ASR with constant memory and latency for 24/7 use
_akhaliq · x · 2026-09-30
akhaliq highlights Audio8 ASR Infinite, a streaming speech recognition model:
- Native streaming architecture decodes 12.5 times per second
- A rolling KV Cache keeps memory and latency constant, enabling true 24/7 operation
- One text token per clock step (12.5/8.3/6.25 decisions/sec) balances perception granularity and resource cost
- An ML intern at HuggingChat set up a Gradio demo for trying it out, and the model weights are available
More from Models
- Meituan LongCat unveils DeepResearch multi-agent system scoring 55.25 on DeepResearchBench — Meituan LongCat Team · 2026-09-30
- Google study: GPT-5.5 flags planted negative results in only 2/200 reports unless told 'be honest' — google · 2026-09-30
- PlaylistEval: frontier video-language judges hit only 75.4% accuracy on 100-hour playlists — Shayekh Bin Islam · 2026-09-30
- Google's TabFM: 400M-param tabular foundation model beats tuned AutoML zero-shot — google · 2026-09-30
- Codex tip: ride out your last $200/mo month before upgrading to the $500 plan — ctjlewis · 2026-09-30
- "Everyday task" benchmarks are meaningless — can we get real data on what models actually build? — Benhamish-WH-Allen · 2026-09-30