Hojo-ASR-Multi-V1 model card: Qwen3 decoder architecture, Apache 2.0 open weights on HF

rohanpaul_ai · x · 2026-09-04

The Hugging Face model card for Hojo-ASR-Multi-V1 reveals technical details and usage: an Encoder-Adapter-LLM architecture built on the Qwen3 LLM decoder with multi-frame acoustic fusion, optimized via multi-stage modular training and RL, performing well in noisy and conversational conditions across major languages. Apache 2.0 licensed, installable via the hojo-asr pip package for quick inference.

Related event: Hojo-ASR Tops Open Multilingual ASR Leaderboard with 3.54% WER(2 posts)→

Original post →

More from Multimodal

Multimodal channel →