Hojo-ASR-Multi-V1 model card: Qwen3 decoder architecture, Apache 2.0 open weights on HF
rohanpaul_ai · x · 2026-09-04
The Hugging Face model card for Hojo-ASR-Multi-V1 reveals technical details and usage: an Encoder-Adapter-LLM architecture built on the Qwen3 LLM decoder with multi-frame acoustic fusion, optimized via multi-stage modular training and RL, performing well in noisy and conversational conditions across major languages. Apache 2.0 licensed, installable via the hojo-asr pip package for quick inference.
Related event: Hojo-ASR Tops Open Multilingual ASR Leaderboard with 3.54% WER(2 posts)→
More from Multimodal
- OmniVoice open-sources diffusion-based TTS with zero-shot voice cloning in 600+ languages — tom_doerr · 2026-09-04
- sanoTTS: 294k-param TTS stack runs on a $3 microcontroller, mid model beats rivals 3-10x its size — Affectionate_Hat_585 · 2026-09-04
- User claims GPT-6 Astra is so good at 3D modeling he opens Blender daily — flavioAd · 2026-09-04
- Open-source ComfyUI toolkit turns an idea into a finished MiniMax Music 3 song with enhanced audio — Vivid_Promise1700 · 2026-09-04
- Makepad flow: a Rust single-executable ComfyUI alternative built in one day — anselm · 2026-09-04
- RTX 4060 Ti 16GB runs MiniMax video model locally: 768p in 3 minutes — aziib · 2026-09-04