Tencent Releases Embodied Multimodal Model Hy-Embodied-VLM-1.0
_akhaliq · x · 2026-07-15
Tencent has released Hy-Embodied-VLM-1.0 on Hugging Face, an efficient MoE vision-language model designed for embodied agents.
Key details include:
- Activates only 3B parameters per token
- Achieves SOTA on 19 out of 38 benchmarks
- Primarily targets embodied agents scenarios
This release leans heavily toward multimodal applications while showcasing significant research attributes.
Related event: Tencent Hunyuan Open-Sources Two Embodied AI Foundation Models(3 posts)→
More from Multimodal
- Seedance 2.0 demo turns ketchup on spaghetti in Rome into an AI reaction meme — azed_ai · 2026-07-21
- A reusable “Lunar Eclipse Dreamscape” prompt comes with multiple example renders — LudovicCreator · 2026-07-21
- Midjourney 8.2 preview shows a double-exposure prompt with strong style control — michaelrabone · 2026-07-21
- Travel MCP Server adds flight, hotel, weather and budget tools for agents — modelcontextprotocol · 2026-07-21
- Douyin Video Analysis MCP turns share links into structured video summaries — modelcontextprotocol · 2026-07-21
- Synthesia launches Dubbing 2.0 with 130+ languages and lip-sync video translation — synthesiaIO · 2026-07-21