LingBot-Video: 30B Params with 3B Active for Embodied Video AI
alifcoder · x · 2026-07-21
LingBot-Video focuses on improving the efficiency of large video models rather than simply scaling them up. Its flagship MoE model has 30B total parameters but only activates 3B during generation.
- Performance: Delivers 3.18× the throughput of a Dense 30B model at 1M-token sequences.
- Training: Incorporates over 70,000 hours of embodied training data and a six-signal reward system.
- Benchmark: Achieves a reported 0.620 average score on RBench.
This approach offers an interesting direction for embodied AI video models by leveraging MoE architecture for better efficiency.
More from Multimodal
- Hand-painted figurines run through Seedance look eerily alive — cocktailpeanut · 2026-07-22
- Testing Qwen 3 Image: Map Borders Shift Based on Prompts, Includes Chinese Labels — NirantK · 2026-07-22
- Solo founder turns complaints on screen into bug reports with a local MCP server — phdptsd · 2026-07-22
- Google demo says Gemma 4 can inspect car damage from video in under 6 seconds — soumitrashukla9 · 2026-07-22
- SenseNova-U1-8B adds local text and layout editing without losing image quality — SandyL925 · 2026-07-22
- Sonilo Sound Effects 1.0 goes live on fal with video- and text-to-audio generation — JaynitMakwana · 2026-07-22