Tencent Benchmarks Hybrid-Thinking MLLMs for Response Alignment
tencent · hf · 2026-08-24
Tencent addresses response-pattern misalignment in hybrid-thinking MLLMs between thinking and non-thinking modes. The work introduces a diagnostic benchmark and applies pattern-specific reinforcement learning penalties to align model behaviors.
More from Research
- CWoMP accepted to EMNLP 2026: Interpretable retrieval-based glossing for endangered languages — fredahshi · 2026-08-24
- SemiAnalysis Open Sources $3M AgentX Benchmark for Agentic Coding Workloads — AccBalanced · 2026-08-24
- Vinci2 Agent Outperforms GPT-5-mini in Proactive Assistance Benchmark — jiqizhixin · 2026-08-24
- OpenAI hiring for Economics of Transformative AI, MATS fellowship applications open — Astral Codex Ten · 2026-08-24
- New "Discovery Episode" Framework Measures AI Scientists by Real Research Cycles — 量子位 · 2026-08-24
- AI Claims Breakthrough on Erdős Problem Transcendence — inductionheads · 2026-08-24