PatternEval and PatternRL: Aligning response patterns in hybrid-thinking MLLMs

YouJiacheng · x · 2026-08-26

The original post introduces PatternEval, a 2,415-prompt benchmark for detecting CoT leakage, repetition, and contradictions. PatternRL adds penalties during RL to improve consistency. The quote questions why this general research direction is confined to 'multimodal'.

Original post →

More from Multimodal

Multimodal channel →