DeepSeek V4F-0731 Underperforms on EQ-Bench v4: Does Heavy RL Hurt Model Personality?

xeophon · x · 2026-08-01

TNG Tech evaluated DeepSeek's V4F-0731 model on the new EQ-Bench v4 benchmark, yielding okayish but unremarkable results.

This raises an interesting systemic question regarding the model's post-training: does heavy reinforcement learning (RL) negatively impact a model's personality? The findings suggest that aggressive RL alignment could potentially be detrimental to the model's performance on specific qualitative benchmarks.

Original post →

More from Models

Models channel →