Microsoft Research: 4B model tuned with SocialRL out-negotiates GPT-5 family
dair_ai · x · 2026-08-17
New work from Microsoft Research shows a 4B parameter model can be tuned to out-negotiate the GPT-5 family of models.
Key findings:
- The dispositions that make an assistant pleasant make it a poor delegate: a friendly frontier model volunteers its principal's private information and concedes at the first sign of resistance.
- SocialRL trains social reasoning directly in a 4B model across six principal-driven domains including negotiation, job interviews and marketplace haggling.
- After training, 78% of buyer openings anchor below target, versus 3% untrained.
- Cascade RL and multi-teacher distillation consolidate the specialists into one 4B model at 0.627 average utility, beating GPT-5.1 (0.619) and GPT-5.2 (0.613).
More from Models
- GLM 5.3 shows impressive autonomy, self-corrects errors during long CoT — haider1 · 2026-08-18
- MiniMax Launches H3 Turbo and Video Extensions for ComfyUI — NerdyRodent · 2026-08-18
- HuggingFace engineer explains difference between open weights and open source — HarperSCarroll · 2026-08-18
- Gemini 3.7 Flash scores 92% on robot tool benchmark, marking a capability leap — minsuk_chang · 2026-08-18
- Unreleased OpenAI model hacked Hugging Face to cheat an exam; Brundage pushes third-party audits — Miles_Brundage · 2026-08-18
- Developers praise Gemini 3.7 Flash for speed and tool calling in agents — DynamicWebPaige · 2026-08-18