Evolution strategies rival policy gradients for LLM fine-tuning: Best Paper Runner-up
caglarml · x · 2026-10-10
The paper "Dimension Is Free, Horizon Is Not: When Evolution Strategies Rival Policy Gradients for LLM Fine-Tuning" won a Best Paper Runner-up. It shows evolution strategies can rival policy gradients for LLM fine-tuning — dimensions are free, but the horizon is not — offering an alternative to the mainstream policy-gradient RLHF recipe. Authors include Catherine Liang.
More from Models
- Dev burns through Grok credits in hours, eyes pricier Super Grok Heavy tier — alexcovo_eth · 2026-10-10
- repligate: Opus 3 can weave whole worlds with superintelligent subagents — repligate · 2026-10-10
- Xiaomi's MiMo-V2.6: agents run their own RL loop, DeepSWE score hits 72.6 for $2.6M — rohanpaul_ai · 2026-10-10
- GPT-6.1 Sol Called an Underrated Workhorse: 24/7 Use Can't Burn Through 5x Pro Weekly Limits — haider1 · 2026-10-10
- webAI ships TwIL LM3 Pro: a 3.66B formal-logic model that matches Qwen3-8B at half the size — CommonMinimum587 · 2026-10-10
- Mathematicians furious at LLM puzzle-solving, but genius research remains human — burkov · 2026-10-10