Multi-agent evals still undecided, but colocated async RL training is catching on
stochasticchasm · x · 2026-09-11
A technical take: multi-agent evals are the most interesting part of a recent dense paper — the optimal multi-agent harness design is still very much undecided, but training for it seems straightforwardly beneficial and scales better than single-agent. The author also notes colocated async RL gaining popularity, with k3 doing the same.
More from Models
- GPT-6 Astra drives a robot arm on first try via physical ICL; Ken Goldberg touts Agentic Robotics — zhaoran_wang · 2026-09-11
- OpenAI's internal model claims a Navier-Stokes millennium prize proof, says analyst — QuintinPope5 · 2026-09-11
- DeepSeek 4.1 flash reportedly uses large ngram embeddings, echoing Qwen4 architecture — ccerrato147 · 2026-09-11
- ValsAI launches RSI Index, first third-party benchmark measuring how close AI is to self-improvement — JenniferHli · 2026-09-11
- Assistant Benchmark goes live: 61 assistants scored across 15 real-use dimensions — Scobleizer · 2026-09-11
- Devin's New Model Verdict: Not a Benchmaxxer, a 'Killer Execution Model' at $20/Month — brandon_galang · 2026-09-11