Multi-agent evals still undecided, but colocated async RL training is catching on

stochasticchasm · x · 2026-09-11

A technical take: multi-agent evals are the most interesting part of a recent dense paper — the optimal multi-agent harness design is still very much undecided, but training for it seems straightforwardly beneficial and scales better than single-agent. The author also notes colocated async RL gaining popularity, with k3 doing the same.

Original post →

More from Models

Models channel →