RSIAgent beats GPT-6 Astra on OSWorld 2.0 without updating model weights
jiqizhixin · x · 2026-09-23
Aether AI presents RSIAgent, claiming 78.98% on OSWorld 2.0 vs GPT-6 Astra's 72.60% and 84.82% on Agents' Last Exam vs 82.26% — with zero model parameter updates, on open-source Kimi-K3 and GLM-5.3.
The idea is to scale experience instead of the model:
- The agent decides what to learn, explores the environment, and verifies results
- Stable action-condition-outcome causal relationships are deposited into Memory for reuse
- A recursive loop: curriculum generation → execution → verification → memory evolution
- New environments (enterprise software, private workflows) require no data collection, fine-tuning, or RL
More from coding & agent
- Microsoft Ships Foundry Dev Pack: One-Command Install for Hosted Agent Development — lee_stott · 2026-09-23
- Rabbit R1 Turned Into a Pocket Portal for Self-Hosted Local AI Models via rabbitOS 3 — SimonBalmain · 2026-09-23
- Opus 5.5 Read 347 Ranking Pages, Found 131 Untouched SEO Angles for $6.35 — alex_verem · 2026-09-23
- Opus 5.5 Cut Prices 40% and Within a Day It Was Porting C to Rust: 8 Use Cases — alex_verem · 2026-09-23
- CodeMidas: Turning Raw Source Code into Executable RL Environments for Coding Agents — burny_tech · 2026-09-23
- Supply-chain attack hits MemOS: malicious Golang binaries hidden in PyPI and npm packages — cyb3rops · 2026-09-23