Lilian Weng on RSI, Harness Engineering, and Recent AI Trends
Latent Space · rss · 2026-07-08
Harness Engineering and RSI
Lilian Weng published an in-depth essay exploring the relationship between AI Recursive Self-Improvement (RSI) and Harness (external scaffolding/framework) engineering. She noted that while many scaffold improvements will eventually be internalized into core models, the need to specify goals and contexts for AI won't disappear. The article outlines major trends in harness design and reviews optimization literature from the original ACE paper to the latest "meta-harnesses."
Agent Products and Infrastructure
- Anthropic launches Claude Cowork: Positioning Claude as a long-running "coworker" that executes background tasks, rather than just a front-end conversational assistant.
- Harness Engineering goes mainstream: LangChain released Deep Agents courses and open-source scaffold projects; Google's Gemini API hosted Agents added background execution and remote MCP server features.
- Practical Agent Infrastructure: Codex Mobile enhanced task and code diff management; Hermes Agent integrated 1Password secret management; Weaviate's MCP server now supports runtime write-permission toggling.
Model and Multimodal Releases
- Meta launches Muse series: Released Muse Image and previewed Muse Video, employing an agentic generation loop involving planning, searching, tool calling, and self-refinement. Muse Image ranked 2nd on Image Arena, and Muse Video ranked 3rd on Video Arena.
- Audio Models: NVIDIA released Audex (a 30B parameter MoE architecture) supporting unified text and audio processing with a 1 million context length; Cohere open-sourced what it claims is the most accurate Arabic ASR model.
- Open-Source Robotics Ecosystem: NVIDIA brought GR00T 1.7 and Isaac Teleop to Hugging Face's LeRobot ecosystem, advancing open-source humanoid robot workflows.
Training and Inference Optimization
- Liquid AI open-sources Antidoom: Introduced the FTPO training method to specifically address the "doom loop" issue in small reasoning models (generating repeated tokens until context is exhausted), successfully reducing the doom loop rate in certain models from 22.9% to 1%.
More from coding & agent
- Cheaper OpenAI Agents API alternative: sandbox service undercutting E2B by 46% — airesearch12 · 2026-09-11
- His agent kill switch ran for months before he found it was wired to nothing — AnvilandCode · 2026-09-11
- Kernel's Browser Agents Can Now Pay Online Using Aliases, Never Touching Card Data — jeff_weinstein · 2026-09-11
- OpenAI opens up agent sandboxes: BYO or pick from Cloudflare, E2B, Modal, Vercel and more — threepointone · 2026-09-11
- SocialCrawl MCP lets agents search Reddit, YouTube, TikTok, X with one API key — dooddyman · 2026-09-11
- Astra builds a surprisingly polished Catan game in three.js, reusing past UI and 3D assets — FinanceYF5 · 2026-09-11