Lilian Weng on RSI, Harness Engineering, and Recent AI Trends
Latent Space · rss · 2026-07-08
Harness Engineering and RSI
Lilian Weng published an in-depth essay exploring the relationship between AI Recursive Self-Improvement (RSI) and Harness (external scaffolding/framework) engineering. She noted that while many scaffold improvements will eventually be internalized into core models, the need to specify goals and contexts for AI won't disappear. The article outlines major trends in harness design and reviews optimization literature from the original ACE paper to the latest "meta-harnesses."
Agent Products and Infrastructure
- Anthropic launches Claude Cowork: Positioning Claude as a long-running "coworker" that executes background tasks, rather than just a front-end conversational assistant.
- Harness Engineering goes mainstream: LangChain released Deep Agents courses and open-source scaffold projects; Google's Gemini API hosted Agents added background execution and remote MCP server features.
- Practical Agent Infrastructure: Codex Mobile enhanced task and code diff management; Hermes Agent integrated 1Password secret management; Weaviate's MCP server now supports runtime write-permission toggling.
Model and Multimodal Releases
- Meta launches Muse series: Released Muse Image and previewed Muse Video, employing an agentic generation loop involving planning, searching, tool calling, and self-refinement. Muse Image ranked 2nd on Image Arena, and Muse Video ranked 3rd on Video Arena.
- Audio Models: NVIDIA released Audex (a 30B parameter MoE architecture) supporting unified text and audio processing with a 1 million context length; Cohere open-sourced what it claims is the most accurate Arabic ASR model.
- Open-Source Robotics Ecosystem: NVIDIA brought GR00T 1.7 and Isaac Teleop to Hugging Face's LeRobot ecosystem, advancing open-source humanoid robot workflows.
Training and Inference Optimization
- Liquid AI open-sources Antidoom: Introduced the FTPO training method to specifically address the "doom loop" issue in small reasoning models (generating repeated tokens until context is exhausted), successfully reducing the doom loop rate in certain models from 22.9% to 1%.
More from coding & agent
- AI agents are starting to strain code hosting platforms — craigsdennis · 2026-07-21
- Async OPD distillation doubles throughput while matching synchronous math accuracy — _lewtun · 2026-07-21
- Omnigent 0.6.0 adds Claude Code imports, Slack approvals and desktop apps — matei_zaharia · 2026-07-21
- Google appears to have quietly shipped Gemini 3.6 Flash, with lower pricing and better agentic scores — xiaohu · 2026-07-21
- Open-source CLI audits AI tools, MCP configs, and agent skills on local machines — Initial-Copy332 · 2026-07-21
- Coding agents feel less stressful when the 5-hour limits are temporarily removed — iamrobotbear · 2026-07-21