AI Weekly: Execution, Quantization, and World Models
Latent Space · rss · 2026-07-15
This AINews/Latent Space recap is packed with AI industry progress from July 13–14:
- Coding agents / harness: OpenAI's Codex and ChatGPT Work usage is surging; JetBrains made Codex its recommended agent; authors discussed the importance of harness quality, observability, and eval environments.
- Local inference & quantization: PrismML's Bonsai 27B, Tencent Hunyuan's 1-bit/4-bit Hy3, and Gemma series NVFP4 show that aggressive compression is bringing stronger models to consumer devices and single-GPU setups.
- Multimodality & world models: Projects like MOSS-VL-Realtime, OmniAgent, and LingBot-World are pushing continuous video understanding, active frame fetching, and real-time interactive generation.
- Eval & research infrastructure: Perplexity open-sourced WANDR, emphasizing realistic evaluations for dynamic web research; it also noted more realistic and adversarial agent eval designs.
- Physical AI: Sakana's Smart Cellular Bricks demonstrated distributed self-repairing/reconfiguring systems.
The overarching theme: the AI ecosystem is shifting from "chatting" to "execution," while quantization, local deployment, evaluation, and multimodal systems are all leveling up simultaneously.
More from coding & agent
- Hermes Agent adds built-in Word, Excel, PDF and PowerPoint support — Teknium · 2026-07-21
- Marker 2 claims better quality than MinerU and docling while hitting 27 pages/sec — VikParuchuri · 2026-07-21
- Cursor publishes a strong deep dive on agent swarms for coding workflows — Sam_Witteveen · 2026-07-21
- A deleted NanoGPT PR still propagated into later world-record submissions — yacineMTB · 2026-07-21
- Soft Clamp cuts tool-call overuse in multi-teacher distillation, from 13.7% to 9.0% — antgroup · 2026-07-21
- Agent harness memory loss and compaction are still a major usability problem — adityaag · 2026-07-21