Shifts in Models and Agents Over the Past Two Months
downingARK · x · 2026-07-13
This post summarizes the intense updates from major US AI companies in the model and agent layers over the past two months, offering a snapshot of the current landscape.
Model Layer: Leaderboard Reshuffle
- Google launched Gemini 3.5 Flash at Google I/O, highlighting its high cost-effectiveness relative to intelligence.
- Meta then released models with similar or slightly higher intelligence at lower costs; Grok 4.5 also stands out in its "cost/intelligence ratio."
- Result: According to Artificial Analysis benchmarks, Google was briefly pushed to fifth place.
- In terms of frontier intelligence, Anthropic and OpenAI are still considered tier-one, but at higher prices.
- The author also notes that OpenAI's 5.6 Sol offers similar intelligence to fable at less than half the price, making it a better deal.
Agent Layer: Accelerating Product Cadence
- Claude Cowork is now available on mobile and web.
- OpenAI launched ChatGPT Work powered by Codex.
- The new ChatGPT Voice can delegate tasks to frontier models for advanced reasoning and tool usage, delivering a "generational leap" in experience.
More from coding & agent
- Soft Clamp cuts tool-call overuse in multi-teacher distillation, from 13.7% to 9.0% — antgroup · 2026-07-21
- Agent harness memory loss and compaction are still a major usability problem — adityaag · 2026-07-21
- SpecJudge runs locally on Ollama to pick the right-sized AI model for your project — jokiruiz · 2026-07-21
- A developer maps out six design rules for CLIs that humans and AI agents can both use — yujiezha · 2026-07-21
- A coding-agent skill that forces ADHD-friendly, answer-first output — ayghri · 2026-07-21
- A set of agent skills for CAD, robotics, and hardware design — earthtojake · 2026-07-21