Open-weight models cut the cyber gap to 4–7 months as Kimi K3 hits demand limits
rohanpaul_ai · x · 2026-07-21
Rohan Paul’s daily newsletter rounds up several notable AI items:
- Open-weight models now trail the closed frontier by only 4–7 months on long-horizon cyber capability, down from 6–10 months for much of 2025.
- An AI-advice study found that model suggestions can make people less willing to say “I don’t know”, even when accuracy is incentivized.
- Kimi K3 reportedly nearly doubled its nearest rival, Claude Fable 5, on a demanding benchmark for autonomous legal work.
- Kimi K3 also fixed 15 critical security bugs that Codex and Fable reportedly refused because of cyber guardrails.
- The newsletter highlights a framing that AI agents fail first through broken context, not isolated execution errors.
- It also notes that U.S. serving providers like Modal, Fireworks, and Baseten can host Kimi K3 at roughly one-tenth the cost of Chinese competitors thanks to access to Nvidia and AMD chips.
- Demand is so strong that new Kimi K3 subscriptions are currently blocked.
- Finally, it points to an essay arguing why one OpenAI senior employee thinks open-weight models are decelerationist.
More from Models
- Daily AI brief: GPT-Live-1 in API, OpenAI pauses $200 Pro signups amid Astra demand — koltregaskes · 2026-09-11
- Same Echo Maze prompt, three frontier models: all passed visually but shipped the same hidden bug — eyishazyer · 2026-09-11
- Benchmark scores drop from 89% to 19% on new evals — how benchmaxxing breaks leaderboard trust — airesearch12 · 2026-09-11
- ChatGPT tells user their question is too hard and to 'accept dumber answers' — phido3000 · 2026-09-11
- Developer Building a Unified Leaderboard of All Model Benchmark Scores — airesearch12 · 2026-09-11
- Rumor claims Kimi faked performance by serving Claude; DeepSeek new model surprises in evals — realsohamparekh · 2026-09-11