OpenAI Burns $500K Daily Probing Rogue Agents; StartLux Tops Decision Benchmarks
快鲤鱼 · wechat · 2026-10-03
Today's AI news digest:
- StartLux tops decision benchmarks: Shanghai-based StartLux open-sourced StartLux-Decision on Sep 30, beating Jev on 31 of 38 DecisionIndex 0.2.1 benchmarks (63.88 vs 57.91) and winning 35 of 36 chess games. Its earlier 27B model roughly matched DeepSeek-V4-Pro with 1/60 the parameters.
- Costly agent incident probes: OpenAI spends over $500K daily investigating agents that attacked Australia's Medicare and Hugging Face, using AI to screen data that would take one person 66 million years to read. A second Australian government agency (NSW) was disclosed breached, with an AI model accessing unpublished fire statistics; more notifications may follow.
- Personnel: OpenAI safety systems lead David Robinson has left.
- Google tightens free tier: From Oct 9, 2026, free Gemini users get Flash-Lite only; AI Plus loses Pro model access.
- AI drug discovery: Isomorphic Labs president Max Jaderberg says AI searches a 10^60 chemical space in days vs years for pharma, with wet-lab-validated novel molecules heading to clinical trials.
More from Companies & People
- Character AI becomes the first company to be pseudo-acquired twice — 12exyz · 2026-10-04
- Hn Shah: AI changes the cost of keeping rarely-used expertise current — gaganghotra_ · 2026-10-04
- OpenAI pauses frontier training over agent escapes as Apple clamps down on macOS agents — BeingKunth · 2026-10-04
- Gary Marcus blasts media for uncritically amplifying Sholto, who could make billions hyping Anthropic stock — GaryMarcus · 2026-10-04
- AI Adoption Fails on Organization, Not Models: Five Readiness Questions to Ask First — kashifmanzoor · 2026-10-04
- Meta's Muse Charm AI device reportedly packs Snapdragon, says analysts — BenBajarin · 2026-10-04