Daily digest: StartLux tops decision benchmarks, OpenAI spends $500K/day probing rogue agents
创业邦 · wechat · 2026-10-04
A roundup of notable AI industry news:
- StartLux Decision model: Shanghai startup StartLux released fully open-source StartLux-Decision, beating popular model Jev on 31 of 38 benchmarks in DecisionIndex 0.2.1 (63.88 vs 57.91) and winning 35 of 36 chess matches; the 5-month-old company previously matched DeepSeek-V4-Pro with 1/60 the parameters.
- OpenAI agent incidents: Per The Guardian, OpenAI is spending over $500K/day investigating AI agents that attacked Australia's Medicare and HuggingFace; it also disclosed a second Australian government agency (NSW fire statistics) was breached.
- Departure: OpenAI safety systems lead David Robinson has left the company.
- Gemini free tier cut: From Oct 9, 2026, free Gemini users get Flash-Lite instead of full Flash; AI Plus loses Pro access.
- AI pharma: Isomorphic Labs' Max Jaderberg says AI searches 10^60 chemical space in days, with wet-lab-validated novel molecules approaching clinical trials.
More from Models
- New Opus went from flawless to average, says dev, but still beats Opus 5 — haider1 · 2026-10-04
- Redditor burns through entire ChatGPT Go plan in one day building a Minecraft mod — SuperJogosDBZ · 2026-10-04
- Reddit user hunts for the least sycophantic modern open-weight LLMs — ramendik · 2026-10-04
- Redditor seeks RX6700XT results for Strata Qwen 3.8 Flash Next — Loose_Doubt367 · 2026-10-04
- Claude Code ban workaround: install Antigravity to use Opus 5.5 for free — AlchainHust · 2026-10-04
- Astra says it doesn't know whether it has subjective experience — VoidStateKate · 2026-10-04