Daily AI brief: OpenAI's cyber-critical Astra, Grok 4.7 dated, Fable 5.1 live
testingcatalog · x · 2026-09-02
Sept 2 roundup of dense AI news:
- OpenAI: Official "Path to Astra" post; Astra is the first model designated Critical for cybersecurity under the Preparedness Framework, scored 100% on ExploitBench and found 2 zero-days in evals; new codenames vega-alpha and ultima-alpha in testing.
- Google: Gemini 3.8 Flash already answering; WSJ says Google engineers preferred it to Opus for coding. Agentic video understanding rolls out on 3.7/3.6/3.5 Flash — up to 88% fewer tokens, 66% cheaper, 7% better accuracy on long video.
- Anthropic: Claude Fable 5.1 live — 52.6% Terminal-Bench-Science 0.1, 55.8% vs 42.0% on Terminal-Bench 4.0; same price, 75% cheaper cache reads, up to 45% cheaper on heavy agent runs.
- Meta: Muse Voice Transcribe launched — first real-time audio perception model, SOTA streaming STT with native diarization (20+ speakers).
- xAI: Elon says Grok 4.7 arrives in 10 days (Sept 12).
- Alibaba: Qwen3.8-Max-0902 live at $2/$6 per 1M tokens, #1 on Code Arena: WebDev at 1691, 3 pts above Claude Opus 5 (Max).
- World Labs: Fei-Fei Li's lab shipped Atlas, an omni world model — few photos to camera control, up to 1 min of 1440p video plus 3D reconstruction; early access only.
More from Models
- User Test: Fable 5.1 Outperforms GPT 5.6 on Complex Profiling Task — remilouf · 2026-09-02
- RTX 5090 laptop owner asks: best abliterated Qwen 3.6 vs 3.8 27B quants for 24GB VRAM — ThomasAger · 2026-09-02
- Merge Gateway Adds Claude Fable 5.1 with 75% Cheaper Cache Reads — shensi · 2026-09-02
- OpenAI's 'Astra' rumored to launch this week with new recurrent depth architecture — mark_k · 2026-09-02
- iFlytek's Spark-X2.5-4B trends on Hugging Face — XHToken · 2026-09-02
- OpenAI Model Tested: Provides Crime Plans But They Are Nonsense — xeophon · 2026-09-02