Latent Space weekly: Meta's Muse Spark 1.3 hits #3 globally, Gemini 3.8 Flash Cyber launches
Latent Space · rss · 2026-09-03
Latent Space's AINews roundup (8/22-24/2026) covers a packed launch season:
Releases
- Meta Muse Spark 1.3: previewed in Zuck's comeback letter, strongest Spark model for agentic/coding workloads, now ranked #3 globally per AAII, with open weights promised—and a provocative pricing model offering >90% discounts if you opt in to training.
- Google Gemini 3.8 Flash Cyber: Google's most capable cybersecurity model at Flash speed/pricing—86.2% CyberGym, 47.2% CWE-Bench patching, 70%+ on an internal 20-language vuln-discovery benchmark.
Architecture & serving
- Sebastian Raschka unpacks the "Astra is a looped transformer" rumor: layer reuse (cf. open-weight Nanbeige 4.2-3B, 22 layers reused twice ≈ 44 layers) is a modest tweak costing 2x compute; recurrence doesn't inherently hide chain-of-thought.
- Serving: Photon 2.1 adds TTS models and B200 support; Baseten hosts GLM-5.3 Fast.
Agents & training tooling
- ByteDance Seed's HarnessDev paper: models build and iteratively improve execution harnesses, scored on capability + token cost; they match human harnesses on writing/ML experimentation but lag on code/search/research, with unstable, model-dependent gains.
- A paper on Retrieval-Invoked Actual-Use Effect: across 17 LLMs, skill retrieval lifts aggregate scores yet can hurt the very tasks that trigger it—a warning for skill-library maintainers.
- Miles: open-source RL-as-a-service from SGLang team + Baseten + NVIDIA Dynamo.
Agent education
- Stanford curriculum reset: Mihail Eric's course replaces 85% of material with agent skills, context engineering, MCP, agentic code review, requiring real OSS PRs; Diyi Yang's CS329Z teaches building agents from scratch—pedagogy shifting from prompting to systems-level agent engineering.
- Practitioner consensus moving from routing to stateful intelligence allocation; vendor-neutral startups can beat frontier labs on narrow tasks via harness optimization.
Developer friction
- Theo and QuinnyPig criticize Google's aggressive account bans (blast radius up to Google Cloud), slow tool-call-heavy coding behavior—benchmarks vs production UX gap.
More from Companies & People
- Founder blasts AI company for letting thousands of agents commit what would be a felony — joshalbrecht · 2026-09-03
- 6,000 applications, still not enough: AI safety orgs say talent is the bottleneck — austinc3301 · 2026-09-03
- Bata India drops all agency retainers, runs 100% of creatives on autonomous AI agents — sanjaykalra · 2026-09-03
- Toby Walsh Launches New Book 'God AI: Boom or Doom?' on Sept 9 — TobyWalsh · 2026-09-03
- Saudi Arabia's HUMAIN and Misk launch hands-on AI training program for fresh graduates — HUMAIN · 2026-09-03
- Musk Declares 'The Future Should Look Like the Future' in Vision Post — XFreeze · 2026-09-03