Gemini 4 Argon launches with 1M-token output, tops 13 of 19 benchmarks at $2/$10 intro pricing
Latent Space · rss · 2026-10-01
Gemini 4 Argon: Google Returns to the Frontier
- Launch & access: GDM's first frontier-class model since 3.1 Pro is SOTA on 13 of 19 credible benchmarks, initially gated to government users and trusted cyber defenders via the Fairwind Program.
- Output & pricing: Industry-first 1M-token output via new Long Decode Continuation API; $4/$20 per 1M tokens with an open-ended 50% intro discount ($2/$10), 95% cached-input discount.
- Internal deployments: Google says Argon agents freed 300+ TiB of data-center memory and are migrating 800K lines of C/C++ kernel code to Rust; a video decoder rewrite made a Rust port 2.7x faster. The team also claims internal agent loops helped complete the CK conjecture.
- Third-party evals: Artificial Analysis Intelligence Index 53, matching GPT-6 Astra. At intro pricing it costs $1.99/task vs Astra's $3.26 — savings come from price, not efficiency (62K vs 27K output tokens per task). Hallucination rate 15% vs Astra's 51%, but accuracy 50% vs 63%. Terminal-Bench 4 jumped from 19% to 57.6%; perfect IOI 2024–2026. Skeptics note a 19.6% on Harvey's legal benchmark (trailing Muse Spark 1.2's 25.42%) and possible preference-data benchmaxxing.
OpenAI: GPT-6.1 Sol and DevDay
- GPT-6.1 Sol is the new MathArena #1 and ranks #3 on WebDev at $2/$10; $0.72 per task vs Astra's $3.26. A Luna image-encoding bug fix added 1 Intelligence Index point.
- Up to 300 tok/s: generation is 8x faster but end-to-end agent tasks only 2–4x (tool latency dominates); computer-use gains are largest.
- DevDay brought dots (persistent agents with their own cloud computers), a Decisions API, and computer use; ChatGPT Sites can host MCP servers as installable plugins. Users report cut usage limits.
Other Releases
- Perplexity open-sourced pplx-embed-v2-context-9b-preview: encodes whole documents then pools chunk vectors; new SOTA on ConTEB, +14.4 points over voyage-context-4 with 1 KB int8 vectors.
- Cohere Embed 5: Pro/Fast tiers in one shared embedding space; Fast claims +6 points over fast-tier rivals at a third less cost than Pro.
- Ideogram 4.5: multi-turn editing keeps 94–99% of untouched content over ten consecutive edits; open weights promised.
- AA-Video-T2V v2.0: Wan 3.0 leads at $12/min; Seedance 2.5 second; MiniMax H3 statistically tied at $4.80/min.
- Open/small models: Ling-3.1-flash (500B) reportedly near GPT-5.6 Sol and Opus 5; Runway open-sourced world-action model Praxis-1; Upstage Solar Mini 4 (35B/3B active) is cheap but only 48% cache-hit rate makes it 5x Luna per task.
- Meta research: Context Language Models treat context as an editable file with context-management policies learned in the weights.
More from Models
- Sol 6.1 called a strong answer to Opus 5.5, arguably beating Astra in some ways — teortaxesTex · 2026-10-01
- No, open vs closed model usage didn't flip 80:20 in 12 weeks — analyst fact-checks viral claim — AccBalanced · 2026-10-01
- OpenAI's $300 Ultrafast mode hits 70tps while rivals match it at $20-$50 — NandaVegg · 2026-10-01
- First look at Opus 5.5 building an SCP-096 game in one go — imjustnewatai · 2026-10-01
- Singapore punches above its weight in AI as the world's No.4, but Claude's web integration lags OpenAI — teortaxesTex · 2026-10-01
- Hidden Dates in System Prompts Swing LLM Eval Scores by Up to 14% — Mario Sanz-Guerrero · 2026-10-01