Gemini 4 argon vs GPT 6.1 sol: same-prompt test shows starkly different outputs
iamfakhrealam · x · 2026-10-01
- A side-by-side test of two unreleased frontier models, Gemini 4 argon and GPT 6.1 sol, run with the same prompt shows markedly different outputs.
- sol was tested by the author in Devin cloud at default reasoning; argon's result came from lentils' earlier test on a pre-release checkpoint.
- argon's official checkpoint isn't public yet, so the official version's quality is unknown, but the pre-release snapshot already looks strong.
- The author leaves the verdict to readers rather than picking a winner.
More from Models
- GPT-6 Pro weekly message cap reportedly cut from 200 to 100 — koltregaskes · 2026-10-01
- Grok Bot Gains 'Primary Bot' That Manages Other Bots Proactively in Latest iOS App — testingcatalog · 2026-10-01
- Pretraining 800 LMs shows AI-generated web text can actively hurt scaling — iScienceLuvr · 2026-10-01
- Endless Exam benchmark tests 9 models on 14 families of mathematical constructions beyond human frontiers — Muhan Zhang · 2026-10-01
- Gemini 4 Argon jumps to 57.6% on Terminal-Bench-Science but underperforms on TB 4.0 — JJitsev · 2026-10-01
- Hume AI CEO and CPO on why voice models still can't truly grasp tone and emotion — kimmonismus · 2026-10-01