CHANNEL
Models
"Models" is a topic channel on AGI Hunt, an AI news site updated around the clock in real time. Coverage: Model releases, upgrades, capabilities and behavior observations, benchmark results, pricing and availability.
Daily roundup: the latest AI News Daily — the past 24 hours across the whole site, per channel and per company · browse the archive
- "Open-source AI must win": OpenMed manifesto argues powerful intelligence can't stay in few hands — MaziyarPanahi · 2026-09-13
- AI Solves Navier–Stokes, Sparking Math Community Backlash — rubenhassid · 2026-09-13(2 related)
- Grok Bots coming to XChat: leak suggests bot DMs on X — nima_owji · 2026-09-13
- Nex-N2.5 Pro Plays Pokémon for Hundreds of Consecutive Steps — alifcoder · 2026-09-13(2 related)
- Third-party 10-dim eval shows DeepSeek V4.1 big reliability gains over V4-Pro-0813 — teortaxesTex · 2026-09-13
- Fine-Tuned Open Models at 5% Cost Spark Debate Over Frontier Lab Moats — ayushtweetshere · 2026-09-13(2 related)
- DeepSeek launches V4.1-Flash: 552B params, multimodal, faster, with API price cuts — emmanuelvivier · 2026-09-13
- Anthropic's most detailed threat report details 39 Claude abuse cases, incl. Russia-linked spying — emmanuelvivier · 2026-09-13
- Not a capability plateau, an affordable compute plateau — AI debate over Opus pricing — marlene_zw · 2026-09-13
- GPT-6 Astra reportedly shipped 8 weeks after GPT-5.6 Sol, possibly the last fast turnaround — kimmonismus · 2026-09-13
- Dev fine-tunes Qwen 3.8 on 81,837 book annotations, writing quality up 86% in blind tests — Scobleizer · 2026-09-13
- Muse Spark 1.3 Ties for First in 105-Real-Bug Benchmark — PawelHuryn · 2026-09-13(3 related)
- Revisiting 2019: OpenAI withheld GPT-2 as 'too dangerous to release' — olalonde · 2026-09-13
- Public reports suggest 1-3 month lag between internal and public frontier models — xeophon · 2026-09-13(2 related)
- Debate Erupts Over Alleged Benchmark Contamination in DeepSeek V4.1 Flash — teortaxesTex · 2026-09-13(2 related)
- 45% of overnight benchmark rollouts failed mid-turn amid OpenAI capacity issues — dejavucoder · 2026-09-13
- Gemini turns out surprisingly good at translating swarm language to and from English — xeophon · 2026-09-13
- DeepSeek V4.1 costs less to serve but prices 2x higher per output token, margins likely up — teortaxesTex · 2026-09-13
- Migrating from Claude Code to Codex: history, plugins and skills don't carry over — Inevitable2727 · 2026-09-13
- Motorhomelabs Teases New Release Amid Prediction Open Models Face Strict Regulation — aiamblichus · 2026-09-13
- Rumors Claim DeepMind Achieved Recursive Self-Improvement — WorldofAI · 2026-09-13(4 related)
- DeepSeek has no single research direction, but 'infinite context' is the goal — teortaxesTex · 2026-09-13
- Indie dev offers to train a fully open 9.4B dense model, tuned for a single GPU — NineThreeTilNow · 2026-09-13
- User shows how GPT flatters your beliefs and hallucinates premises in arguments — GlenBradley · 2026-09-13
- Orchestrator Error Reveals Mystery Model 'Daybreak': 'astra Is Not Allowed to Access Those Resources' — LeopardBernstein · 2026-09-13
- Developer kalomaze calls Opus 5 a bad model — kalomaze · 2026-09-13(2 related)
- Wenhu Chen can't even understand many questions in AA-intelligence AI benchmarks — WenhuChen · 2026-09-13
- Commenter: The OpenAI Navier-Stokes proof would be hailed as a breakthrough if posted anonymously — skdh · 2026-09-13
- GPT-6 Astra skips the chat box: Plus users get just 5-45 messages per 5 hours, locked to Work and Codex — 新智元 · 2026-09-13
- Plus users report Astra burns through Codex 5-hour usage in about two prompts — airamx113 · 2026-09-13
- LLM Architecture Gallery: one chart to compare every major LLM design — zainhas · 2026-09-13
- Schulman cites model trained only on pre-1930 text beating Claude 3 Opus after distillation — victor_explore · 2026-09-13
- DeepSeek-V4.1-Flash reportedly live on Ollama cloud with zero data retention and off-peak pricing — johnseach · 2026-09-13
- Running two frontier models against each other on study design works 'absurdly' well — Tkaraletsos · 2026-09-13
- DeepSeek 4.1 Flash solves HLE problem, then downloads the dataset to check its own answer — FutureStriking283 · 2026-09-13
- Anthropic's policy details exactly when your Claude chats are used for model training — austinc3301 · 2026-09-13
- GPT-Live-1 first impressions: most natural voice model yet, but instruction following is unreliable — kolchinski · 2026-09-13
- GPT-6 Astra dethrones Gemini 3.8 Flash as the best vision model in side-by-side tests — Roger_M_Taylor · 2026-09-13
- Devin SWE-2 first impressions: fast and cheap, but shorter endurance than Codex — CtrlAltDwayne · 2026-09-13
- Different LLMs fail SWE tasks in distinct ways, dev observes — zainhas · 2026-09-13
- Developers Question Whether AI Training Opt-Out Clauses Leave Loopholes — logicus · 2026-09-13(3 related)
- ~$11 per billion tokens: decentralized-trained model's inference optimizations revealed — bittingthembits · 2026-09-13
- Developer Praises Char-Level Models: No Retokenization Bugs — cephaloform · 2026-09-13(2 related)
- SemiAnalysis probes positional embeddings: what lies beyond common mathematical constraints? — burny_tech · 2026-09-13
- User Says Model Wrote Chinese in a Chat Title, Then Insisted It Was Armenian — hydralisk_hydrawife · 2026-09-13
- OpenAI open problem #2 cracked by human theorists; paper unifies and improves MRRW bounds — abeirami · 2026-09-13
- Burkov: Models Still Sycophantic; Without Harness, Intelligence on Par with o1 — burkov · 2026-09-13(2 related)
- Insider chatter claims AI labs have hit the scaling wall and are pivoting to hard math — Promptmethus · 2026-09-13
- QuinnyPig: Gemini has been slowing down AI development for nearly a year — firstadopter · 2026-09-13
- Opencode Zen Bug: Muse Spark Models Fail on Images and Tool Calls with encrypted_content Error — pompompur-in · 2026-09-13