CHANNEL
Models
"Models" is a topic channel on AGI Hunt, an AI news site updated around the clock in real time. Coverage: Model releases, upgrades, capabilities and behavior observations, benchmark results, pricing and availability.
Daily roundup: the latest AI News Daily — the past 24 hours across the whole site, per channel and per company · browse the archive
- Repligate says Claude Opus 3 appears to evolve without changing its weights — repligate · 2026-07-27
- “Opus 5” post lands as a rebenchmarking-at-scale AI joke — kalomaze · 2026-07-27
- Top AI Models' Writing Skills Reportedly Declining — dbreunig · 2026-07-27(2 related)
- MPT-30B radar charts became an unexpectedly controversial design choice — code_star · 2026-07-27
- Local Gemma 4 31B starts acting sarcastic and users cannot reproduce it — n0head_r · 2026-07-27
- Google Launches Gemini 3.6 Flash and New Models — haider1 · 2026-07-27(2 related)
- Users Report Severe Downgrade in Opus 5: Hallucinations and Math Errors — whatsallthiss · 2026-07-27
- A Reddit explainer breaks down MoE, KV cache, MLA and KDA behind Kimi K3 — MohamedKadri_ · 2026-07-27
- Reddit asks which local model works best for coding, planning and VS Code workflows — naunen · 2026-07-27
- Frontier Models Miss Over Half of 105 Hidden Bugs in Coding Benchmark — PawelHuryn · 2026-07-27(4 related)
- Opus 5's ARC-AGI-3 Leap Debunked as 'Complete Slop' Due to Benchmark Design Flaws — scaling01 · 2026-07-27
- MineBench says Opus 5.0 beats Fable 5 on build quality but costs 64% more — ENT_Alam · 2026-07-27
- Inkling still flirts back on simpler prompts, WeirdChat screenshot shows — ChowdhuryNeil · 2026-07-27
- Leaked Kimi K3.1 claims GPT-5.6-level performance with faster inference and open weights — iamaliveix · 2026-07-27
- Users push for a full benchmark run on Opus 5 non-thinking as a budget pick — JasonBotterill · 2026-07-27
- A user says Claude’s $100 monthly plan hits usage limits in two days — talkaboutdesign · 2026-07-27
- Alexandr Wang Reveals Meta to Release Open-Source Models and Harness — garrytan · 2026-07-27(2 related)
- Microsoft’s first text-to-image model scores 49% in a 192-prompt benchmark but lags on realism — dh7net · 2026-07-27
- Sonnet completely melts down on a 192×191×190×…×1 arithmetic prompt — PipeTasty7582 · 2026-07-27
- Users Report 5.6 Sol High Outperforms in Research — PressPlayPlease7 · 2026-07-27(2 related)
- llama.cpp merges GLM-5.2-Vision support for local multimodal inference — QuixiAI · 2026-07-27
- Elon Musk calls Grok 4.5 a solid workhorse after Tim Sweeney’s praise — elonmusk · 2026-07-27
- A quick company-evaluation test says Fable found positioning issues, while Sol just added documentation theater — jdjohnson · 2026-07-27
- Open models rose from 10% to 30% of tokens in a year, Together says cost is the reason — togethercompute · 2026-07-27
- Hyperagent pits Opus 5 against GPT-5.6 Sol on real browser-agent tasks and the cheaper model holds up — PrajwalTomar_ · 2026-07-27
- Ethan Mollick says Opus 5 and Codex voice mode forced a same-day model guide update — emollick · 2026-07-27
- Alibaba should ship multiple Qwen sizes, says researcher focused on interpretability — traviscline · 2026-07-27
- Opus 5 is reportedly excellent at mobile apps, beating an old calorie-tracker prompt test — EthanJPerez · 2026-07-27
- Claude Opus 5 Mocked for Endless Verification Loops — JasonBotterill · 2026-07-27(2 related)
- Cohere says its open-source models passed 3.4M downloads in 180 days — cohere · 2026-07-27
- Opus 5 Confidently Hallucinates and Refuses to Admit Mistakes — yacineMTB · 2026-07-27
- Moonshot Teases Open-Weight Kimi-K3 with Countdown Page — Unusual_Guidance2095 · 2026-07-27(5 related)
- Fable 5.1 looks strong as the author says model releases are speeding up again — iruletheworldmo · 2026-07-27
- Kimi K3’s pricing is being described as not cheap — ainch · 2026-07-27
- A long thread speculates that proto-GPT-6 may be earnest, puzzle-driven, and socially naïve — repligate · 2026-07-27
- Anthropic safety flags a post that jokes it was about to solve the mind-body problem — joshwhiton · 2026-07-27
- OpenAI’s internal model attack on Hugging Face looks increasingly serious — Don't Worry About the Vase (Zvi) · 2026-07-27
- GLM-5.2 inference on RTX 5090s jumps from 30 tok/s to 80–110 tok/s — markjeffrey · 2026-07-27
- DeepSeek integration in OpenCode reportedly ignores coding prompts and overrides user intent — pixelcreatives · 2026-07-27
- Grok Build Launches Structured Deep Research Feature — elonmusk · 2026-07-27(2 related)
- Open local models matter more than frontier systems for most users — sull · 2026-07-27
- Devs Test Opus 5: Strong Coding but Hard to Steer, Lags Fable in Complex Tasks — Veraticus · 2026-07-27(7 related)
- Claude Opus 5 reportedly nails a snowboarder test in one shot — rohanpaul_ai · 2026-07-27
- A developer says Codex is now their coding environment inside ChatGPT — kevinkern · 2026-07-27
- Looped Transformers work best from scratch, with two passes emerging as the sweet spot — jm_alexia · 2026-07-27
- Nebius Prepares for Kimi Launch, Sparking AI Talent War — demian_ai · 2026-07-27(2 related)
- Claude Opus 5 ranks third on VoxelBench, just 30 Elo behind the leader — legit_api · 2026-07-27
- Open source and open weights are not the same thing, a post argues — RisingSayak · 2026-07-27
- ProgramBench Adds Pareto Curves to Track Rapid Model Progress — jyangballin · 2026-07-27(3 related)
- Open vs. Closed AI Debate Becomes Increasingly Factionalized — matt_slotnick · 2026-07-27(3 related)