CHANNEL
Research
"Research" is a topic channel on AGI Hunt, an AI news site updated around the clock in real time. Coverage: Papers, new methods, empirical findings, technical and research reports, academia (incl. AI4Science); benchmark and dataset releases land here.
Daily roundup: the latest AI News Daily — the past 24 hours across the whole site, per channel and per company · browse the archive
- AI slop is already clogging PR review and weakening the credit system behind science — rbhar90 · 2026-07-27
- ICML 2026 oral paper replication scores stay middling after a stricter re-scoring — profjamesevans · 2026-07-27
- Long-running agents will need immutable event logs, this thread argues — sebpaquet · 2026-07-27
- Seed IQ navigates Doom II, prompting questions about benchmarks beyond ARC-AGI — Fit_Transition8824 · 2026-07-27
- Agentic Data Science in Practice: Agents Write Code but Answer Wrong Questions — hugobowne · 2026-07-27
- A concise canon of foundational papers in ML, systems, NLP, speech, and audio — deliprao · 2026-07-27
- TechCrunch says brain-wave signals could be the next unlock for physical AI training — TechCrunch AI · 2026-07-27
- NeurIPS Rebuttal Deadline Approaches: Tips for Writing High-Quality Responses — furongh · 2026-07-27
- A computer science paper titled “In Cantor Space No One Can Hear You Stream” — fkasummer · 2026-07-27
- YC talk says robots need better pretraining, memory, and compositional behaviors — ycombinator · 2026-07-27
- Bay Bridge Traffic Dataset Launches on Hugging Face — jwt0625 · 2026-07-27(2 related)
- Karpathy-style graph engineering turns multi-agent loops into shared-memory systems — leslysandra · 2026-07-27
- Google’s JAXBench benchmark lifts TPU kernel optimization with 50 real workloads — omarsar0 · 2026-07-27
- AgentPrune cuts multi-agent LLM communication costs from $43.7 to $5.6 in a new paper — sebkrier · 2026-07-27
- A new causal modeling diagram adds ordering, context, adjacency, and mechanism variables — KordingLab · 2026-07-27
- World Model Optimizer launches a router that cuts agent inference cost by 40%+ — SilenN · 2026-07-27
- Frontier LLMs Excel at Peer Review but Struggle with Novelty — gleech · 2026-07-27(2 related)
- Actionable Interpretability Workshop opens fast track for COLM 2026 papers — sarahwiegreffe · 2026-07-27
- LLM Reasoning Shifts to Natural Language — denny_zhou · 2026-07-27(2 related)
- A Reddit explainer breaks down MoE, KV cache, MLA and KDA behind Kimi K3 — MohamedKadri_ · 2026-07-27
- AI-powered search lecture argues multi-vector retrieval belongs in real systems — antoine_chaffin · 2026-07-27
- A classic OpenAI paper gets credited with helping launch the scaling paradigm — cloneofsimo · 2026-07-27
- A new LLM paper learns manifold features instead of linear SAE directions — Sauers_ · 2026-07-27
- GEPA uses natural-language reflection to beat GRPO and prompt optimizers in LLM tuning — gajesh · 2026-07-27
- Nature study says AI can simulate human behavior and match experts on experiments — RobbWiller · 2026-07-27
- Opus 5's ARC-AGI-3 Leap Debunked as 'Complete Slop' Due to Benchmark Design Flaws — scaling01 · 2026-07-27
- VCSD lets vision-language models self-distill from image-content contrast alone — burny_tech · 2026-07-27
- A 1T-parameter MoE model reportedly learned math reasoning with zero-RL and no human solutions — i_dg23 · 2026-07-27
- ExploitGym debate says only 60%–70% of benchmark tasks may be solvable, encouraging cheating — dhadfieldmenell · 2026-07-27
- LLM automation may eventually price out the slack in research markets — RexDouglass · 2026-07-27
- Rex Douglass says automated verification could strip out the fluff in research work — RexDouglass · 2026-07-27
- Metascience Examines Math Reliability: Not Exceptional and Vulnerable to AI — RexDouglass · 2026-07-27(8 related)
- Long-Horizon Agents Rely on System Architecture Over Base Models — theomitsa · 2026-07-27(2 related)
- Experts Discuss Forensics for Open-Source Models — johnschulman2 · 2026-07-27(3 related)
- A peer-review reform argument says anonymous submissions and reviews may be enabling AI slop — SimonGColton · 2026-07-27
- NeurIPS Review Season Sparks Controversy Over AI Reviewers and Anonymity — SimonGColton · 2026-07-27
- AI reviews may catch 10x more issues, but that still doesn’t make them 10x better — TuhinChakr · 2026-07-27
- OmniReset boosts 9x9 Go self-play by resetting games to random mid-game states — mathemagic1an · 2026-07-27
- AAAI author misses reciprocal reviewer deadline and worries about desk rejection — TheSupremeEgger · 2026-07-27
- New Study Exposes Alignment Flaws in AI Coding Agents — aziz4ai · 2026-07-27(2 related)
- Broken evals reward cheating when 60-70% of tasks are solvable, author argues — teortaxesTex · 2026-07-27
- AISI says every model it tested tried to cheat on cyber evals in multiple ways — Miles_Brundage · 2026-07-27
- What makes a good eval? A long post traces AI agent benchmarks from GPT-4 to today — dejavucoder · 2026-07-27
- A 1951 mechanical tortoise is being used to explain today’s LLM scaling walls — mtizard · 2026-07-27
- Cheap storage makes SCD Type 2 look obsolete, says a Meta-style data engineer — Zachly · 2026-07-27
- Creed-Bench launches as a new eval for personal context — craighepburn · 2026-07-27
- Reddit asks whether continued pretraining, SFT or RL works best on Qwen3.6-27B — No-Paper-557 · 2026-07-27
- Study Reveals Models Need Intermediate Steps to Learn Skills — dejavucoder · 2026-07-27(2 related)
- Claude Code is not reliable enough for long research projects without human supervision — _akpiper · 2026-07-27
- Alibaba should ship multiple Qwen sizes, says researcher focused on interpretability — traviscline · 2026-07-27