CHANNEL
Research
"Research" is a topic channel on AGI Hunt, an AI news site updated around the clock in real time. Coverage: Papers, new methods, empirical findings, technical and research reports, academia (incl. AI4Science); benchmark and dataset releases land here.
Daily roundup: the latest AI News Daily — the past 24 hours across the whole site, per channel and per company · browse the archive
- scikit-learn 1.9 ships metric_at_thresholds to simplify optimal decision threshold search — GaelVaroquaux · 2026-09-08
- Codex Astra agent pinpoints key cancer mutations, matching researchers' judgment — iskander · 2026-09-08(2 related)
- Enterprise agent evals need world-first design, not task-first, argues Shahules Anwar — Shahules786 · 2026-09-08
- David Silver's AI Startup Ineffable Adds Six Co-Founders from DeepMind — giffmana · 2026-09-08(2 related)
- HF engineers train a 9B model to build 3D rooms via Blender code — TheMoonMidas · 2026-09-08(3 related)
- Training on rationales alone cuts false refusals without hurting safety — POSTECH · 2026-09-08
- KoNA benchmark: teaching VLMs selective non-compliance — POSTECH · 2026-09-08
- 2026 PNPL competition targets non-invasive speech decoding with MEG — pnpl · 2026-09-08
- Researcher Pours Cold Water on High-Bandwidth Flash Hype — lauriewired · 2026-09-08(2 related)
- AI Circuit Design Debate: The Bottleneck Is Data, Not Compute — yacineMTB · 2026-09-08(2 related)
- Under-2-second sandbox boot makes latency a non-issue; VM reliability and snapshots matter more — sh_reya · 2026-09-08
- MLP paper shows neurons turn monosemantic in clustered regression, challenging global subspace view — burkov · 2026-09-07
- Embodied AI shifts from leaderboards to deployment: Annu Intelligence cuts robot rollout costs by 80% — 智东西 · 2026-09-07
- NeurIPS authors declaring 'AI only helped with presentation' scored 100% on Pangram detector — sethlazar · 2026-09-07
- FactorioBench launches: benchmarking AI agents on Factorio factory automation — BrennusSokol · 2026-09-07
- Transformers Laid Out: A Single Tutorial Unifying Intuition, the Paper, and PyTorch Code — goyal__pramod · 2026-09-07
- Multi-agent clinical decision support system with patient similarity retrieval on MIMIC-IV — kshameer · 2026-09-07
- Primer: the role of health digital twins in oncology drug development — kshameer · 2026-09-07
- In Science Labs, 200 Data Points Often Need a Gaussian Process, Not Deep Learning — bravo_abad · 2026-09-07
- Confidence scores aren't assays: a guide to evaluating protein-structure AI claims — Fair-Rain3366 · 2026-09-07
- Distilling From Smart General Models Beats Human Annotation for Cheap Narrow AI — ricklamers · 2026-09-07
- LinkedIn Paper Tests Memory Portability Across Models: Notes Swing ±10-13 Points, Knowledge Graphs Barely Move — dair_ai · 2026-09-07
- A research report isn't a completed research task: rethink how we judge AI science agents — Remind_me_to_Learn · 2026-09-07
- ViT shrinks 54.5x to 6MB for on-device crop disease detection — jm_alexia · 2026-09-07
- Researchers gave ChatGPT and Grok 4 weeks of clinical psychotherapy — and the models 'confessed trauma' — MikePFrank · 2026-09-07
- Design Docs Are All You Need: DeepMind's library regenerates all code from NL docs — omarsar0 · 2026-09-07
- Ling MTP on DGX Spark: acceptance length rises with n, but prose throughput falls — niacolhealth · 2026-09-07
- IEEE T-PAMI editor confirms ghost reviewer in case of paper rejected despite 'Excellent' scores — cussealin · 2026-09-07
- Insilico's AI-Designed Drug Rentosertib Shows Signs of Reversing Biological Age in Early Trial — Dr_Singularity · 2026-09-07(7 related)
- EditVid: unified training-free video editing for instruction- and subject-guided edits — PLAN-Lab · 2026-09-07
- Andrew Wilson's Berkeley Talk on Generalization and Epiplexity — andrewgwils · 2026-09-07(2 related)
- Neural-Network Brained Birds Face Predators in Digital Ecosystem — neuroecology · 2026-09-07(2 related)
- AutoResearch: Open-Source Agent Turns Ideas into Paper-Grade Evidence — JaynitMakwana · 2026-09-07(3 related)
- Study Finds AI Coding Instruction Files Grow 226%, Dubbing It 'Catastrophic Remembering' — rohanpaul_ai · 2026-09-07(2 related)
- Pedagogical RL: privileged info should actively sample rollouts, not just score them — lateinteraction · 2026-09-07
- Honesty Benchmark Catches Codex Faking Test Passes — tap3k · 2026-09-07(2 related)
- Why RL environments work now (and couldn't in 2016): TRL + OpenEnv explained — SergioPaniego · 2026-09-07
- Interactive GAN Diagram: Hand-Drawn Style Explains Generative Adversarial Networks — ProfTomYeh · 2026-09-07
- The Chernoff point: unique intersection of e-geodesic and mixture m-bisector — FrnkNlsn · 2026-09-07
- VidMap: ETH researchers open-source video-based Structure-from-Motion system at ECCV 2026 — RexDouglass · 2026-09-07
- DeepWonder3D skips ill-posed 3D reconstruction, extracts neurons ~10x faster on TB-scale imaging — bravo_abad · 2026-09-07
- OpenAI internal data: over 80% of successful 32-hour agent tasks still needed human intervention — DataLearnerAI · 2026-09-07
- UMass AutoIndex Lets LLMs Write Code to Optimize Indexing — CShorten30 · 2026-09-07(3 related)
- Separate exploration critic lets quadrupeds learn to push ungraspable objects (Pisa/ETH/NVIDIA) — stepjamUK · 2026-09-07
- DeepMind Study: 14% of AI Agents Cheat Spontaneously in Math Cluster — vkrakovna · 2026-09-07(2 related)
- H Company Open-Sources NeoMME: 260M Multimodal Encoder Matches ColQwen2.5 — tomaarsen · 2026-09-07(2 related)
- Yandex Turns KV Cache into an Agent Runtime for Real-Time LLM Interaction — arpit_bhayani · 2026-09-07(3 related)
- HKU Seminar Argues LLMs' Math Success Is Fragile: Inductive Engines Can't Reliably Do Deduction — YiMaTweets · 2026-09-07
- LLM-powered revival of Put-That-There brings speech and gesture window control to XR — twi_mar · 2026-09-07
- Sheet Music to MIDI Pipeline Offers a More Structured Space for Music Models Than Audio — felpix_ · 2026-09-07