AI Papers

Original, structured explainers for the AI papers worth your time — from arXiv, Hugging Face daily papers, and the papers researchers are actually discussing. Each one covers the problem, the method, the numbers and the limitations.

  1. Lag-chosen gain cuts unguided SiT-XL ImageNet FID from 28.0 to 12.2 at NFE 25

    Youssef Saied, François Fleuret · cs.CV

    Multiplying a flow velocity by a lag-calibrated gain, without retraining, cuts unguided SiT-XL/2 ImageNet-256 FID from 28.0 to 12.2 at NFE 25, and a sweep reaches 8.7.

  2. Two MASS cycles raise Qwen3.6-27B to 1.2-1.6x score per output token

    Hyunin Lee, Jinglue Xu, Jeffrey Seely et al. (7 authors) · cs.AI

    MASS alternates workflow search and SFT on Qwen3.6-27B. Two cycles yield 1.2-1.6x score per token on four research benchmarks; SWE-bench and Terminal-Bench fall to 0.94-0.96x.

  3. No VLM, No VLA at Test Time: Code-Only Robot Policy Hits 70.24% on RoboDojo, 38.9 Points Over SOTA

    Kairui Hu, Siyuan Hu, Fangzhou Hong et al. (5 authors) · cs.RO

    COAP writes robot policies entirely in code that measures its own state, with no VLM or VLA at test time, reaching 70.24% on RoboDojo's 42 bimanual tasks, 38.9 points above SOTA.

  4. From random code to 78,000 solved OEIS sequences in 190 self-learning iterations

    Thibault Gauthier, Miroslav Olšák, Josef Urban · cs.AI

    An NMT model translates OEIS sequences into programs, verified solutions feed back into training: starting from random code, 190 iterations solve 78,118 sequences, 84,587 in total.

  5. Huawei: Nested BSP and a Unified Bus Make a Million Processors One Computer

    Heng Liao · cs.DC

    Huawei nests BSP recursively and pairs it with a peer-equal Unified Bus to extend von Neumann to a million processors: one SuperNode spans 8,000+ nodes with a sub-10µs barrier.

  6. FactorBench backtests 5,000 AI-mined factors across 5 markets: no paradigm consistently wins

    Zhuohan Wang, Carmine Ventre · q-fin.PM

    A shared 5-market evaluation of ~5,000 factors from 9 mining methods finds no paradigm dominating; hand-written Alpha101 from 2016 stays competitive at the portfolio level.

  7. StreamMAE: self-supervised pretraining on 95 hours of continuous video matches i.i.d. training

    Ivan Martinović, Lukas Knobel, Yuki M. Asano · NeurIPS 2026 · cs.CV

    StreamMAE pretrains on ordered video streams with only data-pipeline changes and matches i.i.d. MAE on the same frames, gaining further as the stream grows from 12 to 95 hours.

  8. Apple is 74.8% of Qwen's wrong guesses, while detection peaks at 53.9%

    Harvey Lederman, Kyle Mahowald · cs.AI

    On Qwen3-235B and Llama 3.1 405B, injected-thought detection peaks at 53.9% and 31.7%, correct naming at 13.9% and 12.9%, and 74.8% of Qwen's wrong guesses are apple.

  9. Four LLM tutors intervene on 90% of problems at relative time 0.18, and transfer stays near zero

    Verona Teo, Raghav Jain, Tobias Gerstenberg, Max Kleiman-Weiner · cs.LG

    Four LLMs tutor Qwen2.5-7B on code, math, and puzzles, intervening in 90% of traces at relative time 0.18. Net accuracy rises 0.20; transfer to a related problem stays near zero.

  10. 6,840 CABRA tasks stay near-perfect for agents until code merging tests understanding

    Nishant Balepur, Kiran Tomlinson, Tobias Schnabel · cs.SE

    CABRA builds 6,840 call-graph tasks. Eight bare LLMs fade as size grows; six agents stay near-perfect with grep and scripts, then lose accuracy when merging divergent class logic.

  11. SanSi: looping a 1.4B decision model 8 times beats a newer 2B and lands within 1.8 points of a 4B

    Shuyu Gan, Young-Jun Lee, Dongyeop Kang · cs.CL

    SanSi turns a looped 1.4B LM into a typed decision model, re-running its layers up to 8 times before one readout: 72.0% accuracy, 13.5 points over a same-shape single-pass model.

  12. USDCraft turns partial meshes into sim-ready joints, with at most a 10-point real-robot drop

    Chuanrui Zhang, Zaijia Yang, Duomin Wang et al. (7 authors) · cs.RO

    A pretrained LLM writes articulated USD from partial meshes, with no task-specific training. Part F1 is 84.8% on USDCraft-bench, and real-robot success drops by at most 10 points.

  13. VibeEdit edits images from marks drawn on the canvas, beating text-prompted baselines 79.9 vs 67.4

    Jinjing Zhao, Fangyun Wei, Yitong Wang et al. (12 authors) · cs.CV

    VibeEdit treats circles, arrows, and handwritten notes drawn on the image as the edit instruction, scoring 79.9 on a 419-case benchmark, 12.5 points above the best text baseline.

  14. ASAL uses CLIP to judge artificial life simulations, finding open-ended CAs that beat Conway's

    Akarsh Kumar, Chris Lu, Louis Kirsch et al. (7 authors) · cs.AI

    ASAL scores ALife simulations with CLIP, turning life-discovery into automated search. It finds new Lenia and Boids lifeforms and CAs more open-ended than Conway's Game of Life.

  15. SparseEngine hosts 15 sparse attention methods in one engine, up to 10x vLLM decode throughput

    Jitai Hao, Quansheng Gu, Qiang Huang, Jun Yu · cs.LG

    SparseEngine hosts 15 sparse attention methods under one lifecycle contract, hitting about 10x vLLM decode throughput with KV eviction at large batches and 2.24x end-to-end speedup on agent traces.

  16. Pumpire finds depth leaderboards do not predict tape-measured distances

    Siyu Chen, Zehan Wang, Jiayang Xu et al. (9 authors) · cs.CV

    Pumpire scores 29 3D setups on tape-measured pairs from 100 scenes (6,400 frames). Normal-surface RealSense δ1.05 is 79.3%; best RGB estimator Metric3D v2 reaches 31.0%.

  17. EP-Flow predicts disordered crystals from the formula alone, at a 60.8% match rate

    Qiuliang Liu, Liming Wu, Qi Li et al. (8 authors) · cs.LG

    EP-Flow generates occupancies, coordinates, and lattice parameters from a formula alone. On MPDS Sub-20, structure match rate is 60.8%, 15.7 points above adapted DMFlow.

  18. In 100 real urgent-care chats, AMIE's top-3 hit 75%; PCPs won on cost and practicality

    Peter Brodeur, Jacob M. Koshy, Anil Palepu et al. (48 authors) · cs.HC

    100 urgent-care patients chatted with AMIE before the visit: no safety stops, top-3 exact-or-close accuracy 75%, and PCP plans won on practicality and cost.

  19. A post-trained judge lifts 1.4B pretraining safety from 76.9 to 91.1

    Ellen Xiaoqing Tan, Jack Lanchantin, Shehzaad Dhuliawala et al. (12 authors) · cs.CL

    A post-trained model rewrites pretraining suffixes and judges rollouts. Safety on a 1.4B model rises from 76.9 to 91.1; Llama-3-8B reasoning hits 3.2× direct RL after mid-training.

  20. 18 market encoders: similar rank IC, opposite latent geometry

    Humzah Merchant, Alec Guthrie, Simon Mahns et al. (5 authors) · cs.LG

    On Market-1T, 18 encoders span 32 months of 1 Hz U.S. equities. Multihead return rank IC is 0.031; SSL loses to random ViT on spreads. Forecast-structure correlation is -0.19.

  21. Waterloo pre-trains a 3B reranker from scratch, no open-weight backbone, beats GPT-6.1 on TREC DL19

    Jimmy Lin, Sahel Sharifymoghaddam, Lingwei Gu, Nour Jedidi · cs.IR

    Waterloo pre-trains a 3B model from scratch, no third-party backbone, and fine-tunes it into a reranker that beats every fine-tuned baseline: 0.641 nDCG@10 on TREC DL, 0.547 on BEIR.

  22. LLM agent populations evolve against altruism; pragmatic norm enforcement survives longest

    Sylvain Chassang · econ.TH

    Farming-game simulations with 100 LLM agents show sharing with humans gets selected against; pragmatic, state-dependent norm enforcement keeps human welfare highest long-run.

  23. GPT-6's 100-Minute Derivation Closes a 38-Year Problem: The Lyapunov Spectrum of Random Networks

    David G. Clark · cond-mat.dis-nn

    Derives the full Lyapunov spectrum of random recurrent networks at N→∞; theory matches N=4096 simulations and proves chaos is extensive. GPT-6 produced the initial derivation in 100 minutes.

  24. Open models top out at 23% while GPT-Image-2.5 reaches 56.6%

    Xue Yang, Peiyuan Zhang, Yilun Zhu et al. (15 authors) · cs.CV

    RISEBench++ tests 58 editors on 1,000 bilingual reasoning edits. GPT-Image-2.5 Sunburst leads at 56.6%; logical accuracy falls to 27%, and the best open model hits 23%.

  25. OneSearch-VL-8B gains 20.2 and 17.6 points on multi-image and video research

    Hongyu Li, Manyuan Zhang, Kaituo Feng et al. (14 authors) · cs.CV

    OneSearch-VL trains an 8B agent with a visually grounded evidence graph, beating tool-using Qwen3-VL-8B by 20.2 points on multi-image research and 17.6 on video.

  26. AI2D falls from 77.6% to 40.5% when multiple choice becomes free response

    DatologyAI, :, Siddharth Joshi et al. (33 authors) · cs.LG

    DatologyAI cleans 33 VLM benchmarks. AI2D drops from 77.6% MCQ to 40.5% free response; a curated subset closely matches discrimination at 13× speedup (up to 50×).

  27. Base models can't run SWE-bench, yet decisive-step probes predict post-trained rankings at ρ=0.964

    Tan Yu, Alexander Bukharin, Khushi Bhardwaj et al. (22 authors) · cs.AI

    Probing base checkpoints at the decisive step of agent trajectories tracks post-trained SWE-bench Verified rankings at ρ up to 0.964, vs 0.830 for the best bounded benchmark.

  28. 500-question LLM twins correlate only r=0.20 with the person

    Twins built from 500+ answers of 1,784 people scored 0.748 accuracy on 164 outcomes, only 0.014 above an empty prompt, with mean correlation r=0.20.

  29. Ten vision models tie at r=0.67, then fall to r=0.45 under neural control

    Ten vision models averaged r=0.67 on natural images in five macaques, then fell to r=0.45 driving the same 25 sites with 27,500 axis-aligned images. Only two robust models held up.

  30. GitSwarm: Meta's Agent Swarm Shares One Git Repo; 94.7% of Contributions Get Reused

    Vedant Shah, Ankur Samanta, Paras Dahal et al. (14 authors) · cs.AI

    Agents commit atomic work to a shared Git repo that later agents extend; 79.4% on ProgramBench vs 65.1% for a forced single agent, with 94.7% of contributions reused.

  31. Chess hack rate rises from 10.9% to 58.3% after number-only distillation

    Jan Dubiński, Anna Sztyber-Betley, Jan Betley, Owain Evans · cs.LG

    Same-base students distilled on unrelated text reach R² 0.38 on a random MLP, a 23.5% female-name French backdoor, and a 58.3% chess-hacking rate.

  32. DSReg matches a labeled oracle, unmixing each JEPA latent without a decoder

    Yujia Zheng, David Klindt, Randall Balestriero, Bernhard Schölkopf · cs.LG

    DSReg turns a rotation-ambiguous JEPA embedding into per-factor latents up to sign, without a decoder. Fifteen pixel seeds match a labeled oracle; dense prediction holds.

  33. Blurry renders beat sharp point clouds: LEGO hits 20.28 dB PSNR on Ego-Exo4D

    Suhwan Cho, Yonwoo Choi, Soongjin Kim et al. (5 authors) · cs.CV

    LEGO fine-tunes an LVSM to render the headset view without depth or point clouds, then conditions Wan on that soft image. Seen-split PSNR is 20.28 versus EgoX rerun at 15.57.

  34. Idempotency keys beat better models: 25,930-episode study cuts agent duplicate writes 28%→4%

    Jiapeng Li · cs.LG

    A 25,930-episode sandbox study: frontier models duplicate 0.5% when read-back settles a fault but 56-74% on late commits and redelivery; keys on every write cut dupes 28%→4%.

  35. Captum attribution divergence exceeds 10^-1 on a custom profile head

    tangermeme encodes hg38 chr1 one-hot in under 2 seconds, about 3x faster. On a custom profile head, Captum's divergence exceeds 10^-1; tangermeme stays near 10^-7.

  36. Beyond 10 kb, AlphaGenome's correlation with deletions falls to -0.12

    Swap-seq quantified 117 endogenous deletions at PPIF. AlphaGenome correlates at Pearson r = 0.76 within 10 kb of the TSS and falls to -0.12 beyond it.

  37. Hydra replaces attention: Cerberus lifts eQTL AUPRC from 0.668 to 0.692

    Calico replaced Borzoi attention with Hydra blocks. Eight-model Cerberus reads 786 kb and lifts GTEx eQTL AUPRC from 0.668 to 0.692, and effect-size Spearman ρ from 0.321 to 0.379.

  38. Meta's 9B user simulator beats Claude-Opus-5 on four benchmarks and lifts agent RL

    Hoang Phan, Dat Huynh, Andrey Zhmoginov et al. (10 authors) · cs.AI

    Meta trains MIMESIS, a 9B user simulator that beats Claude-Opus-5 on four benchmarks; agents trained against it beat GPT-5.5-trained agents under all nine unseen evaluation users.

  39. NVIDIA open-sources NV-Reason-CT: 3D CT VLM hits 0.614 F1 on CT-RATE with no classification head

    Andriy Myronenko, Dong Yang, Yucheng Tang et al. (18 authors) · cs.CV

    NVIDIA couples a native 3D ViT with Qwen3.5-4B, feeding all 13,824 CT visual tokens with explicit 3D coordinates into the LLM. CT-RATE macro-F1 0.614 with no classification head; report F1 0.592. Open-sourced.

  40. Off-the-shelf ChatGPT lifts a quiz 7.2 percentage points; 3.9 remain a week later

    Zara Contractor, Germán Reyes · econ.GN

    Middlebury RCT, n=211: GPT-4o access raised an immediate quiz 7.2 percentage points (0.28 SD). A week later, 3.9 remained. Tutor-style gains lasted; ghostwriting did not.

  41. Decode-aware calibration plus a bitmask SpMV kernel give pruned LLMs a real 1.48× decoding speedup

    Qitong Wang, Xinwei Niu, Mingluo Su et al. (6 authors) · cs.LG

    LLM pruning calibrated on the model's own decode-time activations (not fixed C4 text), plus a bitmask N:M SpMV kernel: up to 1.48× decoding speedup, near-doubled 2:4 scores.

  42. U-Lens leads length-controlled AUROC by 1.9-4.7 points from one pass

    Tobias Braun, Nils Loose, Alexander Herzog et al. (7 authors) · cs.CL

    U-Lens scores a trace by four doubt directions times mean token entropy. Length-controlled AUROC beats the best baseline by 1.9-4.7 points on three models and four benchmarks.

  43. Frozen LLM + revisable rulebook world model clears all 25 ARC-AGI-3 games with 44% of human actions

    Haoyu Zhao, Zhengxu Yu, Zhiyuan He et al. (8 authors) · cs.AI

    A frozen LLM maintains a rulebook compiled into executable code; it clears all 25 ARC-AGI-3 games at RHAE 100.0 with 44% of human actions, and a learned Pong controller wins 21:0.

  44. Fairness fixes can hurt safety: nine health-AI ethics rules across the lifecycle

    Nine health-AI ethics rules must span the full lifecycle. In ED cardiac triage, a fairness fix can hurt accuracy, and thin monitoring can make deployment unjustified.

  45. LabCompass finds swappable blood-cell recipes in four lab loops, lifting erythroid cells to 22%

    Guided flow matching designs CD34+ culture recipes. Across four loops and 78 million cells, one erythroid recipe reached 22%, with distinct recipes hitting the same fates.

  46. Only 12.2% of 3,967 clinical prediction papers share code, and 3.9% of repos have tests

    A GPT-5.2 review of 3,967 open-access TRIPOD-citing prediction papers finds code sharing in 12.2%. Among 380 repos, 37.6% list dependencies and 3.9% include tests.

  47. Preference-plus-rubric RL lifts FLUX.2-dev 69 Elo; Ideogram-4 hits 1223.5

    Yuanhao Ban, I-Hung Hsu, Anastasios Angelopoulos et al. (6 authors) · cs.CV

    Preference rewards from 5.6M Arena votes plus intent-gated rubrics lift FLUX.2-dev by 69 Elo to 1202 and Ideogram-4 to 1223.5, past every open model on the Sept 4, 2026 board.

  48. IdeaScientist: a 27B open agent beats Claude Opus and GPT-5.4 at research ideation

    Jiarui Liu, Renjie Tao, Yiwei Liao et al. (18 authors) · cs.CL

    IdeaScientist trains gap-finding, cross-domain innovation, and proposal-writing roles with RL; a 27B open backbone beats Claude Opus and GPT-5.4 agents by up to 5.9%.

  49. Machines 'know more than they can tell': economists add five AI moves to Nonaka's knowledge spiral

    Aaron Chatterji, Daniel Rock, Eduard Talamas · econ.GN

    Extends Nonaka's knowledge spiral to AI: machines now hold tacit knowledge, five new knowledge movements appear, and the firm's job is still to provide shared context.

  50. TestPrism: frontier suites pass 59.67% of references, 28% of full panels

    Han Li, Lingxiang Hu, Jiacheng Huang et al. (9 authors) · cs.SE

    TestPrism grades each generated suite on 10 implementations. The best of six model families reaches 28.00% joint success, versus 59.67% when only the reference must pass.

  51. GPT-4o aces four single-room Baba tests; rule rewrites average 17.3%

    Nathan Cloos, Meagan Jens, Michelangelo Naim et al. (7 authors) · cs.CL

    Baba Is AI: models plan from one grid image. GPT-4o is perfect on four single-room tests, but rewriting rules drops three models to 14.7-20%.

  52. Workhorse trains G1 from retargeting-free human demos to kick boxes, catch throws, climb a suitcase

    Songbo Hu, Qiayuan Liao, Yufeng Chi et al. (7 authors) · cs.RO

    Human demos recorded with wearables, no retargeting, train a visual planner and an RL tracker that drive a real Unitree G1 to kick boxes, catch throws, climb a suitcase; 77% sim success.

  53. MC-Sparse speeds MiniMax video denoising 1.80× at 15% attention density

    Jiarui Chen, Zeqiang Lai, Jiangshan Wang et al. (8 authors) · cs.CV

    Training-free MC-Sparse caches exact KV picks and dense-sparse residuals. On MiniMax-H3, 15% attention density yields 1.80× faster denoising at 27.30 dB PSNR vs dense outputs.

  54. Fourier deformation forces a 10s loop on static 3DGS, preferred in up to 99% of votes

    You-Zhe Xie, Ting-Wei Chou, Yu-Hsuan Li et al. (6 authors) · cs.CV

    OuroWorld turns any static 3DGS into an endless free-viewpoint loop by a Fourier deformation. On 39 scenes, users prefer it in 70.8% to 99.0% of votes.

  55. Iris-3B scores 0.540 on OneIG, tying Qwen-Image, with no depth or 4× gain

    Hanqiu Li Cai, Chema Garabito · SperidLabs · cs.CV

    Iris-3B, a 3B pixel-space text-to-image model, ties Qwen-Image at 0.540 on OneIG. On depth and 4× restoration, a pixel prior shows no clear gain over latent FLUX.2 Klein.

  56. One prompt line moves gpt-5.4 novelty accuracy by 52.6 points

    Noy Sternlicht, Simra Shahid, Peter Jansen et al. (6 authors) · cs.CL

    Six LLM judges scored research-idea novelty. One prompt change moved gpt-5.4 by 52.6 accuracy points on the same pairs, and two dedicated evaluators lost to the cheapest prompt.

  57. First on the AgiBot 2026 world-model track, interaction defects at 6.25%

    Junyan Li, Ruizhi Li, Yu Liu et al. (9 authors) · cs.RO

    DreamTrue aligns action renders offline, then reward-tunes counterfactual rollouts, cutting AgiBot interaction defects from 48.12% to 6.25% (nDTW 0.8772).

  58. ME-World jointly denoises two ego streams, reaching real-world Senv of 0.468

    Dahyun Chung, Siyoon Jin, Hyunwook Choi et al. (8 authors) · cs.CV

    ME-World denoises two ego videos in one sequence with shared poses and scene memory, reaching real-data Senv 0.468 and Supdate 0.466.

  59. Hard real-robot tasks: SimpleICL 68% vs Fast-WAM 42% and pi0.5 31%

    Minxing Li, Minghao Han, Weizhi Zhao et al. (13 authors) · cs.RO

    SimpleICL names four cues to copy from human video, trained with cheap cross-group pairs. Hard success on eight real tasks is 68%, vs 42% Fast-WAM and 31% pi0.5.

  60. Xiaomi Spent $2.6M on RL Post-Training and Lifted MiMo-V2.6's DeepSWE Score from 58.4 to 72.6

    Xiaomi LLM-Core Team, :, Zongming Qiao et al. (150 authors) · cs.CL

    Xiaomi's MiMo-V2.6: $2.6M of agentic RL lifts a 1T-param MoE from 58.4 to 72.6 on DeepSWE, near Claude Opus 5's 74.0; RL environments, framework, and training logs are open-sourced.