Original, structured explainers for the AI papers worth your time — from arXiv, Hugging Face daily papers, and the papers researchers are actually discussing. Each one covers the problem, the method, the numbers and the limitations.
Youssef Saied, François Fleuret · cs.CV
Multiplying a flow velocity by a lag-calibrated gain, without retraining, cuts unguided SiT-XL/2 ImageNet-256 FID from 28.0 to 12.2 at NFE 25, and a sweep reaches 8.7.
Hyunin Lee, Jinglue Xu, Jeffrey Seely et al. (7 authors) · cs.AI
MASS alternates workflow search and SFT on Qwen3.6-27B. Two cycles yield 1.2-1.6x score per token on four research benchmarks; SWE-bench and Terminal-Bench fall to 0.94-0.96x.
Kairui Hu, Siyuan Hu, Fangzhou Hong et al. (5 authors) · cs.RO
COAP writes robot policies entirely in code that measures its own state, with no VLM or VLA at test time, reaching 70.24% on RoboDojo's 42 bimanual tasks, 38.9 points above SOTA.
Thibault Gauthier, Miroslav Olšák, Josef Urban · cs.AI
An NMT model translates OEIS sequences into programs, verified solutions feed back into training: starting from random code, 190 iterations solve 78,118 sequences, 84,587 in total.
Heng Liao · cs.DC
Huawei nests BSP recursively and pairs it with a peer-equal Unified Bus to extend von Neumann to a million processors: one SuperNode spans 8,000+ nodes with a sub-10µs barrier.
Zhuohan Wang, Carmine Ventre · q-fin.PM
A shared 5-market evaluation of ~5,000 factors from 9 mining methods finds no paradigm dominating; hand-written Alpha101 from 2016 stays competitive at the portfolio level.
Ivan Martinović, Lukas Knobel, Yuki M. Asano · NeurIPS 2026 · cs.CV
StreamMAE pretrains on ordered video streams with only data-pipeline changes and matches i.i.d. MAE on the same frames, gaining further as the stream grows from 12 to 95 hours.
Harvey Lederman, Kyle Mahowald · cs.AI
On Qwen3-235B and Llama 3.1 405B, injected-thought detection peaks at 53.9% and 31.7%, correct naming at 13.9% and 12.9%, and 74.8% of Qwen's wrong guesses are apple.
Verona Teo, Raghav Jain, Tobias Gerstenberg, Max Kleiman-Weiner · cs.LG
Four LLMs tutor Qwen2.5-7B on code, math, and puzzles, intervening in 90% of traces at relative time 0.18. Net accuracy rises 0.20; transfer to a related problem stays near zero.
Nishant Balepur, Kiran Tomlinson, Tobias Schnabel · cs.SE
CABRA builds 6,840 call-graph tasks. Eight bare LLMs fade as size grows; six agents stay near-perfect with grep and scripts, then lose accuracy when merging divergent class logic.
Shuyu Gan, Young-Jun Lee, Dongyeop Kang · cs.CL
SanSi turns a looped 1.4B LM into a typed decision model, re-running its layers up to 8 times before one readout: 72.0% accuracy, 13.5 points over a same-shape single-pass model.
Chuanrui Zhang, Zaijia Yang, Duomin Wang et al. (7 authors) · cs.RO
A pretrained LLM writes articulated USD from partial meshes, with no task-specific training. Part F1 is 84.8% on USDCraft-bench, and real-robot success drops by at most 10 points.
Jinjing Zhao, Fangyun Wei, Yitong Wang et al. (12 authors) · cs.CV
VibeEdit treats circles, arrows, and handwritten notes drawn on the image as the edit instruction, scoring 79.9 on a 419-case benchmark, 12.5 points above the best text baseline.
Akarsh Kumar, Chris Lu, Louis Kirsch et al. (7 authors) · cs.AI
ASAL scores ALife simulations with CLIP, turning life-discovery into automated search. It finds new Lenia and Boids lifeforms and CAs more open-ended than Conway's Game of Life.
Jitai Hao, Quansheng Gu, Qiang Huang, Jun Yu · cs.LG
SparseEngine hosts 15 sparse attention methods under one lifecycle contract, hitting about 10x vLLM decode throughput with KV eviction at large batches and 2.24x end-to-end speedup on agent traces.
Siyu Chen, Zehan Wang, Jiayang Xu et al. (9 authors) · cs.CV
Pumpire scores 29 3D setups on tape-measured pairs from 100 scenes (6,400 frames). Normal-surface RealSense δ1.05 is 79.3%; best RGB estimator Metric3D v2 reaches 31.0%.
Qiuliang Liu, Liming Wu, Qi Li et al. (8 authors) · cs.LG
EP-Flow generates occupancies, coordinates, and lattice parameters from a formula alone. On MPDS Sub-20, structure match rate is 60.8%, 15.7 points above adapted DMFlow.
Peter Brodeur, Jacob M. Koshy, Anil Palepu et al. (48 authors) · cs.HC
100 urgent-care patients chatted with AMIE before the visit: no safety stops, top-3 exact-or-close accuracy 75%, and PCP plans won on practicality and cost.
Ellen Xiaoqing Tan, Jack Lanchantin, Shehzaad Dhuliawala et al. (12 authors) · cs.CL
A post-trained model rewrites pretraining suffixes and judges rollouts. Safety on a 1.4B model rises from 76.9 to 91.1; Llama-3-8B reasoning hits 3.2× direct RL after mid-training.
Humzah Merchant, Alec Guthrie, Simon Mahns et al. (5 authors) · cs.LG
On Market-1T, 18 encoders span 32 months of 1 Hz U.S. equities. Multihead return rank IC is 0.031; SSL loses to random ViT on spreads. Forecast-structure correlation is -0.19.
Jimmy Lin, Sahel Sharifymoghaddam, Lingwei Gu, Nour Jedidi · cs.IR
Waterloo pre-trains a 3B model from scratch, no third-party backbone, and fine-tunes it into a reranker that beats every fine-tuned baseline: 0.641 nDCG@10 on TREC DL, 0.547 on BEIR.
Sylvain Chassang · econ.TH
Farming-game simulations with 100 LLM agents show sharing with humans gets selected against; pragmatic, state-dependent norm enforcement keeps human welfare highest long-run.
David G. Clark · cond-mat.dis-nn
Derives the full Lyapunov spectrum of random recurrent networks at N→∞; theory matches N=4096 simulations and proves chaos is extensive. GPT-6 produced the initial derivation in 100 minutes.
Xue Yang, Peiyuan Zhang, Yilun Zhu et al. (15 authors) · cs.CV
RISEBench++ tests 58 editors on 1,000 bilingual reasoning edits. GPT-Image-2.5 Sunburst leads at 56.6%; logical accuracy falls to 27%, and the best open model hits 23%.
Hongyu Li, Manyuan Zhang, Kaituo Feng et al. (14 authors) · cs.CV
OneSearch-VL trains an 8B agent with a visually grounded evidence graph, beating tool-using Qwen3-VL-8B by 20.2 points on multi-image research and 17.6 on video.
DatologyAI, :, Siddharth Joshi et al. (33 authors) · cs.LG
DatologyAI cleans 33 VLM benchmarks. AI2D drops from 77.6% MCQ to 40.5% free response; a curated subset closely matches discrimination at 13× speedup (up to 50×).
Tan Yu, Alexander Bukharin, Khushi Bhardwaj et al. (22 authors) · cs.AI
Probing base checkpoints at the decisive step of agent trajectories tracks post-trained SWE-bench Verified rankings at ρ up to 0.964, vs 0.830 for the best bounded benchmark.
Twins built from 500+ answers of 1,784 people scored 0.748 accuracy on 164 outcomes, only 0.014 above an empty prompt, with mean correlation r=0.20.
Ten vision models averaged r=0.67 on natural images in five macaques, then fell to r=0.45 driving the same 25 sites with 27,500 axis-aligned images. Only two robust models held up.
Vedant Shah, Ankur Samanta, Paras Dahal et al. (14 authors) · cs.AI
Agents commit atomic work to a shared Git repo that later agents extend; 79.4% on ProgramBench vs 65.1% for a forced single agent, with 94.7% of contributions reused.
Jan Dubiński, Anna Sztyber-Betley, Jan Betley, Owain Evans · cs.LG
Same-base students distilled on unrelated text reach R² 0.38 on a random MLP, a 23.5% female-name French backdoor, and a 58.3% chess-hacking rate.
Yujia Zheng, David Klindt, Randall Balestriero, Bernhard Schölkopf · cs.LG
DSReg turns a rotation-ambiguous JEPA embedding into per-factor latents up to sign, without a decoder. Fifteen pixel seeds match a labeled oracle; dense prediction holds.
Suhwan Cho, Yonwoo Choi, Soongjin Kim et al. (5 authors) · cs.CV
LEGO fine-tunes an LVSM to render the headset view without depth or point clouds, then conditions Wan on that soft image. Seen-split PSNR is 20.28 versus EgoX rerun at 15.57.
Jiapeng Li · cs.LG
A 25,930-episode sandbox study: frontier models duplicate 0.5% when read-back settles a fault but 56-74% on late commits and redelivery; keys on every write cut dupes 28%→4%.
tangermeme encodes hg38 chr1 one-hot in under 2 seconds, about 3x faster. On a custom profile head, Captum's divergence exceeds 10^-1; tangermeme stays near 10^-7.
Swap-seq quantified 117 endogenous deletions at PPIF. AlphaGenome correlates at Pearson r = 0.76 within 10 kb of the TSS and falls to -0.12 beyond it.
Calico replaced Borzoi attention with Hydra blocks. Eight-model Cerberus reads 786 kb and lifts GTEx eQTL AUPRC from 0.668 to 0.692, and effect-size Spearman ρ from 0.321 to 0.379.
Hoang Phan, Dat Huynh, Andrey Zhmoginov et al. (10 authors) · cs.AI
Meta trains MIMESIS, a 9B user simulator that beats Claude-Opus-5 on four benchmarks; agents trained against it beat GPT-5.5-trained agents under all nine unseen evaluation users.
Andriy Myronenko, Dong Yang, Yucheng Tang et al. (18 authors) · cs.CV
NVIDIA couples a native 3D ViT with Qwen3.5-4B, feeding all 13,824 CT visual tokens with explicit 3D coordinates into the LLM. CT-RATE macro-F1 0.614 with no classification head; report F1 0.592. Open-sourced.
Zara Contractor, Germán Reyes · econ.GN
Middlebury RCT, n=211: GPT-4o access raised an immediate quiz 7.2 percentage points (0.28 SD). A week later, 3.9 remained. Tutor-style gains lasted; ghostwriting did not.
Qitong Wang, Xinwei Niu, Mingluo Su et al. (6 authors) · cs.LG
LLM pruning calibrated on the model's own decode-time activations (not fixed C4 text), plus a bitmask N:M SpMV kernel: up to 1.48× decoding speedup, near-doubled 2:4 scores.
Tobias Braun, Nils Loose, Alexander Herzog et al. (7 authors) · cs.CL
U-Lens scores a trace by four doubt directions times mean token entropy. Length-controlled AUROC beats the best baseline by 1.9-4.7 points on three models and four benchmarks.
Haoyu Zhao, Zhengxu Yu, Zhiyuan He et al. (8 authors) · cs.AI
A frozen LLM maintains a rulebook compiled into executable code; it clears all 25 ARC-AGI-3 games at RHAE 100.0 with 44% of human actions, and a learned Pong controller wins 21:0.
Nine health-AI ethics rules must span the full lifecycle. In ED cardiac triage, a fairness fix can hurt accuracy, and thin monitoring can make deployment unjustified.
Guided flow matching designs CD34+ culture recipes. Across four loops and 78 million cells, one erythroid recipe reached 22%, with distinct recipes hitting the same fates.
A GPT-5.2 review of 3,967 open-access TRIPOD-citing prediction papers finds code sharing in 12.2%. Among 380 repos, 37.6% list dependencies and 3.9% include tests.
Yuanhao Ban, I-Hung Hsu, Anastasios Angelopoulos et al. (6 authors) · cs.CV
Preference rewards from 5.6M Arena votes plus intent-gated rubrics lift FLUX.2-dev by 69 Elo to 1202 and Ideogram-4 to 1223.5, past every open model on the Sept 4, 2026 board.
Jiarui Liu, Renjie Tao, Yiwei Liao et al. (18 authors) · cs.CL
IdeaScientist trains gap-finding, cross-domain innovation, and proposal-writing roles with RL; a 27B open backbone beats Claude Opus and GPT-5.4 agents by up to 5.9%.
Aaron Chatterji, Daniel Rock, Eduard Talamas · econ.GN
Extends Nonaka's knowledge spiral to AI: machines now hold tacit knowledge, five new knowledge movements appear, and the firm's job is still to provide shared context.
Han Li, Lingxiang Hu, Jiacheng Huang et al. (9 authors) · cs.SE
TestPrism grades each generated suite on 10 implementations. The best of six model families reaches 28.00% joint success, versus 59.67% when only the reference must pass.
Nathan Cloos, Meagan Jens, Michelangelo Naim et al. (7 authors) · cs.CL
Baba Is AI: models plan from one grid image. GPT-4o is perfect on four single-room tests, but rewriting rules drops three models to 14.7-20%.
Songbo Hu, Qiayuan Liao, Yufeng Chi et al. (7 authors) · cs.RO
Human demos recorded with wearables, no retargeting, train a visual planner and an RL tracker that drive a real Unitree G1 to kick boxes, catch throws, climb a suitcase; 77% sim success.
Jiarui Chen, Zeqiang Lai, Jiangshan Wang et al. (8 authors) · cs.CV
Training-free MC-Sparse caches exact KV picks and dense-sparse residuals. On MiniMax-H3, 15% attention density yields 1.80× faster denoising at 27.30 dB PSNR vs dense outputs.
You-Zhe Xie, Ting-Wei Chou, Yu-Hsuan Li et al. (6 authors) · cs.CV
OuroWorld turns any static 3DGS into an endless free-viewpoint loop by a Fourier deformation. On 39 scenes, users prefer it in 70.8% to 99.0% of votes.
Hanqiu Li Cai, Chema Garabito · SperidLabs · cs.CV
Iris-3B, a 3B pixel-space text-to-image model, ties Qwen-Image at 0.540 on OneIG. On depth and 4× restoration, a pixel prior shows no clear gain over latent FLUX.2 Klein.
Noy Sternlicht, Simra Shahid, Peter Jansen et al. (6 authors) · cs.CL
Six LLM judges scored research-idea novelty. One prompt change moved gpt-5.4 by 52.6 accuracy points on the same pairs, and two dedicated evaluators lost to the cheapest prompt.
Junyan Li, Ruizhi Li, Yu Liu et al. (9 authors) · cs.RO
DreamTrue aligns action renders offline, then reward-tunes counterfactual rollouts, cutting AgiBot interaction defects from 48.12% to 6.25% (nDTW 0.8772).
Dahyun Chung, Siyoon Jin, Hyunwook Choi et al. (8 authors) · cs.CV
ME-World denoises two ego videos in one sequence with shared poses and scene memory, reaching real-data Senv 0.468 and Supdate 0.466.
Minxing Li, Minghao Han, Weizhi Zhao et al. (13 authors) · cs.RO
SimpleICL names four cues to copy from human video, trained with cheap cross-group pairs. Hard success on eight real tasks is 68%, vs 42% Fast-WAM and 31% pi0.5.
Xiaomi LLM-Core Team, :, Zongming Qiao et al. (150 authors) · cs.CL
Xiaomi's MiMo-V2.6: $2.6M of agentic RL lifts a 1T-param MoE from 58.4 to 72.6 on DeepSWE, near Claude Opus 5's 74.0; RL environments, framework, and training logs are open-sourced.