Deep Learning Weekly #475: GPT-6.1 Sol, LLM-as-a-Judge, JIT memory for agents
skdh · x · 2026-10-02
Deep Learning Weekly Issue #475 is live, covering the introduction of GPT-6.1 Sol and Luna, an AI evals guide, Jev vs. LLM-as-a-Judge, Just-in-Time Memory for Agents, and a paper on RRSI (Regularized Recursive Self-Improvement of Agents), among other items.
More from Research
- Marin's 535B MoE hero run hits ~27% MFU on 11 NVL72 racks: expert parallelism deep dive — dlwh · 2026-10-03
- SWE-sweep benchmark tests whether AI agents can find bugs before users hit them — klieret · 2026-10-02
- Harvard-led team unveils brain imaging 60x faster than fMRI, tracking activity in ~100ms — melnykowycz · 2026-10-02
- RExBench: best coding agent implements research extensions only 33% of the time — najoungkim · 2026-10-02
- Retrieval-augmented episodic memory narrows LLMs' lexical frequency gap in syntax tests — najoungkim · 2026-10-02
- EMPIRIC teaches robots missing physics as code, solving all 25 tasks where baselines get 14-16 — tomssilver · 2026-10-02