Microsoft Trains a 'Night Science' Agent with RL, Expanding Research Directions 27.8%
MicrosoftResearch · hf · 2026-09-30
Microsoft Research introduces AI Night-Scientist, an agentic framework that tackles the homogeneity and predictability of LLM outputs in open-ended scientific ideation by using reinforcement learning to teach models when and how to depart from predictable reasoning.
- Grounded in cognitive science, creativity is modeled along three axes: action (what to do and how creatively), process (when to explore vs. exploit), and outcome (novelty and usefulness of the idea).
- Models are trained with GRPO across these axes, exposing them to varying degrees and forms of creativity during training.
- Results: substantially more diverse scientific proposals, expanding research directions by 27.8% and contribution types by 14.9%; predicted citation impact up to 32.0 percentage points and originality up 66.2 points.
- Key finding: these gains cannot be reproduced by simply raising decoding temperature; semantic guidance specifying what kind of creativity to pursue is critical.
The authors conclude creativity is a learnable, multi-level ability that can be shaped to help researchers reach ideas beyond those typically explored by LLMs.
More from AGI Musings
- Agentic traffic wave: under 1% of consumers use agents today — and almost no business is ready — omooretweets · 2026-09-30
- Chollet: The Test for Human-Level AGI Is Passing ARC-AGI-(n+1) on Release — fchollet · 2026-09-30
- The Id, the Ego and the Superintelligence: NYT on AI and Human Judgment — nytopinion · 2026-09-30
- The cruel cycle: labs RL a skill, wrapper startups pile in, labs ship the killer — generativist · 2026-09-30
- Fortune Argues AI Writing Disclosure Should Be a Spectrum, Not Yes/No — shdw_0x0 · 2026-09-30
- Top interpretability expert Neel Nanda: don't rely on us to save you on the current trajectory — burny_tech · 2026-09-30