One prompt line moves gpt-5.4 novelty accuracy by 52.6 points

Old Ideas, Novel Problems: The Instability of LLM-Based Novelty Evaluation

Noy Sternlicht, Simra Shahid, Peter Jansen, Daniel S. Weld, Pao Siangliulue, Tom Hope

cs.CL

2026-10-02

Six LLM judges scored research-idea novelty. One prompt change moved gpt-5.4 by 52.6 accuracy points on the same pairs, and two dedicated evaluators lost to the cheapest prompt.

What problem this solves

Systems that generate research ideas, from end-to-end AI scientists to multi-agent hypothesis editors, almost all report one number: the ideas are more novel. That number usually comes from an LLM judge, a prompted model acting as a reviewer. Which model, whether it can search, and whether it scores ideas singly or in pairs, travel with the system and are rarely tested on their own.

When a check exists, it tends to use a small labeled set, data from before the model's knowledge cutoff, or human-written papers. That distribution is not the generated ideas the judge is later asked to score. Acceptance is a poor stand-in for novelty. A paper can be accepted while reviewers still call the contribution incremental. Several venues no longer publish a novelty score at all.

If the ruler moves when the prompt moves, novelty gains cannot be compared across systems.

Method

Labels come from ICLR 2026 reviews on OpenReview, not from accept or reject decisions. claude-opus-4-6 kept only sentences that explicitly call the contribution original, or explicitly call it unoriginal. Generic praise is dropped, as is any line where the reviewer says they cannot tell. The high-novelty pool has 154 accepted papers: mean rating in the top 10% of the primary area and at least 6, mean contribution score at least 3, a majority of positive novelty signals, and no negative signal. The low-novelty human pool has 145 rejected papers with every condition flipped: bottom 10% and at most 3.5, mean contribution at most 2, majority negative, zero positive. Metrics and baseline wins are stripped from abstracts so a judge cannot shortcut through experimental numbers.

The generated pool is one pass from claude-sonnet-4-5, given only the area name, with no papers and no tools. The model is more than a year old and was chosen because it is not a frontier system, which keeps the weak label conservative. That label is a pool-level assumption: every generated idea is less novel than every high-novelty human idea. A professor blindly judged 30 pairs, 28 of them errors by opus-4-6 under the reference setup. In about five hours the professor picked the human idea all 30 times. The check asks only whether the assumption fails in an obvious way.

Both setups share the 154 high-novelty ideas. Human-Only uses the 145 human papers on the low side. Human+Generated uses 154 generated ideas. Pointwise judging is a binary novel-or-not call. Pairwise judging asks which idea is more novel inside the same ICLR area. Each setup has 154 pairs, and 299 versus 308 pointwise items.

The six judges are gpt-5.1, gpt-5.2, gpt-5.4, sonnet-4-5, opus-4-5, and opus-4-6. Training cutoffs predate the submission deadline. Reviews became public on 2025-11-12, about two months later, so the novelty verdicts sit outside the training data. The reference setup uses a high reasoning budget and no retrieval. Pointwise decisions take a majority of 3 votes. Pairwise decisions take 3 votes in each presentation order, and an equal score is a tie. Strict accuracy counts a tie as wrong. Soft accuracy counts it as half a point.

Each run changes one knob. P1 drops the guardrail against treating a mashup of existing modules as novel. P2 drops the novelty definition. P3 tells the judge that top-conference reviewers found exactly one idea novel and asks which one. P4 deletes that premise and asks which idea is novel. P5 keeps the premise and asks which was judged more novel. Retrieval adds five related abstracts, cut off at 2025-03-01, about six months before the ICLR 2026 abstract deadline of 2025-09-19, so the idea's own preprint stays out. Other single changes are a low reasoning budget, a two-field purpose-and-mechanism format, no vote aggregation, and regenerating the low side with opus-4-5, gpt-5.1, or gpt-5.4. Purpose-built comparators are Idea Novelty Checker and the AI-Scientist retrieval agent, plus gpt-5.6-sol searching arXiv with no query cap at reasoning effort xhigh.

Results

Reference pairwise strict accuracy on human papers sits between 74.0 and 83.8. Even this easy split, opposite ends of the ICLR rating range, stays far from perfect.

ModelHuman-OnlyHuman+Generated
sonnet-4-574.039.6
opus-4-579.263.0
opus-4-682.586.4
gpt-5.183.863.0
gpt-5.283.876.0
gpt-5.483.177.9

Generated ideas pull most judges down. sonnet-4-5 falls from 74.0 to 39.6, below chance. opus-4-5 and gpt-5.1 both fall to 63.0. opus-4-6 is the exception at 86.4.

Prompt wording moves the number more than the choice of model. On generated pairs, gpt-5.4 scores 93.5 with P3 and 40.9 with P4, a gap of 52.6 points. The prompts see the same pairs, so they disagree on at least 52.6% of them. Opposite verdicts cannot both be right, and the gap does not depend on the weak labels. The same swap on human pairs is about 7 points (83.1 versus 76.0). P3 raises all six judges, and opus-4-6 reaches 100.0. P1 hurts all six on generated pairs and drops sonnet-4-5 to 18.2. On human pairs that edit costs sonnet-4-5 9.7 points, and the other models move by at most 2.

Retrieval adds little. The largest significant gain is +15.6 for opus-4-5 (63.0 to 78.6), still smaller than switching to P3, while gpt-5.4 drops 5.2. Unbounded arXiv search with gpt-5.6-sol reaches a pointwise macro-F1 of 0.78, about 11 points behind opus-4-6 given five pre-retrieved abstracts (0.889), at about 10 times the cost per decision. Retrieved papers can also flip a correct call, recasting a combination reviewers treated as novel as a stitching of known ingredients.

Lowering the reasoning budget from high to low costs at most about 7 points, and some models improve. Ties are common. opus-4-6 ties on 35% of pairs under some setups, sonnet-4-5 on more than half. Those two tie about twice as often on generated pairs as on human pairs. Dropping vote aggregation can more than double the tie rate. Strict accuracy and soft accuracy do not rank systems in the same order.

The dedicated pipelines are less accurate. With opus-4-6, the cheapest prompted setup beats Idea Novelty Checker and AI-Scientist by 8 to 29 macro-F1 points and costs up to about 32 times less. With gpt-5.4 every prompted setup also wins, at about 6 times lower cost. Rewriting every idea into the same two-field plan does not calm the swings. Some edits flip more than two thirds of the verdicts. Regenerating the low side with stronger models still leaves some judges below chance.

Why it matters

Novelty gains in ideation papers sit on this kind of judge. One sentence moves gpt-5.4 by 52.6 points and can leave another backbone almost unchanged. A check on human papers shows only about 7 points of that movement. Generated ideas, the setting where the judge is actually used, show the 52.6.

Search, longer reasoning, custom novelty models, and agentic retrieval do not stabilize the call. On this data the cheapest prompt is the more accurate one. An LLM novelty number needs a sweep over prompts and backbones on generated ideas, with ties reported on their own. As published, these scores do not support comparisons between ideation systems.

Limitations

The generated side is not item-level human annotation. The 30-pair check uses one expert, and 28 pairs were picked because the strongest judge got them wrong. Absolute accuracies have to be read with that pool-level assumption. When two configs disagree on the same pair, the disagreement does not need the label. That is the firmer result.

P3 tells the model that exactly one idea was judged novel. opus-4-6 then scores 100.0 on generated pairs. The perfect score may be recognizing a conference abstract against a one-pass generation. The shared plan format removes some style shortcuts, and the instability remains, but the paper does not show whether that perfect score survives the rewrite. P3 and P4 are different instructions. The claim is that the gap is large enough to change the conclusion.

Reviewers read full papers. Judges read abstracts with the numbers removed. The benchmark keeps only the top and bottom deciles where reviewers agree, so close calls between similar ideas are unmeasured. Instability on the easy extremes makes the middle look harder, but the accuracies here do not transfer to that comparison. Signal extraction also uses claude-opus-4-6. The venue is only ICLR 2026, and the judges come from two closed vendors.

Terms

Source

What people are saying

Related papers

All paper explainers