GPT-5.6 Reportedly Found a BH Counterexample

Multiple reposts, along with a Chinese article by Xinzhiyuan, said GPT-5.6 helped advance a long-standing open problem in statistics by constructing a key counterexample. The question concerns whether the Benjamini-Hochberg (BH) procedure still always controls the false discovery rate (FDR) under correlated two-sided Gaussian tests, making the claim notable because BH is a basic tool across statistics and genomics.

Key details

The posts recap that Benjamini and Hochberg introduced the BH method in 1995 and proved FDR control under independence, which is why it became widely used. The open problem discussed here is more specific: whether BH continues to control FDR in the correlated two-sided Gaussian setting. Xinzhiyuan wrote that discussions involving Wharton statistician Edgar Dobriban said GPT-5.6 constructed a counterexample in about 90 minutes. Several reposts also added that the classic BH paper has been cited more than 130,000 times.

Interpretations and caveats

Many of the posts framed this as an example of AI contributing to math and statistics research. One repost said the underlying author described it as one of the most interesting long-unsolved problems in that area of statistics. Another repost summarized the argument that LLMs may be more likely than humans to combine known concepts in nonstandard ways and thus generate valid counterexamples humans would not try first.

At the same time, the material in this cluster is largely second-hand. Based on the posts alone, the safest characterization is that GPT-5.6 was described as helping push forward the resolution of this statistical open problem, rather than treating every stronger claim as independently verified.

2026-07-15 ~ 2026-07-16 · 8 related posts

1 near-duplicate retellings: romainhuet