Academic Non-Profit Consortium Finds the Outer Limits of AI – Humans Beat AI Badly in Complex Problem Solving – “It’s Discovery 101”
Humans solved 70 out of 70 puzzles in study. World’s best AI model solved 50 As problems got more difficult, AI
Press Release Disclaimer: This is a press release distributed through the XPR Media network. It has not been independently verified by our newsroom.

![]()
A new study by a consortium of some of the world’s leading neuroscientists and academics found that the people were twice as good at complex problem solving than AI and five times better on the really difficult problems – a finding that has major implications for AI’s utility in solving complex medical and scientific problems.
During the DiG-Bench competition conducted in July 2026 DiG-Bench asked humans and AI to solved 70 complex, real world text-based puzzles. The humans solved 70 out of 70, but the best AI frontier model tested solved only 50 when they were confronted by puzzles without having problem-solving rules provided to the AI models.
As the problems got successively harder through seven tiers of increasing difficulty, AI performed even more poorly. While humans were able to solve all 20 of the most challenging problems, AI could only solve eight.
The players, whether human or machine, were only able to discover the rules and solve the problems by running experiments and revising their hypotheses when they failed.
Opus 5 from Anthropic solved 50 of these environments. Of the other models measured, GPT-5.5 solved 18, Gemini 3.1 Pro solved 16 and DeepSeek V4 solved four.
“These results challenge some industry claims that current AI models are capable of autonomous scientific discovery in complex, real-world settings,” Dr. James C.R Whittington, one of the two co-leaders of the study, said. “The research demonstrates that frontier models struggle to discover rules even in simple text-based games that humans easily solve within an hour—despite text being language models’ natural domain.
“Human Intelligence could do this much better than the Artificial Intelligence models. They could adapt, improvise, test, and overcome,” he said. “Even the best models couldn’t match the people.”
“Accelerating science with AI, particularly AI that can autonomously come up with and test good hypotheses, is one of the few unequivocal places AI can do good,” said Dr. Ruairidh M. Battleday, the other lead author on the published study, DiG: Discovery in Games
“This has really important scientific implications,” he said. “If we are to find the Rosetta Stone to solve the thorniest scientific and medical problems we face, AI isn’t there yet. Our competition revealed that current frontier models fail at this even in simple text environments, which should theoretically be well within their domain. This is Discovery 101.”
DiG-Bench asked competitors to solve 70 environments in total, but only 21 of them are public and can be played on the website by anyone with access to API , hosted at digbench.ai/. The other 49 have not been released; no model or human can be trained on them and therefore no result can be produced without going through the same experimental procedure.
About the Study
The complete methodology has been published alongside the results in DiG: Discovery in Games (2026), a study led by Drs. Battleday and Whittington and its authors include influential AI and cognitive science researchers Drs. Jürgen Schmidhuber, Joshua Tenenbaum and Thomas L. Griffiths, alongside other researchers from Thinking About Thinking, the University of Oxford, Inria, MIT and Princeton University. DIG.bench and the accompanying paper are available at https://digbench.ai/.
View source version on businesswire.com: https://www.businesswire.com/news/home/20260922483083/en/
Media gallery

