Can AI Read Scientific Papers Like a Scientist? New Benchmark Reveals LLM Weaknesses

by Anika Shah - Technology
0 comments

Can AI Read Papers Like a Scientist? New Benchmarks Reveal LLM Limitations

To stay up to date and work forward in their fields, scientists must have at their fingertips and in their minds thousands of published studies. Large language models (LLMs) demonstrate promise as a tool for exploring the vast scientific literature, but are they trustworthy when it comes to providing full and scientifically accurate answers to complex questions in specialized fields?

Putting Language Models to the Test

Researchers at Cornell University and Google have evaluated the ability of six LLM systems—including ChatGPT, Claude, and others—to understand scientific literature at the level of a specialist. The study focused on the field of high-temperature cuprates, a class of superconducting materials, as a case study. The findings, published in the Proceedings of the National Academy of Sciences on March 10, 2026, reveal both strengths and gaps in current LLM capabilities.

Benchmark Design and Methodology

The researchers created a database of 1,726 scientific papers curated by human experts covering the history of high-temperature cuprates. They also developed a set of 67 questions, written by a larger group of experts, designed to probe deep understanding of the literature. Four LLMs were examined: ChatGPT-4, Claude 3.5, Perplexity, and Gemini Advanced Pro 1.5. They tested NotebookLM, a Google product designed to answer questions based on provided documents, and a custom retrieval-augmented generation (RAG) system capable of retrieving both text and images from the curated documents.

Which AI Tools Performed Best?

The systems that utilized curated information—Google’s NotebookLM and the custom RAG system—performed the best. According to Haoyu Guo, Bethe/KIC postdoctoral fellow with Cornell’s Laboratory of Atomic and Solid State Physics (LAASP), “LLMs operating on trusted data sources—papers we collected ourselves, not from the LLM searching the Internet—tend to perform better. Among these, NotebookLM performs better when I have a set of papers that I desire to understand better.”

Strengths and Weaknesses Identified

All LLMs demonstrated surprising proficiency in extracting text-based information. However, they were “totally incapable” of effectively engaging with data visualization, according to Eun-Ah Kim, the Hans A. Bethe Professor of physics in the College of Arts and Sciences (A&S) at Cornell. The custom RAG model, with its ability to retrieve images, showed significantly improved performance in understanding data visualization.

Future Improvements for LLMs

Researchers identified several areas for improvement in future LLM development. These include more accurate attribution of claims (reducing instances of fabricated references), enhanced ability to synthesize complex information, and improved comprehension of plots and figures. Guo noted that while models have improved in many aspects over the past year, visual reasoning remains underdeveloped.

Implications for Scientific Research

Kim suggests that trustworthy LLM systems could be valuable tools for young researchers, helping them explore literature and generate creative ideas. She emphasizes that the ability to think creatively and approach problems from new angles is becoming more important than simply memorizing facts. “Knowing the facts used to be brandished as a ticket to the table. Holding a fact in your head should not be the ticket. The ticket should be: Do you understand how to think in a creative way? Can you approach problems from a creative angle?”

Citation: Can AI read papers like a scientist? A new benchmark shows where LLMs fail (2026, March 10) retrieved 10 March 2026 from https://techxplore.com/news/2026-03-ai-papers-scientist-benchmark-llms.html

Related Posts

Leave a Comment