A new multi-institution study has found that today's most advanced AI agents can handle the routine mechanics of research — but they can't produce original work that would get accepted at a top AI conference. The study tested frontier AI systems on the full research pipeline, from literature review to paper writing.
What the AI got right
The AI agents successfully performed tasks like summarizing existing papers, generating code for experiments, and formatting results. They could mimic the structure of a research paper and even produce plausible-sounding citations. The mechanics of research, the study found, are within reach.
Where the AI fell short
But when it came to generating novel ideas or making a genuine contribution to the field, the AI failed. The study's evaluators judged the AI-generated papers as lacking the originality required for acceptance at a top-tier AI conference. The agents could not produce work that advanced the state of the art.
What the study means
The findings suggest that while AI can be a powerful tool for researchers, it is not yet capable of independent scientific discovery. The study was conducted by researchers from multiple institutions, though the specific names were not disclosed. The results add to a growing debate about the role of AI in science.
The study highlights a clear gap between AI's ability to execute and its ability to innovate.

