A new study found today's frontier AI agents could complete many of the engineering tasks required for AI research but failed to produce original work worthy of acceptance at a top machine learning conference.
“Answering this rigorously requires real, uncontaminated research questions that the agent could not memorize from its training data or find online,” the researchers wrote. “To satisfy these requirements, we rely on high-quality AI research that was not public at the time we conducted the experiments.”
The agents completed much of the engineering required for research, conducting literature reviews, debugging software, running experiments, managing GPU resources, and producing complete academic papers without human intervention. But reviewers concluded the systems failed to generate original scientific contributions worthy of publication at a top machine learning conference.
The authors said their evaluation better measures scientific reasoning than previous benchmarks because it tests open-ended research problems rather than predefined tasks.
The authors cautioned that the study examined only two research projects and acknowledged limitations, including the small sample size and the fact that the original researchers evaluated the AI-generated papers. They said the results suggest current frontier AI agents can automate many of the engineering tasks involved in research but continue to struggle with generating original scientific work.
The study comes as researchers continue to uncover surprising and sometimes risky behaviors in increasingly autonomous AI agents.


















