AI agents can run experiments and draft papers but fail to produce publishable research, study finds

·

Frontier AI agents were able to handle much of the mechanics of AI research in two test cases — writing code, running experiments and drafting papers — but failed to produce publishable advances on the underlying research questions, according to a new arXiv preprint.

That finding matters because many claims about fast-moving AI progress rely on the idea that AI systems will not just help with routine programming work, but also automate AI research itself and speed up further development. The new paper suggests current systems may be better at executing research workflows than at generating the kind of original insight needed for top-tier research.

The preprint, titled “Can AI agents conduct open-ended AI research? Early evidence from two case studies,” was posted to arXiv on July 29, with a PDF dated July 30. It was produced under the CRUX project, with Peter Kirgis and Sayash Kapoor listed as equal-contributing lead authors among a larger multi-institution group that includes Princeton University affiliates.

The authors tested what they call “shadow evaluations,” meant to measure AI on open-ended research rather than narrow benchmark tasks with clear right answers. In the method, an AI agent receives the central question from a high-quality unpublished paper, then works independently with web and computing access. Afterward, the original paper’s authors grade the AI-generated result as if it were a conference submission. The researchers present that as a middle ground between standard benchmarks, which can miss the messiness of real research, and submitting AI-written papers to blind peer review, which they argue can be noisy and difficult to interpret.

The team ran the evaluation on two unpublished NeurIPS 2026 submissions. In the main runs, they used the OpenClaw scaffold with Opus 4.8. Each run got about six days — described in the paper as a 120-hour deadline plus a 24-hour extension — along with $3,000 in API credits per paper, plus GPU and virtual-machine access.

According to the paper, the agents completed major engineering tasks without human help, including literature review, environment and GPU debugging, running hundreds of experiments, compiling LaTeX papers, and producing code and reproducibility scripts. As the abstract puts it, the systems “can do the engineering of AI research, but struggle with critical parts of the research lifecycle.”

Those critical parts were the difference between a competent workflow and a publishable result. The paper says the agents “did not make substantial progress” on the central research questions. The original authors reviewed the resulting papers on a NeurIPS-like 1-to-6 scale and judged both to be rejections, with reported overall scores of 2 out of 6 for one paper and 1 out of 6 for the other. The study describes both as “unambiguous rejections.”

The authors identified five recurring failure modes: poor judgment about the bar for publishable research, uncreative responses to shortcomings in the research design, ineffective backtracking from dead ends, poor resource awareness and instruction drift. A robustness check using a different setup — OpenAI Codex with GPT-5.6 Sol — reportedly reproduced nearly all of the same failure modes.

The paper is explicitly framed as early evidence, not a final verdict. It is a preprint, meaning it has not been peer-reviewed. It also covers just two case studies, and the grading was not blinded because the original authors knew they were reviewing AI-generated work. Still, the researchers say they are releasing the expert reviews, survey responses, agent repositories and run logs, which could allow outside researchers to scrutinize the results and the proposed evaluation method.

Tags: #ai, #research, #machinelearning, #neurips