UCLA-Affiliated Paper Says AI Agent Solved 10 Research Math Problems — But Proofs Were Accepted Only by an LLM Verifier

·

A UCLA-affiliated author group has posted an arXiv paper claiming its AI math agent, called Ansatz, solved all 10 problems in the First Proof — Second Batch benchmark and several named open problems. But the authors’ own paper and code release say those results were accepted by an LLM-based verifier, not independently confirmed by outside mathematicians or checked in a formal proof system such as Lean or Coq.

The preprint, “Continual Graph Memory for Mathematical Research Agents,” was posted to arXiv as 2610.02945v1 on Oct. 2. It lists Junyi Zhang, Jinxi Yu, Eric Hanchen Jiang and numerous coauthors, including mathematician Terence Tao, with the University of California, Los Angeles listed as the affiliation in the paper. The claim is notable because First Proof — Second Batch, introduced in a June 10, 2026, report, is a benchmark built around 10 research-level math problems whose proofs were chosen so they were not publicly discoverable online. A genuine 10-for-10 result there would mark a significant milestone for AI-assisted mathematical research.

The paper describes Ansatz as a system built around “Continual Graph Memory,” a graph-based memory method meant to store intermediate proof artifacts, retrieve them later and reuse them across long-running proof searches. In experiments, the paper reports “closure on all ten research tasks” in First Proof — Second Batch. It also claims fully autonomous solutions to the Jamison caterpillar conjecture and Erdős Problems 289, 348 and 488, along with partial progress on other open problems.

The authors also released code and supporting materials in a GitHub repository, uclanlp/continual-graph-memory. The release includes implementation details, an annotated result index, run materials and 13 manuscripts intended to support the reported outcomes.

The central caveat is verification. In mathematics, stronger evidence usually comes either from independent expert refereeing or from a machine-checkable formal proof in software such as Lean or Coq. The paper explicitly says it is not making that kind of claim. “Throughout the paper, ‘accepted’ means accepted by an LLM verifier, not checked by a proof assistant,” the authors wrote in the arXiv paper.

The GitHub repository makes the same point in even broader terms. “Here, accepted means accepted by the configured model-based review process. It does not imply a formal proof-assistant certificate or independent expert validation,” the README says. The repository also says, “No new mathematical referee review was performed for this release.”

That distinction matters because a benchmark result or claimed theorem solution can look much stronger than it is if “solved” means only that another language model judged the proof acceptable. An LLM-based verifier may provide a useful internal check, but it is a materially weaker standard than outside peer review or formal certification in a proof assistant.

The released materials also include a reproducibility caveat. While the current paper reports a 10-for-10 result on First Proof — Second Batch, the repository says an earlier archival paper PDF and run snapshot included in the release recorded a 9-for-10 local-verifier result, with task 03 still open and a different model configuration. The release notes state: “The current paper source reports 10/10 outcomes on First Proof Second Batch. The supplied archival paper PDF and supporting run snapshot record an earlier 9/10 local-verifier outcome, with task 03 open and a different recorded model configuration. These are distinct records; this release does not establish that the newer claim has been reproduced from the older archive.”

Taken together, the paper and repository make a set of public claims that, if later validated, would be significant for AI-driven mathematical research. But the same materials also state that the reported acceptances come from a model-based review process, not from Lean- or Coq-style formal proof checking, and that no independent mathematical referee review has yet been provided.

Tags: #ai, #mathematics, #arxiv, #theorem-proving