Revised preprint from Caltech, Google and MIT claims exponential quantum advantage for classical datasets — scrutiny continues
A team spanning Caltech, Google Quantum AI, MIT and Oratomic has posted a revised arXiv preprint claiming an exponential quantum advantage for machine-learning tasks on massive classical datasets, a notable assertion in a field where broad, practical quantum advantage has been difficult to establish. But the claim is still unsettled: The earlier version drew a detailed public technical audit, and the paper has not yet been peer-reviewed.
The preprint, “Exponential quantum advantage in processing massive classical data,” is listed on arXiv as arXiv:2604.07639. It was first submitted April 8, 2026, and a revised version, v2, appeared Oct. 1. The authors — Haimeng Zhao, Alexander Zlokapa, Hartmut Neven, Ryan Babbush, John Preskill, Jarrod R. McClean and Hsin-Yuan Huang — say the result would apply to ordinary classical data rather than a narrowly quantum-native problem. In the abstract, they write that “we prove that a small quantum computer of polylogarithmic size can perform large-scale classification and dimension reduction on massive classical data.” They also claim “four to six orders of magnitude reduction in size with fewer than 60 logical qubits.”
That would be a striking result if it holds. The paper says a comparatively small quantum machine could process data samples as they arrive and carry out two common machine-learning tasks — classification, which assigns inputs to categories, and dimension reduction, which compresses data while preserving useful structure. Comparable classical machines, the authors argue, would need exponentially larger size for the same job. They attribute the advantage to a method they call quantum oracle sketching, used with classical shadows, which they say sidesteps the usual quantum-machine-learning bottlenecks around loading large classical datasets into quantum form and reading the results back out.
Those bottlenecks are a big reason this paper is getting attention. In quantum machine learning, many earlier speedup claims ran into a basic problem: Even if a quantum algorithm looked fast on paper, the cost of feeding ordinary data into the system could wipe out the benefit. This work matters because it explicitly claims to bypass that issue.
The authors say they found supporting evidence on real datasets, including PBMC68k, a single-cell RNA sequencing dataset, and IMDb movie-review sentiment analysis. They have also released a public GitHub repository, haimengzhao/quantum-oracle-sketching, with code, notebooks and scripts for the figures and dataset experiments described in the paper. Zhao published a plain-language blog post on Quantum Frontiers on April 9 summarizing the work and pointing readers to the code.
Still, the public record around the paper is already unusually contested for such a high-profile claim. Carmelo Vellón Gascón of GatePhys published a public audit of the original v1 manuscript, arguing that it contained proof and interface problems. The GatePhys repository README says that “its central result is an explicit finite counterexample to the sufficient threshold printed in Theorem D.16 / Eq. D.99.” That critique was directed at v1, not necessarily at the newly posted v2. But the timing matters: v2 was posted after months of public criticism of v1, and as of now there is no documented independent public assessment of v2 in the audit repository.
That kind of scrutiny is standard in this corner of the field. Quantum machine learning has a history of bold claims later being narrowed, or challenged by so-called quantum-inspired classical methods and dequantization results that reproduce the supposed advantage without a quantum computer. Because of that track record, researchers pay close attention to the assumptions behind how algorithms access data and what resources they actually require.
So the state of play is unusual but clear. A heavyweight team has released a revised preprint claiming an exponential quantum advantage for processing massive classical datasets, along with code and concrete examples. At the same time, the earlier version faced a detailed public challenge, the revised paper is not yet peer-reviewed, and independent vetting of the new version has not yet fully played out in public. For a field long constrained by data-loading debates, that makes this a consequential claim — and one still very much under examination.