Audit Finds Google AI Overviews Cite Sources Outside Top Results, 11% of Claims Unsupported
A large independent audit of Google’s AI Overviews found that the AI-generated search summaries often draw on sources users do not see in the ordinary first page of results, and that about 11% of the summaries’ atomic factual claims were not supported by the pages they cited. The findings matter in part because Google has said AI Overviews now reach more than 2 billion users.
The paper, “Measuring Google AI Overviews: Activation, Source Quality, Claim Fidelity, and Publisher Impact,” was written by Haofei Xu, Umar Iqbal and Jacob M. Montgomery of Washington University in St. Louis. It was first posted to arXiv on May 13, 2026, revised Oct. 1, and is marked as accepted to the 2026 ACM Internet Measurement Conference, a leading venue for internet-measurement research, scheduled for Oct. 12-16. In the paper’s abstract, the authors describe AI Overviews as “arguably the most widely encountered deployment of generative AI.”
AI Overviews are Google’s AI-written summaries that appear above some search results, replacing the older pattern in which users mainly saw ranked links and had to assess sources themselves. Google has said the feature is “built to only surface information that is backed up by top web results.” The new study set out to test that claim using 55,393 trending Google queries across 19 topical categories collected over 40 days, from March 13 through April 21, 2026.
Of those queries, 7,583 triggered an AI Overview, for an overall activation rate of 13.7%. Query format made a large difference: question-form searches triggered AI Overviews 64.7% of the time, compared with 9.5% for non-question searches. Activation also varied sharply by topic, with especially low rates in politically sensitive areas such as Politics, at 7.5%, and Law & Government, at 9.6%.
One of the study’s clearest findings was that AI Overviews were not simply summarizing the same links users saw in the top search rankings. Only 31.5% of AI Overview reference domains overlapped with the top five first-page results for the same query. The overlap rose to 70.2% when measured against the full first page, meaning 29.8% of cited domains did not appear anywhere on that page.
That finding cuts both ways. On one hand, it suggests users cannot assume the summary is based only on the links placed in front of them in ordinary search results. On the other, the authors found that AI Overview-cited domains scored, on average, as more credible than the co-displayed first-page result pool, based on the study’s credibility metrics.
The paper also examined whether the summaries’ claims were actually supported by the cited sources. The researchers broke 7,491 verifiable AI Overviews into 98,020 atomic claims — individual factual assertions that could be checked one by one. Their verification pipeline, which was LLM-based and validated against human annotators, labeled 89.0% of claims as consistent with the cited pages and 11.0% as inconsistent.
That inconsistent category does not mean all of those claims were flatly false. The largest failure mode was omission: about 6.98% of claims were not addressed by the cited pages at all. Another 2.66% were labeled incorrect because the cited source contradicted the claim, while 1.39% were marked ambiguous. The authors said their automated system showed high agreement with human reviewers; after adjudication, human review matched the pipeline on 98 of 100 sampled verdicts.
The paper also points to potential effects on publishers. At least 50.63% of pages cited in AI Overviews carried display advertising, a figure the authors describe as a lower bound because some pages could not be crawled or may have used monetization methods the study could not detect. If AI Overviews answer users’ questions without a click, the paper argues, publishers can lose traffic and ad revenue even when their reporting or reference material is used in the summary.
That publisher question is already under scrutiny in Europe and the U.K., where regulators have examined Google’s use of publisher content and controls around AI-generated search summaries.
The authors include several important caveats. Their sample was limited to trending queries, not all Google searches, so the 13.7% activation rate should not be generalized across the full search universe. The factuality checks relied on an automated pipeline, though one the paper says was validated against human annotators. And while the project links to a GitHub repository, the README says the dataset “is currently being prepared and will be released in a future update,” limiting immediate reproducibility.
Stocks: GOOGL