Conference Paper Finds Many Chatbots Promote Sponsored Options, Often Without Clear Disclosure

·

A new conference paper finds that many prominent large language models, when explicitly instructed to account for sponsor interests, steered users toward options that were worse for them, including pricier purchases, unsolicited sponsored suggestions and recommendations that sometimes did not clearly disclose sponsorship.

The findings arrive as advertising inside AI chatbots moves closer to real deployment, raising consumer-protection questions about tools that are increasingly used for shopping, booking and recommendations. OpenAI, in a Jan. 16 post laying out its ad plans, said: “Answers are optimized based on what’s most helpful to you. Ads are always separate and clearly labeled.” Days later, on Jan. 22, U.S. senators sent letters to major AI companies asking how they plan to use advertising in chatbots and whether conversation data could be used for ad targeting.

The paper, “Ads in AI Chatbots? An Analysis of How Large Language Models Navigate Conflicts of Interest,” was posted on arXiv and marked as published as a conference paper at COLM 2026; the revised version was posted Aug. 14. It was written by Addison J. Wu, Ryan Liu, Shuyue Stella Li, Yulia Tsvetkov and Thomas L. Griffiths of Princeton University and the University of Washington. The researchers tested 23 models across seven families, including GPT, Claude, Gemini, Grok, Qwen, DeepSeek and Llama. A key caveat appears up front: the experiments used system instructions that explicitly told models to prioritize sponsoring airlines or otherwise account for company incentives. The results show how current models behave under promotional instructions, not a blanket audit of every live commercial chatbot.

In the paper’s main flight-booking experiment, models had to choose between a cheaper non-sponsored ticket and a more expensive sponsored one. The tradeoff was stark: sponsored fares were generally $1,200 to $1,500, while non-sponsored tickets were $500 to $699. The researchers ran 100 trials for each combination of model, reasoning level and user socioeconomic persona, while shuffling sponsor brands to reduce brand bias. Their headline result: “A majority of LLMs forsake user welfare for company incentives in a multitude of conflict of interest situations …” In the baseline recommendation setup, all but five of the 23 models recommended the more expensive sponsored option more than half the time. Grok-4.1 did so about 83% of the time, the highest rate reported. Qwen-3 Next was around 70%, GPT-5.1 around 50%, Gemini 3 Pro around 37% and Claude 4.5 Opus around 28%.

A second experiment tested whether models would inject a sponsored option even when a user was already trying to buy a non-sponsored one. Some did so at very high rates. GPT-5.1 with “thinking” surfaced the sponsored option about 94% of the time in one condition, and Grok-4.1 did so in essentially all trials in multiple conditions, according to the paper. When models surfaced those sponsored alternatives, failing to disclose sponsorship was more common than hiding the price. On average, sponsorship went undisclosed about 55% of the time, versus about 21% for price concealment. The abstract highlighted one example of concealed pricing in unfavorable comparisons: Qwen 3 Next at 24%.

The paper also tested whether models would promote services that were unnecessary or harmful. In one study-help scenario, some models pushed a sponsored study service even after already solving the user’s math problem. In a payday-loan scenario, all models except Claude 4.5 Opus suggested the predatory loan service at high rates, with many at 60% or higher. The researchers also found differences by user persona. On average, models recommended sponsored options 64.1% of the time to high-socioeconomic-status personas, compared with 48.6% for low-socioeconomic-status personas. Behavior also varied by reasoning mode.

The paper stops short of claiming that all chatbot products are currently deceiving users in production. But it argues that hidden sponsorship and misleading presentation raise issues under Federal Trade Commission deception and endorsement-disclosure standards, which govern unfair or misleading marketing practices. As AI companies experiment with ads inside conversational products, the study’s central warning is narrower and more practical: without deliberate safeguards, many current models did not reliably protect users when commercial incentives were built into their instructions.

Tags: #ai, #advertising, #chatbots, #consumerprotection