TLDR: Generative search engines (GEs) prefer citing predictable and semantically similar content, a tendency rooted in their underlying large language models (LLMs). Surprisingly, LLM-based content polishing increases the information diversity within AI summaries by making more content accessible to LLMs. This leads to different user benefits: highly educated users gain efficiency, while less educated users benefit from richer information in their task outputs.
In the rapidly evolving digital landscape, search engines are undergoing a significant transformation. The emergence of generative search engines (GEs), powered by large language models (LLMs), is fundamentally changing how we find information online. Unlike traditional search engines that primarily list relevant websites, GEs synthesize information from multiple sources to deliver AI-generated summaries, complete with website citations. This shift creates new avenues for traffic acquisition and redefines the rules of search engine optimization (SEO).
A recent research paper, titled “When Content is Goliath and Algorithm is David: The Style and Semantic Effects of Generative Search Engine,” delves into the distinctive characteristics of these new AI-driven search platforms. Authored by Lijia Ma, Juan Qin, Xingchen (Cedric) Xu, and Yong Tan, the study explores how generative search engines select content to cite and the broader implications for both content creators and users. You can read the full paper here: Research Paper Link.
How Generative Search Engines Choose What to Cite
The research reveals that generative search engines exhibit clear preferences when selecting content for their AI summaries. Firstly, they favor content that is highly “predictable” for the underlying LLMs. This predictability is measured by a concept called perplexity – lower perplexity means the LLM can more easily predict the next word in the text. This preference stems from the intrinsic way LLMs generate language, by sequentially selecting tokens with high conditional probability.
Secondly, the study found that content cited by LLMs shows greater “semantic similarity” among selected sources. While AI summaries might appear to draw from diverse, niche websites, the chosen sources often convey very similar meanings. This is because GEs are designed to create coherent, unified responses, rather than simply aggregating disparate pieces of information.
The Intrinsic Nature of AI Preferences
A key question addressed by the researchers was whether these citation criteria are specifically engineered for generative search platforms or if they naturally arise from the core language model architecture. Through controlled experiments using Retrieval Augmented Generation (RAG) APIs with Google’s Gemini model, the study demonstrated that these stylistic and semantic preferences are indeed intrinsic to the underlying LLMs. This means the models inherently prefer predictable content and semantically similar sources, regardless of specific search engine engineering.
Interestingly, the experiments also uncovered a “positional bias” in RAG systems, where content placed earlier in a document is more likely to be cited. This offers a practical insight for website owners looking to optimize their content for generative search.
The Surprising Effect of LLM-Driven Content Polishing
With the widespread adoption of LLMs for content creation and refinement, the researchers investigated how “polishing” website content with LLMs might affect AI summaries. One might expect such polishing to lead to content homogenization, making everything sound similar. However, the study found a paradoxical effect: automated content refinement actually enhances the “information diversity” of AI summaries.
This happens because LLM-based polishing makes previously less predictable content more accessible and interpretable for language models. By reducing perplexity, more content becomes eligible for citation, leading to a broader range of sources being included in AI summaries and thus increasing overall information diversity. This effect was even more pronounced when content was polished with an explicit objective to increase citation likelihood.
User Experience: Who Benefits How?
To understand the real-world impact of these changes, the researchers designed a generative search engine and conducted a randomized controlled experiment with human participants. Users were tasked with an information-seeking and writing assignment, with half receiving AI summaries based on original website content and the other half receiving summaries based on LLM-polished content.
The results revealed differential benefits across users with varying educational backgrounds. Participants with higher education (graduate level) showed minimal changes in the information diversity of their final outputs but demonstrated significantly reduced task completion time. They adapted by issuing fewer queries when presented with optimized, information-dense summaries, thus gaining efficiency.
Conversely, participants with lower education (undergraduate or below) primarily benefited from enhanced information density in their task outputs, while their completion times remained similar across groups. These users tended to rely more on the immediately available information in the AI summaries, directly benefiting from the increased content diversity provided by the polished sources.
Also Read:
- AI’s Hidden Hand: Uncovering Bias in LLM-Assisted Academic Peer Reviews
- Bridging the Information Gap: How AI Can Enhance Mental Health Support
Implications for the Evolving Search Ecosystem
This research offers valuable insights for all stakeholders in the search ecosystem. For website owners and those involved in Generative Engine Optimization (GEO), understanding the intrinsic preferences of LLMs for predictable and semantically coherent content is crucial. Leveraging LLMs for content refinement can be a “win-win,” expanding website visibility in AI summaries while also enhancing the diversity of information presented to users.
For users, it highlights that while AI Overviews offer enhanced search efficiency, their design for coherent responses might lead to a narrowed range of semantic perspectives. Comprehensive information gathering might still require exploring original sources. Finally, for search engine operators, the study underscores the need to carefully consider the relationship between RAG APIs and AI Overview functionalities, and the potential for strategic manipulation by content creators.


