Summary
An observational study examined around 2 million citations from four commercial large language models across 10,000 web pages to identify features influencing citation frequency. The research covers data from ChatGPT, Claude, Google AI, and Gemini over six months.
AI-assisted summary based on the listed source.
What happened
Production large language models retrieve and cite web pages alongside generated answers, yet the page-level features that predict citation frequency remain poorly characterised. We present an observational study of approximately 2 million LLM citations from four commercial engines (ChatGPT, Claude, Google AI,...
Why it matters
Understanding what drives citation frequency helps clarify how LLMs source and prioritize information, which is crucial for evaluating their reliability and transparency. This insight can guide improvements in AI-generated content attribution.
Signal Intelligence
Signal Strength 95%
Technical label SOURCE-BACKED
Public Interest 42
Category RESEARCH
Reader Depth TECHNICAL
Signal Strength reflects source quality, relevance, freshness and evidence. Public Interest helps organize discovery; it is not proof of truth.
Public Interest components
Recognizable Entity Score 68
Practical Impact Score 20
Novelty Interest Score 70
Consequence Score 18
Curiosity Score 0
Shareability Score 40