Summary
This study evaluates three large language model chatbots—Claude Sonnet 5, Gemini 3.1 Pro, and ChatGPT GPT-5.5—on their ability to retrieve relevant clinical studies for medical questions. It addresses a gap in prior research by focusing on the quality of retrieved studies rather than citation fabri...
AI-assisted summary based on the listed source.
What happened
Large language model (LLM) chatbots are increasingly used to answer clinical questions with citations to relevant clinical studies. Prior research has largely focused on citation fabrication, leaving a gap in evaluating the quality of retrieved studies and the factors driving their selection. In this study, we...
Why it matters
Understanding how well AI chatbots retrieve expert-level clinical studies is crucial for their reliable use in medical decision-making. This research helps clarify factors influencing study selection by different LLMs, informing their deployment in healthcare.
Signal Intelligence
Signal Strength 95%
Technical label SOURCE-BACKED
Public Interest 32
Category RESEARCH
Reader Depth TECHNICAL
Signal Strength reflects source quality, relevance, freshness and evidence. Public Interest helps organize discovery; it is not proof of truth.
Public Interest components
Recognizable Entity Score 68
Practical Impact Score 0
Novelty Interest Score 48
Consequence Score 18
Curiosity Score 0
Shareability Score 32