Large language models (LLMs) serving as physician recommendation tools prioritize a doctor's reputation over demographic characteristics, according to an algorithm audit published on arXiv cs.CL. The research indicates that factors like patient ratings significantly influence which doctors these AI systems suggest, a dynamic that could profoundly affect physician visibility and patient choice. As patients increasingly consult LLMs for healthcare decisions, these AI systems are becoming central to how individuals select their medical providers.

The study, titled "Whose doctor does the AI recommend? An algorithm audit of reputation and demographic signals in large language model-assisted physician choice," utilized a randomized algorithm audit methodology. Researchers evaluated seven different LLMs, including GPT-4o-mini and six open-weight models. Each model was presented with 3,024 unique choice sets, each containing five synthetic family medicine physician profiles. These profiles had independently randomized attributes, including signals for gender and ethnicity conveyed through names, following established correspondence-audit practices. The audit also incorporated three distinct patient personas, nine prompt paraphrases, and nine experimental arms, resulting in a total of 40,068 scored responses.

The primary finding was that reputation signals exerted the strongest influence on the LLMs' recommendations. For instance, raising a physician's rating by one point led to a substantial increase in the likelihood of that physician being recommended. In contrast, demographic signals, such as the perceived gender or ethnicity of the physician, had a considerably smaller impact on the recommendation outcomes. This suggests that while LLMs are designed to process and respond to a wide array of information, their decision-making process in this context heavily leans on quantitative measures of professional standing.

The authors characterize LLMs in this role as "AI infomediaries." This term highlights the algorithms' function in intermediating a patient's choice among various healthcare providers, thereby silently influencing which physicians gain visibility and, consequently, patient流量. This mediation occurs at scale, potentially affecting a large number of patient-provider matches without explicit human oversight in the recommendation logic.

The increasing reliance on AI tools for healthcare decisions is a broader trend. A 2025 report by rater8, which surveyed over 1,000 U.S. adults, found that 70% of patients are open to or already using AI tools to research physicians. Among these, 26% reported that AI recommendations directly influenced their decision, a figure nearly comparable to the influence of primary care referrals (28%) and healthcare review sites (29%). The same report also noted a growing trust in AI-generated results, with one-third of respondents trusting AI search outcomes as much as traditional search engines.

However, the authenticity of information remains a key concern for patients. When asked about valued information in AI summaries, 40% of respondents cited verified patient reviews, surpassing physician credentials or convenience factors. This emphasis on verified reviews aligns with the audit's finding that reputation signals are paramount in LLM recommendations.

Previous research has also explored how AI can assist in doctor selection, though with different focuses. One study in 2025 introduced an algorithm called MDE-HYB, designed to appraise doctors based on their track record of accurate diagnoses, particularly when success rates are not publicly available. This algorithm aimed to reduce misdiagnosis rates by 41% compared to other selection methods. Another paper in May 2026 examined the ethical values embedded in LLMs providing medical advice, noting that while models discuss competing values, their individual decisions are often deterministic, potentially underweighting patient autonomy.

The current audit's findings underscore the importance of understanding the underlying mechanisms of AI recommendations, particularly as these systems become more integrated into critical decision-making processes like healthcare provider selection. While reputation signals may appear to be a neutral criterion, the methodologies for collecting and presenting these signals can themselves carry biases or limitations. The study did not specify the exact nature of the "reputation signals" used beyond "rating," leaving open questions about the specific components of reputation that LLMs prioritize.

The researchers did not elaborate on whether the LLMs demonstrated any biases in how they interpreted or weighted different aspects of reputation. Further investigation could explore the granularity of these reputation signals and how they are synthesized by the AI to form a recommendation. As LLMs continue to evolve as "AI infomediaries," transparency in their recommendation algorithms will become increasingly important for both patients and healthcare providers.