Human experts remain essential for checking clinical AI outputs

AI systems delivered consistent, low-cost ratings, but local clinicians still identified important concerns that automated evaluators overlooked.

Study: Human evaluators vs. LLM-as-a-Judge: toward scalable evaluation of GenAI in global health. Image Credit: elenabsl / Shutterstock

Study: Human evaluators vs. LLM-as-a-Judge: toward scalable evaluation of GenAI in global health. Image Credit: elenabsl / Shutterstock

In a recent 'Article in Press' in the journal npj Digital Medicine, researchers investigated the benefits and limitations of automated "LLM-as-a-judge" evaluation frameworks. The study benchmarked the performance of these frameworks against local clinician ratings of AI- and human-generated clinical decision-support responses in Rwanda to evaluate their reliability for scalable assessment of clinical AI outputs in resource-constrained environments.

The study dataset comprised 524 query-response pairs selected from a larger dataset of queries submitted by Rwandan community health workers simulating requests for clinical decision support. Study analyses revealed that while AI judges demonstrated high internal consistency, even the top-performing model matched these local ratings on only 4 of 11 evaluation criteria.

Critically, all evaluated AI judges and juries did not match clinician ratings on the "Potential for Demographic Bias" criterion. Virtually all AI ratings were perfect, whereas local clinicians identified potential demographic bias in some cases. Consequently, the authors conclude that while automated judging offers a 75-fold reduction in evaluation costs, it may be useful for initial screening but is not yet justified for completely replacing human medical experts.

Background

Deploying generative artificial intelligence (AI) tools in resource-constrained medical settings may offer opportunities for clinical decision support.

However, because large language models (LLMs) generate probabilistic outputs, validating an algorithm prior to deployment does not guarantee sustained safety, as a safe response at one time does not guarantee that a future response will also be safe.

Traditionally, senior medical professionals perform these quality evaluations. However, a growing body of evidence indicates that human oversight is limited by cost and significant inter-clinician variability. In this study, human evaluation cost about $9.17 per query.

To address these scalability limits, developers and policymakers are increasingly turning to "LLM-as-a-judge" paradigms, in which secondary models evaluate primary clinical outputs.

Unfortunately, while preliminary trials have suggested that AI evaluators could approximate physician judgment, whether these models can accurately gauge local clinician ratings and handle linguistic and cultural context in low- and middle-income countries has remained unestablished.

About the study

The present study aimed to address this knowledge gap by evaluating clinical decision-support responses tailored for frontline healthcare contexts in Rwanda. The experimental sample included 524 question-and-answer pairs (416 in English and 108 in Kinyarwanda) derived from queries submitted by Rwandan Community Health Workers (CHWs).

Each response pair was systematically rated across 11 key clinical criteria including medical consensus alignment, reasoning validity, potential for harm, local context accuracy, and demographic bias.

Study evaluations were conducted independently across two arms:

  1. The human panel comprised six experienced, bilingual Rwandan general practitioners, working in two groups of three. Two clinicians scored each response and discussed ratings that differed by more than one point with a supervising clinician.
  2. AI judges and juries, which comprised five flagship LLMs, namely GPT-5, Gemini-2.5-Pro, Claude-4.1-Opus, MedGemma-20B, and GPT-OSS-70B. These LLMs were prompted via a shared evaluation rubric with four few-shot examples, alongside weighted ensemble combinations ("AI juries") designed to balance individual model biases.

Statistical analyses compared AI ratings with local clinician ratings across the 11 criteria and tested whether the differences were small enough to be practically meaningful.

Study findings

The study’s statistical analyses showed that AI judges produced more consistent ratings than clinician pairs, but this consistency did not imply that their judgments reliably matched local clinician ratings.

Notably, the findings revealed that no single LLM achieved global equivalence across all 11 criteria. Claude-4.1-Opus achieved the highest parity, but matched local clinician ratings on only four criteria. Conversely, Gemini-2.5-Pro tended to give more favorable ratings than clinicians, whereas GPT-5 tended to score responses more harshly.

Furthermore, every single AI judge and jury failed to match clinicians on the "Potential for Demographic Bias" metric, rating virtually all responses as flawless, whereas local clinicians identified potential bias in some responses. This shared weakness did not extend to the separate measure of local context.

The best-performing weighted AI jury matched clinicians on five of the 11 criteria, a modest improvement over Claude-4.1-Opus alone. AI judges displayed a greater tendency than clinicians to favor longer responses. Conversely, human clinicians exhibited in-group bias toward human-written answers.

Finally, transitioning the evaluation language from English to Kinyarwanda degraded AI agreement with clinician ratings for several models, particularly MedGemma. However, GPT-OSS improved in Kinyarwanda, and statistically combining model scores into AI juries largely reduced these language-related differences. Economically, AI judging was estimated to cost up to $0.12 per response, compared with $9.17 for human evaluation, resulting in a 75-fold cost reduction.

Stacked posterior predicted category probabilities by evaluator, ordered (within each criterion) by expected score. These four criteria were selected because Inclusion of Irrelevant Content and Potential for Demographic Bias were the worst-performing criteria across all AI judges/juries, and Omission of Important Information and Understanding of Local Context capture related components of these two poorly-evaluated criteria.

Stacked posterior predicted category probabilities by evaluator, ordered (within each criterion) by expected score. These four criteria were selected because Inclusion of Irrelevant Content and Potential for Demographic Bias were the worst-performing criteria across all AI judges/juries, and Omission of Important Information and Understanding of Local Context capture related components of these two poorly-evaluated criteria.

Conclusions

This study demonstrates that while automated AI judging offers scalability and cost-efficiency for initial screening, replacing human medical experts is not yet justified. Current LLM judges exhibit critical blind spots, particularly in detecting demographic prejudice and in processing nuances of underrepresented languages such as Kinyarwanda.

The authors conclude that AI juries may be appropriate for screening out clearly inappropriate systems, even when some imprecision is acceptable. However, until automated evaluators can reliably navigate localized equity and regional contexts, the complete phase-out of human medical experts is not yet supported.

Journal reference:
  • Williams, G., Rutunda, S., Nzabakira, F., & Mateen, B. A. (2026). Human evaluators vs. LLM-as-a-Judge: toward scalable evaluation of GenAI in global health. npj Digital Medicine (Article in Press). DOI: 10.1038/s41746-026-02992-w. https://www.nature.com/articles/s41746-026-02992-w
Hugo Francisco de Souza

Written by

Hugo Francisco de Souza

Hugo Francisco de Souza is a scientific writer based in Bangalore, Karnataka, India. His academic passions lie in biogeography, evolutionary biology, and herpetology. He is currently pursuing his Ph.D. from the Centre for Ecological Sciences, Indian Institute of Science, where he studies the origins, dispersal, and speciation of wetland-associated snakes. Hugo has received, amongst others, the DST-INSPIRE fellowship for his doctoral research and the Gold Medal from Pondicherry University for academic excellence during his Masters. His research has been published in high-impact peer-reviewed journals, including PLOS Neglected Tropical Diseases and Systematic Biology. When not working or writing, Hugo can be found consuming copious amounts of anime and manga, composing and making music with his bass guitar, shredding trails on his MTB, playing video games (he prefers the term ‘gaming’), or tinkering with all things tech.

Citations

Please use one of the following formats to cite this article in your essay, paper or report:

  • APA

    Francisco de Souza, Hugo. (2026, July 21). Human experts remain essential for checking clinical AI outputs. News-Medical. Retrieved on July 22, 2026 from https://www.news-medical.net/news/20260721/Human-experts-remain-essential-for-checking-clinical-AI-outputs.aspx.

  • MLA

    Francisco de Souza, Hugo. "Human experts remain essential for checking clinical AI outputs". News-Medical. 22 July 2026. <https://www.news-medical.net/news/20260721/Human-experts-remain-essential-for-checking-clinical-AI-outputs.aspx>.

  • Chicago

    Francisco de Souza, Hugo. "Human experts remain essential for checking clinical AI outputs". News-Medical. https://www.news-medical.net/news/20260721/Human-experts-remain-essential-for-checking-clinical-AI-outputs.aspx. (accessed July 22, 2026).

  • Harvard

    Francisco de Souza, Hugo. 2026. Human experts remain essential for checking clinical AI outputs. News-Medical, viewed 22 July 2026, https://www.news-medical.net/news/20260721/Human-experts-remain-essential-for-checking-clinical-AI-outputs.aspx.

Comments

The opinions expressed here are the views of the writer and do not necessarily reflect the views and opinions of News Medical.
Post a new comment
Post

While we only use edited and approved content for Azthena answers, it may on occasions provide incorrect responses. Please confirm any data provided with the related suppliers or authors. We do not provide medical advice, if you search for medical information you must always consult a medical professional before acting on any information provided.

Your questions, but not your email details will be shared with OpenAI and retained for 30 days in accordance with their privacy principles.

Please do not ask questions that use sensitive or confidential information.

Read the full Terms & Conditions.

You might also like...
Gut-friendly diet linked to lower mortality risk in coronary heart disease