# AI Verifier Disagreements Highlight Risks in Clinical Evaluations

> The study reveals alarming levels of disagreement among AI systems tasked with evaluating clinical reasoning, challenging the reliability of the popular 'LLM-as-a-judge' approach in medical AI. With disagreement rates reaching up to 74.3%, this research calls for a reevaluation of AI's role in clinical assessments.

**Source**: bioengineer.org | **Published**: 2026-09-07 | **Type**: research

## Key Facts

- AI verifiers showed 62.2%-74.3% disagreement, indicating unreliable evaluation methods in healthcare.
- Fleiss’ kappa values (0.087-0.223) reveal significant vulnerabilities in AI's clinical reasoning assessments.
- Discrepancies among models highlight the risk of erroneous diagnoses due to flawed AI judgment.
- Current reliance on single AI models for evaluation poses financial risks in costly healthcare settings.
- Need for multi-verifier systems suggests strategic shifts toward human oversight in AI evaluations.

## Summary

Recent research published in the Journal of Medical Systems reveals significant shortcomings in the reliability of artificial intelligence (AI) systems tasked with evaluating clinical reasoning in other AI models. The study, led by researchers from Yonsei University, found that three advanced large language models (LLMs) exhibited alarming levels of disagreement when assessing the soundness of clinical conclusions drawn by their peers. This finding raises critical concerns about the validity of the increasingly popular "LLM-as-a-judge" approach in medical AI evaluation.

The study specifically challenged the assumption that independent AI verifiers would agree on the coherence of clinical reasoning presented by generative models. The results were striking: using Fleiss’ kappa, a standard measure of agreement, the researchers recorded values between 0.087 and 0.223, indicating poor reliability. Disagreement rates among the verifiers ranged from 62.2% to 74.3%, suggesting that in most instances, these AI judges could not reach consensus on whether a piece of clinical reasoning was valid. This lack of agreement poses a significant risk in clinical settings, where inaccurate evaluations could lead to unsafe diagnostic conclusions being accepted or flagged for review.

To arrive at these conclusions, the researchers implemented a comprehensive evaluation framework that assessed the outputs of three generator models against 1,000 hospital-stay cases from the MIMIC-IV dataset. The evaluation criteria included medical concept grounding, semantic similarity, semantic uncertainty, and the pivotal measure of evidence–conclusion coherence. While the first three axes have established methodologies, the coherence measure was particularly innovative, focusing on whether the rationale provided by a model genuinely supported its diagnosis. The findings indicated that models could perform well on traditional metrics while still producing logically flawed conclusions.

The implications of this research are profound. As healthcare organizations increasingly adopt AI systems for tasks such as screening hospital discharge summaries and evaluating clinical notes, the reliance on a single model's judgment becomes statistically untenable. The study suggests a potential solution: implementing a unanimous-agreement tier, where automated evaluations are only permitted when all verifiers concur, leaving ambiguous cases for human oversight. However, the authors caution that even this approach requires further validation through larger studies involving multiple clinicians.

The study also highlights a broader methodological concern within the field of medical AI evaluation. Traditional semantic metrics, such as BLEU scores and embedding similarities, fail to capture internal logical inconsistencies within model outputs. The Yonsei team's multi-axis framework emphasizes the need to prioritize coherence in evaluations, signaling a shift away from reliance on single-metric assessments. Yet, the irony remains that the very tools required to audit medical AI—other LLMs—are themselves unreliable in this role, necessitating human oversight that the technology was initially intended to replace.

As the landscape of medical AI continues to evolve, this research underscores the necessity for more robust evaluation frameworks that incorporate multiple independent judgments and structured clinician adjudication. The findings suggest that while automation can enhance efficiency, it cannot replace the critical human element in ensuring the safety and reliability of AI-driven clinical decision-making. Organizations must navigate this complexity as they integrate AI into healthcare, balancing the benefits of automation with the imperative for rigorous oversight.

## Entities

- **Products**: HuatuoGPT-o1-8B, Meta-Llama-3.1-8B-Instruct, Meta-Llama-3.3-70B-Instruct, Claude Sonnet 4.6, Gemini 2.5 Pro, GPT-5.4 mini
- **Technologies**: large language models, semantic similarity, semantic uncertainty, evidence–conclusion coherence
- **People**: Hyunjung Byun, Beakcheol Jang, Dahyoun Lee, Munyoung Jung, Won Hwi Kim
- **Organizations**: Yonsei University, MIT, PhysioNet

## Key Concepts

AI verifiers, clinical reasoning, inter-verifier disagreement, LLM-as-a-judge, evaluation methodology, coherence, MIMIC-IV, Fleiss’ kappa

## Definitions

- **LLM-as-a-judge**: A framework where large language models are used to evaluate the reasoning of other AI systems.
- **Fleiss’ kappa**: A statistical measure used to assess the agreement among multiple raters.
- **semantic uncertainty**: A measure of how much a model's outputs vary when the same question is asked multiple times.
- **evidence–conclusion coherence**: The degree to which a model's justification supports its diagnosis.
- **medical concept grounding**: The evaluation of whether a model's output is based on recognized biomedical terminology.

## Use Cases

- Evaluating AI models in clinical settings
- Assessing hospital discharge summaries
- Quality assessment of medical notes
- Automated evaluation of clinical reasoning
- Multi-verifier panels for AI assessment
- Structured human oversight in AI evaluations

## Frequently Asked Questions

**What is the main finding of the study?**

The study found that AI verifiers often disagree on clinical reasoning, indicating that no single LLM verifier is reliable enough to serve as a stand-alone judge.

**Why is inter-verifier disagreement a concern?**

Inter-verifier disagreement raises questions about the reliability of AI evaluations in clinical settings, where inconsistent judgments could lead to unsafe diagnostic conclusions.

**What metrics were used to evaluate the AI models?**

The evaluation used metrics such as medical concept grounding, semantic similarity, semantic uncertainty, and evidence–conclusion coherence to assess the AI models.

**How does the study suggest improving AI evaluations?**

The authors suggest using a unanimous-agreement tier for automated judgments and combining multi-verifier panels with structured clinician adjudication to enhance reliability.

**What implications does this study have for healthcare AI?**

The findings imply that relying solely on automated evaluations without human oversight is statistically indefensible, highlighting the need for careful integration of AI in clinical decision-making.

## Links

- [Read on Welcome.AI](https://welcome.ai/content/ai-verifier-disagreements-highlight-risks-in-clinical-evaluations)
- [Original source](https://bioengineer.org/new-multi-axis-study-probes-how-ai-verifiers-disagree-on-clinical-reasoning/)

---

Source: Welcome.AI | https://welcome.ai/content/ai-verifier-disagreements-highlight-risks-in-clinical-evaluations