Stable answers help medical AI flag diagnoses clinicians can trust
· Medical Xpressby Dresden University of Technology
edited by Gaby Clark, reviewed by Robert Egan
Gaby Clark
Scientific Editor
Meet our editorial team
Behind our editorial process
Robert Egan
Senior Editor
Meet our editorial team
Behind our editorial process Editors' notes
This article has been reviewed according to Science X's editorial process and policies. Editors have highlighted the following attributes while ensuring the content's credibility:
fact-checked
peer-reviewed publication
trusted source
proofread
The GIST Add as preferred source
AI agents could reliably support diagnoses and clinical decision-making in the future—provided sensitive health data are protected and clinicians can assess the reliability of individual AI-generated results. Researchers at the Else Kröner Fresenius Center (EKFZ) for Digital Health at TU Dresden and Dresden University Hospital have developed an on-premises medical AI system that addresses both challenges. Their findings have been published in the journal Nature Medicine.
Large language models are increasingly capable of handling complex medical tasks. However, their use in clinical practice faces two fundamental challenges: Sensitive patient data need to remain under the control of the respective institution rather than being shared with external services, and clinicians need to be able to assess the reliability of AI-generated diagnoses.
A research team led by Jakob N. Kather in Dresden developed a fully on-premises diagnostic AI system and evaluated it in a simulation environment in which two AI agents interact with each other: One takes on the role of the physician, the other that of the patient. Using this setup, the researchers investigated how reliably medical AI systems can make diagnostic decisions. The system is built on the MIRA AI agent introduced in June 2026.
High diagnostic accuracy in standardized tests
The AI agent was tested using standardized clinical cases covering a range of conditions, including appendicitis, cholecystitis, pneumonia, pulmonary embolism and urinary tract infections. The cases were based on anonymized electronic patient data, including diagnoses, laboratory values, medications and physical examination findings.
The physician AI agent could ask follow-up questions and request clinical findings and laboratory results. Based on this information, it generated a diagnosis and an accompanying rationale. The best locally operated AI model reached the correct diagnosis in approximately 90% of cases in one benchmark and 84% in the other. For 181 randomly selected cases, physicians additionally reviewed the diagnoses. The automated evaluation and the physicians' consensus agreed in more than 90% of cases.
Consistent answers are the strongest indicator of a correct diagnosis
One of the study's key questions was how to determine whether an AI-generated decision can be trusted. To address this, the researchers investigated several indicators of diagnostic correctness. The most informative was whether the AI arrived at the same diagnosis when processing the same case repeatedly. The more stable the diagnosis remained across multiple runs, the more likely it was to be correct.
The model's internal likelihood scores were less useful for predicting whether a diagnosis was correct. This was also evident in a stress test: When the researchers removed reliable information from the system, diagnostic accuracy dropped substantially. Although the AI diagnoses became more frequently incorrect, the model's probability-based signals did not reflect this.
Medical AI agents to support clinical practice
Based on their findings, the researchers propose a possible framework for collaboration between clinicians and AI. Rather than either fully trusting an AI system or requiring every decision to be verified, cases with stronger reliability signals could be distinguished from more uncertain ones. Uncertain cases would always be deferred to medical professionals for review.
Kather, senior author of the study and professor of clinical artificial intelligence at TU Dresden, explains, "Our goal is an AI agent with selective autonomy. These systems should support clinicians in decision-making, but never take over completely. It is therefore essential that AI-generated results are understandable to clinicians and that the systems clearly indicate when their outputs are uncertain. Responsibility for diagnosis and treatment will always remain with humans."
Local deployment enables institutional control
A second focus of the study is technical control over the AI system. The AI models and all data processing ran entirely within locally operated infrastructure. This gives medical institutions greater control over where data are processed and which AI models are used. On-premises infrastructure provides the technical basis for managing data protection, model versions, access rights and monitoring processes within the respective institution.
Kather also sees this as a prerequisite for advancing medical AI responsibly in Europe: "Calls to slow down AI development are, in my view, heading in the wrong direction. In medicine, we are still at the very beginning: AI systems such as our agents are only just starting to show what is possible, and patients have so far seen very little benefit.
"We need to move faster, not slower. What matters is that we develop these systems in a way that is safe, transparent and preserves data sovereignty—for example, by running them entirely within the hospital. But for this to work, Europe also needs to play a role in the foundational technology itself, including the language models, rather than simply building on technologies developed elsewhere."
Next steps toward clinical use
The study also highlights questions that need to be addressed before any potential clinical deployment. The researchers found indications of differences in diagnostic accuracy between patient groups. It remains unclear why simulated cases involving older patients in particular showed poorer results, and this will require further investigation.
In addition, the findings are based on retrospective simulations using existing clinical data rather than on testing during routine clinical care. "Our results show that we have taken an important step toward more reliable medical AI agents. Next, we want to investigate how well this approach performs under realistic conditions, with clinicians involved in the process and across a broader range of diseases and clinical data," says Li Zhang, first author of the publication and a researcher on Kather's team.
Another focus is efficiency. Because estimating reliability currently requires repeated runs for each case, the team is working to reduce the computational cost without compromising the quality of the results. In addition to the Dresden researchers, scientists from the National Center for Tumor Diseases (NCT) Heidelberg at Heidelberg University Hospital also contributed to the study.
Publication details
Li Zhang et al, On-premise medical AI agents for reliable clinical decision-making, Nature Medicine (2026). DOI: 10.1038/s41591-026-04609-x
Journal information: Nature Medicine
Key medical concepts
Clinical categories
Hospital medicine Provided by Dresden University of Technology Who's behind this story?
Gaby Clark
MA in English, copy editor since 2021 with experience in higher education and health content. Dedicated to trustworthy science news. Full profile →
Robert Egan
Bachelor's in mathematical biology, Master's in creative writing. Well-traveled with unique perspectives on science and language. Full profile →
Citation: Stable answers help medical AI flag diagnoses clinicians can trust (2026, September 15) retrieved 15 September 2026 from https://medicalxpress.com/news/2026-09-stable-medical-ai-flag-clinicians.html This document is subject to copyright. Apart from any fair dealing for the purpose of private study or research, no part may be reproduced without the written permission. The content is provided for information purposes only.