Language models identify psychiatric symptoms nearly as accurately as early-career clinicians
· Medical Xpressby Central Institute of Mental Health
edited by Sadie Harley, reviewed by Robert Egan
Sadie Harley
Scientific Editor
Meet our editorial team
Behind our editorial process
Robert Egan
Senior Editor
Meet our editorial team
Behind our editorial process Editors' notes
This article has been reviewed according to Science X's editorial process and policies. Editors have highlighted the following attributes while ensuring the content's credibility:
fact-checked
peer-reviewed publication
trusted source
proofread
The GIST Add as preferred source
A study led by the Central Institute of Mental Health (CIMH) provides initial evidence that large language models can identify complex psychopathological findings in transcripts of psychiatric interviews with accuracy comparable to that of predominantly young clinicians. The study examined 10 language models and 108 practicing clinicians from three psychiatric clinics.
The study is published in the journal npj Digital Medicine.
The results demonstrate the potential for AI-supported decision-making tools, but they do not suggest that these tools can replace medical judgment. Further studies involving real patients are required before any potential clinical application.
From the psychiatric interview to a structured diagnosis
Psychiatric diagnosis usually begins with an interview. During this process, professionals systematically assess a person's mental state. From information that is often ambiguous and unstructured, a psychopathological assessment is derived, which serves as the most important basis for diagnosis and treatment.
It was precisely this step in the process that the researchers examined. While many previous AI studies have focused on diagnoses or standardized questionnaires, this study is the first to systematically examine how large language models perform in the complex task of psychopathological assessment compared with practicing clinicians.
A comparison of 10 language models with 108 clinicians
For the study, 10 large language models evaluated transcripts of three simulated psychiatric interviews on depression, mania and schizophrenia. Their task was to assess all 100 features of the AMDP system. The AMDP system is a method widely used in German-speaking psychiatry to systematically describe psychiatric symptoms according to established criteria.
For comparison, 108 practicing physicians and psychologists from three psychiatric clinics conducted the same evaluations. Most of the participants were still early in their clinical careers. While the language models were provided exclusively with the written transcripts, the clinicians were able to review the complete video and audio recordings to ensure the most realistic assessment possible.
The best language models fell within the range of the clinical control group
The two highest-performing models, GPT-5.1 and Gemini-3-Pro-Preview, fell within the range of the clinical comparison group overall. Across all three interviews, they achieved an average accuracy of 72%, while the clinical comparison group achieved 68%. These figures refer exclusively to the assessment of individual psychopathological characteristics and not to the establishment of a psychiatric diagnosis.
Due to the small number of experienced specialists in the sample, the study does not allow for a reliable comparison with psychiatrists who have many years of specialized experience.
Clinical professionals and AI make different kinds of mistakes
What was particularly revealing was not so much the average accuracy as the nature of the errors. Clinicians were more likely to infer the presence or absence of a characteristic based on incomplete information. The language model, on the other hand, more frequently classified such characteristics as "unassessable."
This difference was particularly pronounced for symptoms where nonverbal or visual cues are important for assessment. The sometimes contrasting error patterns suggest that human clinical experience and machine-based evaluation—which is strictly guided by the available text—could complement one another.
"We deliberately chose not to limit ourselves to having a language model evaluate diagnoses or questionnaires. In everyday psychiatric practice, structured psychopathological findings must be derived from complex and often ambiguous conversations," says lead author Dr. Esra Lenz, a researcher at the Hector Institute for Artificial Intelligence in Psychiatry (HITKIP) at the CIMH.
"Our results provide initial evidence that AI could support this process in the future. However, it does not replace either clinical experience or the direct assessment by a specialist."
Potential as an additional decision-making tool
In a retrospective simulation, the researchers also investigated whether a language model could serve as an additional decision-making aid in cases of conflicting clinical assessments. When the model's evaluation was used to resolve such disagreements, the results were more accurate than when a random choice was made between the two clinical assessments. Similar improvements were observed in a simulated specialist supervision scenario.
However, these results are derived exclusively from a statistical simulation and do not yet allow for any conclusions regarding whether AI actually improves diagnostic assessment in everyday clinical practice.
Confirmation with actual patients is required
The researchers emphasize the exploratory nature of the study. Only three simulated interviews were analyzed—one each on depression, mania and schizophrenia. No real patients were involved. The results therefore cannot be generalized to psychiatric interviews in general or to specific disorders.
Future studies must examine whether the findings can be confirmed in a larger number of different interviews, involving real patients and conducted under real-world clinical conditions. In doing so, strict ethical and data protection requirements must be taken into account.
"The key question is not whether AI will replace mental health professionals, but whether it can provide meaningful support where psychiatric diagnosis actually begins—namely, in clinical interviews and psychopathological assessment," says Dr. Emanuel Schwarz, the study's last author and director of the Hector Institute for Artificial Intelligence in Psychiatry at the CIMH.
"Our study deliberately shifts the focus away from diagnoses and questionnaires and toward clinical interviews and psychopathological assessment. Before AI systems can be used in this field, we must carefully examine their strengths and limitations under real-world conditions," adds Dr. Tobias Gradinger, also a co-author and research associate at HITKIP.
Publication details
Esra Lenz et al, Benchmarking large language models against practicing clinicians on psychopathological assessment, npj Digital Medicine (2026). DOI: 10.1038/s41746-026-02852-7
Journal information: npj Digital Medicine
Key medical concepts
Clinical categories
PsychiatryPsychology & Mental health Provided by Central Institute of Mental Health Who's behind this story?
Sadie Harley
BSc Life Sciences & Ecology. Microbiology lab background with pharmaceutical news experience in oil, gas, and renewable industries. Full profile →
Robert Egan
Bachelor's in mathematical biology, Master's in creative writing. Well-traveled with unique perspectives on science and language. Full profile →
Citation: Language models identify psychiatric symptoms nearly as accurately as early-career clinicians (2026, August 5) retrieved 5 August 2026 from https://medicalxpress.com/news/2026-08-language-psychiatric-symptoms-accurately-early.html This document is subject to copyright. Apart from any fair dealing for the purpose of private study or research, no part may be reproduced without the written permission. The content is provided for information purposes only.