Repository logo
Research Outputs
Projects
People
Statistics
  1. Home
  2. HSG CRIS
  3. HSG Publications
  4. Can Conversational AI Reliably Assess Depressive Symptoms? A Feasibility, Acceptability and Bias-Aware Validation Study of MADRS-Chat
Details

Can Conversational AI Reliably Assess Depressive Symptoms? A Feasibility, Acceptability and Bias-Aware Validation Study of MADRS-Chat

Journal
SSRN
Type
working paper
Date Issued
2026-02-27
Author(s)
Samantha Weber
;
Barbora Provaznikova
;
Nikolas Psathakis
;
Andrea Casanova
;
Lucian Amthor
;
Lara Kirchhofer
;
José Antonio García Macías
;
Tobias Welt
;
Golo Kronenberg
;
Tobias Kowatsch  
;
Birgit Kleim
;
Sebastian Olbrich
DOI
10.2139/ssrn.6306520
Abstract
Background: Structured clinical interviews such as the Montgomery-Åsberg Depression Rating Scale (MADRS) are central to depression assessment but are time-intensive and inconsistently applied in routine care. Large language models (LLMs) offer new opportunities to support standardized symptom assessment through conversational interfaces, yet their clinical feasibility, validity, and potential biases remain insufficiently studied.

Methods: We evaluated MADRS-Chat, an LLM-assisted text-based chatbot designed to administer the MADRS interview and generate symptom-level scores using a clinician-trained scoring model. In a prospective observational study involving psychiatric in- and outpatients, we evaluated feasibility and acceptability of a conversational chatbot and its agreement with clinician-rated MADRS scores and patient self-rated (MADRS-S) scores.

Findings: MADRS-Chat was well accepted by patients and clinicians and chatbot-derived MADRS scores showed meaningful agreement with human ratings. Qualitative feedback indicated that the anonymous nature of the written interaction facilitated more open and detailed symptom disclosure. However, chatbot-administered interviews systematically yielded higher severity scores, likely reflecting interaction-induced shifts in response style. This bias was effectively addressed through lightweight post-hoc calibration.

Interpretation: Our findings demonstrate that conversational AI can support structured psychiatric assessment in real-world clinical settings when used within clearly defined clinical boundaries. Moreover, text-based conversational interactions can alter symptom reporting, making explicit validation and calibration essential before deployment.
HSG Classification
contribution to practical use / society
Refereed
No
Publisher
Elsevier BV
URL
https://www.alexandria.unisg.ch/handle/20.500.14171/125209
Subject(s)

health sciences

Division(s)

MED - School of Medic...

Support
HSG researchers can find instructions here for adding or importing publications (DOI, ORCID). Please send questions to alexandria@unisg.ch

Built with DSpace-CRIS software - Extension maintained and optimized by 4Science

  • Accessibility settings
  • Privacy policy
  • End User Agreement
  • Send Feedback
Repository logo COAR Notify