Figures
Abstract
Background
Major depressive disorder is a severe, recurrent and disabling condition. Although diagnosis and clinical monitoring are based on medical interviews and validated rating scales, voice and speech analysis may provide complementary digital biomarkers reflecting depressive severity and clinical evolution. However, current evidence remains limited by methodological heterogeneity, predominantly cross-sectional designs, limited longitudinal data and underrepresentation of non-English-speaking clinical populations.
Objective
The aim of the VOICE-DEP study is to develop and formalize a standardized, reproducible and clinically grounded protocol for the multimodal analysis of voice and speech during medical interviews as a tool to support the diagnosis of depressive disorder and to assess whether speech-derived digital biomarkers change over time in parallel with clinical severity measures.
Methods
VOICE-DEP is an observational, prospective, longitudinal pilot study of patients with major depressive disorder with a healthy control group, conducted in a hospital-based clinical setting in Spain. The study will include 25 adult patients with moderate or severe unipolar depression, with or without psychotic symptoms, and 50 healthy controls without a personal history of psychiatric disorders. Patients will be assessed at five time points: baseline (V0) and four monthly follow-up visits at 30, 60, 90 and 120 days. Healthy controls will be assessed once at baseline. The planned dataset comprises 175 voice recordings: 125 from patients and 50 from controls. At each assessment, the Montgomery-Asberg Depression Rating Scale related part of the medical interview, lasting approximately 10–30 minutes and including an initial free-speech segment, will be recorded using a standardized audio protocol. Acoustic (e.g., pitch, intensity), paralinguistic (e.g., speed rate, prosodic range) and linguistic (e.g., sentiment polarity, lexical diversity) features will be extracted and analyzed in relation to clinician-rated severity measures and self-reported symptoms.
Citation: Zabalza-Zudaire M, Sayar-Beristain O, Fructos P, Núñez FE, Carpio FF, García E, et al. (2026) The VOICE-DEP study protocol: Multimodal analysis of voice and speech during medical interviews to support diagnosis and longitudinal monitoring of major depressive disorder. PLoS One 21(9): e0353827. https://doi.org/10.1371/journal.pone.0353827
Editor: Jihua Dong, Shandong University, CHINA
Received: July 2, 2026; Accepted: August 24, 2026; Published: September 10, 2026
Copyright: © 2026 Zabalza-Zudaire et al. This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.
Data Availability: No datasets were generated or analysed during the current study. The entire data of this Protocol will be made freely accessible if our manuscript is accepted for publication. This means that the data associated with this work (namely: the full Original protocol approved by our local research ethics board, in Spanish, and its translation to English) have been attached as Supporting information to be freely accessible upon publication.
Funding: The author(s) received no specific funding for this work.
Competing interests: We have read the journal’s policy and the authors of this manuscript have the following competing interests: PM reports (all outside the current work) having received research grants from the Ministry of Education (Spain), the Government of Navarra (Spain), the Spanish Foundation of Psychiatry and Mental Health and AstraZeneca; he has been a clinical consultant for MedAvanteProPhase and Worldwide Clinical Trials Limited and has received lecture honoraria from or has been a consultant for AB-Biotics, Adept Field Solutions, Dialectica, Guidepoint, Janssen, Novumed, Roland Berger, and Scienta, received travel support for taking part in scientific meetings in the last 3 years (air/ground tickets + hotel) from Boston Scientific and Janssen, and has been the principal investigator of several studies supported by Janssen and Novartis about the efficacy and safety of novel pharmacological treatments for depression. The other authors declare no conflicts of interests. This does not alter our adherence to PLOS ONE policies on sharing data and materials.
1. Introduction
Major depressive disorder (MDD) is one of the leading contributors to global disability, affecting more than 280 million people worldwide and constituting a major public health challenge [1]. Despite decades of research, clinical assessment and monitoring of depression still rely predominantly on structured interviews and rating scales, which are inherently subjective and susceptible to recall bias and inter-rater variability [2,3].
In recent years, the identification of objective, scalable and non-invasive digital biomarkers has become a central goal of precision psychiatry [4]. Among behavioral signals, human speech has emerged as a promising source of information because speech production integrates motor control, cognition, affective regulation and autonomic function [5]. Alterations in prosody, speech rate, pauses, intensity and voice quality have been consistently associated with depressive states [6,7].
Recent advances in artificial intelligence have accelerated research on voice-based depression assessment. Machine learning, natural language processing and large language models have shown promising results using acoustic, paralinguistic and linguistic features extracted from speech recordings [8,9]. In parallel, recent studies have increasingly adopted multimodal approaches integrating audio and text information. Representative studies in the field are summarized in Table 1.
However, systematic reviews continue to highlight substantial methodological heterogeneity across this literature [26]. Differences in recording conditions, speech tasks, feature extraction pipelines, validation strategies and clinical reference standards limit reproducibility and comparability across studies. Moreover, the current literature remains dominated by cross-sectional designs and English-speaking datasets, while longitudinal studies in clinically characterized Spanish-speaking populations remain scarce.
Another important limitation is that many studies rely primarily on self-report questionnaires as reference labels for model training and evaluation. In contrast, clinician-rated instruments such as the Montgomery–Asberg Depression Rating Scale (MADRS) may provide a complementary, clinically grounded severity reference given its demonstrated validity and inter-rater reliability [27]. Longitudinal designs are also particularly relevant because they allow evaluation of whether speech-derived digital biomarkers evolve in parallel with changes in depressive severity over time [28].
Recent evidence in Spanish-speaking clinical populations further supports the relevance of integrating linguistic information into speech-based depression assessment [29]. Nevertheless, important challenges remain unresolved, including the standardization of recording conditions, harmonization of clinical anchors, definition of speech elicitation procedures and longitudinal evaluation of within-subject changes. These limitations highlight the need for transparent and reproducible protocol-driven studies designed specifically for clinical translation.
The present study protocol addresses these challenges by proposing a longitudinal, hospital-based framework for the collection and analysis of speech data in Spanish-speaking patients with depressive disorder and healthy controls. By integrating acoustic, paralinguistic and linguistic features and anchoring them to validated clinician-administered scales, the VOICE-DEP study aims to advance voice-based depression research toward clinically interpretable and reproducible methodology.
2. Materials and methods
2.1. Aim of the study, design and setting
The overall aim of the VOICE-DEP study is to develop and formalize a standardized, reproducible and clinically grounded protocol for the multimodal analysis of voice and speech during medical interviews as a tool to support the diagnosis and longitudinal monitoring of depressive disorder.
The working hypothesis is that major depressive disorder and its clinical evolution are associated with changes in acoustic, paralinguistic and linguistic characteristics of voice and speech, and that these changes can be analyzed in a standardized and reproducible manner in a routine clinical setting.
Primary objective. The primary objective is to establish a reproducible clinical framework for extracting and analyzing acoustic, paralinguistic and linguistic features from semi-structured medical interviews, allowing discrimination between patients with a medical diagnosis of depressive disorder and healthy controls using voice-based models anchored to validated clinician-rated clinical reference standards.
Secondary objectives. The secondary objectives are: i) To construct a longitudinal Spanish-language clinical speech corpus including repeated voice recordings from patients with depressive disorder and single-session recordings from healthy controls; ii) To characterize the association between speech-derived digital biomarkers and clinician-rated measures of depressive severity, including their evolution over time; iii) To evaluate the added value of multimodal models combining acoustic, paralinguistic and linguistic features compared with unimodal models; iv) To assess intra-individual vocal changes across repeated assessments and their association with changes in clinical severity; v) To provide a transparent methodological framework that may facilitate replication, cross-study comparison and future external validation in other clinical presentations, diagnoses, languages and sociocultural settings; vi) To explore whether specific treatment modalities are associated with greater changes in voice and speech in parallel with clinical evolution.
Study design and setting. This is an observational, prospective, longitudinal pilot study with a healthy control group. The study will be conducted in a hospital-based clinical setting, including both outpatient consultations and inpatient care at the Department of Psychiatry of a university hospital. The expected study duration is 24 months.
2.2. Study population and participant characteristics
The study will include two groups: (1) patients with a medical diagnosis of moderate or severe unipolar depression and (2) healthy control participants without a personal history of depression or any other psychiatric disorder.
Inclusion criteria for patients are: Age 18 years or older; Native Spanish speakers or individuals with sufficient fluency in Spanish to maintain a clinical interview; Capacity to understand study procedures and provide written informed consent, according to the investigator’s judgement; Medical diagnosis of moderate or severe unipolar depression, with or without psychotic symptoms, according to ICD-10 categories: moderate depressive episode (F32.1), severe depressive episode without psychotic symptoms (F32.2), severe depressive episode with psychotic symptoms (F32.3), recurrent depressive disorder, current episode moderate (F33.1), recurrent depressive disorder, current episode severe without psychotic symptoms (F33.2), or recurrent depressive disorder, current episode severe with psychotic symptoms (F33.3), regardless of whether the patient receives psychotherapeutic, pharmacological or neurostimulation treatment, or chooses not to receive active treatment while remaining under clinical follow-up.
Exclusion criteria for patients are: Severe neurological disorders affecting speech production or cognitive function, such as severe dementia; Severe voice or speech disorders unrelated to depression that prevent completion of the clinical interview through verbal language; State of intoxication due to substance use at the time of assessment; Other severe medical conditions interfering with study participation or speech production.
Inclusion criteria for healthy controls are: Age 18 years or older. Exclusion criteria for healthy controls are: Personal history of depression or any other psychiatric disorder; Severe neurological disorders affecting speech production or cognitive function, such as severe dementia; Severe voice or speech disorders unrelated to depression that prevent completion of the clinical interview through verbal language; State of intoxication due to substance use at the time of assessment; Other severe medical conditions interfering with study participation or speech production.
2.3. Study procedures and visit schedule
Participants will be recruited consecutively according to eligibility criteria. Patients will be invited by physicians from the research team during routine clinical practice. Healthy controls will be invited by members of the research team and, if they agree to participate, will be referred to a physician from the research team for baseline assessment. As this is an observational study, no randomization or blinding procedures will be applied.
At the baseline visit (V0), all participants will receive verbal and written information about the study objectives, methodology, non-invasive nature of the procedures, confidentiality safeguards, and exclusive research use of the collected data. Written informed consent will be obtained before any study-related procedure. A clinical interview will then be conducted to collect baseline sociodemographic and clinical information, followed by psychometric assessment using the clinical instruments specified in the protocol. During the same session, voice recordings will be obtained under standardized conditions.
Only patients with depressive disorder will undergo longitudinal follow-up. Patients will attend four additional visits coinciding, whenever possible, with routine psychiatric follow-up visits: V1 at 30 days, V2 at 60 days, V3 at 90 days and V4 at 120 days, each with an allowed window of ±7 days. At each follow-up visit, the same recording and clinical assessment procedures applied at baseline will be repeated. Clinician-rated and self-reported measures of depressive severity will be updated, and relevant clinical changes, including treatment adjustments or adverse events, will be documented. Healthy controls will complete only the baseline visit under the same recording conditions.
In all participants, data collection will be integrated into the clinical interview. Only the part of the medical interview corresponding to the MADRS assessment will be audio recorded. In order to combine open ended questions by clinicians that allow patients to describe their experiences in their own words, allowing spontaneous and elaborated responses, while adopting a reliable method of exploration, the instructions of Montgomery and Asberg for the MADRS will be adopted, moving from broadly phrased questions about each symptom to more detailed ones which allow a precise rating of severity [27]. The recorded material will comprise the complete MADRS-related clinical interaction, including an initial free-speech segment immediately followed by the semi-structured questions used to complete the scale. The free-speech segment will be an introductory part of the recorded MADRS assessment (namely, patients will be asked to describe, in their own words, how they have been feeling in the past week). This recorded segment is expected to last approximately 10–30 minutes including the initial free-speech segment, allowing voice and speech to be collected within a clinically meaningful interaction while maintaining a standardized procedure.
Voice recordings will be assigned a pseudonymized audio identifier linked to participant type, participant code and visit number. Any relevant recording incident, including interruption due to patient need, clinical need or technical problems, will be documented in the case report form and considered in data quality assessment.
2.4. Sample size and sampling strategy
The planned study sample comprises 75 participants: 25 patients with a clinical diagnosis of moderate or severe unipolar depression and 50 healthy controls. Patients will be assessed at five time points: baseline (V0) and four monthly follow-up visits at 30, 60, 90 and 120 days, each with an allowed window of ±7 days. Assuming complete follow-up, this will generate 125 longitudinal voice recordings in the clinical group. Healthy controls will undergo a single baseline assessment, generating 50 voice recordings. The planned dataset therefore comprises 175 voice recordings.
The sample size was defined pragmatically as a feasible recruitment target for a hospital-based feasibility and methodological pilot study conducted in a real-world clinical setting. The study is not intended as a definitive diagnostic accuracy trial powered to estimate the performance of a final clinical classifier. Rather, it is designed to evaluate the feasibility, standardization, reproducibility and analytical validity of a multimodal voice- and speech-based methodology for depression assessment and longitudinal monitoring.
As a quantitative sensitivity analysis, a baseline comparison of 25 patients and 50 healthy controls, with a two-sided significance level of 0.05 and 80% power, would detect a standardized mean difference of approximately Cohen’s d = 0.70. The study is therefore primarily sensitive to moderate-to-large effects. This is presented as a sensitivity analysis because reliable prospective effect-size estimates are not available for the multiple exploratory speech-derived variables included.
The sampling strategy will be consecutive and non-probabilistic. Eligible patients will be recruited during routine clinical care in the Department of Psychiatry, including both outpatient consultations and inpatient care. Healthy controls will be invited by members of the research team and assessed after confirmation of eligibility.
Although a balanced baseline comparison between patients and healthy controls would be methodologically desirable, the final allocation was defined according to feasibility considerations. Healthy controls will be individually matched to patients by age and sex whenever feasible. If individual matching cannot be achieved, control recruitment will be guided by the age and sex distribution of the patient group to obtain group-level frequency matching as closely as possible. Recruitment of clinically characterized patients with moderate or severe depressive disorder is expected to be more demanding because it requires repeated clinical assessments and longitudinal follow-up. In contrast, healthy controls require only a single baseline visit, allowing the inclusion of a larger control group without substantially increasing study burden. Therefore, the study will include 25 patients and 50 healthy controls. This design is expected to strengthen the baseline characterization of non-depressed speech patterns while preserving the longitudinal focus of the clinical cohort.
At participant level, baseline analyses will compare patients and healthy controls, whereas longitudinal analyses will focus on within-patient changes across repeated assessments. For predictive modeling, the imbalance between groups will be explicitly considered through appropriate analytical strategies, such as participant-level data splitting, balanced resampling, class weighting or sensitivity analyses when applicable.
Given the pilot nature of the study and the moderate sample size, classifier performance estimates will be interpreted as exploratory and hypothesis-generating. The study will prioritize transparent feature extraction, reproducible preprocessing, clinically interpretable modeling and longitudinal within-subject analysis over the development of a definitive deployable diagnostic model. Model complexity will be limited in relation to the available sample size, and classifier performance estimates will be accompanied by confidence intervals whenever feasible.
2.5. Voice recording protocol
Voice recordings will be obtained using a Jabra Speak 510 USB omnidirectional microphone, positioned frontally at a distance of no more than 20 cm from the participant’s mouth. Recordings will be conducted in a quiet hospital room, with minimized background noise, closed doors and windows whenever feasible, silenced electronic devices, and no overlapping speech or external interruptions whenever possible.
The Jabra Speak 510 microphone was selected for pragmatic and translational reasons: it is affordable, widely available, easy to deploy and compatible with routine hospital use. This choice prioritizes ecological validity over laboratory-grade audio fidelity, which may reduce the precision of signal-sensitive measures such as jitter, shimmer, HNR, and intensity. This potential variability will be reduced by maintaining the same microphone, configuration, placement, and recording conditions throughout the study. This approach is consistent with previous studies that have used non-specialized recording devices, including tablet microphones, smartphone microphones and mobile phones, to extract clinically relevant speech and voice features [19,24,25,30].
Repeated assessments will be scheduled within a similar time-of-day window for each participant whenever feasible. Because recordings are integrated into routine clinical visits, exact consistency cannot be guaranteed; recording time will therefore be documented and considered in sensitivity analyses. Audio recordings will be stored in uncompressed WAV format to preserve signal integrity during feature extraction. All files will be labeled using pseudonymized alphanumeric identifiers. Speaker diarization will identify participant and interviewer segments, followed by manual review. Overlapping or uncertain segments will be excluded, and the original recording will be preserved unchanged. For acoustic and paralinguistic analyses, only participant speech will be retained. The complete interaction will be automatically transcribed and manually reviewed to correct transcription and speaker-attribution errors. Linguistic variables and contextual text embeddings will be extracted from the reviewed participant transcripts, while interviewer turns will be retained as conversational context.
For each recording, total recording duration, cumulative valid participant-speech duration, silence-to-speech proportion, number of intelligibly transcribed words, and relevant recording incidents will be documented. A minimum of two minutes of cumulative valid participant speech will be required for the primary multimodal analysis. consistent with previous studies using short samples [11,19]. Recordings below this threshold will be considered non-evaluable for that analysis, but the participant will not be excluded from the study.
2.6. Variables
All study variables will be recorded in a pseudonymized case report form developed specifically for the VOICE-DEP study. Variables will be classified into: (i) baseline variables collected only at V0, and (ii) longitudinal visit variables collected at V0, V1, V2, V3 and V4 in patients with depressive disorder.
Exposure variable: The main exposure variable is the medical diagnosis of moderate or severe unipolar depressive disorder, with or without psychotic symptoms, according to the predefined ICD-10 categories included in the protocol.
Primary outcome variables: The primary outcome variables are the standardized voice recording and the multimodal speech- and speech-derived variables extracted from the recorded clinical interview. These variables will be grouped into acoustic, paralinguistic and linguistic domains.
Clinician-rated and self-reported symptom variables: At each study visit, depressive symptom severity and global clinical status will be assessed concurrently with the voice recording using the MADRS, the Clinical Global Impression scale (CGI), and the Patient Health Questionnaire-9 (PHQ-9). These variables will provide the clinical reference outcomes for cross-sectional and longitudinal analyses. Psychotic depression, operationally defined by an ICD-10 diagnosis of F32.3 or F33.3, will be recorded as a binary clinical variable. The same clinician-rated and self-reported assessments will also be conducted in healthy controls at baseline (V0) to identify possible subthreshold symptoms.
Baseline and longitudinal variables: Baseline variables collected at V0 will include demographic, anthropometric, clinical and contextual information, including psychiatric and medical history, substance use, stressful life events, and previous depressive episodes. Longitudinal variables collected across visits will include recording traceability variables, symptom scales, pharmacological treatment variables, psychotherapy variables and neurostimulation-related variables.
Metadata related to the recording process, including pseudonymized audio identifiers and recording incidents, will also be documented for traceability and quality control purposes.
The complete list of study variables operationalized in the VOICE-DEP case report form is summarized in Table 2. The list of multimodal voice- and speech-derived variables is summarized in Table 3.
2.7. Statistical and analytical plan
The analytical strategy of the VOICE-DEP study is designed to support the primary and secondary objectives of the protocol while ensuring methodological transparency, reproducibility, and robustness. In accordance with the study design, the analysis will address three complementary dimensions: (i) descriptive characterization of the study sample and extracted variables, (ii) exploratory cross-sectional association and discrimination analyses between patients with depressive disorder and healthy controls, and (iii) longitudinal intra-individual modeling of changes in voice- and speech-derived variables in relation to clinical evolution.
First, a descriptive analysis will be performed for all study variables recorded in the case report form. Quantitative variables, including acoustic, paralinguistic, and linguistic features, will be summarized using means and standard deviations or medians and interquartile ranges, as appropriate according to their distribution. Qualitative variables will be described using absolute frequencies and percentages. This descriptive phase will include both baseline variables collected at V0 and longitudinal variables collected across repeated visits in patients. Age, sex, educational level and available socioeconomic-related variables will be compared between groups and considered as covariates in adjusted or sensitivity analyses when relevant imbalance is observed and the sample size permits stable estimation.
Second, the association between continuous voice- and speech-derived digital biomarkers and clinical severity measures will be explored using Pearson or Spearman correlation coefficients, depending on distributional assumptions. These analyses will focus on the relationship between multimodal voice-derived variables and clinician-rated depressive symptom severity measured with the MADRS, global clinical status measured with the CGI, and self-reported depressive symptom burden measured with the Patient Health Questionnaire-9 (PHQ-9).
For diagnostic discrimination between patients and healthy controls, classification models will be developed and evaluated. The area under the receiver operating characteristic curve (AUC) will be used as the principal performance metric, and additional measures such as odds ratios and their corresponding 95% confidence intervals will be reported where applicable. These analyses will be interpreted as exploratory because the study is a pilot methodological validation study rather than a definitive diagnostic accuracy trial.
Third, a systematic comparison between unimodal and multimodal approaches will be conducted. Unimodal models will include models based exclusively on audio-derived features or text-derived features, whereas multimodal models will combine acoustic, paralinguistic, and linguistic information. This comparison is intended to quantify the added value of multimodal integration beyond individual feature domains and is directly aligned with one of the predefined secondary objectives of the study.
Embedding-based acoustic features and contextual text embeddings will be evaluated exploratorily because previous studies suggest that they may capture clinically relevant patterns beyond predefined features [8,10,12,13]. Given their high dimensionality and limited interpretability, they will not constitute the primary analytical approach. Model specifications and extraction procedures will be fully documented to ensure transparency and reproducibility.
Longitudinal analyses will be used to evaluate the evolution of voice- and speech-derived digital biomarkers over time and their association with changes in depressive symptom severity and global clinical status. Because repeated assessments are obtained in the patient group at V0, V1, V2, V3, and V4, these analyses will account for within-subject correlation. Mixed-effects models and, where appropriate, finite-difference or change-score approaches will be used to examine intra-individual trajectories and their relationship with variation in MADRS, CGI, and PHQ-9 scores, as well as with relevant treatment-related variables. Treatment exposure and changes during follow-up will be included, where feasible, as time-varying covariates in the mixed-effects models. Treatment-specific analyses will be considered exploratory and will not be interpreted causally.
Psychotic status will be considered as a potential confounder and included as a binary covariate when the sample permits stable estimation. Sensitivity analyses excluding patients with psychotic depression will be performed, while subgroup comparisons will be considered exploratory.
Regularized logistic regression will be the primary classification model, with a linear support vector machine and random forest models used for comparison. Missing data, including their frequency, pattern, and documented reasons, will be reported by variable and visit. Mixed-effects models will include a participant-specific random intercept, time as a fixed effect, and selected time-varying covariates; a random slope for time will be considered if supported by the data. The Benjamini–Hochberg procedure will control the false discovery rate within each feature family, and feature selection and hyperparameter tuning will be performed within each training partition.
Because repeated recordings from the same patient are not independent, all longitudinal analyses, predictive modeling and validation procedures will use participant-level data splitting and account for within subject correlation, in order to preserve the independence of training and test data. Specifically, all recordings from the same participant will be assigned exclusively to either the training or the test partition, in order to reduce optimistic bias due to repeated observations from the same individual and avoid information leakage. In this context, Leave-One-Group-Out cross-validation will be used as an appropriate validation framework for grouped repeated-measures data.
To improve clinical interpretability, feature attribution analyses will be performed using Shapley-based methods, including SHAP values where applicable. These analyses will help identify which acoustic, paralinguistic, and linguistic signals contribute most strongly to model predictions and will support the interpretation of the relative importance of different multimodal domains. No imputation for missing data will be applied. To maximize the data collection, quality control reviews will be conducted after every visit. A flexible time-window for visits and periodic reminder phone calls will be planned to reduce the risk of follow-up losses.
All analyses will be performed using Python v3.x or equivalent validated statistical software. Statistical significance will be considered at a two-sided threshold of p < 0.05.
2.8. Safety considerations
Participants may pause or stop the recording at any time, without providing a reason and without any negative consequence for their medical care. If fatigue, emotional distress, discomfort, or any other difficulty occurs during the interview, the recording may be interrupted or discontinued according to the participant’s preference and the investigator’s clinical judgement.
All assessments will be conducted by qualified clinical staff in a hospital-based psychiatric setting. If clinically relevant worsening, active suicidal ideation, suicidal behavior, severe anxiety, psychotic symptoms, or any other urgent clinical concern is detected, clinical care will take priority over study procedures and the situation will be managed according to usual clinical practice.
2.9. Ethical considerations and data protection
This protocol has been reviewed and approved by the local Research Ethics Committee, which complies with the established legal requirements for biomedical research, accredited by the Government of Navarra (University of Navarra Research Ethics Committee; reference code: 2026.110). The study will be conducted in accordance with the Declaration of Helsinki and with applicable National and European data protection regulations.
The VOICE-DEP study is observational and non-invasive, and does not modify clinical decision-making, treatment prescription, or usual psychiatric follow-up. The only study-specific procedure consists of audio recording a predefined section of the medical interview. It includes review of clinical variables collected in routine practice and administration of the clinical scales specified in the protocol. No pharmacological or interventional procedure is introduced by the study. All participants will receive verbal and written information about the study and will provide written informed consent before any study-related procedure. Participants will be informed that participation is voluntary and that they may withdraw at any time without any negative consequence for their subsequent medical care. The informed consent form will include authorization for access to clinical records when required for study objectives.
Participants will be identified using sequential study codes. Case report forms, reports and study communications will be identified using these codes. Only the sponsor, the investigator and authorized members of the research team, the Research Ethics Committee and competent health authorities, where applicable, will have access to identifiable study material.
Voice recordings and clinical data will be pseudonymized, stored securely and processed in accordance with applicable Spanish and European data protection regulations. Pseudonymized voice recordings will be transferred without distortion to secure NNBi servers. Recordings will be stored until 31 December 2029 and destroyed no later than 1 January 2030. NNBi servers are certified according to the Spanish National Security Framework at High level and ISO 27001. Technical and organizational safeguards will include encryption at rest and in transit, role-based access control, the principle of least privilege, robust authentication, environment segregation, access logging, backup encryption, recovery procedures and secure deletion at the end of the retention period.
Results will be reported in aggregate form. Individual-level voice recordings and transcriptions will not be publicly released because of privacy and re-identification risks.
2.10. Status and timeline of the study
This study was approved on 2026/05/08. The start date of the recruitment period for this study was 2026/05/21. Recruitment and data collection are ongoing and planned to continue until December 2027. Analysis and interpretation of the data are planned for February 2028. Final results report is planned for September 2028.
3. Discussion
The present manuscript describes a protocol for a longitudinal, hospital-based study aimed at developing and conducting an exploratory evaluation of speech-derived candidate digital biomarkers for the detection and monitoring of major depressive disorder. As a protocol paper, this discussion focuses on methodological considerations, anticipated challenges in study implementation, limitations inherent to the design, and plans for dissemination and protocol management, rather than on empirical findings.
3.1. Limitations and methodological considerations
Several limitations and anticipated challenges of the proposed protocol should be acknowledged, particularly those inherent to voice-based clinical studies that seek to balance standardization, ecological validity, and translational feasibility.
Although speech recordings are obtained under controlled hospital conditions, substantial inter-individual variability in speaking style, emotional expressiveness, and engagement during clinical interviews is expected. To balance standardization with clinical realism, a semi-structured, clinician-administered interview protocol is used to elicit naturalistic speech. This improves ecological validity and clinical relevance, but may introduce variability in speech content, duration, and topic coverage, which can in turn affect both acoustic and linguistic measurements.
Longitudinal data collection poses practical challenges. Repeated assessments increase participant burden and may lead to missed visits or incomplete follow-up data. To mitigate this risk, follow-up visits are scheduled within flexible time windows and all procedures are designed to be brief and minimally invasive. Nevertheless, some degree of attrition is anticipated and will be explicitly addressed in the analytical strategy.
The planned sample size remains moderate. While the longitudinal structure provides valuable information on intra-individual vocal trajectories, a moderate sample may limit statistical power, increase uncertainty in performance estimates, constrain the exploration of complex interaction effects, and raise the risk of overfitting, particularly for higher-dimensional or multimodal models.
Technical variability in recording conditions such as background noise or microphone positioning, may influence acoustic measurements. The protocol addresses this by using a standardized recording device and placement, while intentionally avoiding laboratory-grade equipment to maximize clinical applicability. This trade-off prioritizes real-world deployability over maximal signal fidelity and may reduce the signal-to-noise ratio for certain acoustic digital biomarkers.
The study is conducted in a single clinical center and focuses on Spanish-speaking participants, which may induce selection bias and restrict external validity across healthcare systems, dialects, or linguistic contexts. Dialectal variation across Spanish-speaking regions may influence acoustic, prosodic, and linguistic characteristics. Therefore, the findings will be interpreted within the linguistic and geographical context of the study population and will require future external validation in samples from other Spanish-speaking regions and sociolinguistic settings. However, this focus also addresses a relevant gap in the literature, which has been predominantly centered on English-language datasets.
Because of the observational design and treatment heterogeneity, changes in speech cannot be fully disentangled from the effects of specific treatments. Treatment-related analyses will therefore remain exploratory and will not support causal inference.
Residual variability related to the recording time of the day will be acknowledged as a study limitation.
Finally, clinician-administered scales are used as the primary reference standards, but they are themselves subject to measurement variability and do not constitute an objective ground truth. In addition, symptom-focused MADRS questions may influence some linguistic variables, particularly sentiment-related measures, which will therefore be interpreted cautiously. The present protocol does not aim to replace clinical judgment but rather to complement it by identifying speech-derived markers that correlate with established measures of depressive severity.
In summary, the formal clinical diagnosis will define the study groups, whereas the MADRS, CGI, and PHQ-9 will be used to quantify symptom severity and longitudinal change. Speech-derived measures will therefore be interpreted as complementary correlates of clinical assessment rather than as independent objective standards, under the premise of supporting, but not replacing, the clinical judgment.
3.2. Dissemination and data sharing
The results derived from this study are intended to be disseminated through peer-reviewed scientific publications and presentations at national and/or international conferences focused on psychiatry, digital health, and biomedical signal processing. Findings will be reported in accordance with relevant reporting guidelines, and emphasis will be placed on transparent description of methods and analytical pipelines.
Where feasible and in compliance with ethical and legal constraints, aggregated results and methodological details will be shared to facilitate replication and comparison with future studies.. Raw voice recordings and transcripts will not be publicly shared because of their identifiable nature. A de-identified matrix of derived numerical features will be deposited in a suitable open-access repository upon publication, subject to approval through the required ethical, legal and data-protection review procedures.
The shared dataset will consist of a de-identified, recording-level feature matrix containing exclusively derived acoustic, paralinguistic, and linguistic summary variables. Each row will correspond to one study assessment, with repeated assessments linked through a newly generated public study identifier that will not correspond to the identifiers used in the clinical database or audio repository. The public representation is expected to include a restricted set of interpretable acoustic features, such as summary measures of fundamental frequency, intensity, harmonic-to-noise ratio, jitter, shimmer, selected spectral descriptors, formant statistics, and aggregated MFCC coefficients; paralinguistic measures such as speech and articulation rate, pause duration, pause ratio, and prosodic variability; and aggregated linguistic measures such as lexical diversity, word and sentence counts, sentiment-related measures, first-person pronoun frequency, and summary measures of lexical coherence. The publicly deposited dataset will not contain raw audio, transcripts, frame-level acoustic representations, timestamps, high-dimensional speech or semantic embeddings, or intermediate representations from which the original signal or linguistic content could potentially be inferred. Model-specific transformations, including feature scaling or normalization, will be performed within the corresponding analytical training partitions and documented as part of the analytical pipeline, rather than incorporated into the shared raw feature matrix. Only the minimum clinical and study variables required to interpret and reproduce the reported analyses will accompany the derived features (for example, study group, assessment visit, and the clinical outcome variables analyzed in the corresponding publication). Detailed demographic, treatment, and clinical-history information will not be included in the public matrix unless specifically required for reproducibility and considered compatible with the re-identification risk assessment. The final set and granularity of variables released will be determined after a formal re-identification risk assessment and the corresponding ethical and data-protection review.
Supporting information
S2 File. English translation of the original protocol.
https://doi.org/10.1371/journal.pone.0353827.s002
(DOCX)
S3 File. STROBE checklist of cohort studies (applicable items for a study protocol).
https://doi.org/10.1371/journal.pone.0353827.s003
(DOC)
References
- 1.
World Health Organization. Depression. WHO Fact Sheets. 2023. [cited 2025 Jan 1]. Available from: https://www.who.int/news-room/fact-sheets/detail/depression
- 2. Goldberg D. The heterogeneity of “major depression”. World Psychiatry. 2011;10(3):226–8. pmid:21991283
- 3. Zimmerman M, McGlinchey JB. Why don’t psychiatrists use scales to measure outcome when treating depressed patients?. J Clin Psychiatry. 2008;69(12):1916–9. pmid:19192467
- 4. Insel TR. Digital Phenotyping: technology for a new science of behavior. JAMA. 2017;318(13):1215–6. pmid:28973224
- 5. Cummins N, Scherer S, Krajewski J, Schnieder S, Epps J, Quatieri TF. A review of depression and suicide risk assessment using speech analysis. Speech Commun. 2015;71:10–49.
- 6. Darby JK, Hollien H. Vocal and speech patterns of depressive patients. Folia Phoniatr (Basel). 1977;29(4):279–91. pmid:604242
- 7. Chandrasekaran B, Van Engen K, Xie Z, Beevers CG, Maddox WT. Influence of depressive symptoms on speech perception in adverse listening conditions. Cogn Emot. 2015;29(5):900–9. pmid:25090306
- 8. Wu W, Zhang C, Woodland PC. Self-Supervised Representations in Speech-Based Depression Detection. In: ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). 2023. pp. 1–5.
- 9.
Zhang X, Liu H, Xu K, Zhang Q, Liu D, Ahmed B. When LLMs meet acoustic landmarks: Integrating speech into language models for depression detection. In: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2024. pp. 146–58.
- 10. Jin Y, Chen X, Hong X, Wang M, Niu W, Liu A, et al. Depression screening with textual and audio features based on large language models and machine learning. J Affect Disord. 2026;395(Pt A):120644. pmid:41253255
- 11. Li S, Xie Z, Naqvi SM. Efficient Long Speech Sequence Modelling for Time-Domain Depression Level Estimation. In: ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). 2025. pp. 1–5.
- 12. Xu Z, Gao Y, Wang F, Zhang L, Zhang L, Wang J, et al. Depression detection methods based on multimodal fusion of voice and text. Sci Rep. 2025;15(1):21907. pmid:40593978
- 13. Ali M, Lucasius C, Patel T, Aitken M, Vorstman J, Szatmari P, et al. Speech as a Multimodal Digital Phenotype for Multi-Task LLM-based Mental Health Prediction. 2025.
- 14. Mazur A, Costantino H, Tom P, Wilson MP, Thompson RG. Evaluation of an AI-based voice biomarker tool to detect signals consistent with moderate to severe depression. Ann Fam Med. 2025;23(1):60–5. pmid:39805690
- 15. Ghosh D, Karande H, Gite S, Pradhan B. Psychological disorder detection: a multimodal approach using a transformer-based hybrid model. MethodsX. 2024;13:102976. pmid:39430783
- 16. Huang X, Wang F, Gao Y, Liao Y, Zhang W, Zhang L, et al. Depression recognition using voice-based pre-training model. Sci Rep. 2024;14(1):12734. pmid:38830969
- 17. Sadeghi M, Richer R, Egger B, Schindler-Gmelch L, Rupp LH, Rahimi F, et al. Harnessing multimodal approaches for depression detection using large language models and facial expressions. Npj Ment Health Res. 2024;3(1):66. pmid:39715786
- 18. Anand A, Tank C, Pol S, Katoch V, Mehta S, Shah R. Depression Detection and Analysis using Large Language Models on Textual and Audio-Visual Modalities. 2024.
- 19. Menne F, Dörr F, Schräder J, Tröger J, Habel U, König A, et al. The voice of depression: speech features as biomarkers for major depressive disorder. BMC Psychiatry. 2024;24(1):794. pmid:39533239
- 20. Ronneberg CR, Lv N, Ajilore OA, Kannampallil T, Smyth J, Kumar V, et al. Study of a PST-trained voice-enabled artificial intelligence counselor for adults with emotional distress (SPEAC-2): Design and methods. Contemp Clin Trials. 2024;142:107574. pmid:38763307
- 21. Cansel N, Faruk Alcin Ö, Furkan Yılmaz Ö, Ari A, Akan M, Ucuz İ. A new artificial intelligence-based clinical decision support system for diagnosis of major psychiatric diseases based on voice analysis. Psychiatr Danub. 2023;35(4):489–99. pmid:37992093
- 22. Berardi M, Brosch K, Pfarr J-K, Schneider K, Sültmann A, Thomas-Odenthal F, et al. Relative importance of speech and voice features in the classification of schizophrenia and depression. Transl Psychiatry. 2023;13(1):298. pmid:37726285
- 23. Li N, Feng L, Hu J, Jiang L, Wang J, Han J, et al. Using deeply time-series semantics to assess depressive symptoms based on clinical interview speech. Front Psychiatry. 2023;14:1104190. pmid:36865077
- 24. Kim AY, Jang EH, Lee S-H, Choi K-Y, Park JG, Shin H-C. Automatic depression detection using smartphone-based text-dependent speech signals: deep convolutional neural network approach. J Med Internet Res. 2023;25:e34474. pmid:36696160
- 25. Wasserzug Y, Degani Y, Bar-Shaked M, Binyamin M, Klein A, Hershko S, et al. Development and validation of a machine learning-based vocal predictive model for major depressive disorder. J Affect Disord. 2023;325:627–32. pmid:36586600
- 26. Briganti G, Lechien JR. Speech and voice quality as digital biomarkers in depression: a systematic review. J Voice. 2025. pmid:40410060
- 27. Davidson J, Turnbull CD, Strickland R, Miller R, Graves K. The Montgomery–˚Asberg Depression Scale: reliability and validity. Acta Psychiatrica Scandinavica. 1986;73:544–8.
- 28. Almaghrabi SA, Clark SR, Baumert M. Bio-acoustic features of depression: a review. Biomed Signal Process Control. 2023;85:105020.
- 29. Maran PL, Gabirondo P, Vlaic A, Alonzo-Castillo MT, Rojo T, Zaldua C, et al. Beyond acoustic features: incorporating linguistic variables in automatic speech analysis for depression detection. J Affect Disord. 2026;405:121563. pmid:41796775
- 30. Lin D, Nazreen T, Rutowski T, Lu Y, Harati A, Shriberg E. Feasibility of a machine learning-based smartphone application in detecting depression and anxiety in a generally senior population. Front Psychol. 2022;13(4).