Skip to main content
Advertisement
Browse Subject Areas
?

Click through the PLOS taxonomy to find articles in your field.

For more information about PLOS Subject Areas, click here.

  • Loading metrics

Uncertainty-aware personalized estimation of Parkinson’s disease severity from longitudinal speech

Abstract

Parkinson’s disease (PD) is a progressive neurological disorder characterized by motor impairments whose severity is commonly assessed using the Unified Parkinson’s Disease Rating Scale (UPDRS). Although clinically established, UPDRS assessment is inherently subjective, requiring in-person evaluation by trained specialists, limiting its suitability for frequent monitoring. Speech production is affected early in PD and provides a non-invasive modality for remote symptom assessment. In this study, an uncertainty-aware personalized framework is proposed for estimating PD severity from speech signals. The approach integrates longitudinal temporal modeling of longitudinal speech recordings with patient-specific representations and a probabilistic latent disease state. Continuous motor UPDRS scores are jointly estimated with data-driven ordinal disease severity stages, enabling both fine-grained regression and auxiliary ordinal prediction. Predictive uncertainty is explicitly quantified to characterize predictive variability within the proposed framework. The method is evaluated on a longitudinal speech dataset using a strict patient-wise split, ensuring that all test subjects are unseen during training. On the held-out test set, the proposed model achieves promising predictive accuracy (mean absolute error 0.56 UPDRS points, root mean squared error 0.74, and coefficient of determination R2 = 0.99) for motor UPDRS estimation. Ordinal severity classification attained an accuracy of 0.92 across three stages. Comparative experiments against classical machine learning methods and global temporal baselines demonstrate consistent performance improvements. These results demonstrate the potential of personalized, uncertainty-aware speech modeling for longitudinal PD severity estimation.

1 Introduction

Parkinson’s disease (PD) is a chronic, progressive neurodegenerative disorder that affects motor control, speech, and quality of life for millions of individuals worldwide [13]. As the disease advances, patients experience gradually worsening motor symptoms such as bradykinesia, rigidity, tremor, and speech impairment [4,5]. Accurate assessment of symptom severity is therefore central to clinical management, treatment adjustment, and the evaluation of disease progression [1,6].

The Unified Parkinson’s Disease Rating Scale (UPDRS) remains the most widely used clinical instrument for quantifying PD severity [79]. Despite its clinical acceptance, UPDRS assessment has several well-recognized limitations. First, it requires the physical presence of the patient in a clinical setting, which can be burdensome for individuals with mobility impairments and costly for healthcare systems. Second, assessments are typically performed at infrequent intervals, often every several months, providing only sparse snapshots of a disease process that evolves continuously and exhibits substantial short-term fluctuations. Third, UPDRS scoring is inherently subjective, relying on expert judgment and clinical experience, which introduces inter-rater variability and limits reproducibility across settings and practitioners [10,11]. These limitations have motivated growing interest in telemonitoring approaches that enable remote, frequent, and objective assessment of PD symptoms. Advances in digital health technologies have facilitated collecting patient data in non-clinical environments, reducing logistical barriers while increasing temporal resolution [12,13]. Among the available modalities, speech recordings are particularly attractive: speech production is strongly affected by motor impairment in PD, voice acquisition is non-invasive, and sustained phonations can be reliably collected using simple, self-administered protocols [1416]. Prior work has demonstrated that acoustic features extracted from speech signals carry clinically relevant information related to disease severity [1719].

However, much of the existing literature on speech-based PD assessment adopts a cross-sectional perspective, treating individual recordings as independent samples. This assumption neglects the longitudinal structure of telemonitoring data, where repeated observations are collected from the same individual over time [20]. Ignoring temporal dependencies risks conflating inter-patient variability with disease progression and may lead to overly optimistic performance estimates. Moreover, global models that pool all patients implicitly assume a shared disease trajectory, despite well-documented heterogeneity in PD onset, progression rate, and symptom manifestation [21,22]. From a modeling perspective, PD severity is better viewed as a continuously evolving latent disease state that is imperfectly observed through clinical scores and behavioral signals, such as speech. So, the UPDRS score represents a noisy measurement of this latent state rather than an exact ground truth [11]. Consequently, approaches that produce deterministic point estimates without accounting for uncertainty fail to reflect the subjective and variable nature of clinical assessment. This limitation is particularly important in telemonitoring contexts, where remotely generated predictions may inform clinical decision-making. An additional limitation of existing approaches is the disconnect between continuous severity estimation and clinically meaningful severity categories. Although UPDRS is reported as a numerical score, clinicians often interpret disease status in terms of ordered severity stages (e.g., mild, moderate, severe) [9]. However, the telemonitoring dataset used in this study contains only continuous motor UPDRS scores and does not provide clinically validated severity-stage labels. Therefore, a data-driven ordinal categorization is used as an auxiliary modeling strategy to exploit the inherent ordering of disease severity. Models that ignore this ordinal structure may yield numerically accurate predictions that are nonetheless clinically inconsistent. Taken together, these considerations highlight a theoretical gap at the intersection of telemonitoring, disease progression modeling, and clinical interpretability. There is a need for methods that (i) explicitly model longitudinal disease dynamics, (ii) account for patient-specific variability, (iii) quantify uncertainty arising from subjective clinical measurements, and (iv) jointly reconcile continuous severity estimation with ordinal clinical staging.

In this work, an uncertainty-aware personalized temporal modeling framework is proposed for speech-based estimation of PD severity. By modeling disease progression as a latent probabilistic process inferred from longitudinal speech measurements, the framework aims to better capture temporal variability in PD progression. The proposed approach jointly predicts continuous motor UPDRS scores and data-driven severity stages while incorporating predictive uncertainty estimation. Extensive analyses, including temporal extrapolation, sparse-history evaluation, robustness assessment, and uncertainty calibration experiments, are conducted to investigate the behavior and limitations of the framework under realistic longitudinal telemonitoring conditions. The results suggest that personalized temporal modeling has the potential to improve longitudinal tracking of PD severity from speech-derived measurements while highlighting important challenges related to temporal dependence, generalization, and uncertainty in reliability.

2 Related work

2.1 Speech-based analysis for Parkinson’s disease

Speech impairment is a well-established manifestation of PD, and acoustic analysis of voice signals has emerged as a promising non-invasive biomarker for disease assessment [23,24]. Early studies demonstrated that sustained vowel phonations contain discriminative information capable of distinguishing individuals with PD from healthy controls using handcrafted acoustic features and classical machine learning techniques [15,25,26]. In particular, Little et al. [27] and Tsanas et al. [10] showed that nonlinear dysphonia measures combined with regression and classification models can accurately estimate UPDRS scores from voice recordings. Subsequent work expanded this line of research by investigating feature selection strategies, nonlinear modeling approaches, and robustness to recording conditions. Sajal et al. [28] proposed a smartphone-based telemonitoring system that integrates voice and tremor analysis with ensemble learning methods for remote symptom assessment. Ramezani et al. [29] demonstrated that PD progression can be tracked from speech by selecting informative acoustic features using mRMRC and estimating motor UPDRS with regression models, highlighting the importance of spectral, hoarseness-related, and variability-based features. Oliveira et al. [30] showed that decomposing PD voice severity into multiple binary speech classification tasks using DDK features enables moderate multiclass discrimination of disease stages. These studies established speech as a clinically meaningful modality for PD assessment and laid the foundation for speech-based telemonitoring. However, most approaches treated speech recordings as independent observations and focused primarily on population-level inference.

2.2 Telemonitoring and longitudinal Parkinson’s disease studies

To overcome the limitations of clinic-based assessments, telemonitoring systems have been developed to enable frequent, at-home data collection. Palacios et al. [31] proposed Teca-Park, an integrated, contact-free telemonitoring framework that combines speech-based PD monitoring (MonParLoc) with acoustic neurostimulation (AcousticPar), implemented via a mobile app and scorecard, to support longitudinal remote assessment, symptom tracking, and patient management. Platforms such as the At-Home Testing Device (AHTD) demonstrated the feasibility of collecting longitudinal speech and motor data remotely and estimating symptom severity over time. Tsanas et al. [11] showed that self-administered speech recordings can accurately track PD progression by mapping dysphonia features to UPDRS scores with clinically useful precision using the Oxford Parkinson’s Disease Telemonitoring Dataset [32]. Longitudinal studies have further highlighted the importance of modeling disease progression over time. Azuma et al. [33] reported progressive cognitive decline in Parkinson’s patients, even among cognitively normal individuals, emphasizing the value of baseline behavioral markers for predicting future deterioration. Salmanpour et al. [34] demonstrated that combining longitudinal clinical data with DAT-SPECT radiomics and hybrid machine-learning models can uncover distinct progression trajectories and enable early prediction of these trajectories. These studies illustrate the potential of telemonitoring to capture disease dynamics at a finer temporal resolution than conventional clinical practice. Nevertheless, many telemonitoring approaches rely on static or weakly temporal models that do not explicitly represent disease progression, and patient-specific variability is often treated as noise rather than an informative signal.

2.3 Machine learning limitations and theoretical gaps

Recent advances in machine learning, including deep and recurrent architectures, have enabled more powerful modeling of complex biomedical time series [35]. Despite this progress, most speech-based PD models remain deterministic and provide point estimates of disease severity without quantifying uncertainty. Comparative studies have shown that classical regression models can achieve strong predictive performance. Eskidere et al. [36] reported that Least Squares Support Vector Machines (LSSVM) outperformed SVMs, neural networks, and prior state-of-the-art methods for remote UPDRS estimation. Yoon et al. [37] introduced a positive transfer learning (TL)framework that constructs patient-specific models by selectively transferring beneficial information from other subjects, significantly improving prediction accuracy while mitigating negative transfer. Chandrabhatla et al. [38] provided a comprehensive overview of how advances in sensing technologies and machine learning have enabled a shift from subjective, clinic-based assessment toward data-driven, in-home monitoring of motor symptoms.

Despite these advances, existing methods typically decouple continuous severity estimation from clinically meaningful severity staging, despite the ordinal nature of PD progression. Moreover, uncertainty arising from subjective clinical ratings and heterogeneous disease trajectories is rarely modeled explicitly. In summary, prior work has established the feasibility of speech-based PD assessment and telemonitoring but has not fully addressed the combined challenges of longitudinal disease modeling, patient-specific personalization, uncertainty quantification, and ordinal clinical interpretation. The present work seeks to bridge this gap by introducing an uncertainty-aware, personalized temporal framework that more closely reflects the clinical reality of Parkinson’s disease progression.

3 Theoretical motivation

From a modeling perspective, UPDRS scores can be regarded as noisy observations of an underlying latent disease severity process. Let denote the unobserved disease severity of subject i at time t, which evolves according to patient-specific progression dynamics. The observed clinical score can then be expressed as

(1)

where is an unknown measurement function and captures rater variability, contextual effects, and measurement noise. This formulation reflects two key properties of PD assessment: disease severity is continuous, and clinical observations provide imperfect and noisy measurements of this latent state.

Speech production offers a non-invasive behavioral proxy for motor impairment [39] and can be represented as a sequence of acoustic feature vectors extracted from sustained phonations. Given the longitudinal nature of PD, severity estimation should explicitly account for temporal dependencies in speech. Accordingly, the latent disease state may be modeled as a function of the subject’s historical speech observations,

(2)

where denotes a subject-specific temporal mapping. This formulation captures both within-subject temporal structure and inter-subject heterogeneity in disease progression. In clinical practice, PD severity is often interpreted not only as a continuous score but also in terms of ordered severity stages (e.g., mild, moderate, severe) [30]. These stages impose ordinal constraints on the latent disease state and can be naturally modeled via a threshold-based (ordinal) formulation,

(3)

where are ordered thresholds and denotes the ordinal severity stage. This representation motivates joint modeling of continuous severity estimation and ordinal classification, ensuring consistency between numerical predictions and clinically interpretable categories. Finally, both disease progression and clinical assessment are intrinsically uncertain. Representing the latent disease state as a probability distribution rather than a point estimate provides a principled framework for uncertainty quantification [40]. Specifically, it is modeled as,

(4)

where denotes the estimated disease severity and captures epistemic uncertainty arising from limited data as well as observational uncertainty induced by measurement noise.

4 Methodology

4.1 Dataset description

4.1.1 Subjects.

This study uses the Parkinson’s disease telemonitoring dataset [32] collected in a multi-center longitudinal trial and publicly available through the UCI Machine Learning Repository. The original cohort consisted of 52 individuals with idiopathic PD recruited across six U.S. medical centers [41]. Following the exclusion of early dropouts and subjects with insufficient recordings, data from 42 participants (28 males) were retained, each contributing at least 20 valid sessions. All subjects were recently diagnosed, remained unmedicated throughout the six-month study period, and were clinically evaluated using motor and total UPDRS at baseline, three months, and six months.

4.1.2 Data acquisition and features.

Voice data were collected weekly in participants’ homes using the Intel ATHD with a head-mounted microphone sampled at 24 kHz and 16-bit resolution. Each session comprised six sustained phonations of the vowel /a/ recorded under controlled pitch and loudness conditions. After quality screening, a total of 5,923 phonations were retained. Each recording was represented using a set of linear and nonlinear dysphonia measures, producing scalar features used to predict motor and total UPDRS scores. A summary of the dataset characteristics is provided in Table 1.

thumbnail
Table 1. Summary of the Parkinson’s telemonitoring dataset [32].

https://doi.org/10.1371/journal.pone.0343191.t001

4.2 Exploratory data analysis

An Exploratory Dataset Analysis (EDA) was conducted to characterize the longitudinal structure of the dataset and to assess variability in both observation frequency and disease severity across subjects. Fig 1 summarizes key properties of the data at the patient, temporal, and population levels.

thumbnail
Fig 1. Exploratory analysis of the Parkinson’s telemonitoring dataset.

(A) Distribution of the number of visits per subject. (B) Distribution of motor UPDRS scores across all visits. (C) Individual longitudinal motor UPDRS trajectories plotted against time since recruitment. (D) Population-level scatter showing a weak global temporal trend.

https://doi.org/10.1371/journal.pone.0343191.g001

Fig 1(A) shows the distribution of the number of recording sessions per subject. While all included participants contributed a sufficient number of visits, notable variability in visit counts is observed, reflecting irregular and subject-specific sampling schedules. This heterogeneity motivates modeling approaches that can accommodate unequal temporal resolution and leverage subject-wise history rather than assuming uniformly sampled trajectories. Fig 1(B) illustrates the distribution of motor UPDRS scores across all visits. The scores span a wide range of disease severity and exhibit a multimodal structure, suggesting the presence of distinct severity regimes. This observation supports modeling disease severity as a continuous latent variable while also motivating the use of ordinal stratification into clinically meaningful stages.

Individual disease trajectories are shown in Fig 1(C), where motor UPDRS is plotted against time since recruitment for each subject. Substantial inter-individual variability is evident in both baseline severity and progression patterns. Some subjects exhibit gradual monotonic worsening, while others show fluctuating or weakly increasing trends. This diversity supports the assumption of patient-specific progression dynamics rather than a shared global temporal model. At the population level, Fig 1(D) aggregates all observations and reveals only a weak average temporal trend in motor UPDRS over time. The absence of a strong global progression pattern indicates that population-level temporal models are insufficient to capture disease dynamics and further motivates personalized temporal representations conditioned on individual histories.

Overall, the exploratory findings support the theoretical assumptions that PD’s severity evolves heterogeneously across individuals, that observed UPDRS scores represent noisy measurements of an underlying latent disease state, and that uncertainty-aware, patient-specific temporal modeling is required for accurate severity estimation.

4.3 Data preprocessing

Let denote the acoustic feature vector extracted from the t-th voice recording of subject i, where d = 16 corresponds to the selected dysphonia measures. All recordings were temporally ordered by test time for each subject. Each acoustic dimension was standardized using z-score normalization for feature comparability,

(5)

where and denote the mean and standard deviation of feature j computed over the training data.

Motor UPDRS scores were used as continuous regression targets. In addition, ordinal disease severity stages were derived via quantile-based discretization of the motor UPDRS distribution. Let q0.33 and q0.66 denote the 33rd and 66th percentiles, respectively. The ordinal stage label was defined as

(6)

Fixed-length sliding windows were constructed independently for each subject to capture temporal dependencies. For a window size W, the model input at time t was defined as

(7)

with the corresponding regression and ordinal targets given by and . Windows were generated without overlap across subjects, preserving subject-wise temporal ordering.

4.4 Model architecture

The proposed model estimates PD severity using a personalized stochastic latent representation derived from longitudinal speech features. For each subject i, a sequence of standardized acoustic feature windows is provided as input.

Temporal dependencies within each window are modeled using a recurrent encoder. Specifically, an LSTM processes the input sequence and produces a hidden state corresponding to the final time step,

(8)

Then, each subject is associated with a learnable embedding vector , which is concatenated with the temporal representation,

(9)

These embeddings are learned only for training subjects and remain fixed during inference; they are not adapted online for previously unseen patients. The combined representation is mapped to the parameters of a latent disease state distribution. The latent variable is modeled as a Gaussian random variable with mean and diagonal covariance ,

(10)

where and are linear transformations. Sampling is performed via the reparameterization trick,

(11)

Continuous motor UPDRS estimation is obtained through a linear regression head,

(12)

where denotes a fully connected layer. In parallel, ordinal disease severity is estimated using an ordinal regression head with K ordered stages. The probability that the latent state exceeds threshold is given by

(13)

where is the logistic sigmoid function, is a shared projection vector, and are learnable ordered thresholds.

The model outputs the continuous severity estimate , ordinal stage probabilities, and the parameters of the latent disease state distribution for each input window.

4.5 Training objective

Model parameters are learned by minimizing a composite loss function that jointly accounts for continuous severity estimation, ordinal stage prediction, latent state regularization, and cross-task consistency. For continuous motor UPDRS prediction, a mean squared error (MSE) loss is used,

(14)

where and denote the true and predicted motor UPDRS scores, respectively. Ordinal disease severity is modeled using a cumulative link formulation. Let denote the predicted probability that the latent disease state exceeds ordinal threshold k. The ordinal loss is defined as

(15)

where denotes binary cross-entropy and is the indicator function. To regularize the stochastic latent disease state, a Kullback–Leibler divergence [42] term is included,

(16)

encouraging the approximate posterior to remain close to a standard normal prior.

To enforce consistency between continuous predictions and ordinal stage assignments, a constraint-based loss is applied. For each ordinal stage k with interval , the consistency loss [43] is defined as

(17)

The total training objective is given by a weighted sum of the individual components,

(18)

where , , and are fixed hyperparameters. In all experiments, these weights were set to , , and . The proposed framework is illustrated in Fig 2.

thumbnail
Fig 2. Overview of the proposed Parkinson’s disease progression framework using longitudinal voice biomarkers.

https://doi.org/10.1371/journal.pone.0343191.g002

4.6 Experimental setup

All experiments were conducted using a strict patient-wise data split to prevent subject leakage between training and evaluation. Let denote the set of all subjects. A subset comprising 25% of the subjects was randomly selected and held out for testing, while the remaining subjects formed the training set. All recordings from a given subject were assigned exclusively to either the training or test set. Summary statistics for the resulting splits are reported in Table 2.

thumbnail
Table 2. Summary statistics of the patient-wise training and test splits.

https://doi.org/10.1371/journal.pone.0343191.t002

The distributions of motor UPDRS scores in the training and test sets were examined to investigate a systematic shift in disease severity. Fig 3 shows that both subsets exhibit comparable coverage across the severity range, with overlapping density profiles and no pronounced distributional bias. Sliding windows of fixed length W = 10 were constructed independently for each subject, as described in the preprocessing stage. Models were trained using mini-batch stochastic optimization with a batch size of 32 and the Adam optimizer. Training was performed for 100 epochs with a fixed learning rate of 10−3.

thumbnail
Fig 3. Distribution of motor UPDRS scores for training and test sets under the patient-wise split.

Histogram and kernel density estimates indicate comparable severity coverage and overlapping distributions, supporting the validity of the evaluation protocol.

https://doi.org/10.1371/journal.pone.0343191.g003

4.7 Evaluation metrics

Model evaluation was performed exclusively on the held-out test subjects defined in the patient-wise split. All reported metrics were computed using predictions generated without gradient updates. The proposed framework is compared with several naive baselines, including a global mean predictor, subject-specific mean prediction, and a last-observation-carried-forward (LOCF) [44] strategy to better understand the influence of temporal continuity.

Regression performance.

Continuous motor UPDRS estimation was evaluated using mean absolute error (MAE), root mean squared error (RMSE), and the coefficient of determination (R2),

(19)(20)(21)

where and denote the true and predicted motor UPDRS scores, respectively, and is the mean of the ground truth scores.

Patient-level robustness.

MAE was computed separately for each test subject and summarized by the mean and standard deviation across subjects. This analysis evaluates the consistency of model performance under heterogeneous disease trajectories.

Ordinal severity prediction.

Ordinal disease severity prediction was evaluated using classification accuracy,

(22)

where and denote the true and predicted severity stages, respectively. Predicted stages were obtained by counting the number of ordinal thresholds exceeded by the model’s cumulative probabilities.

In addition to global metrics, patient-level robustness was evaluated by computing the mean absolute error separately for each test subject and summarizing the distribution of these errors across subjects. To assess robustness beyond short-term temporal continuity, temporal extrapolation experiments were conducted by training the model on earlier segments of each patient trajectory and evaluating performance on later visits. Additional experiments were performed using reduced temporal window sizes consisting of one or two prior recordings. Predictive uncertainty was evaluated using prediction interval coverage probability (PICP) [45] and reliability calibration analysis after variance scaling calibration. Coverage probabilities for calibrated and confidence intervals were compared against expected Gaussian confidence levels to assess global calibration behavior.

5 Results

5.1 Regression analysis results

First, regression performance was evaluated on the held-out test set, where the proposed model achieved an MAE of 0.561 motor UPDRS points, an RMSE of 0.740, and an R2 of 0.989, indicating accurate estimation of continuous disease severity across the full range of scores. The high (R2) should be interpreted in the context of the dataset and task formulation. The dataset consists of dense longitudinal recordings with repeated measurements per subject and relatively low short-term variability in motor UPDRS, particularly over the six-month study period. As a result, a substantial proportion of the variance is explained by subject-specific baseline severity and short-term temporal continuity. Similar levels of explained variance have been reported in prior work on the same dataset when using longitudinal or subject-aware models [10,46]. Table 3 summarizes the comparison of the proposed framework with several naive baselines.

thumbnail
Table 3. Comparison with naive temporal baselines.

https://doi.org/10.1371/journal.pone.0343191.t003

The LOCF baseline achieved comparatively stronger performance, indicating substantial short-term temporal autocorrelation within the longitudinal UPDRS trajectories. Nevertheless, the proposed framework provides several additional capabilities beyond simple temporal persistence, including uncertainty quantification, latent disease representation learning, and joint continuous-ordinal severity modeling.

Fig 4(A) shows predicted versus ground truth motor UPDRS values for the test set. Predictions closely align with the identity line, with low dispersion across mild to severe severity levels, demonstrating consistent accuracy and absence of systematic bias. Performance remains stable across the severity spectrum, suggesting effective modeling of both low and high disease burden. Uncertainty-aware regression results are illustrated in Fig 4(B), where predictions are sorted by true severity. The predicted trajectory closely follows the ground truth progression, while the estimated uncertainty bands adapt to local variability in the data. Regions exhibiting greater prediction variability are associated with wider uncertainty intervals, reflecting the model’s capacity to express confidence in its estimates.

thumbnail
Fig 4. Regression performance on the test set.

(A) Predicted versus ground truth motor UPDRS scores for unseen patients. (B) Uncertainty-aware regression with predictive intervals, with samples sorted by severity.

https://doi.org/10.1371/journal.pone.0343191.g004

Fig 5 presents representative longitudinal trajectories for two unseen patients. The model accurately tracks individual disease progression over time, capturing both gradual trends and short-term fluctuations. Predicted trajectories closely match observed motor UPDRS scores, demonstrating the personalized temporal model’s ability to generalize to new subjects while preserving subject-specific progression patterns.

thumbnail
Fig 5. Longitudinal motor UPDRS trajectories for two representative unseen patients.

Predicted trajectories closely follow ground truth measurements across time.

https://doi.org/10.1371/journal.pone.0343191.g005

Under this substantially challenging setting of temporal extrapolation experiment, the proposed framework achieved an MAE of 2.994, RMSE of 3.946, and R2 = 0.774. Although performance decreased compared to the original longitudinal evaluation, the model retained reasonably strong predictive capability, suggesting that the framework captures meaningful longer-range temporal disease structure beyond simple local persistence effects.

5.2 Ordinal classification results

Ordinal disease severity prediction was also evaluated on the held-out test set consisting of unseen patients. The model achieved an overall classification accuracy of 0.916 across the three severity stages. Table 4 reports precision, recall, and F1-score for each class, along with macro-averaged and weighted-averaged performance. Performance is balanced across classes, with high recall for mild and severe stages and slightly lower recall for the moderate stage.

thumbnail
Table 4. Ordinal classification performance on the test set.

https://doi.org/10.1371/journal.pone.0343191.t004

Fig 6(A) shows the confusion matrix. Most predictions lie on the diagonal, indicating correct stage assignment, while misclassifications primarily occur between adjacent severity levels. Fig 6(B) visualizes the learned latent disease representations projected using t-SNE, where samples are colored according to true severity stage. The latent space exhibits structured separation between stages, with smooth transitions between neighboring severity levels.

thumbnail
Fig 6. Ordinal classification results on the test set.

(A) Confusion matrix for three-stage disease severity classification. (B) Two-dimensional t-SNE projection of the learned latent disease states, colored by true severity stage.

https://doi.org/10.1371/journal.pone.0343191.g006

5.3 Feature importance analysis

A permutation-based feature importance analysis was conducted to assess the relative contribution of individual acoustic features to prediction. For each feature, values were randomly permuted across samples while keeping all other features unchanged, and the resulting increase in MAE was recorded.

Fig 7 shows the change in MAE induced by permuting each feature, averaged across permutations. Larger increases in MAE indicate greater importance for prediction. Features related to harmonicity and nonlinear vocal dynamics, including HNR, RPDE, and PPE, exhibit the largest impact on performance. Measures of jitter and shimmer also contribute substantially, indicating sensitivity to both frequency and amplitude perturbations in sustained phonation.

thumbnail
Fig 7. Permutation-based global feature importance analysis.

Bars indicate the increase in MAE after permuting each acoustic feature, with larger values corresponding to greater importance for motor UPDRS prediction.

https://doi.org/10.1371/journal.pone.0343191.g007

5.4 Model analysis and uncertainty characterization

Further analysis of the proposed personalized latent framework focuses on its sensitivity to key architectural hyperparameters and on the properties of the learned latent disease representation and associated predictive uncertainty.

Hyperparameter sensitivity.

Fig 8 summarizes model performance as a function of temporal window length, latent state dimensionality, and patient embedding size. Increasing the temporal window initially yields substantial performance gains, reflecting the benefit of incorporating short-term longitudinal context, after which performance saturates for window lengths beyond approximately 10 recordings. A similar trend is observed for the latent state dimensionality and patient embedding size, where performance improves up to an intermediate capacity and then plateaus. These results indicate that the proposed model is robust to hyperparameter selection and does not rely on excessively large representations to achieve strong performance.

thumbnail
Fig 8. Hyperparameter sensitivity analysis.

(A) Temporal window length. (B) Latent state dimensionality. (C) Patient embedding size. Performance is reported in terms of MAE and R2 on the held-out test set.

https://doi.org/10.1371/journal.pone.0343191.g008

Uncertainty-aware error analysis.

To characterize the behavior of predictive uncertainty, test samples were stratified into uncertainty levels based on the posterior standard deviation of the latent disease state. Fig 9(A) shows the distribution of absolute prediction error across low, medium, and high uncertainty groups. While predictions associated with higher uncertainty exhibit somewhat broader error distributions, the median prediction error remains relatively similar across groups. This observation is consistent with the quantitative calibration analysis, which showed only a weak association between predictive uncertainty and sample-wise prediction error. Consequently, the current uncertainty estimates should be interpreted as reflecting overall predictive variability rather than as validated indicators of individual prediction reliability.

thumbnail
Fig 9. Properties of the learned latent disease representation.

(A) Distribution of absolute prediction error across uncertainty levels. (B) Relationship between latent disease severity and motor UPDRS scores.

https://doi.org/10.1371/journal.pone.0343191.g009

Clinical interpretability of the latent disease state.

Fig 9(B) illustrates the relationship between the learned latent representation and motor UPDRS scores. A monotonic association is observed, suggesting that the latent space captures information related to disease severity, and deviation from linearity likely reflects the heterogeneous progression of PD and measurement variability in clinical scoring.

5.5 Uncertainty calibration analysis

Uncertainty calibration was further evaluated using PICP and reliability calibration analysis. Table 5 summarizes the quantitative calibration results.

thumbnail
Table 5. Quantitative uncertainty calibration results after variance scaling calibration.

https://doi.org/10.1371/journal.pone.0343191.t005

Fig 10 illustrates the calibration behavior of the uncertainty estimates after post-hoc uncertainty scaling. Fig 10(A) compares expected and observed prediction interval coverage probabilities for and confidence intervals. The calibrated intervals achieved empirical coverage probabilities reasonably close to the expected Gaussian confidence levels, although mild undercoverage remained at higher confidence intervals. Fig 10(B) presents the reliability calibration curve across multiple nominal confidence levels. Observed coverage followed the ideal calibration trend consistently but remained slightly below the diagonal, indicating residual underestimation of predictive uncertainty after calibration. While the scaling procedure improved global interval calibration, the near-zero correlation between predictive uncertainty and sample-wise prediction error () suggests that the current latent uncertainty representation remains weakly associated with local prediction difficulty.

thumbnail
Fig 10. Uncertainty calibration analysis after variance scaling calibration.

(A) Expected versus observed prediction interval coverage probabilities for and intervals. (B) Reliability calibration curve comparing expected and empirical coverage across multiple confidence levels.

https://doi.org/10.1371/journal.pone.0343191.g010

5.6 Robustness to synthetic label noise

To evaluate robustness against noisy clinical annotations, an additional experiment was conducted in which synthetic Gaussian noise was injected into the motor UPDRS labels during training. Noise levels with standard deviations ranging from to were evaluated, while testing was performed on clean labels. Table 6 summarizes the results. Predictive performance degraded gradually as label corruption increased, indicating that the proposed framework remained relatively stable under moderate annotation noise. Even under substantial perturbation (), the model retained strong predictive capability with R2 = 0.963. In addition, predictive uncertainty increased moderately with larger noise levels, suggesting that the probabilistic latent formulation partially reflects increased ambiguity in the supervision signal.

thumbnail
Table 6. Robustness analysis under synthetic Gaussian label noise added to motor UPDRS targets during training.

https://doi.org/10.1371/journal.pone.0343191.t006

5.7 Limited-history prediction analysis

Experiments were conducted using temporal windows consisting of only one or two prior recordings to assess model performance under a limited longitudinal context. With a single-recording window, the proposed framework achieved an MAE of 1.710, an RMSE of 2.292, and an R2 of 0.920. Using two recordings further improved performance, yielding an MAE of 1.552, an RMSE of 2.087, an R2 = 0.934, and an average predictive uncertainty of 0.112. Although performance improved with longer temporal context, as shown in Fig 8, the model remained reasonably accurate even under extremely limited history, suggesting usefulness in practical real-world settings where prior recordings are scarce, such as newly enrolled patients or sparse telemonitoring scenarios.

5.8 Ablation Analysis

An ablation study was conducted to evaluate the contribution of each major component of the proposed framework, including ordinal supervision, probabilistic latent inference, patient-specific embeddings, and temporal modeling. Results are summarized in Table 7.

thumbnail
Table 7. Ablation analysis of the proposed framework.

https://doi.org/10.1371/journal.pone.0343191.t007

Removing any individual component resulted in performance degradation relative to the full model, indicating that each module contributes positively to overall prediction quality. The largest performance decrease was observed when temporal modeling was replaced with a non-temporal MLP encoder, highlighting the importance of longitudinal temporal structure for disease severity estimation. Removal of patient embeddings also reduced performance, suggesting that personalization captures subject-specific disease characteristics. In contrast, removing ordinal supervision or probabilistic latent inference produced smaller but consistent decreases in accuracy, indicating that both components provide complementary regularization and improve the stability of the longitudinal representation.

5.9 Comparison with existing works

Table 8 compares the proposed framework with prior speech-based approaches for PD severity estimation on the telemonitoring dataset [32]. Early studies by Tsanas et al. [10,11] employed classical machine learning models such as CART and LASSO, reporting MAE in the range of 5–7 UPDRS points. Subsequent work by Eskidere et al. [36] improved predictive accuracy using LSSVM, achieving an MAE of 4.87. These approaches primarily relied on cross-sectional modeling and did not explicitly account for longitudinal structure or patient-specific variability. More recent studies have explored nonlinear and hybrid learning paradigms. Nilashi et al. [46] reported strong performance using an adaptive neuro-fuzzy inference system (ANFIS), achieving a low MAE of 0.491 and an R2 of 0.959. Mohammadi et al. [47] employed decision tree–based regression (REPTree), while Tang et al. [48] proposed the NoRo framework, both yielding moderate reductions in prediction error. However, these methods generally focused on point estimation and did not provide explicit uncertainty quantification or jointly model continuous and ordinal disease severity.

thumbnail
Table 8. Comparison with existing speech-based PD severity estimation methods on the telemonitoring dataset.

https://doi.org/10.1371/journal.pone.0343191.t008

In comparison, the proposed framework achieves competitive predictive performance (MAE = 0.56, R2 = 0.99) while incorporating personalized temporal modeling, probabilistic latent inference, ordinal severity supervision, and uncertainty-aware prediction within a unified formulation. Although direct numerical comparison should be interpreted cautiously due to differences in evaluation protocols and data splits, the results suggest that integrating temporal context and personalization can improve longitudinal speech-based assessment of Parkinson’s disease severity.

6 Discussion

This study presents a personalized, uncertainty-aware framework for estimating PD severity from longitudinal speech recordings. By jointly modeling temporal dynamics, patient-specific representations, and a stochastic latent disease state, the proposed approach enables simultaneous prediction of continuous motor UPDRS scores and data-driven ordinal severity stages in a remote telemonitoring setting.

The regression results demonstrate accurate estimation of motor UPDRS on held-out test patients, with low prediction error and high explained variance. Predictive performance was generally consistent across the observed range of motor UPDRS scores, suggesting that the model did not exhibit an obvious bias toward a particular severity range. Consistent patient-level performance further suggests that incorporating subject-specific embeddings and temporal context may help capture aspects of inter-individual heterogeneity in disease progression, although additional validation on larger and more diverse cohorts is needed to assess generalizability. The auxiliary ordinal prediction task complements continuous regression by providing data-driven ordinal severity categories derived from the motor UPDRS distribution. Most misclassifications occurred between adjacent categories, which is consistent with the ordered nature of the labels and the gradual progression of disease severity. The agreement between continuous severity estimates and the corresponding ordinal categories suggests that the joint learning framework learns internally consistent continuous and ordinal representations.

A key contribution of this work is the explicit modeling of predictive uncertainty through a stochastic latent disease representation. Uncertainty-aware analyses show that predicted confidence intervals vary across samples and provide reasonably calibrated coverage after post-hoc variance scaling. However, the observed weak association between predictive uncertainty and sample-wise prediction error suggests that the current uncertainty estimates should not be interpreted as validated indicators of individual prediction reliability. Instead, they provide a probabilistic characterization of model predictions that may be useful for understanding overall predictive variability while requiring further validation for clinical applications. Analysis of the learned latent space suggests a structured relationship with disease severity, with gradual transitions across data-driven ordinal categories. This behavior suggests that the latent representation captures a continuous disease manifold rather than discrete class boundaries, consistent with the gradual and heterogeneous progression of PD (see Fig 9(B)). The observed relationship between latent severity and motor UPDRS further establishes a promising goal towards the clinical interpretability of the latent disease state.

Permutation-based feature importance analysis showed that both linear and nonlinear dysphonia measures contribute meaningfully to prediction, consistent with prior evidence linking vocal instability and reduced harmonic structure to Parkinsonian speech impairment [10,36]. Temporal extrapolation and limited-history experiments further demonstrated that the model retained stable predictive capability under more challenging longitudinal conditions, including sparse recording histories. Robustness analysis under synthetic label noise revealed gradual performance degradation with increasing perturbation levels, while ablation experiments confirmed the importance of temporal modeling and patient-specific embeddings. Finally, uncertainty calibration analysis showed reasonably consistent global coverage behavior after calibration, although uncertainty remained weakly associated with sample-wise prediction difficulty.

Several limitations should be acknowledged. The dataset spans a relatively short six-month period and includes only unmedicated patients, which restricts the observable range of disease progression. In addition, the high explained variance observed in this study is influenced by the dense longitudinal structure and limited temporal span of the dataset and may not directly translate to settings with sparser sampling or longer-term disease evolution. Longer-term datasets and medication-state variability may introduce additional dynamics not captured in this study. Future work should investigate adaptation strategies for truly unseen patients [49], including meta-learning [50], adaptive embedding initialization [51], and online personalization mechanisms for realistic longitudinal telemonitoring deployment. In addition, ordinal severity stages were derived using quantile-based thresholds, which, while data-driven, may not correspond exactly to clinically defined stage boundaries, such as Hoehn-Yahr scaling [52]. A further limitation is that patient-specific embeddings are learned only for training subjects and are not updated for unseen patients. Consequently, the current framework does not address cold-start personalization for newly enrolled individuals.

Future work should investigate clinically established staging criteria on a larger longitudinal dataset and extend this framework by incorporating multimodal signals, such as accelerometry or handwriting data, and by explicitly modeling medication effects. Integrating clinician-in-the-loop calibration [53] or longitudinal Bayesian updating [54] may further enhance reliability in real-world telemonitoring applications. Future work should also investigate online adaptation [55] and meta-learning strategies to progressively refine patient representations. Consequently, the findings may not fully generalize to broader clinical populations with heterogeneous disease stages, medication conditions, and recording environments.

7 Conclusion

This work presents a personalized, uncertainty-aware framework for estimating PD severity from longitudinal speech recordings by modeling disease progression as a stochastic latent process and incorporating patient-specific temporal representations. The proposed framework jointly predicts continuous motor UPDRS scores and data-driven ordinal severity categories from remotely collected voice data. Experimental results on a public telemonitoring dataset suggest accurate severity estimation on unseen patients, consistent patient-level performance, and the potential benefits of integrating temporal modeling with patient-specific representations. The probabilistic latent formulation provides calibrated uncertainty estimates that characterize predictive variability, although further validation is required before these estimates can be considered reliable indicators of individual prediction confidence. Analysis of the learned latent representation suggests that it captures information related to disease severity, while feature importance analysis identifies several dysphonia measures that contribute substantially to prediction. Overall, this study provides a unified framework for temporal modeling, personalization, ordinal learning, and uncertainty quantification and establishes a promising direction for future validation on larger and more diverse longitudinal cohorts, as well as multimodal extensions for PD telemonitoring.

Acknowledgments

The author gratefully acknowledges Bangladesh University of Engineering and Technology (BUET) for providing institutional support and facilities that enabled the completion of this research.

References

  1. 1. Ben-Shlomo Y, Darweesh S, Llibre-Guerra J, Marras C, San Luciano M, Tanner C. The epidemiology of Parkinson’s disease. Lancet. 2024;403(10423):283–92. pmid:38245248
  2. 2. Bloem BR, Okun MS, Klein C. Parkinson’s disease. Lancet. 2021;397(10291):2284–303. pmid:33848468
  3. 3. Marsden CD. Parkinson’s disease. J Neurol Neurosurg Psychiatry. 1994;57(6):672–81. pmid:7755681
  4. 4. Movement Disorder Society Task Force on Rating Scales for Parkinson’s Disease. The unified Parkinson’s disease rating scale (UPDRS): status and recommendations. Movement Disorders. 2003;18(7):738–50.
  5. 5. Sveinbjornsdottir S. The clinical symptoms of Parkinson’s disease. J Neurochem. 2016;139 Suppl 1:318–24. https://doi.org/10.1111/jnc.13691 pmid:27401947
  6. 6. Shahed J, Jankovic J. Motor symptoms in Parkinson’s disease. Handb Clin Neurol. 2007;83:329–42. pmid:18808920
  7. 7. Visser M, Marinus J, Bloem BR, Kisjes H, van den Berg BM, van Hilten JJ. Clinical tests for the evaluation of postural instability in patients with parkinson’s disease. Arch Phys Med Rehabil. 2003;84(11):1669–74. pmid:14639568
  8. 8. Shulman LM, Gruber-Baldini AL, Anderson KE, Fishman PS, Reich SG, Weiner WJ. The clinically important difference on the unified Parkinson’s disease rating scale. Arch Neurol. 2010;67(1):64–70. pmid:20065131
  9. 9. Martínez-Martín P, Rodríguez-Blázquez C, Mario Alvarez, Arakaki T, Arillo VC, Chaná P, et al. Parkinson’s disease severity levels and MDS-Unified Parkinson’s Disease Rating Scale. Parkinsonism Relat Disord. 2015;21(1):50–4. pmid:25466406
  10. 10. Tsanas A, Little M, McSharry P, Ramig L. Accurate telemonitoring of Parkinson’s disease progression by non-invasive speech tests. Nat Prec. 2009. https://doi.org/10.1038/npre.2009.3920.1
  11. 11. Tsanas A, Little MA, McSharry PE, Ramig LO. Enhanced classical dysphonia measures and sparse regression for telemonitoring of Parkinson’s disease progression. In: 2010 IEEE International Conference on Acoustics, Speech and Signal Processing, 2010. 594–7. https://doi.org/10.1109/icassp.2010.5495554
  12. 12. Achey M, Aldred JL, Aljehani N, Bloem BR, Biglan KM, Chan P, et al. The past, present, and future of telemedicine for Parkinson’s disease. Mov Disord. 2014;29(7):871–83. pmid:24838316
  13. 13. Konitsiotis S, Alexoudi A, Zikos P, Sidiropoulos C, Tagaris G, Xiromerisiou G, et al. Paradigm shift in Parkinson’s disease: using continuous telemonitoring to improve symptoms control. Results from a 2-years journey. Front Neurol. 2024;15:1415970. pmid:38903169
  14. 14. Amato F, Borzi L, Olmo G, Artusi CA, Imbalzano G, Lopiano L. Speech Impairment in Parkinson’s Disease: Acoustic Analysis of Unvoiced Consonants in Italian Native Speakers. IEEE Access. 2021;9:166370–81.
  15. 15. Almeida JS, Rebouças Filho PP, Carneiro T, Wei W, Damaševičius R, Maskeliūnas R, et al. Detecting Parkinson’s disease with sustained phonation and speech signals using machine learning techniques. Pattern Recognition Letters. 2019;125:55–62.
  16. 16. Suppa A, Costantini G, Asci F, Di Leo P, Al-Wardat MS, Di Lazzaro G, et al. Voice in Parkinson’s Disease: A Machine Learning Study. Front Neurol. 2022;13:831428. pmid:35242101
  17. 17. Asgari M, Shafran I. Predicting severity of Parkinson’s disease from speech. Annu Int Conf IEEE Eng Med Biol Soc. 2010;2010:5201–4. pmid:21095825
  18. 18. Bayestehtashk A, Asgari M, Shafran I, McNames J. Fully Automated Assessment of the Severity of Parkinson’s Disease from Speech. Comput Speech Lang. 2015;29(1):172–85. pmid:25382935
  19. 19. Rábano-Suárez P, Del Campo N, Benatru I, Moreau C, Desjardins C, Sánchez-Ferro Á, et al. Digital Outcomes as Biomarkers of Disease Progression in Early Parkinson’s Disease: A Systematic Review. Mov Disord. 2025;40(2):184–203. pmid:39613480
  20. 20. Lewinski AA, Walsh C, Rushton S, Soliman D, Carlson SM, Luedke MW, et al. Telehealth for the Longitudinal Management of Chronic Conditions: Systematic Review. J Med Internet Res. 2022;24(8):e37100. pmid:36018711
  21. 21. Wüllner U, Borghammer P, Choe C-U, Csoti I, Falkenburger B, Gasser T, et al. The heterogeneity of Parkinson’s disease. J Neural Transm (Vienna). 2023;130(6):827–38. pmid:37169935
  22. 22. Greenland JC, Williams-Gray CH, Barker RA. The clinical heterogeneity of Parkinson’s disease and its therapeutic implications. Eur J Neurosci. 2019;49(3):328–38. pmid:30059179
  23. 23. Ho AK, Iansek R, Marigliani C, Bradshaw JL, Gates S. Speech impairment in a large sample of patients with Parkinson’s disease. Behav Neurol. 1999;11(3):131–7. pmid:22387592
  24. 24. Smith KM, Caplan DN. Communication impairment in Parkinson’s disease: Impact of motor and cognitive symptoms on speech and language. Brain Lang. 2018;185:38–46. pmid:30092448
  25. 25. Vaiciukynas E, Verikas A, Gelzinis A, Bacauskiene M. Detecting Parkinson’s disease from sustained phonation and speech signals. PLoS One. 2017;12(10):e0185613. pmid:28982171
  26. 26. Arias-Vergara T, Vásquez-Correa JC, Orozco-Arroyave JR. Parkinson’s Disease and Aging: Analysis of Their Effect in Phonation and Articulation of Speech. Cogn Comput. 2017;9(6):731–48.
  27. 27. Little M, McSharry P, Roberts S, Costello D, Moroz I. Exploiting Nonlinear Recurrence and Fractal Scaling Properties for Voice Disorder Detection. Nat Prec. 2007. https://doi.org/10.1038/npre.2007.326.1
  28. 28. Sajal MSR, Ehsan MT, Vaidyanathan R, Wang S, Aziz T, Mamun KAA. Telemonitoring Parkinson’s disease using machine learning by combining tremor and voice analysis. Brain Inform. 2020;7(1):12. pmid:33090328
  29. 29. Ramezani H, Khaki H, Erzin E, Akan OB. Speech features for telemonitoring of Parkinson’s disease symptoms. Annu Int Conf IEEE Eng Med Biol Soc. 2017;2017:3801–5. pmid:29060726
  30. 30. Oliveira GC, Pah ND, Ngo QC, Yoshida A, Gomes NB, Papa JP, et al. A pilot study for speech assessment to detect the severity of Parkinson’s disease: An ensemble approach. Comput Biol Med. 2025;185:109565. pmid:39709867
  31. 31. Palacios-Alonso D, Melendez-Morales G, Lopez-Arribas A, Lazaro-Carrascosa C, Gomez-Rodellar A, Gomez-Vilda P. MonParLoc: A Speech-Based System for Parkinson’s Disease Analysis and Monitoring. IEEE Access. 2020;8:188243–55.
  32. 32. Tsanas A, Little M. Parkinsons Telemonitoring. UCI Machine Learning Repository. 2009. https://doi.org/10.24432/C5ZS3N
  33. 33. Azuma T, Cruz RF, Bayles KA, Tomoeda CK, Montgomery EB Jr. A longitudinal study of neuropsychological change in individuals with Parkinson’s disease. Int J Geriatr Psychiatry. 2003;18(11):1043–9. pmid:14618557
  34. 34. Salmanpour MR, Shamsaei M, Hajianfar G, Soltanian-Zadeh H, Rahmim A. Longitudinal clustering analysis and prediction of Parkinson’s disease progression using radiomics and hybrid machine learning. Quant Imaging Med Surg. 2022;12(2):906–19. pmid:35111593
  35. 35. Bock C, Moor M, Jutzeler CR, Borgwardt K. Machine Learning for Biomedical Time Series Classification: From Shapelets to Deep Learning. Methods Mol Biol. 2021;2190:33–71. pmid:32804360
  36. 36. Eskidere Ö, Ertaş F, Hanilçi C. A comparison of regression methods for remote tracking of Parkinson’s disease progression. Expert Systems with Applications. 2012;39(5):5523–8.
  37. 37. Yoon H, Li J. A Novel Positive Transfer Learning Approach for Telemonitoring of Parkinson’s Disease. IEEE Trans Automat Sci Eng. 2019;16(1):180–91.
  38. 38. Chandrabhatla AS, Pomeraniec IJ, Ksendzovsky A. Co-evolution of machine learning and digital technologies to improve monitoring of Parkinson’s disease motor symptoms. NPJ Digit Med. 2022;5(1):32. pmid:35304579
  39. 39. Quatieri TF, Williamson JR, Lambert A, Heaton KJ, Palmer JS. Noninvasive biomarkers of neurobehavioral performance. Linc Lab J. 2020;24:28–59.
  40. 40. Gray A, Wimbush A, de Angelis M, Hristov PO, Calleja D, Miralles-Dolz E, et al. From inference to design: A comprehensive framework for uncertainty quantification in engineering with limited information. Mechanical Systems and Signal Processing. 2022;165:108210.
  41. 41. Goetz CG, Stebbins GT, Wolff D, DeLeeuw W, Bronte-Stewart H, Elble R, et al. Testing objective measures of motor impairment in early Parkinson’s disease: Feasibility study of an at-home testing device. Mov Disord. 2009;24(4):551–6. pmid:19086085
  42. 42. van Erven T, Harremoes P. Rényi Divergence and Kullback-Leibler Divergence. IEEE Trans Inform Theory. 2014;60(7):3797–820.
  43. 43. Shahriar K. Lightweight Resolution-Aware Audio Deepfake Detection via Cross-Scale Attention and Consistency Learning. arXiv preprint arXiv:260106560. 2026.
  44. 44. Lachin JM. Fallacies of last observation carried forward analyses. Clin Trials. 2016;13(2):161–8. pmid:26400875
  45. 45. Khosravi A, Mazloumi E, Nahavandi S, Creighton D, van Lint JWC. Prediction Intervals to Account for Uncertainties in Travel Time Prediction. IEEE Trans Intell Transport Syst. 2011;12(2):537–47.
  46. 46. Nilashi M, Ibrahim O, Samad S, Ahmadi H, Shahmoradi L, Akbari E. An analytical method for measuring the Parkinson’s disease progression: A case on a Parkinson’s telemonitoring dataset. Measurement. 2019;136:545–57.
  47. 47. Mohammadi P, Hatamlou A, Masdari M. A comparative study on remote tracking of Parkinsons disease progression using data mining methods. arXiv preprint arXiv:13122140. 2013.
  48. 48. Tang Z, Hou C, Zhang T, Tian B, Wang J, Lv H. Enhancing Noise Robustness of Parkinson’s Disease Telemonitoring via Contrastive Feature Augmentation. arXiv preprint arXiv:251001588. 2025.
  49. 49. Bleeker SE, Moll HA, Steyerberg EW, Donders ART, Derksen-Lubsen G, Grobbee DE, et al. External validation is necessary in prediction research: a clinical example. J Clin Epidemiol. 2003;56(9):826–32. pmid:14505766
  50. 50. Jia Z, Shi Y, Hu J. Personalized Neural Network for Patient-Specific Health Monitoring in IoT: A Metalearning Approach. IEEE Trans Comput-Aided Des Integr Circuits Syst. 2022;41(12):5394–407.
  51. 51. Yang Y, Sun Z-Q, Zhu H, Fu Y, Zhou Y, Xiong H, et al. Learning Adaptive Embedding Considering Incremental Class. IEEE Trans Knowl Data Eng. 2021;:1–1. https://doi.org/10.1109/tkde.2021.3109131
  52. 52. Goetz CG, Poewe W, Rascol O, Sampaio C, Stebbins GT, Counsell C, et al. Movement Disorder Society Task Force report on the Hoehn and Yahr staging scale: status and recommendations the Movement Disorder Society Task Force on rating scales for Parkinson’s disease. Movement disorders. 2004;19(9):1020–8.
  53. 53. Tang S, Modi A, Sjoding M, Wiens J. Clinician-in-the-loop decision making: Reinforcement learning with near-optimal set-valued policies. In: International Conference on Machine Learning. PMLR. 2020. 9387–96.
  54. 54. Tucker A, Li Y, Garway-Heath D. Updating Markov models to integrate cross-sectional and longitudinal studies. Artif Intell Med. 2017;77:23–30. pmid:28545609
  55. 55. Marceglia S, Rossi E, Rosa M, Cogiamanian F, Rossi L, Bertolasi L, et al. Web-based telemonitoring and delivery of caregiver support for patients with Parkinson disease after deep brain stimulation: protocol. JMIR Res Protoc. 2015;4(1):e30. pmid:25803512