Figures
Abstract
Background
Most conventional diagnostic systems rely on fixed thresholds to differentiate disease states from normal. However, early pathological changes may begin before these thresholds are crossed. Therefore, a system that works in this pre-threshold state can potentially lead to earlier diagnosis.
Objective
To propose and evaluate a geometric framework that models early disease as a directional drift from a physiological plane to a pathological plane, allowing for pre-threshold detection using biologically interpretable variables.
Methods
In this modeling study on synthetic data derived from published clinical trends, two clinically meaningful variables were used to define a 2D feature space. The model was applied to a synthetic dataset of 4000 eyes divided into four phenotypes: normal stable (NS), early disease stable (ED_S), early disease progressive (ED_P), and pre-threshold progressive (PT_P). A physiological plane was constructed using range-normalized values from the NS group. A canonical disease vector was derived from the ED_P group. Each subject's follow-up data was transformed into a subject-specific drift vector, and the Composite Drift Score (CDS) was calculated as the product of directional alignment (Directional Emphasis Multiplier, DEM) and a Magnitude-to-Noise Ratio (MNR).
Results
CDS increased significantly over follow-ups in both ED_P and PT_P, distinguishing them from the two stable cohorts (p < 0.001). DEM and MNR components showed consistent trends, with progressive cases exhibiting higher alignment with the disease vector and supra-noise magnitude of change. Visual and statistical analyses confirmed early drift detection even within numerically normal ranges.
Citation: Prakash G (2026) Directional drift in biologically meaningful vector planes: A proposed geometric framework for early detection of subthreshold disease. PLoS One 21(7): e0353723. https://doi.org/10.1371/journal.pone.0353723
Editor: Zeheng Wang, Commonwealth Scientific and Industrial Research Organisation, AUSTRALIA
Received: February 21, 2026; Accepted: June 26, 2026; Published: July 30, 2026
Copyright: © 2026 Gaurav Prakash. This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.
Data Availability: All relevant data are within the manuscript and its Supporting Information files.
Funding: The author(s) received no specific funding for this work.
Competing interests: NO authors have competing interests.
Introduction
Over the past few decades, there have been significant improvements in diagnostic technology, particularly in imaging and laboratory methods. This has allowed clinicians to detect increasingly subtle physiological variations, even in individuals without overt disease [1–5]. Despite these capabilities, many diagnostic systems continue to depend on fixed thresholds to define disease. This approach can overlook earlier physiological changes that precede those cutoffs [6–14].
Due to their underlying design, these threshold-based, cut-off style tests tend to perform well in structured datasets where subjects are clearly categorized as either normal or diseased. However, in real-world clinical settings, normal and diseased cases often share overlapping characteristics across multiple parameters. [6,7,13,14].
Futhermore, oversights with multivariable threshold style testing can lead to redundancy and overfitting and reduced diagnostic accuracy [15–23].
While threshold-based tests are useful for categorizing disease at a single point in time, they cannot account for how a subject reached that point (Fig 1A, 1B). For example, in Fig 1A, patient P1 remains stable and above the threshold (disease zone) at both and
, while P2 worsens but only crosses the threshold at the second time point. Despite their very different disease dynamics (P1 has stable disease and P2 has worsening disease), both appear similar if only the current value is considered. Also, deviation away from normal state and progression of worsening often begins well before any statistical cutoffs are breached. For example, Fig 1B shows two more patients, P3 and P4, coming for their first-ever visit at time t. Both are within the so-called normal range currently. However, P3 has been drifting steadily toward the disease state from before, while P4 is stable. A static threshold cannot distinguish between them, failing to detect the early directional trend in P3. Early pathological drift begins when the subject or the organ is still numerically within the accepted values for ‘normal’ population distribution. This phenomenon has been documented in many chronic progressive diseases [24–28].
1A. Two patients in the disease zone may appear clinically similar at a single timepoint despite very different trajectories leading up to disease. P1 is stable and P2 is worsening. 1B. Two subjects in the normal zone may be on very different paths. P4 is stable but P3 was worsening.
Clinicians have long recognized that normal range, sub-threshold rising values (more than the reference change value), can be concerning [29–31]. However, it has been challenging to formally characterize or quantify this type of change using traditional tools.
Worsening this issue is the assumption that physiological parameters follow Gaussian distributions, which is often not accurate [7,32–34]. Cross-sectional cutoffs are further complicated by changes in how these scores, or numerical thresholds are defined: as clinical panels revise cutoffs, the decision boundary itself moves one way or the other [30,35].
One way to circumvent these issues is to detect drift, a directional deviation away from the normal state.
To explore the potential of using drift as an alternate vantage point, we should first review natural fluctuations (to due homeostasis) and observation-related noise and formalize our definitions with reference to this study.
Healthy physiological systems maintain a stable state through active regulation resulting in natural fluctuations. Homeostasis is the organism’s tendency to maintain this state. Physiological variation can therefore be defined as homeostatic self-correction oscillations occur within bounded ranges. [36–38]. Observation variation, on the other hand, arises due to multiple reasons within the device-operator-interpreter system. This occurs whenever a measurement is attempted on a biological system or structure [39–46]. Observation variation, like physiological variation, is typically non-directional and self-correcting. For example, diurnal variations in serum glucose or intraocular pressure (IOP), or postural changes in blood pressure (BP), are expected and healthy [36–38]. Similarly, when measuring IOP, the intra-measurement variation arising from the tonometer, operator experience, and protocol introduces an observation effect that is random and self-correcting, not directional [39–46].
Physiological and observational variation can interact stochastically. There, may not be a clear rule to discriminate them from each other. Therefore, it may be mathematically useful to club them into a single entity, ‘noise’.
This is different from an actual movement aligned with disease -like worsening, which we term Drift. Many chronic progressive diseases involve continued cellular deterioration long before diagnostic thresholds are crossed [47–55]. Let us assume that a patient presents to the clinic for annual health checkup and has an HbA1c level of 5.5%, This will be classified as non-diabetic and therefore reassuring. However, if this value rose from 4.5% over a single year, the rate of change may exceed expected biological variability despite remaining below the diagnostic cutoff. Similarly, in keratoconus, an increase in maximum keratometry (Kmax) from 42.0 to 44.5 Diopters may not raise concern because it has not yet crossed the conventional threshold of 47.2D, yet the directional trend is clinically meaningful. This concept extends into multiple clinical domains: in glaucoma, retinal nerve fibre layer thinning may precede visual field loss by years while intraocular pressure remains within the normal range [36]; in chronic kidney disease, directional changes in eGFR and albuminuria signal progressive nephron loss before staging thresholds are crossed [49]; and in Alzheimer's disease, amyloid and tau biomarker trajectories diverge from normal years before cognitive thresholds are breached [24,25]. Across these domains, the common thread is that disease progression is directional and accumulative, and therefore potentially detectable, long before any single numerical cutoff is crossed.
Therefore, Drift can be defined as a slow, directional deviation away from this physiological corridor aligned with the expected direction of pathology progression. This occurs before conventional diagnostic thresholds are crossed. It is not defined by a single measurement but by a trajectory over time. This is often still within the so-called “normal” numerical range. [47–55].
It can be argued that medicine, in essence, is systems biology viewed through the lens of pathology. In systems biology, it is well recognized that transitions in system behavior are often preceded by early warning signals. These signs include increased variance, autocorrelation, or critical slowing down, where recovery from oscillations becomes delayed [56–59]. These signals suggest that a system is approaching a tipping point, after which a spontaneous return to the original state becomes unlikely.In our keratoconus example above, the subject trending from 42.0 to 44.5 Diopters over follow-up was demonstrating this Drift, numerically normal, yet directionally meaningful. This concept of drift from a new post-tipping baseline extends beyond keratoconus; in post-transplant settings, directional changes in endothelial cell density and corneal thickness have been shown to predict outcomes such as graft survival [54,55].
Therefore, evidence, as outlined above, suggests homeostasis and drift state are not similar, following different trajectories and presentations. These two states therefore appear to be governed by different sets of rules which can be potentially utilized to identify them. The simulated temporal and trajectory-based behavior of 6 representative cases (including one ‘stable normal’) illustrates this further in Fig 2A-2D.
Temporal patterns of directional drift in simulated subjects. 2A. Six representative cases showing variable progression and stability. 2B. Comparison of noise corridor vs disease trajectory. 2C. Timepoint-overlaid cross-sectional views demonstrating missed progression (in the distribution plots, N = normal, S = suspect, D = Disease) 2D. Conceptual “window of opportunity” for early detection and intervention.
In this visualization, we show six representative simulated cases: one stable normal and five with varying disease trajectories (Fig 2A). These illustrate the diverse ways patients can drift from physiological equilibrium. Some progress gradually, some abruptly, and some not at all. Highlighting the contrast between normal noise (green zone) and true disease drift (blue curve) can suggest how directional movement outside the noise corridor marks meaningful change even before the threshold breach (Fig 2B). Cross-sectional views can obscure these temporal patterns. Fig 2C demonstrates this by comparing overlaid distribution plots at t₀ to t₅. Despite clear progression in several cases, static thresholds at single time points miss the bigger picture. This creates a zone of clinical opportunity: a region of directional drift where early intervention could alter trajectory (Fig 2D). Case 5 in particular shows how an early flag could precede and prevent threshold-defined diagnosis. This case becomes stable after early pre-conventional threshold diagnosis and treatment.
As we have discussed earlier, traditional diagnostic frameworks assume that normal and disease lie along a shared linear axis. This implies that progression is merely a matter of movement along a single continuum [6,7,60,61]. It is a mathematical requirement for comparison and is therefore convenient [60–62]. Threshold-based models assume that a variable’s behavior is interpretable within a known distribution. However, when variables come from fundamentally different biological states, such as a regulated homeostatic system compared to a decompensating one such as the drift, this assumption becomes sub-optimal.
We propose a two plane hypotheses and suggests that normal and disease states are better represented as existing on distinct geometric planes. The physiological plane is constrained by homeostasis. Values fluctuate but remain bounded and self-correcting. The disease plane has different behavior, what we term as ‘intentional directionality’ to differentiate from noise (which on measurement shows as non-intentional, non-directional movement). The tipping point marks the origin of the disease plane (Fig 3).
Conceptual model of the two-plane hypothesis. Normal subjects remain within the physiological plane, while progressive cases drift onto the pathological plane after the tipping point.
This transition is illustrated in the Fig 3, where normal subjects remain within the physiological plane with the passage of time, while progressively worsening pre-threshold and diagnosed cases drift on the pathological plane after the tipping event. The overlay of threshold-based distribution panel on the left of the image illustrates how classical models may produce overlapping decision regions, whereas the geometric model separates populations through angular divergence between planes. The exact moment of tipping may not be directly observable when intermittent sampling is being done compared to real-time tracking. In a clinical setting, this means that the tipping point may occur between follow-ups. However, the behavior before and after the tipping point can be traced back to a putative inflection point, giving mathematical support to this concept. Once tipping occurs, the subject may still appear within the normal range numerically. However, its new trajectory begins to follow the rules of the disease plane (Fig 3). The parameters continue to change and eventually cross conventional diagnostic thresholds, progressing further into disease.
The concept of the two planes can be extended beyond qualitative or illustrative value by mathematically defining them. This can help in creating a framework that documents the changes occurring on these planes around the tipping point. This quantification may help into extending it into clinical implications.
In the next section, we present a vector-based geometric framework that captures this transition in feature space.
Methods
For this exploratory study, we first developed a geometric framework in which normal and disease planes, and their derived vector products, are defined. We then formalized the rules governing their interactions and applied this framework to a simulated dataset built on biological behavior described in the literature. The following is a narrative summary of the framework; complete mathematical derivations and a detailed commentary are provided in S1 Appendix.
A. Framework Design
1. Variable selection and feature space definition.
Two biologically meaningful variables, x and y, are selected to define the feature space. To ensure clinical interpretability and framework generalizability, selected variables should satisfy the following criteria: (i) established clinical relevance to the disease process; (ii) a biologically plausible directional relationship with disease progression; (iii) wide availability and routine use in clinical practice; (iv) preference for simple, direct measurements or low-complexity derived values; and (v) low inter-variable correlation in healthy subjects. In the current study, steepest keratometry (Kmax, diopters) and thinnest corneal thickness (TCT, microns) were selected as x and y respectively, based on their established roles in keratoconus progression [63,64].
2. Normalization, scaling, and plane construction.
Variables are range-normalised using the physiological (normal) population as the reference. For a variable whose value increases with disease progression, the normalized value is:
where m is the raw value, m_min is the minimum value in the physiological population, and is the physiological range [Eq 2a, Supplementary S1 Appendix]. For variables that decrease with disease progression, the symmetric transformation applies [Eq. 2c, S1 Appendix]. This directional alignment ensures that worsening always maps in the positive direction for both variables, regardless of their raw units.
Normalized physiological subjects define a bounded reference plane (Eq 3a, Supplementary S1 Appendix). Disease subjects are scaled using the same physiological range (Eqs. 4a–4b, Supplementary S1 Appendix), placing them in a comparable but unbounded pathological plane
, where
represent the scaled end-state disease maxima. Since this study focuses on early drift near the physiological boundary, end-state values serve only to define the upper limit of the pathological plane and are not modelled directly. Full derivations of both planes are provided in Supplementary S1 Appendix (Eqs. 2–5).
3. Physiological noise scalar (η).
Physiological and observational variability are pooled into a single noise scalar η, derived from the within-subject standard deviation (Sw) of repeated measurements in the stable normal group. The final noise magnitude is defined as:
where is the joint within-subject standard deviation for physiological subject j, and the factor 2.77 corresponds to the 95% coefficient of repeatability (CR =
) [65]. This extends the univariate CR concept (used previously by us for statistical threshold estimation in keratoconus progression) into a radial 95% confidence boundary in the joint feature space, serving as the denominator in the signal-to-noise framework described below [66]. The full per-subject Sw derivation, including the covariance term accounting for correlation between x and y, is provided in Supplementary S1 Appendix (Eqs. 6 a-f).
4. Age-related physiological reference drift (
).
Many physiological parameters change gradually with age, and this background drift must be distinguished from pathological change [67–69]. A linear age-related drift vector
,
) was incorporated as a correction term in all subsequent vector calculations. Over the 2.5-year observation window of this exploratory study, age-related drift was assumed to be negligible; the correction terms were retained as placeholders for future iterations with longer follow-up. The full formulation is provided in Supplementary S1 Appendix (Eqs. 6g-j).
5. Canonical disease vector (
).
The direction of disease progression is captured as a pooled disease vector D, derived from the mean baseline-to-follow-up change across all confirmed progressive disease subjects (ED_P group), adjusted for age-related drift. For each pathological subject k, the age-adjusted change at each follow-up visit is computed and then averaged across visits and subjects to yield a single canonical direction vector in the normalised feature space. This vector serves as the reference direction against which individual subject drift is compared. The complete derivation, including per-subject trajectory vectors and the pooling procedure, is provided in Supplementary S1 Appendix (Eqs. 6j–6q).
6. Subject-specific drift vector
.
For any subject under evaluation (denoted u), the drift vector is computed as the age-adjusted change from baseline to follow-up visit t, using the same scaling principles applied to pathological subjects. The magnitude
quantifies how far the subject has moved in the feature space, and the angle
defines the direction of that movement. Full expressions for
,
and
are provided in Supplementary S1 Appendix (Eqs. 6r–6t).
7. Derived metrics: MNR, DEM, and Composite Drift Score (CDS).
Magnitude-to-Noise Ratio (MNR). A subject's drift is meaningful only if it exceeds the expected physiological noise. Unit-dependent absolute change thresholds are poorly generalisable across variables and clinical domains. We therefore define a dimensionless signal-to-noise ratio:
An MNR > 1 indicates that the subject's displacement exceeds the 95% physiological noise boundary. Just as a radio-based system uses SNR to deduce a meaningful signal over background static, the MNR intends to deduce subject-level change over the expected noise.
Directional Emphasis Multiplier (DEM). Drift magnitude alone is insufficient, a subject moving perpendicular or opposite to the disease direction should not be flagged. The cosine of the angle between the subject vector and the canonical disease vector
measures angular alignment (cos
= 1 for perfect alignment, 0 for orthogonal, −1 for opposite). To reward strong alignment quadratically and penalise misalignment, we apply a signed cosine-square weighting:
DEM ranges from −1 to +1. Movement precisely aligned with the disease vector yields DEM = 1; movement in the opposite direction yields DEM = −1 (which may prove useful as a treatment-response indicator in future work). The signed-square function is customizable and steeper alternatives such as cos³(ϕ) can be used when tighter directional specificity is required.
Composite Drift Score (CDS). The final metric combines both directional and magnitude components into a single, unitless score:
This scalar, unitless value reflects both the strength and directionality of pathological drift and forms the basis for early detection in this framework. It is intuitively set at a threshold of ≥1.0 (perfect alignment with disease process and change greater than seen with noise, which can vary over multiple case scenarios as we discuss later in results and discussion).
8. Extension to higher dimensions and correlated variables (Mahalanobis approach).
The Euclidean CR-based noise scalar η described above performs well for the 2D, low-covariance setting of this study (Kmax and TCT show minimal correlation in normal eyes). However, for datasets with three or more variables, or where inter-variable covariance is substantial, the Euclidean approach may underestimate the true noise boundary. Mahalanobis distance has been noted to counter this issue [70–72]. Therefore, we derived a Mahalanobis distance (MD)-based alternative noise metric, yielding MNR(MD) and CDS(MD) as covariance-adjusted analogues of MNR and CDS:
The divisor 2.45 corresponds to the square root of the 95th percentile of the chi-square distribution with 2 degrees of freedom (√5.99 ≈ 2.45), providing a multivariate equivalent of the univariate CR threshold. In our synthetic dataset, both CR-based and MD-based metrics showed consistent directional trends across groups, with MD-based scores yielding numerically higher values in progressive cohorts. For this initial 2D study, we retain the CR-based MNR and CDS as primary metrics for their clinical interpretability; the MD framework is reserved for future higher-dimensional implementations. The complete Mahalanobis derivation, including covariance matrix construction and the chi-square reference boundary, is provided in Supplementary S1 Appendix (Eqs. 8a–8i).
B. Application to synthetic subjects: With the above framework established, we applied the above steps 1–8 on a labeled synthetic dataset composed of four distinct subject groups (n = 1000 per group), representing a different disease stage or behavioral phenotype. For this exploratory study, we wanted to explore a disease with domain knowledge before expanding to others in future studies, so that we are aware of the expected physiological and pathological behaviors. Therefore, we modeled normal cornea and early keratoconus for this study and the two variables used were steepest Keratometry (Kmax in diopters) and thinnest corneal thickness (TCT) in microns. The synthetic data was generated using normal and abnormal ranges, rate of progression, and intra-measurement standard deviation provided in the literature [63,64]. The complete data generation methodology, including randomization logic, Gaussian noise parameters, seed values, and parameter distributions, is provided in S2 Appendix. The excel sheet with the final data is provided as S3 Appendix and is available publicly at https://doi.org/10.17605/OSF.IO/2JNQF (CC-BY 4.0 International). An illustrative pdf with the framework of the construction is supplied as S4 Appendix. To prevent any future chances of data redundancy or overfitting, we did not use any of our previous studies or our published data with normative corneal or keratoconus data.
To prevent any future chances of data redundancy or overfitting, we did not use any of our previous studies or our published data with normative corneal or keratoconus data.
1. The 4 subject groups were:
- a. Stable normal cases: Used to define the reference range for physiological noise
and physiological reference drift
.
- b. Stable (early) disease cases: served as anchor points for comparison with progressive disease. These cases belong clearly to the pathological plane (no suspects or borderline case) without directional progression.
- c. Progressive (early) disease cases: These subjects also belonged to the pathological plane and have with similar baselines as (b). They were simulated to reflect realistic but stochastic disease progression. Changes in the individual variables
were adjusted to trend over worsening values consistent with known early disease behavior over time. The disease vector
was calculated from this group.
- d. Progressive normal-range cases: This group represents the primary target of our framework: subjects whose baseline values fell within traditional “normal” ranges (based on Gaussian cutoffs or consensus thresholds), but whose homeostatic compensatory mechanisms have failed, leading them to drift beyond the biological tipping point and thus, by definition, into the pathological plane. Variables
were simulated to reflect stochastic but directionally worsening trends, consistent with known early disease behavior over time. To maintain biological plausibility, the average rate of change in this group was scaled to approximately 80% of that used for the early progressive disease group (Group C). This conservative modeling assumption reflects the current lack of direct empirical data on progression dynamics in this pre-diagnostic window. The drift parameters used here are customizable and can be refined in future models using more granular longitudinal datasets.
2) Computations in the synthetic dataset: noise ( and disease vector (
were computed at the dataset level, and
, DEM, MNR, and CDS were computed at the subject level.
Results
A. Group Characteristics and Internal Validation: The synthetic exploratory dataset included four simulated groups (n = 1000 each): Normal Stable (Norm), Early Disease Stable (ED_S), Early Disease Progressive (ED_P), and Pre-threshold Progressive (PT_P).
As noted previously, in this study, the two variables were assigned values from Kmax and TCT, respectively, such that Kmax
and TCT
. All subsequent analyses refer to these generalized variables
for clarity and broader applicability.
Group means and standard deviations for baseline () and final follow-up (
) values of (
) and normalized variables
confirmed expected characteristics hypothesized while modeling the framework. Both the groups designed to worsen with time (ED_P, PT_P) showed worsening of variables with time, while the groups designed to remain stable (Norm, ED_S) remained relatively unchanged (Tables 1 and 2).
A repeated-measures ANOVA on raw values across time
) revealed a statistically significant main effect of
(p < .001), and of
interaction for all comparison (p < .001). Plotting the raw values demonstrated diverging trajectories in progressive groups (increasing
, decreasing
) while stable and normal groups remained flat. Normalized values (
confirmed that progression occurs along a biologically plausible axis and further highlight the temporal separation between disease and stable behavior (Fig 4).
This ensured that the synthetic dataset followed the planned intention of creating 4 different datasets. The randomization and simulation did result in the trends we wanted to mimic, reassuring internal consistency.
B: Indices’ performance Next, we wanted to evaluate the performance of the indices we had conceptualized. The progressive groups showed a significant increase in the three metrics DEM, MNR, and CDS, whereas the stable groups did not show a significant change (Table 3). Both the sets of CR derived and the Mahalanobis distance (MD) derived metrics (MNR, MNR(MD), CDS, CDS(MD)) showed this trend. MNR(MD) and CDS(MD) both showed higher values for the progressives group.
To formally evaluate agreement between the two approaches, paired t-tests and classification concordance analysis were performed at t5. MNR(MD) was significantly higher than MNR(CR) across all groups (all p < 0.001), with mean differences of 0.13 in NS, 0.19 in ED_S, 0.53 in ED_P, and 0.39 in PT_P, consistent with the MD approach's greater sensitivity to multivariate deviation. However, this did not translate into a meaningful CDS difference in stable groups (NS: mean difference 0.005, p = 0.49; ED_S: mean difference 0.001, p = 0.95). This is consistent with the structure of the framework: CDS = DEM × MNR, where near-zero DEM values in stable subjects (NS mean DEM 0.032; ED_S mean DEM 0.025) suppress any MNR difference at the CDS level, regardless of which noise metric is used. In progressive groups where DEM approaches 1.0, the CDS difference between methods closely tracks the MNR difference (ED_P mean difference 0.52, PT_P mean difference 0.37; both p < 0.001), which is numerically larger but clinically consistent. Classification concordance at CDS ≥ 1.0 was 100% for both progressive groups (kappa = 1.0): every subject flagged by CR-based CDS was identically classified by MD-based CDS and vice versa. In stable groups, concordance was 99.7% (NS) and 98.5% (ED_S), with the small number of discordant cases (3 and 12 respectively) representing MD-only flags in subjects with sub-threshold CR-based values. These findings confirm that while MD-based metrics yield numerically higher scores, clinical classification is identical between the two approaches in this 2D, low-covariance setting, supporting the use of the simpler CR-based approach as the primary metric for this study.
For the purpose of this initial study, we continue to use CR-derived MNR and CDS for the rest of the evaluation and would reserve further exploration with MD derived metrics in a higher dimensional dataset in the future.
As shown in Fig 5, the subject vector angle (top-left panel) remained stable over time, with no significant interaction, reflecting that subject maintained a consistent drift direction in the
space (Repeated-measures ANOVA p > 0.5).
Top-right: Directional Emphasis Multiplier (DEM). Bottom-left: Magnitude-to-Noise Ratio (MNR). Bottom-right: Composite Drift Score (CDS).
In contrast, Directional Emphasis Multiplier (DEM) increased significantly over time in progressive groups (repeated-measures ANOVA, p < 0.001) (top-right panel). This may reflect improving alignment with the disease trajectory and not a large change in direction. This could also suggest the self-correcting nature of the index due to repeated measures from the baseline over follow-up. Again, as this was synthetic data, we do expect more noisiness in real cases and therefore this finding if repeated in those groups, needs to be explored further. Also, with time the pre-threshold progressive cases (PT_P) trended to match the performance of the early progressive disease (ED_P). This again is suggestive of a possible stronger alignment with confirmed disease process with time. This was not a planned or hard coded step but is explainable on disease behavior, further suggesting that even though this data was synthetic, did show patterns we expected.
The Magnitude-to-Noise Ratio (MNR) (bottom-left panel) rose in progressive groups, indicating that observed drift exceeded physiological noise boundaries (repeated-measures ANOVA, p < 0.001).
Finally, the Composite Drift Score (CDS) (bottom-right panel), showed a rise in Early Disease Progressive and Pre-threshold Progressive groups, while remaining flat in stable cohorts (repeated-measures ANOVA, p < 0.001).
Taken together, these vector-based metrics demonstrated strong discriminative behavior across simulated disease states, justifying the potential for further exploration for trajectory analysis and diagnostic comparison in more complicated or real-life datasets.
To visualize how directional drift evolves over time and contributes to the Composite Drift Score, we constructed a four-panel boxplot montage (Fig 6).
Top-right: Directional Emphasis Multiplier (DEM). Bottom-left: Magnitude-to-Noise Ratio (MNR). Bottom-right: Composite Drift Score (CDS).
Subject angle (top-left panel) revealed that progressive groups converged directionally toward the canonical disease vector, while stable and normal groups remained without a clear directional trend. This angular trend was reflected in rising DEM values (top-right panel), which translated into increased alignment. Combined with the rising MNR (bottom-left panel), CDS values (bottom-right panel), rose in progressive cohorts ED_P and PT_P, confirming that directional alignment and supra-noise drift both contributed to the CDS scores getting higher.
C. Subject-level stochasticity visualization: Even though we induced Gaussian noise in our samples as described in the methods, at the level of measures of central tendency it is difficult to see if the outcomes created were artificially smooth without. To assess intra-individual variability visually, we mapped each subject's vector displacement over time. Fig 7 shows directional scatter plots for all four groups across five follow-up time points, with simulated subject vectors (orange) overlaid on the baseline physiological noise cloud (blue).
The stable and normal cohorts remained within the noise cloud over the follow-ups. The progressive groups showed clear but directional trends away from the noise cloud, which varied in speed between subjects. This further gave visual support to the drift varying in quantity but similar in the direction in the subjects where it was intended.
Perhaps, we should discuss the method of noise cloud representation in this visual analogy in more detail: As CR is a positive number, plotting the CR only would have artificially resulted in a data set limited to quadrant I of the cartesian system. In constructing the baseline physiological appearing noise cloud, we aligned the noise distribution with the biological directionality of change observed at the last visit
Where theare the coordinates of the noise cloud for subject j,
are the absolute values of
,
, the change in the normalized
variables at the last follow-up
, and
are subject level intra-measurement standard deviations which when multiplied with 2.77 gives the CR for that subject. Though this approximation is not mathematically equivalent to the CR we used as a pooled value, hopefully, it gives a visual representation of the noise cloud primarily at an intuitional level. This could also be useful as we further deal with more stochastic and subject-level data in future work.
Also, Fig 8 presents heat maps of subject-level evolution for all the follow-ups of all the cases for CDS, the cosine angle between disease and subject vectors, and magnitude of subject vector.
Rows represent individual subjects; warmer colors indicate increasing values.
Progressive cohorts show increasing intensity (green to orange to red) in both alignment and magnitude, consistent with disease-like drift accumulation. In contrast, stable disease and normal groups show lesser changes in the hues from the cooler colors over follow-ups. Together, these visualizations help to demonstrate the biological plausibility and internal variability of the synthetic dataset, and similar methods could be used in future work to compare the internal variability of more complicated datasets.
D. Sampling Frequency Sensitivity Analysis.
As a trajectory-based framework, CDS depends on the availability of serial measurements: fewer follow-up visits mean fewer opportunities to detect drift before conventional thresholds are crossed. To quantify this expected relationship, four clinically relevant sampling cadences were modelled in the PT_P group (n = 1000): 6-, 12-, 18-, and 24-month intervals, using the visits available within the 2.5-year observation window. The detection criterion was CDS ≥ 1.0 at any available visit. Lead time was calculated as the interval between first CDS flag and conventional TCT threshold crossing (TCT ≤ 500 µm, univariate threshold crossing as an example) in the 156 PT_P subjects who crossed the threshold during the observation period. Results are summarized in Table 4.
As expected, detection rate and lead time both declined as sampling frequency decreased. With 6-month follow-up, 99.8% of PT_P subjects were flagged at a mean of 13.4 months, with a mean lead time of 11.3 months ahead of threshold crossing. At 12-month intervals, detection rate was maintained at 98.5% with mean lead time of 10.1 months, a modest reduction consistent with the loss of alternate visits. At 18-month intervals, detection fell to 92.7% and lead time to 7.5 months. At 24-month intervals, detection recovered to 98.5% as the t4 visit captures most progressors by that point, but lead time compressed to 3.5 months, the shortest across all cadences and leaving little clinical window for intervention before threshold crossing.
These results are an expected property of any trajectory-based detection system and are presented here to quantify the trade-off rather than as a novel finding. In this dataset and for the disease model used, the framework's pre-threshold advantage is preserved under annual review and diminishes substantially at intervals beyond 18 months. The dependence on sampling frequency is acknowledged as a limitation of the framework and is discussed accordingly. Future work will explore cadence-specific optimization for individual disease domains and progression rate profiles.
Discussion
This novel framework quantifies directional progression instead of relying solely on value-based thresholds. In this paper, we have conceptualized and then evaluated a geometric framework rooted in biological plausibility and clinical interpretability. The framework is derived from previous work in the fields of medicine, biomedical imaging, and signal processing. It derives inspiration from concepts in ecology, systems biology, and physics [56–59]. This is also an extension from our previous studies of univariate & bivariate range normalization, rate of change assessment and predictive modeling in keratoconus, use of repeatability indices for clinical decision-making, and modeling the role of directional alignment in the disease process [66,73,74]. The goal of this paper is to formalize a clinical intuition many physicians have when population-based cutoffs, which have historically worked well, lag in predicting patient’s unique changes [60–62]. We should also discuss the use of synthetic data in the study. Similar approaches using synthetic or semi-synthetic disease progression models have been used previously [75–79]. There are relatively large sizes of keratoconus, Fuchs dystrophy, and normal eye databases, and we have also reported findings from our databases. However, most of these are cross-sectional [80–83]. Subjects who are clinically normal are often not followed up serially in a systemic fashion. Therefore, we decided to use a more customizable method such as an Excel based dataset generation vs publicly available or pre-generated cross sectional synthetic datasets in this initial explorative study.
Often, correlated biomarkers lead to overfitting and interpretation difficulty [84, 85]. Therefore, the parsimonious design of this model is an attempt to safeguard against overfitting. While there can be exceptions, the initial plan should be to include 2 or 3 parameters that are used for threshold-based screening. The reason to use parameters already being used commonly in threshold-based testing is to keep the clinical interpretability intact due to preexisting familiarity in clinicians for those parameters (directionality and not necessarily the numerical thresholds). Continuing from our example of keratoconus, studies have shown the advantage of corneal wavefront measurement as a more sensitive marker for keratoconus diagnosis, and Brillouin microscopy can detect keratoconus much earlier than conventional testing [86]. However, most ophthalmologists will know that the cornea gets thinner and steeper with keratoconus and will be able to relate to keratometry and pachymetry at a more intuitive level than corneal wavefront [4,28].
As an expansion of scope, we visualize this system to be a framework that, if proven, may be transferable to the more comprehensive practice physicians such as a primary care physician or a comprehensive ophthalmologist who can work along with subspecialists. Therefore, the common interpretability of the parameters becomes important.
The use of relatively more esoteric parameters and complex indices has an undeniable role in severity mapping for advanced diseases, but the current framework is aimed to supplement that by working in a hitherto unexplored area [3,5,74]. Our geometric vector-plane framework complements prior approaches that use information technologies to process multidimensional physiological state-spaces for health monitoring and intelligent decision support [87].
It is worth discussing the place of the CDS framework within the landscape of existing longitudinal monitoring approaches. Reference change values (RCV), derived from within-subject biological variation and analytical imprecision, are the most widely used method for flagging clinically significant interval change in a single variable. [31] CDS extends this concept in two important directions: it operates simultaneously across multiple variables in a shared feature space, and it incorporates directional alignment with a disease-specific trajectory rather than treating any sufficiently large change as meaningful regardless of its direction.
Composite multiparameter indices such as the Belin-Ambrósio Enhanced Ectasia Display in keratoconus, the AGIS score in glaucoma, or KDIGO staging in chronic kidney disease offer powerful classification but are designed to characterize disease severity or cumulative damage rather than to detect directional pre-threshold drift in subjects still classified as normal [88–90]. Machine learning classifiers for early disease detection can achieve high discriminative accuracy in well-powered datasets but typically require large, labelled training cohorts, offer limited interpretability at the individual subject level, and are not inherently designed around the concept of directional biological drift [91, 92]. CDS is not intended to replace any of these approaches. Rather, it is designed to operate upstream of them in the pre-threshold longitudinal window where a subject's measurements remain numerically normal, yet their trajectory has already diverged from physiological noise. In this sense, CDS is less a competitor to existing diagnostic tools and more a temporally earlier layer of the same clinical decision process.
However, one may be concerned if this is an oversimplification (using lower order, linear approximations rather than the complex black box style models). Clinicians and researchers are aware of the complex stochastic process of disease progression after the clinical thresholds are triggered up to complete loss of structure or function or both (as the case may be) [93–95]. Simply put, disease progression starts and stops, accelerates, and decelerates, and multiple genetic and environmental roles need to be modeled when mapping the entire disease process [96, 97].
However, this framework does not aim to model that process but zooms in a small area of that process with an awareness of the start and end points. The starting point of this evaluation is the deviation away from the normal physiology, or as the model conceptualizes, the physiological plane. The endpoint is the crossing over of the diagnostic boundary. When the classic decision boundary is already breached, the clinician has proof of worsening, so the geometric framework does not serve any further diagnostic advantage. The patient can be flagged by either threshold-based or drift-based methods at that point in time. Linear approximation of what is possibly still a small part of the trajectory may be both mathematically acceptable and clinically translatable. Fig 9 demonstrates this concept.
The full curve is nonlinear and stochastic (blue). The dashed red line represents the actual disease trajectory within the noise-to-threshold zone, which can be locally approximated using linear assumptions. This model is not designed to fit the entire disease course but to identify early directional drift (gray arrow) away from the physiological corridor (green band). Later progression may require alternate modeling approaches, such as polynomial fitting (solid red line).
For example, imagine a cornea that starts to develop keratoconus beginning at 45D (when it was structurally normal and stable) and ends at 75D (when there is an advanced, end-stage disease). The change from 45D to hitting the statistical threshold at 47.4 D is a smaller part of that curve and the endpoints can be approximated. A clinician is also interested in knowing the trend (for example, a change from 45 to 46.5 D) on a more practical basis rather than knowing if a complex polynomial equation can map this change. That style of fitting, as we have previously demonstrated, becomes more useful in mapping the entire disease range, from normal through early to advanced [73]. From a more methodological perspective, this idea of linear regression is derived from the practice of segmented regression analysis used to evaluate changes in interrupted time series [97].
The visualization of the start point perhaps needs more explanation: it is the deviation away from normal in a significant, biologically plausible manner. This uses two concepts: normal noise envelope and angular alignment (cosine) with the disease process.
As discussed earlier, the noise in a subject-tester-instrument-interpretation environment can be accounted for using the intra-measurement standard deviation. This extends from the classic work of Bland and Altman [98]. This is a usual practice in many clinical domains, especially ophthalmology. Multiple previous repeatability and reliability studies have used the intra-measurement standard deviation (Sw) and its derived parameter the coefficient of repeatability (CR), especially for normal subjects [44,46,98–104] Again, reference change values (RCV) used for laboratory measurements are conceptually similar [31]. There is an interest in using a log-normal spread for reference change value (RCV) and this model can be customized based on the set of tests it is being applied to in the future [105]. For this study, we retained the CR at , per the Gaussian assumption. Using the Mahalanobis distance-based metric as an alternative is very promising, especially when we would deal with ≥3-dimensional testing, or we are forced to use correlated variables. Our future studies will explore this more in Python based platforms.
When measuring multiple variables, measures of confidence can be treated as orthogonal scalars creating a vector, as has been demonstrated previously [106,107]. This is generally derived from pooled data. However, to incorporate the real-world noisiness we calculated the individual vectors for Sw between the two parameters, accounting for a creeped-in covariance, and later scaled them by 2.77 to construct what we term the “noise fog envelope”. Like the rest of our model, this is also scalable into higher dimensions.
We deliberately avoided pooling noise metrics till later steps. This method retains the option of customization for a specific subject where the noise fog envelope can be calibrated based on their noted inter- measurement standard deviations(Sw) [42]. Once this envelope is modeled, it is supposed to be the upper limit of normal difference in measurement seen and therefore it was used to scale the change noted in the subject being tested. Classically, noise ellipses have been used as comparative tools (e.g., for inter and intra-device agreement). However, we have previously demonstrated their potential as a thresholding tool for clinical decision-making for cross-linking in progressive keratoconus [66]. The current metric is an extension of that concept.
Parameter values can change due to non-pathological reasons or due to a different pathology. For example, for a patient being evaluated for keratoconus, a tight contact lens can create a corneal warpage or scarring can cause corneal thinning. However, warpage will not induce significant corneal stromal thinning (can cause focal epithelial thickening) and corneal scarring will generally cause flattening [4,108]. So, the directionality of the pathological process becomes relevant (In our keratoconus example- corneal thinning and corneal steepening is the expected combination over time).
So, we need a method to note if the change is more than physiological noise AND is in the direction of the progression of pathology (alignment with the disease vector). This is the purpose of the combined metric, the Composite Drift Score (CDS). It combines both the magnitude and the directionality and thereby creates a potential metric to evaluate the change.
A CDS threshold of 1 requires both conditions to be met simultaneously, the subject's drift must exceed the physiological noise boundary (MNR > 1) and must be directionally aligned with the disease vector (DEM > 0). Neither condition alone is sufficient, which is what gives the combined score its specificity. The raw cosine was replaced by the signed cosine-square (DEM = cos ϕ· |cos ϕ|) because it quadratically rewards strong alignment while progressively penalizing misalignment, without requiring an arbitrary cutoff at a particular angle(or degree of misalignment). This function is customizable. Steeper alternatives such as cos³(ϕ) increase directional specificity at the cost of sensitivity, as illustrated in Figs 10a and 10b. The operating properties of different weighting functions and their effect on CDS across the full angular range are shown there and are not re-derived here.
The additional penalty values show the additional reduction each of two functions (DEM and Cubed cosine) imposes relative to Raw cosine. Effect of Angular Weighting on CDS Magnitude. Left: Constant magnitude-to-noise ratio (MNR). Center: Angular weighting functions: Cos(θ), DEM, and Cos3(θ). Right: Product of MNR and Angular function defines CDS magnitude (gray shaded area under curve).
Further optimization of directional weighting, phenotype-specific modeling, and performance tuning across noisy datasets will be addressed in subsequent work.
A practical implementation of the CDS framework would follow a staged pathway. In the calibration stage, a disease-specific reference dataset is assembled from routine clinical records: stable normal subjects provide the noise scalar η and physiological plane, while confirmed progressive early-disease subjects provide the canonical disease vector D. These steps require two routinely measured variables and standard repeatability data already available in most clinical domains, for example, Pentacam-derived Kmax and TCT in corneal practice, or eGFR and urine albumin-to-creatinine ratio in nephrology [63,109]. The calibration dataset can be updated periodically as institutional data accumulates, analogous to how laboratory reference intervals are re-derived over time [35].
At the individual patient level during routine follow-up, the framework produces a single unitless output: CDS below 1.0 indicates that observed change is either within physiological noise or not aligned with the disease direction; CDS at or above 1.0 indicates supra-noise, disease-aligned drift warranting closer surveillance. Critically, CDS functions as a triage layer upstream of existing diagnostic workflows rather than as a standalone diagnostic, analogous to how rising biomarker velocity prompts further evaluation even when the absolute value remains below the diagnostic threshold [110–112]. Conversely, a persistently low CDS in a patient with borderline absolute values could provide quantitative reassurance and support longer surveillance intervals.
Several implementation considerations merit acknowledgement. The framework requires a minimum of two longitudinal visits and cannot operate on a single cross-sectional measurement. Calibration cohorts must be representative of the target population in terms of age, ethnicity, and measurement device. The CDS threshold of 1.0 is a mathematically motivated starting point; optimal thresholds for specific clinical applications will require ROC analysis in real-world datasets. The assumptions and boundary conditions under which this framework may underperform are discussed in below.
This study has several limitations inherent to its exploratory, proof-of-concept design. The framework was evaluated on synthetic data generated under controlled assumptions; performance in real-world datasets with irregular follow-up intervals, missing visits, and population heterogeneity remains to be established in future work. The current implementation uses two variables; while the Mahalanobis extension was derived and showed consistent trends, systematic comparison with the Euclidean approach in higher-dimensional or strongly correlated settings is reserved for subsequent studies. The canonical disease vector assumes a dominant directional trajectory; diseases with distinct phenotypic subtypes following divergent progression paths may benefit from subtype-specific vectors, which the framework can accommodate through separate calibration cohorts. The linear approximation within the noise-to-threshold window is intentional and appropriate for this narrow zone, though its adequacy for diseases with abrupt or stepwise transitions would need to be evaluated individually. The noise scalar η is derived from pooled normal data and assumes broadly comparable measurement variability across subjects; personalized noise calibration, which the framework supports, may improve performance in heterogeneous clinical settings. Finally, while the mathematical structure is domain-agnostic, transferability beyond the keratoconus model used here requires independent calibration and validation in each clinical setting.
Conclusion
We hypothesized a 2-plane disease vs normal model, and conceptualized a mathematical framework based on it. Then we constructed an index system, CDS (= DEM x MNR) to with the eventual target to flag early change within the subthreshold range in synthetic data. This metric and its subcomponents were able to demonstrate similar trends in progressive early disease and progressive pre-threshold groups. This suggests that with more studies, rigorous evaluation, and iterations it may be able to mathematically map its intended overarching goal: a unitless ratio based intuitive method to denote disease directional change in the pre-threshold group and therefore have impact in early detection and management of a subset of chronic diseases.
Threshold based tests are agnostic to pre-threshold movement. Therefore, the next step is to compare these metrics in unlabeled data and evaluate their performance and lead time compared to threshold-based cutoffs. This early-stage simulation study sets up ground to evaluate this metric in our next work on stress testing this on highly stochastic data, then conceptualizing case-based scenarios, leading to possible multicentric/ collaborative work in real life situations.
Supporting information
S1 Appendix. Full mathematical derivations of the geometric drift framework, including complete derivations of normalization equations (Eqs. 2a–2c), physiological and pathological plane construction (Eqs. 3–5), noise scalar (Eqs. 6a–6f), age-related drift (Eqs. 6g–6i), disease vector (Eqs. 6j–6q), subject drift vector (Eqs. 6r–6t), and Mahalanobis distance framework (Eqs. 8a–8i).
https://doi.org/10.1371/journal.pone.0353723.s001
(DOCX)
S2 Appendix. Synthetic dataset generation methodology: group-specific parameter distributions, randomisation logic, Gaussian noise specifications, seed values, and Mahalanobis distance computation pipeline.
https://doi.org/10.1371/journal.pone.0353723.s002
(DOCX)
S3 Appendix. Excel dataset containing the full synthetic data for all four groups (NS, ED_S, ED_P, PT_P; n = 1000 each), including raw variables, normalised variables, computed drift vectors, and CR-based and Mahalanobis-based metrics across all follow-up timepoints. Publicly available at
https://doi.org/10.17605/OSF.IO/2JNQF
https://doi.org/10.1371/journal.pone.0353723.s003
(XLSX)
S4 Appendix. Visual representation of the framework construction steps (panels A–K): raw data scatter, directional alignment, plane construction, normalisation, disease vector derivation, noise envelope, and subject vector evaluation.
https://doi.org/10.1371/journal.pone.0353723.s004
(PDF)
References
- 1. Neumaier M. Diagnostics 4.0: the medical laboratory in digital health. Clin Chem Lab Med. 2019;57(3):343–8. pmid:30530888
- 2. Khatab Z, Yousef GM. Disruptive innovations in the clinical laboratory: catching the wave of precision diagnostics. Crit Rev Clin Lab Sci. 2021;58(8):546–62. pmid:34297653
- 3. Bouchareb Y, AlSaadi A, Zabah J, Jain A, Al-Jabri A, Phiri P, et al. Technological Advances in SPECT and SPECT/CT Imaging. Diagnostics (Basel). 2024;14(13):1431. pmid:39001321
- 4. Fan R, Chan TC, Prakash G, Jhanji V. Applications of corneal topography and tomography: a review. Clin Exp Ophthalmol. 2018;46(2):133–46. pmid:29266624
- 5. Munari E, Scarpa A, Cima L, Pozzi M, Pagni F, Vasuri F, et al. Cutting-edge technology and automation in the pathology laboratory. Virchows Arch. 2024;484(4):555–66. pmid:37930477
- 6. Warner JL, Najarian RM, Tierney LM Jr. Perspective: Uses and misuses of thresholds in diagnostic decision making. Acad Med. 2010;85(3):556–63. pmid:20182138
- 7. Giannoni A, Baruah R, Leong T, Rehman MB, Pastormerlo LE, Harrell FE, et al. Do optimal prognostic thresholds in continuous physiological variables really exist? Analysis of origin of apparent thresholds, with systematic review for peak oxygen consumption, ejection fraction and BNP. PLoS One. 2014;9(1):e81699. pmid:24475020
- 8. Sheldrick RC, Breuer DJ, Hassan R, Chan K, Polk DE, Benneyan J. A system dynamics model of clinical decision thresholds for the detection of developmental-behavioral disorders. Implement Sci. 2016;11(1):156. pmid:27884203
- 9. Patel BS, Steinberg E, Pfohl SR, Shah NH. Learning decision thresholds for risk stratification models from aggregate clinician behavior. J Am Med Inform Assoc. 2021;28(10):2258–64. pmid:34350942
- 10. Nease RF Jr, Bonduelle Y. Solid recommendations from soft numbers: the test/treatment decision. Med Decis Making. 1987;7(4):220–33. pmid:3683109
- 11. Schumacher GE, Barr JT. Bayesian and threshold probabilities in therapeutic drug monitoring: when can serum drug concentrations alter clinical decisions?. Am J Hosp Pharm. 1994;51(3):321–7. pmid:8160684
- 12. Schumacher GE, Barr JT. Using population-based serum drug concentration cutoff values to predict toxicity: test performance and limitations compared with Bayesian interpretation. Clin Pharm. 1990;9(10):788–96. pmid:2242660
- 13. Cooper RV. Avoiding false positives: zones of rarity, the threshold problem, and the DSM clinical significance criterion. Can J Psychiatry. 2013;58(11):606–11. pmid:24246430
- 14. Zimmerman M. Would broadening the diagnostic criteria for bipolar disorder do more harm than good? Implications from longitudinal studies of subthreshold conditions. J Clin Psychiatry. 2012;73(4):437–43. pmid:22579144
- 15. Streiner DL. Best (but oft-forgotten) practices: the multiple problems of multiplicity-whether and how to correct for many statistical tests. Am J Clin Nutr. 2015;102(4):721–8. pmid:26245806
- 16. Wang Z, Dendukuri N, Zar HJ, Joseph L. Modeling conditional dependence among multiple diagnostic tests. Stat Med. 2017;36(30):4843–59. pmid:28875512
- 17. van Walraven C, Austin PC, Jennings A, Forster AJ. Correlation between serial tests made disease probability estimates erroneous. J Clin Epidemiol. 2009;62(12):1301–5. pmid:19716680
- 18. Sonnenberg A. Combining the outcomes of endoscopy, laboratory testing, and professional judgement in gastroenterological decision-making. Eur J Gastroenterol Hepatol. 2017;29(12):1321–6. pmid:29111998
- 19. Meisner A, Carone M, Pepe MS, Kerr KF. Combining biomarkers by maximizing the true positive rate for a fixed false positive rate. Biom J. 2021;63(6):1223–40. pmid:33871887
- 20. Dendukuri N, Joseph L. Bayesian approaches to modeling the conditional dependence between multiple diagnostic tests. Biometrics. 2001;57(1):158–67. pmid:11252592
- 21. Brenner H. How independent are multiple “independent” diagnostic classifications?. Stat Med. 1996;15(13):1377–86. pmid:8841648
- 22. Greenland S, Gustafson P. Accounting for independent nondifferential misclassification does not increase certainty that an observed association is in the correct direction. Am J Epidemiol. 2006;164(1):63–8. pmid:16641307
- 23. Thompson WH, Wright J, Bissett PG, Poldrack RA. Dataset decay and the problem of sequential analyses on open datasets. Elife. 2020;9:e53498. pmid:32425159
- 24. Insel PS, Ossenkoppele R, Gessert D, Jagust W, Landau S, Hansson O, et al. Time to amyloid positivity and preclinical changes in brain metabolism, atrophy, and cognition: evidence for emerging amyloid pathology in Alzheimer’s disease. Front Neurosci. 2017;11:281.
- 25. Jia J, Ning Y, Chen M, Wang S, Yang H, Li F, et al. Biomarker Changes during 20 Years Preceding Alzheimer’s Disease. N Engl J Med. 2024;390(8):712–22.
- 26. Lange P, Celli B, Agustí A, Boje Jensen G, Divo M, Faner R, et al. Lung-Function Trajectories Leading to Chronic Obstructive Pulmonary Disease. N Engl J Med. 2015;373(2):111–22.
- 27. Long MT, Fox CS. The Framingham Heart Study--67 years of discovery in metabolic disease. Nat Rev Endocrinol. 2016;12(3):177–83. pmid:26775764
- 28. Li X, Rabinowitz YS, Rasheed K, Yang H. Longitudinal study of the normal eyes in unilateral keratoconus patients. Ophthalmology. 2004;111(3):440–6. pmid:15019316
- 29. Steinvil A, Shapira I, Ben-Bassat OK, Cohen M, Vered Y, Berliner S, et al. The association of higher levels of within-normal-limits liver enzymes and the prevalence of the metabolic syndrome. Cardiovasc Diabetol. 2010;9:30. pmid:20633271
- 30. Kim S-K, Chung J-W, Lim J, Jeong T-D, Chang J, Seo M, et al. Interpreting changes in consecutive laboratory results: clinician’s perspectives on clinically significant change. Clin Chim Acta. 2023;548:117462. pmid:37390943
- 31. Fraser CG. Reference change values. Clin Chem Lab Med. 2011;50(5):807–12. pmid:21958344
- 32. Limpert E, Stahel WA. Problems with using the normal distribution--and ways to improve quality and efficiency of data analysis. PLoS One. 2011;6(7):e21403. pmid:21779325
- 33. Fossion R, Rivera AL, Estañol B. A physicist’s view of homeostasis: how time series of continuous monitoring reflect the function of physiological variables in regulatory mechanisms. Physiol Meas. 2018;39(8):084007. pmid:30088478
- 34. Feldman M, Dickson B. Plasma Electrolyte Distributions in Humans-Normal or Skewed?. Am J Med Sci. 2017;354(5):453–7. pmid:29173354
- 35. Ozarda Y, Sikaris K, Streichert T, Macri J. Distinguishing reference intervals and clinical decision limits - A review by the IFCC Committee on Reference Intervals and Decision Limits. Crit Rev Clin Lab Sci. 2018;55(6):420–31.
- 36. David R, Zangwill L, Briscoe D, Dagan M, Yagev R, Yassur Y. Diurnal intraocular pressure variations: an analysis of 690 diurnal curves. Br J Ophthalmol. 1992;76(5):280–3. pmid:1356429
- 37. Kawano Y. Diurnal blood pressure variation and related behavioral factors. Hypertens Res. 2011;34(3):281–5. pmid:21124325
- 38. Poggiogalle E, Jamshed H, Peterson CM. Circadian regulation of glucose, lipid, and energy metabolism in humans. Metabolism. 2018;84:11–27. pmid:29195759
- 39. Gao F, Liu X, Zhao Q, Pan Y. Comparison of the iCare rebound tonometer and the Goldmann applanation tonometer. Exp Ther Med. 2017;13(5):1912–6. pmid:28565785
- 40. Veríssimo D, Vinhais J, Ivo C, Martins AC, Nunes E Silva J, Passos D, et al. Continuous Glucose Monitoring vs. Capillary Blood Glucose in Hospitalized Type 2 Diabetes Patients. Cureus. 2023;15(8):e43832.
- 41. Myers MG, Kaczorowski J, Dawes M, Godwin M. Automated office blood pressure measurement in primary care. Can Fam Physician. 2014;60(2):127–32. pmid:24522674
- 42. Coşkun A, Sandberg S, Unsal I, Cavusoglu C, Serteser M, Kilercik M, et al. Personalized reference intervals in laboratory medicine: a new model based on within-subject biological variation. Clin Chem. 2021;67(2):374–84.
- 43. Coskun A. Diagnosis Based on Population Data versus Personalized Data: The Evolving Paradigm in Laboratory Medicine. Diagnostics (Basel). 2024;14(19):2135. pmid:39410539
- 44. Prakash G, Agarwal A, Jacob S, Kumar DA, Agarwal A, Banerjee R. Comparison of fourier-domain and time-domain optical coherence tomography for assessment of corneal thickness and intersession repeatability. Am J Ophthalmol. 2009;148(2):282-290.e2. pmid:19442961
- 45. Hernández-Camarena JC, Chirinos-Saldaña P, Navas A, Ramirez-Miranda A, de la Mota A, Jimenez-Corona A, et al. Repeatability, reproducibility, and agreement between three different Scheimpflug systems in measuring corneal and anterior segment biometry. J Refract Surg. 2014;30(9):616–21. pmid:25250418
- 46. Prakash G, Agarwal A, Mazhari AI, Chari M, Kumar DA, Kumar G, et al. Reliability and reproducibility of assessment of corneal epithelial thickness by fourier domain optical coherence tomography. Invest Ophthalmol Vis Sci. 2012;53(6):2580–5. pmid:22427573
- 47. Esser N, Utzschneider KM, Kahn SE. Early beta cell dysfunction vs insulin hypersecretion as the primary event in the pathogenesis of dysglycaemia. Diabetologia. 2020;63(10):2007–21. pmid:32894311
- 48. Cohrs CM, Panzer JK, Drotar DM, Enos SJ, Kipke N, Chen C, et al. Dysfunction of Persisting β Cells Is a Key Feature of Early Type 2 Diabetes Pathogenesis. Cell Rep. 2020;31(1):107469. pmid:32268101
- 49. Avila MN, Luciardi MC, Oldano AV, Aleman MN, Pérez Aguilar RC. Chronic kidney disease prevalence in asymptomatic patients with risk factors-usefulness of serum cystatin C: a cross-sectional study. Porto Biomed J. 2023;8(6):e233.
- 50. Keenan TDL, Cukras CA, Chew EY. Age-related macular degeneration: epidemiology and clinical aspects. Adv Exp Med Biol. 2021;1256:1–31. pmid:33847996
- 51. Zhang J, Patel DV. The pathophysiology of Fuchs’ endothelial dystrophy--a review of molecular and cellular insights. Exp Eye Res. 2015;130:97–105. pmid:25446318
- 52. Méthot S, Proulx S, Brunette I, Rochette PJ. The Presence of Guttae in Fuchs Endothelial Corneal Dystrophy Explants Correlates With Cellular Markers of Disease Progression. Invest Ophthalmol Vis Sci. 2023;64(5):13. pmid:37195656
- 53. Ozgurhan EB, Kara N, Yildirim A, Bozkurt E, Uslu H, Demirok A. Evaluation of corneal microstructure in keratoconus: a confocal microscopy study. Am J Ophthalmol. 2013;156(5):885-893.e2. pmid:23932262
- 54. Benetz BA, Lass JH, Gal RL, Sugar A, Menegay H, Dontchev M, et al. Endothelial morphometric measures to predict endothelial graft failure after penetrating keratoplasty. JAMA Ophthalmol. 2013;131(5):601–8. pmid:23493999
- 55. Eleiwa T, Elsawy A, Ozcan E, Chase C, Feuer W, Yoo SH, et al. Prediction of corneal graft rejection using central endothelium/Descemet’s membrane complex thickness in high-risk corneal transplants. Sci Rep. 2021;11(1):14542.
- 56. Scheffer M, Bascompte J, Brock WA, Brovkin V, Carpenter SR, Dakos V, et al. Early-warning signals for critical transitions. Nature. 2009;461(7260):53–9. pmid:19727193
- 57. Kitano H. Biological robustness. Nat Rev Genet. 2004;5(11):826–37. pmid:15520792
- 58. Clements CF, Ozgul A. Indicators of transitions in biological systems. Ecol Lett. 2018;21(6):905–19. pmid:29601665
- 59. Trefois C, Antony PMA, Goncalves J, Skupin A, Balling R. Critical transitions in chronic disease: transferring concepts from ecology to systems medicine. Curr Opin Biotechnol. 2015;34:48–55. pmid:25498477
- 60. Djulbegovic B, Hozo I, Mayrhofer T, van den Ende J, Guyatt G. The threshold model revisited. J Eval Clin Pract. 2019;25(2):186–95. pmid:30575227
- 61. Pauker SG, Kassirer JP. Therapeutic decision making: a cost-benefit analysis. N Engl J Med. 1975;293(5):229–34. pmid:1143303
- 62. Büttner J. Biological variation and quantification of health: the emergence of the concept of normality. Clin Chem Lab Med. 1998;36(1):69–73. pmid:9594090
- 63. Kreps EO, Jimenez-Garcia M, Issarti I, Claerhout I, Koppen C, Rozema JJ. Repeatability of the pentacam HR in Various Grades of Keratoconus. Am J Ophthalmol. 2020;219:154–62. pmid:32569740
- 64. Kosekahya P, Caglayan M, Koc M, Kiziltoprak H, Tekin K, Atilgan CU. Longitudinal Evaluation of the Progression of Keratoconus Using a Novel Progression Display. Eye Contact Lens. 2019;45(5):324–30. pmid:30724839
- 65.
Bland M. An introduction to medical statistics. 4th ed. Oxford University Press. 2015.
- 66. Prakash G, Philip R, Srivastava D, Bacero R. Evaluation of the Robustness of Current Quantitative Criteria for Keratoconus Progression and Corneal Cross-linking. J Refract Surg. 2016;32(7):465–72. pmid:27400078
- 67. Ferrer-Blasco T, González-Méijome JM, Montés-Micó R. Age-related changes in the human visual system and prevalence of refractive conditions in patients attending an eye clinic. J Cataract Refract Surg. 2008;34(3):424–32. pmid:18299067
- 68. Laing RA, Sanstrom MM, Berrospi AR, Leibowitz HM. Changes in the corneal endothelium as a function of age. Exp Eye Res. 1976;22(6):587–94. pmid:776638
- 69. Mitnitski A, Rockwood K. Aging as a process of deficit accumulation: its utility and origin. Interdiscip Top Gerontol. 2015;40:85–98. pmid:25341515
- 70. Mahalanobis PC. On the generalized distance in statistics. Proc Natl Inst Sci India. 1936;2(1):49–55.
- 71. De Maesschalck R, Jouan-Rimbaud D, Massart DL. The Mahalanobis distance. Chemom Intell Lab Syst. 2000;50(1):1–18.
- 72. Ghorbani H. MAHALANOBIS DISTANCE AND ITS APPLICATION FOR DETECTING MULTIVARIATE OUTLIERS. FU Math Inform. 2019;:583.
- 73. Prakash G, Suhail M, Srivastava D. Predictive Analysis Between Topographic, Pachymetric and Wavefront Parameters in Keratoconus, Suspects and Normal Eyes: Creating Unified Equations to Evaluate Keratoconus. Curr Eye Res. 2016;41(3):334–42. pmid:25803133
- 74. Prakash G, Perera C, Jhanji V. Comparison of machine learning-based algorithms using corneal asymmetry vs. single-metric parameters for keratoconus detection. Graefes Arch Clin Exp Ophthalmol. 2023;261(8):2335–42. pmid:37022493
- 75. Goncalves A, Ray P, Soper B, Stevens J, Coyle L, Sales AP. Generation and evaluation of synthetic patient data. BMC Med Res Methodol. 2020;20(1):108. pmid:32381039
- 76. Tucker A, Wang Z, Rotalinti Y, Myles P. Generating high-fidelity synthetic patient data for assessing machine learning healthcare software. NPJ Digit Med. 2020;3(1):147. pmid:33299100
- 77. Pantanowitz J, Manko CD, Pantanowitz L, Rashidi HH. Synthetic Data and Its Utility in Pathology and Laboratory Medicine. Lab Invest. 2024;104(8):102095. pmid:38925488
- 78. Dixit A, Yohannan J, Boland MV. Assessing Glaucoma Progression Using Machine Learning Trained on Longitudinal Visual Field and Clinical Data. Ophthalmology. 2021;128(7):1016–26. pmid:33359887
- 79. Alexiuk M, Tangri N. Prediction models for earlier stages of chronic kidney disease. Curr Opin Nephrol Hypertens. 2024;33(3):325–30. pmid:38420892
- 80. Ferdi AC, Nguyen V, Gore DM, Allan BD, Rozema JJ, Watson SL. Keratoconus Natural Progression: A Systematic Review and Meta-analysis of 11 529 Eyes. Ophthalmology. 2019;126(7):935–45. pmid:30858022
- 81. Das AV, Chaurasia S. Clinical profile and demographic distribution of Fuchs’ endothelial dystrophy: An electronic medical record-driven big data analytics from an eye care network in India. Indian J Ophthalmol. 2022;70(7):2415–20. pmid:35791122
- 82. Das AV, Chaurasia S. Clinical Profile and Demographic Distribution of Corneal Dystrophies in India: A Study of 4198 Patients. Cornea. 2021;40(5):548–53. pmid:32740009
- 83. Prakash G, Srivastava D, Avadhani K, Thirumalai SM, Choudhuri S. Comparative evaluation of the corneal and anterior chamber parameters derived from scheimpflug imaging in arab and south asian normal eyes. Cornea. 2015;34(11):1447–55. pmid:26203755
- 84. Babyak MA. What you see may not be what you get: a brief, nontechnical introduction to overfitting in regression-type models. Psychosom Med. 2004;66(3):411–21. pmid:15184705
- 85. Subramanian J, Simon R. Overfitting in prediction models - is it a problem only in high dimensions?. Contemp Clin Trials. 2013;36(2):636–41.
- 86. Scarcelli G, Pineda R, Yun SH. Brillouin optical microscopy for corneal biomechanics. Invest Ophthalmol Vis Sci. 2012;53(1):185–90. pmid:22159012
- 87. Babker AMa. Information Technologies of the Processing of the Spaces of the States of a Complex Biophysical Object in the Intellectual Medical System HEALTH. IJATCSE. 2019;8(6):3221–7.
- 88. Belin MW, Jang HS, Borgstrom M. Keratoconus: Diagnosis and Staging. Cornea. 2022;41(1):1–11. pmid:34116536
- 89. Advanced glaucoma intervention study. 2. Visual field test scoring and reliability. Ophthalmology. 1994;101(8):1445–5.
- 90. KDIGO. KDIGO 2024 Clinical Practice Guideline for the Evaluation and Management of Chronic Kidney Disease. Kidney Int. 2024;105(4S):S117–314.
- 91. Rajkomar A, Dean J, Kohane I. Machine Learning in Medicine. N Engl J Med. 2019;380(14):1347–58. pmid:30943338
- 92. Quer G, Arnaout R, Henne M, Arnaout R. Machine Learning and the Future of Cardiovascular Care: JACC State-of-the-Art Review. J Am Coll Cardiol. 2021;77(3):300–13. pmid:33478654
- 93. Jenkins D, Chubb JR, Galea G. Stochastic processes in development and disease. Philos Trans R Soc Lond B Biol Sci. 2024;379(1900):20230043. pmid:38432319
- 94. Becker S, L’Ecuyer Z, Jones BW, Zouache MA, McDonnell FS, Vinberg F. Modeling complex age-related eye disease. Prog Retin Eye Res. 2024;100:101247. pmid:38365085
- 95. Janssen SF, Gorgels TGMF, Ramdas WD, Klaver CCW, van Duijn CM, Jansonius NM, et al. The vast complexity of primary open angle glaucoma: disease genes, risks, molecular mechanisms and pathobiology. Prog Retin Eye Res. 2013;37:31–67. pmid:24055863
- 96. Thalhauser CJ, Komarova NL. Alzheimer’s disease: rapid and slow progression. J R Soc Interface. 2012;9(66):119–26. pmid:21653567
- 97. Wagner AK, Soumerai SB, Zhang F, Ross-Degnan D. Segmented regression analysis of interrupted time series studies in medication use research. J Clin Pharm Ther. 2002;27(4):299–309. pmid:12174032
- 98. Bland JM, Altman DG. Statistical methods for assessing agreement between two methods of clinical measurement. Lancet. 1986;1(8476):307–10. pmid:2868172
- 99. Asmussen A, Smith BS, Møller F, Jakobsen TM. Repeatability and inter-observer variation of choroidal thickness measurements using swept-source optical coherence tomography in myopic danish children aged 6-14 years. Acta Ophthalmol. 2022;100(1):74–81. pmid:34126650
- 100. McCafferty S, Tetrault K, McColgin A, Chue W, Levine J, Muller M. Intraocular Pressure Measurement Accuracy and Repeatability of a Modified Goldmann Prism: Multicenter Randomized Clinical Trial. Am J Ophthalmol. 2018;196:145–53. pmid:30195894
- 101. Kumar M, Shetty R, Jayadev C, Rao HL, Dutta D. Repeatability and agreement of five imaging systems for measuring anterior segment parameters in healthy eyes. Indian J Ophthalmol. 2017;65(4):288–94. pmid:28513492
- 102. García-Resúa C, Pena-Verdeal H, Miñones M, Giráldez MJ, Yebra-Pimentel E. Interobserver and intraobserver repeatability of lipid layer pattern evaluation by two experienced observers. Cont Lens Anterior Eye. 2014;37(6):431–7. pmid:25113047
- 103. Aslam T, Mahmood S, Balaskas K, Patton N, Tanawade RG, Tan SZ, et al. Repeatability of visual function measures in age-related macular degeneration. Graefes Arch Clin Exp Ophthalmol. 2014;252(2):201–6. pmid:23884391
- 104. Salvetat ML, Zeppieri M, Tosoni C, Brusini P. Repeatability and accuracy of applanation resonance tonometry in healthy subjects and patients with glaucoma. Acta Ophthalmol. 2014;92(1):e66-73. pmid:23837834
- 105. Røys EÅ, Viste K, Kellmann R, Guldhaug NA, Alaour B, Sylte MS, et al. Estimating Reference Change Values Using Routine Patient Data: A Novel Pathology Database Approach. Clin Chem. 2025;71(2):307–18. pmid:39492685
- 106. Albert A, Zhang L. A novel definition of the multivariate coefficient of variation. Biom J. 2010;52(5):667–75. pmid:20976696
- 107. Walter SD, Gafni A, Birch S. A geometric confidence ellipse approach to the estimation of the ratio of two variables. Stat Med. 2008;27(28):5956–74. pmid:18720350
- 108. Tang M, Li Y, Chamberlain W, Louie DJ, Schallhorn JM, Huang D. Differentiating Keratoconus and Corneal Warpage by Analyzing Focal Change Patterns in Corneal Topography, Pachymetry, and Epithelial Thickness Maps. Invest Ophthalmol Vis Sci. 2016;57(9):OCT544-9. pmid:27482824
- 109. Levey AS, Grams ME, Inker LA. Uses of GFR and Albuminuria Level in Acute and Chronic Kidney Disease. N Engl J Med. 2022;386(22):2120–8. pmid:35648704
- 110. Carter HB, Pearson JD, Metter EJ, Brant LJ, Chan DW, Andres R, et al. Longitudinal evaluation of prostate-specific antigen levels in men with and without prostate disease. JAMA. 1992;267(16):2215–20. pmid:1372942
- 111. Paran Y, Yablecovitch D, Choshen G, Zeitlin I, Rogowski O, Ben-Ami R, et al. C-reactive protein velocity to distinguish febrile bacterial infections from non-bacterial febrile illnesses in the emergency department. Crit Care. 2009;13(2):R50. pmid:19351421
- 112. Hing JX, Mok CW, Tan PT. Clinical utility of tumour marker velocity of CA 15-3 and CEA in breast cancer surveillance. Breast. 2020;52:95–101.