Skip to main content
Advertisement
Browse Subject Areas
?

Click through the PLOS taxonomy to find articles in your field.

For more information about PLOS Subject Areas, click here.

< Back to Article

Table 1.

Abbreviations and Notation.

More »

Table 1 Expand

Fig 1.

Feature extraction pipeline for estimating P(AKI) during rehospitalization.

We sought to estimate the probability of AKI during rehospitalization given all of a patient’s previous hospitalizations, ), shown by the red arrow. An example of three hospitalizations (H1, H2, H3) is shown. Here H1 and H2 are used to estimate P(AKI) during H3. The EHR captures raw data (shown in boxes closest to the time series tracings) of which our dataset contains N + 1 features. At each level, data is aggregated via domain-expertise-informed functions F and G. The pipeline produces a single, fixed-length representation of all previous hospitalizations to serve as input to a learning algorithm. Measurements from each hospitalization, and series of hospitalizations, are treated as sequences, denoted with operator s.

More »

Fig 1 Expand

Fig 2.

Cohort selection.

On the top left, the selection procedure used to obtain the rehospitalization cohort is shown. On the top right, the distribution of the 197,046 hospitalizations not preceded by a diagnosis of ESRD is shown. On the bottom, a schematic of predictor/target generation is shown for an example patient with n hospitalizations from which n − 1 training cases were derived. For each target rehospitalization, y, data from all prior hospitalizations, X, are used as predictors. Multiple prior hospitalizations were aggregated using G as described above.

More »

Fig 2 Expand

Table 2.

AKI diagnosis distribution.

More »

Table 2 Expand

Table 3.

Cohort demographics.

Statistics are computed per hospitalization. There are a total of 124,518 hospitalizations from 34,505 patients, each with more than one hospitalization. These are therefore all hospitalizations generated by patients in the final cohort (including the first hospitalization from each patient, for which AKI is not predicted).

More »

Table 3 Expand

Fig 3.

GBC evaluation.

ROC, Calibration, and PR curves for 50 iterations of 5-fold CV (250 lines shown; each of 50 iterations has 5 lines corresponding to the 5 outer folds of CV). The black diagonal line represents chance for the ROC curve and ideal for the calibration curve. Results are reported per hospitalization, not patient. Alpha level = 0.5, line weight = 0.5.

More »

Fig 3 Expand

Table 4.

Predictive performance.

ROC = Receiver Operating Characteristic, PR = Precision Recall, ALR1 = Anscombe LR1, GBC = Gradient Boosting Classifier, LR1 = l1-Penalized Logistic Regression, LSTM = Long Short-term Memory, HP = Highly Penalized, W = Weighted, S = Sampled, R = Recent (for GBC) or Randomized (for LR1, HPLR1), M = Medication, N = Noise.

More »

Table 4 Expand

Table 5.

Predictive performance comparison.

ROC = Receiver Operating Characteristic, PR = Precision Recall, ALR1 = Anscombe LR1, GBC = Gradient Boosting Classifier, LR1 = l1-Penalized Logistic Regression, LSTM = Long Short-term Memory, HP = Highly Penalized, W = Weighted, S = Sampled, R = Recent (for GBC) or Randomized (for LR1, HPLR1), M = Medication, N = Noise.

More »

Table 5 Expand

Fig 4.

GBC hospitalization- and patient-specific risk distributions.

Observed hospitalization-level risk is plotted against predicted risk (top row) and patient-level mean observed risk against mean predicted risk (bottom row). Distributions of predictions PP are shown at the hospitalization and patient level. At patient level, distributions that are difficult to discern from the scatter plot are shown. In the scatter plots, alpha level is 0.05 and the red calibration curve corresponds to all hospitalizations or to patients who had either mean risk over hospitalizations of 1 or 0. The calibration curves are computed according to the macro-averaged predicted output per hospitalization or patient over the 50 iterations of 5 fold CV (over 250 total folds). Ideal calibration is the dotted black diagonal. Histograms have 1000 bins to give necessary resolution. PO = observed risk per hospitalization, PP = predicted risk per hospitalization, = mean observed risk over hospitalizations, = mean predicted risk over hospitalizations.

More »

Fig 4 Expand

Fig 5.

GBC prediction variance.

The mean and standard deviation of predicted probabilities are plotted over iterations (per hospitalization). Alpha = 0.01 for all plots.

More »

Fig 5 Expand

Table 6.

Coefficients of features associated with error.

For diagnoses, features correspond to the count assigned in prior hospitalizations. Note that age and diagnosis were fit in separate regressions despite being displayed in the same table.

More »

Table 6 Expand

Fig 6.

GBC error by utilization.

The mean and STD absolute error is shown as a function of the number of hospitalizations. Patients were binned based on the number of hospitalizations in the dataset and then, over bins, the mean error and STD of the predictions were computed. Stratification by outcome is performed since it was earlier established that the hospitalization:patient ratio is higher in cases than in controls.

More »

Fig 6 Expand

Table 7.

Feature importances/coefficients for GBC and LR1.

For laboratory results, the first function is G, aggregation over hospitalizations, and the second is F, aggregation within a hospitalization; e.g., “mean max sCr” is the mean over hospitalizations of the maximum sCr of each hospitalization.

More »

Table 7 Expand

Table 8.

Coefficients of HPLR1.

More »

Table 8 Expand

Table 9.

Feature importances/coefficients for MGBC and MLR1.

Each feature corresponds to the count of administrations of the medication over prior hospitalizations.

More »

Table 9 Expand