Skip to main content
Advertisement
Browse Subject Areas
?

Click through the PLOS taxonomy to find articles in your field.

For more information about PLOS Subject Areas, click here.

  • Loading metrics

Validation of expected objective performance indicators in surgical data analysis

  • Ye Li ,

    Contributed equally to this work with: Ye Li, Max Berniker

    Roles Formal analysis, Methodology, Validation, Writing – original draft, Writing – review & editing

    ye.li@intusurg.com

    Affiliation Advanced Product Development, Intuitive Surgical, Inc., Peachtree Corners, Georgia, United States of America

  • Sreeram Kamabattula,

    Roles Writing – review & editing

    Affiliation Advanced Product Development, Intuitive Surgical, Inc., Peachtree Corners, Georgia, United States of America

  • Kiran Bhattacharyya,

    Roles Conceptualization, Supervision, Writing – review & editing

    Affiliation Advanced Product Development, Intuitive Surgical, Inc., Peachtree Corners, Georgia, United States of America

  • Max Berniker

    Contributed equally to this work with: Ye Li, Max Berniker

    Roles Conceptualization, Formal analysis, Methodology, Writing – original draft, Writing – review & editing

    Affiliation Advanced Product Development, Intuitive Surgical, Inc., Sunnyvale, California, United States of America

Abstract

Objective Performance Indicators (OPIs) quantify surgical activity by measuring critical aspects of a case (e.g., surgeon movements, interaction forces, etc.) at clinically meaningful steps. They are particularly useful in robotic-assisted surgery, where the information required to compute them can be harvested automatically, and early studies have found that they help quantify surgeon skill. However, since surgical procedures can be highly variable and the boundaries between surgical steps often subject to interpretation, OPIs become similarly variable. Yet, OPIs should be insensitive to this variability to remain informative and unambiguous. Previously we proposed a method for computing expected OPI values in a probabilistic framework that has theoretical support for being robust to this variability. Here we validate this method and compare it against the conventional approach on a large clinical dataset. Using 5408 annotated clinical steps from 1,016 clinical cases across seven procedure types, we performed an extensive comparison of this new method alongside the conventional approach. A series of analyses compared the resulting OPI values focusing on potential biases and the variability of their values in the presence of annotator noise. We found that the probabilistic method resulted in OPI values that were less often biased than their conventional OPI counterparts (49.37%, and 56.63% respectively), and the bias magnitudes were smaller. Finally, we found that the probabilistic method resulted in significantly less variable OPI values in 40.56% of comparison, whereas 8.52% showed the opposite. Our analysis suggests that the expected OPIs are more robust to annotator variability, with less bias and smaller variance. We suggest that this probabilistic approach holds promise for the computational analysis, and perhaps the conceptual interpretation, of surgical data.

Introduction

Robotic assisted surgery (RAS) is growing in prominence and prevalence. The number of cases performed each year with RAS is growing at an exponential rate [1], and the types of procedure offered are expanding [2,3]. Importantly, each RAS procedure yields a wealth of data as a matter of course (e.g., video, surgeon motor behavior, tool kinematics, forces, etc.), providing surgeons and scientists with the unique opportunity to analyze surgeries in precise and quantifiable terms. Objective performance indicators (OPIs) are one such means of quantifying surgery.

An OPI is a metric to analyze surgical activity during a specific time window. For example, how quickly and far tools were moved, or the forces used to interact with tissue, or cauterization rates, during a particular window of activity, are all features that can be quantified through OPIs. The windows of activity they examine can be defined with different levels of granularity, such as an entire case, a phase, or what we focus on here, a single surgical step. As evidence of their importance, connections between OPIs measured during critical categories of steps and surgeon skills have already been made [4,5]. Choosing not only which feature to measure, but which step to measure it, may prove crucial for OPI analyses.

Unfortunately, unambiguously identifying steps can prove particularly challenging. For a given surgical case, various techniques, surgeons’ preferred strategies, and patients’ unique anatomies all increase the variability of surgical steps. In addition, annotations are usually made postoperative, without the benefit of the surgeon’s direct input. As such, the boundaries (start and stop times) of steps can be ambiguous and subject to interpretation. This variability, in turn, renders the OPIs variable and potentially biased, limiting their precision in evaluating surgical performance. A method for computing OPIs that is robust to this annotation variability is paramount.

We previously proposed a probabilistic framework for computing OPIs [6]. Rather than computing OPIs with “hard” temporal boundaries (i.e., discrete points in time for the start and stop of tasks), we described annotations as random samples from an underlying distribution and computed the resulting expected OPI value. We demonstrated how the expected OPIs, approximated using a finite sample of annotations, were more accurate and precise (with smaller biases and variances) than the conventional hard-boundary method.

The aim of this study is to verify those theoretical benefits on a rich, real-world dataset. We gathered data from seven procedure types and over a thousand surgical cases. Accompanying each case were the annotations, kinematics, and system event data necessary to compute OPI values. For each annotation we examined five different OPIs, computing both their hard-boundary and expected values. Using empirical estimates of annotation variability, we found that the expected OPI values were more accurate and less variable, verifying the theoretical benefits of the probabilistic approach. We suggest it may be helpful to think of OPIs as random variables that must be estimated.

Materials and methods

Data collection

In this study, we retrospectively collected data from 1,016 de-identified surgical cases performed using the da Vinci® Xi Surgical System. These cases included seven procedure types: lobectomy (N = 238), ventral hernia (N = 76), hysterectomy (N = 137), low anterior resection (N = 56), radical prostatectomy (N = 54), inguinal hernia (N = 211), and cholecystectomy (N = 244). For each case, we captured the tool kinematics (position and orientation) of the system’s four robotic arms, along with hand and camera clutch events, and energy application events (occurring during cauterization).

Surgical procedures can be broken into a series of clinically relevant, mutually exclusive steps. Each procedure type has unique surgical steps based on an ontology derived from consensus recommendations made by surgeons [7], with detailed methodology described elsewhere [8]. For some procedure types, the parameters which defined the start and stop times of steps were modified early in our data collection process to improve annotator alignment.

All annotators first completed a standardized training and calibration process with clinical oversight provided by a trained surgeon, until annotations were within five seconds of agreement. Teams of four to seven annotators were assigned to each procedure type. Following calibration, each case was annotated by a single trained annotator, who identified the start and stop times for the case’s surgical steps.

Retrospective system data collected during surgical procedures were analyzed under an approved IRB protocol (IRB tracking number: 20182083, WCG Clinical), with all patient identifiers removed prior to analysis. De-identified raw system data for this study were accessed between 01/09/2022 and 16/12/2024. Consent form was obtained from surgeons and patients before data was collected.

Objective performance indicators

Objective performance indicators (OPIs) computed over surgical steps can quantify any number of surgeon behaviors or techniques. Typical OPIs include the duration of a step, or the distance travelled by an instrument [9]. These OPIs are usually computed via integrals, or in discrete-time, sums.

For example, suppose quantifies some function, , integrated from to Assume is defined over a set of equally spaced discrete values in time (with a step size, ∆), . Then , from start time, to stop time, , is computed as:

(1)

where . In this study, we refer to OPIs computed by this method as the “hard-boundary” OPIs.

As an alternative, consider the scenario where the start and stop times of a step are not known, but rather we have a belief, or joint distribution, over start and stop times, . The OPI is now a random variable. To compute the expected OPI value we begin with the probability of a step. That is, is the probability of a step at time t, implying and :

(2)

The expected OPI value, can be found as:

(3)

This is our best estimate of when it is a random and uncertain variable. We refer to OPIs computed by this method as “expected” OPIs.

Prior to this investigation, an internal study was conducted to quantify the uncertainty associated with human annotation of surgical steps. Multiple annotators independently annotated the same cases. Using these replicate annotations, we estimated the variance of start and stop annotations and calculated the corresponding standard deviation, which was approximately two seconds. We observed that annotation variability did not differ substantially across steps or procedure types. Based on these findings, we assumed the joint distribution of annotations, , could be approximated with two independent normal distributions, both with standard deviations of two seconds. In this study we chose to limit our analysis to surgical steps that had only a single annotated instance per case, since two or more instances of the same step in quick succession could invalidate our normality assumption for (e.g., if the stop of one instance was less than two seconds before the start of the next instance). For a complete list of surgical steps analyzed see S1 Table.

We computed five different OPIs for the surgical steps in this study, using the two methods described above. These OPIs and their definitions are as follows: (1) tool path length: the total path length traveled by all instruments during a given surgical step, (2) angular path length: the total endowrist angular path length traveled by all instruments during a given surgical step, (3) camera clutch rate: the average rate of endoscope clutches performed during the step, (4) energy use rate: the average rate of energy pedal presses during the step, and (5) hand clutch rate: the average rate of hand controller clutches performed during the step.

Statistical analyses

To verify that the probabilistic framework for computing OPIs is superior to the conventional hard-boundary method, we tested two aspects: 1) whether the expected OPI was more accurate than the hard-boundary OPI, and 2) whether the expected OPI was more robust to annotator variability. To do so we performed the following resampling process. For each single-instance step, we used the annotator derived start and stop times to compute “nominal” hard-boundary and expected OPI values. These nominal values were used as reference OPI values for subsequent comparisons. Then, using a normal distribution centered on these nominal annotations we resampled 500 start and stop time pairs. We used a standard deviation of two seconds, based on an earlier study of the variability of annotations among skilled annotators. For edge instances where either the step start or stop time was the beginning or the end of the procedure respectively, we adjusted the resampling process and only resampled the free unconstrained boundary. For each resampled pair, the five OPIs were calculated using the two methods.

To compare accuracy, we evaluated how closely the randomly sampled OPI values matched their nominal values, and how large the sample bias was. One-sample t-tests were performed to determine whether the sampled OPIs were significantly different from their respective nominal values. We also compared the OPI values from two methods directly, using two-sample t-tests. To evaluate the sample bias, we normalized the mean bias with the nominal OPI values. Then we compared the normalized bias of two OPI methods by the Wilcoxon signed-rank test.

To assess robustness, we examined whether either method produced sample sets with a narrower range of values. We used Levene’s test to evaluate differences in sample variances between the two methods. Similar to the analyses of bias, we also computed a normalized standard deviation using the nominal OPI values. Then we compared the normalized standard deviation of two OPI methods by the Wilcoxon signed-rank test.

The percentage of steps showing significant differences in biases and variances provides insight into the extent to which OPI values may differ when replacing the conventional hard-boundary approach with the novel probabilistic approach. For all tests a significance level of α = 0.05 was applied.

Results

Data from 1,016 surgical cases spanning seven procedure types was collected. In total, 16,124 annotated step instances were identified across all cases. Of these, 10,393 step instances (64.46%) were excluded because they occurred multiple times within a case. The proportion of excluded multi-instance steps varied across procedure types, ranging from 34.6% to 78.4%, reflecting differences in inherent workflow structure and the frequency of iterative steps within certain procedures. 323 additional step instances were excluded due to data quality issues. The remaining 5,408 single-instance surgical steps (ranging from 1 to 14 per case, mean = 5) were included in the analysis. Five nominal OPIs (instrument path length, endowrist angular length, camera clutch rate, energy use rate and hand clutch rate) were computed for each step, along with paired, randomly sampled sets of hard-boundary and expected OPI values. In total this yielded 27,040 paired OPIs for comparison and analysis.

To provide some numerical context for the OPIs, the means for the two methods are displayed below (see Table 1). These are averaged across all steps and cases within each procedure type, so no detailed conclusions should be drawn, but rather the values are displayed to convey their relative magnitudes and their variability. Note that while the OPI values do vary considerably (in both value and nature), their values are relatively similar between methods. This demonstrates that the distinctions between the two methods require a careful analysis.

thumbnail
Table 1. Mean values of hard-boundary and expected OPIs across steps of each procedure type.

https://doi.org/10.1371/journal.pone.0355276.t001

Accuracy of expected OPIs

Our first comparison was in terms of accuracy, and in particular whether one approach provided OPI values that were closer to their nominal value when the annotations were noisy and uncertain. To do so we compared the nominal OPIs values with the means of their randomly resampled counterparts. First, we note that across procedure types and OPIs, many comparisons, indeed often most, are statistically distinct (see Table 2). In theory, the OPIs do not vary linearly with their annotations and adding noise to the annotations will bias their values. This result drives home how in practice, this is clearly the case as well. Regardless, relative to the hard-boundary OPIs, the mean expected OPIs were more often indistinguishable from their nominal values. Averaging across the five OPIs and seven procedure types, the percentage of steps whose sample means were significantly distinct from their nominal OPI values were 56.63% and 49.37%, for the hard-boundary and expected OPIs, respectively.

thumbnail
Table 2. Percentage of steps with significant biases OPI values.

https://doi.org/10.1371/journal.pone.0355276.t002

When comparing the kinematic OPIs (tool path length and angular path length) the expected OPIs showed a consistent and substantial improvement in accuracy across procedure types. Approximately 72.25% of hard-boundary OPI comparisons were statistically distinct, whereas the expected values were distinct in only 49.78% of comparisons. In summary, all five OPIs tested, particularly the kinematic OPIs, were more accurate using the probabilistic approach than those obtained with the hard-boundary approach.

Next, we examined the similarity of the values generated by the two approaches (Table 2, “comparisons”). The frequency and size of these differences would be indicative of hard-boundary biases. Similar to the above observation, the kinematic OPI of angular path length consistently had the largest number of distinct comparisons. When averaging across the five OPIs and seven procedure types, 48.61% of the hard-boundary and expected OPI samples were significantly distinct. These differences indicate the relative fragility of the hard-boundary approach.

Lastly, to provide more context we computed the average biases of the two approaches. Since the OPIs differ in category and magnitude we compared biases as an unsigned percentage of their nominal value (Table 3). As can be seen, the biases are all less than 10%, and the biases of the expected OPIs are smaller than their hard-boundary counterparts in every single example. We tested the significance of the biases (Wilcoxon signed-rank test) and found that in all but three scenarios the expected bias was significantly smaller than that of the hard-boundary bias.

thumbnail
Table 3. Mean normalized percent bias of hard-boundary and expected OPIs and the p-value of their comparisons (Wilcoxon signed-rank test).

https://doi.org/10.1371/journal.pone.0355276.t003

Robustness of hard-boundary OPIs

Next, we compared the robustness of the two approaches, examining the relative effects of annotator noise on the resulting variability in OPI values (independent of their accuracy). We compared the sample variances across the seven procedure types and five OPIs (Levene’s test, see Table 4). Almost all OPIs had a greater number of smaller variance results with the probabilistic method. Angular path length was the exception, where there were a moderate amount of steps with larger expected OPI variance (23.30% of steps, relative to 20.28% of the steps with the hard-boundary approach). However, when averaging across the five OPIs, the probabilistic approach was superior. Finally, a weighted average (taking into account the differing steps per procedure) across OPIs and procedure types found that 40.56% of the steps had larger hard-boundary variance, whereas only 8.52% of the steps had larger expected variances.

thumbnail
Table 4. Percentage of tasks showing higher variance in one OPI method than the other.

https://doi.org/10.1371/journal.pone.0355276.t004

As with the biases, we compared the standard deviations as an unsigned percentage of their nominal OPI value (Table 5). As can be seen, the standard deviations of the expected OPIs were significantly smaller in all conditions.

thumbnail
Table 5. Mean normalized percent of standard deviation based on the OPI values and the p-value of its comparison using Wilcoxon signed-rank test.

https://doi.org/10.1371/journal.pone.0355276.t005

Discussion

In a previous study we proposed a probabilistic approach to OPI computation, outlining its theoretical benefits over the conventional hard-boundary method, and presenting a proof of concept on a limited data set [6]. Here our aim was to determine if these previously reported benefits withstand a rigorous examination with a rich and extensive clinical data set. We collected data from 1,016 surgical cases, across seven procedure types, and 194 step categories. In total this provided 5,584 manually annotated surgical steps. For every annotation we computed five OPIs. This provided a uniquely large dataset of clinically-derived OPIs to validate the probabilistic OPI approach.

Across all seven procedure types and 5,584 surgical steps, we observed that expected OPIs exhibit bias less frequently than the traditional hard-boundary OPIs. This was particularly the case for the kinematic OPIs of tool path length and angular path length. Furthermore, about half of the surgical steps we examined showed distinct OPI values between the hard-boundary and probabilistic approaches, highlighting the potential benefits of transitioning to a probabilistic framework for OPI computation. In terms of variance, expected OPIs consistently demonstrated an advantage across all procedure types with more steps exhibiting significantly higher variance with the hard-boundary method than vice versa. This pattern was observed for all OPIs except angular path length, suggesting that probabilistic OPIs provide more stable measurements and are less sensitive to annotation variability than conventional hard-boundary OPIs.

Probabilistic OPIs demonstrated comparable or superior performance, and adopting them can provide additional benefits as well. Consider that probabilistic OPIs do not require a change in the established methodology for annotating steps. This is crucial since human annotations are usually the most time and resource intensive element of the analysis. There is a nominal increase in computational effort, however, and either the probability of a step or the distribution of annotation times is required. Yet this effort can be viewed as an advantage since it provides a mechanism to incorporate variability and uncertainty in the annotation process. For example, annotation variability may differ across procedures, steps, and human annotators. All these distinct sources of variability can be explicitly addressed both at the level of annotation, but also in subsequent downstream computations.

Here we limited our analysis in three important respects. First, our analysis was restricted to single instance steps. There are conditions under which a surgical step might be completed with two or more distinct portions in quick succession. When the duration between instances is short relative to the variability of an annotation, modeling annotation uncertainty is more complicated. To keep the analysis straightforward we chose to ignore all multi-instance steps regardless of whether or not they violated our assumptions. In the future we hope to explore probabilistic techniques to that are robust to these conditions. Second, we assumed a fixed annotation variability (σ = 2 seconds), derived from an internal study of inter-annotator variability. We did not perform a formal sensitivity analysis to evaluate the robustness of our results with respect to this assumption. While this representative value enabled a clear demonstration of the method, future work could assess the impact of varying levels of annotation uncertainty. Finally, we have assumed annotation uncertainty is fixed and constant across all steps and procedure types. In settings where annotation uncertainty is more likely to change, for example, with complicated workflow, context-dependent uncertainty could be examined.

OPIs are promising for a number of applications such as, evaluating surgeon skill [812], predicting surgical outcomes [1316], and providing directed feedback for surgeon training [17,18]. The same basic approach of addressing variability and uncertainty employed here, could be used in these downstream OPI applications. For example, rather than providing a read-out of surgical skill as a function of OPIs, a probability distribution over skill could be computed. We hope to explore possibilities like this in the future.

Identifying the beginning and ending of surgical segments is an inherently variable process, yet crucial for computing objective performance indicators (OPIs). To mitigate the influence of this variability on OPIs, we propose taking a probabilistic approach that computes their expected value. Our results suggest that this probabilistic approach holds promise for the computational analysis, and perhaps the conceptual interpretation, of surgical data.

Supporting information

S1 Table. Surgical steps included in the analysis.

https://doi.org/10.1371/journal.pone.0355276.s001

(DOCX)

References

  1. 1. Rizzo KR, Grasso S, Ford B, Myers A, Ofstun E, Walker A. Status of robotic assisted surgery (RAS) and the effects of Coronavirus (COVID-19) on RAS in the Department of Defense (DoD). J Robot Surg. 2023;17(2):413–7. pmid:35739435
  2. 2. Intuitive Surgical, Inc. Intuitive Environmental, Social Responsibility, and Governance (ESG) Report. 2022. Available from: https://www.intuitive.com/en-us/-/media/ISI/Intuitive/Pdf/2022-intuitive-esg-report.pdf
  3. 3. Childers CP, Maggard-Gibbons M. Estimation of the acquisition and operating costs for robotic surgery. JAMA. 2018;320(8):835–6.
  4. 4. Reiley CE, Lin HC, Yuh DD, Hager GD. Review of methods for objective surgical skill evaluation. Surg Endosc. 2011;25(2):356–66. pmid:20607563
  5. 5. Oh DS, Ershad M, Wee JO, Sancheti MS, D’Souza DM, Herrera LJ, et al. Comparison of Global Evaluative Assessment of Robotic Surgery with objective performance indicators for the assessment of skill during robotic-assisted thoracic surgery. Surgery. 2023;174(6):1349–55. pmid:37718171
  6. 6. Berniker M, Bhattacharyya KD, Brown KC, Jarc A. A probabilistic approach to surgical tasks and skill metrics. IEEE Trans Biomed Eng. 2022;69(7):2212–9. pmid:34971527
  7. 7. Meireles OR, Rosman G, Altieri MS, Carin L, Hager G, Madani A, et al. SAGES consensus recommendations on an annotation framework for surgical video. Surg Endosc. 2021;35(9):4918–29. pmid:34231065
  8. 8. Mlambo B, Shields M, Bach S, Bauer A, Hung A, Kudsi OY, et al. A standardized temporal segmentation framework and annotation resource library in robotic surgery. Mayo Clin Proc Digit Health. 2025;3(4):100257. pmid:41050187
  9. 9. Brown KC, Bhattacharyya KD, Kulason S, Zia A, Jarc A. How to bring surgery to the next level: interpretable skills assessment in robotic-assisted surgery. Visc Med. 2020;36(6):463–70. pmid:33447602
  10. 10. Lyman WB, Passeri MJ, Murphy K, Siddiqui IA, Khan AS, Iannitti DA, et al. An objective approach to evaluate novice robotic surgeons using a combination of kinematics and stepwise cumulative sum (CUSUM) analyses. Surg Endosc. 2021;35(6):2765–72.
  11. 11. Quinn KM, Chen X, Runge LT, Pieper H, Renton D, Meara M, et al. The robot doesn’t lie: real-life validation of robotic performance metrics. Surg Endosc. 2023;37(7):5547–52. pmid:36266482
  12. 12. Jarc AM, Curet MJ. Viewpoint matters: objective performance metrics for surgeon endoscope control during robot-assisted surgery. Surg Endosc. 2017;31(3):1192–202. pmid:27422247
  13. 13. Lazar JF, Brown K, Yousaf S, Jarc A, Metchik A, Henderson H, et al. Objective performance indicators of cardiothoracic residents are associated with vascular injury during robotic-assisted lobectomy on porcine models. J Robot Surg. 2023;17(2):669–76. pmid:36306102
  14. 14. Kaoukabani G, Gokcal F, Fanta A, Liu X, Shields M, Stricklin C, et al. A multifactorial evaluation of objective performance indicators and video analysis in the context of case complexity and clinical outcomes in robotic-assisted cholecystectomy. Surg Endosc. 2023;37(11):8540–51. pmid:37789179
  15. 15. Maan ZN, Maan IN, Darzi AW, Aggarwal R. Systematic review of predictors of surgical performance. Br J Surg. 2012;99(12):1610–21. pmid:23034658
  16. 16. Curtis NJ, Foster JD, Miskovic D, Brown CSB, Hewett PJ, Abbott S, et al. Association of surgical skill assessment with clinical outcomes in cancer surgery. JAMA Surg. 2020;155(7):590–8. pmid:32374371
  17. 17. Ma R, Lee RS, Nguyen JH, Cowan A, Haque TF, You J, et al. Tailored feedback based on clinically relevant performance metrics expedites the acquisition of robotic suturing skills-an unblinded pilot randomized controlled trial. J Urol. 2022;208(2):414–24. pmid:35394359
  18. 18. Cizmic A, Häberle F, Wise PA, Müller F, Gabel F, Mascagni P, et al. Structured feedback and operative video debriefing with critical view of safety annotation in training of laparoscopic cholecystectomy: a randomized controlled study. Surg Endosc. 2024;38(6):3241–52. pmid:38653899