Skip to main content
Advertisement
Browse Subject Areas
?

Click through the PLOS taxonomy to find articles in your field.

For more information about PLOS Subject Areas, click here.

  • Loading metrics

Optimising the Delphi survey method during core set development: The impact of summarised feedback on stakeholders’ prioritisation of core data items and consensus

  • Katy A. Chalmers ,

    Roles Data curation, Formal analysis, Investigation, Methodology, Project administration, Validation, Visualization, Writing – original draft, Writing – review & editing

    katy.chalmers@bristol.ac.uk

    Affiliations National Institute for Health and Care Research Bristol Biomedical Research Centre, University Hospitals Bristol and Weston NHS Foundation Trust and University of Bristol, Bristol, United Kingdom, Bristol Centre for Surgical Research, University of Bristol, Bristol, United Kingdom

  • Karen Coulman,

    Roles Conceptualization, Data curation, Methodology, Supervision, Validation, Writing – original draft

    Affiliations National Institute for Health and Care Research Bristol Biomedical Research Centre, University Hospitals Bristol and Weston NHS Foundation Trust and University of Bristol, Bristol, United Kingdom, Bristol Centre for Surgical Research, University of Bristol, Bristol, United Kingdom, Obesity and Bariatric Surgery Service, North Bristol NHS Trust, Bristol, United Kingdom

  • Jane M. Blazeby,

    Roles Conceptualization, Funding acquisition, Methodology, Supervision, Writing – original draft, Writing – review & editing

    Affiliations National Institute for Health and Care Research Bristol Biomedical Research Centre, University Hospitals Bristol and Weston NHS Foundation Trust and University of Bristol, Bristol, United Kingdom, Bristol Centre for Surgical Research, University of Bristol, Bristol, United Kingdom

  • John Dixon,

    Roles Investigation, Validation, Writing – review & editing

    Affiliation Iverson Health Innovation Research Institute, Swinburne University of Technology, Melbourne, Australia

  • Lilian Kow,

    Roles Investigation, Project administration, Validation, Writing – review & editing

    Affiliation College of Medicine and Public Health, Flinders University, Adelaide, Australia

  • Ronald Liem,

    Roles Conceptualization, Investigation, Methodology, Validation, Writing – original draft, Writing – review & editing

    Affiliation Department of Surgery, Groene Hart Hospital, Gouda, The Netherlands

  • Dimitri J. Pournaras,

    Roles Data curation, Investigation, Methodology, Project administration, Validation, Writing – original draft, Writing – review & editing

    Affiliation Obesity and Bariatric Surgery Service, North Bristol NHS Trust, Bristol, United Kingdom

  • Johan Ottosson,

    Roles Investigation, Methodology, Writing – review & editing

    Affiliation School of Medical Sciences, Örebro University, Örebro, Sweden

  • Richard Welbourn,

    Roles Investigation, Methodology, Writing – original draft, Writing – review & editing

    Affiliation Department of Upper GI and Bariatric Surgery, Somerset NHS Foundation Trust, Taunton, United Kingdom

  • Wendy Brown,

    Roles Conceptualization, Investigation, Writing – review & editing

    Affiliation Department of Surgery, Monash University, Melbourne, Australia

  • Kerry N. L. Avery

    Roles Conceptualization, Data curation, Funding acquisition, Investigation, Methodology, Project administration, Supervision, Validation, Visualization, Writing – original draft, Writing – review & editing

    ‡ This author are senior author on this work.

    Affiliations National Institute for Health and Care Research Bristol Biomedical Research Centre, University Hospitals Bristol and Weston NHS Foundation Trust and University of Bristol, Bristol, United Kingdom, Bristol Centre for Surgical Research, University of Bristol, Bristol, United Kingdom

Abstract

Background

Core outcome sets are an established method for standardising the collection, measurement and reporting of treatment outcomes in effectiveness trials. Using Delphi survey methodology, core sets are developed by prioritising and re-prioritising data items facilitated by provision of feedback of other stakeholders’ responses. It is unknown how best to provide feedback to ensure that it influences the re-prioritisation of items effectively. This study examined whether informing participants of the top-rated items from the previous survey round may influence the re-prioritisation of data items in a subsequent survey round during the development of core data sets.

Methods

This study was nested in the development of a registry core data set. In round two of the Delphi survey, participants were randomised to receive ‘standard’ or ‘enhanced’ instructions. ‘Enhanced’ instructions included summarised data of the top five data items scored by participants in the previous survey round) in addition to standard feedback (the median round 1 score per item). Items scored 7–9 by ≥70% of participants in round 2 were considered ‘prioritised’. Concordant/discordant items were determined and extent of agreement between groups calculated (kappa statistics).

Results

Both groups prioritised a larger number of items in round 2 than in round 1 and there was little difference in the percentage of respondents prioritising the ‘Top 5’ items in round 2 (mean change in prioritisation of Top 5 items for all four core sets combined – 2.3% increase in standard group and 3.2% increase in enhanced group). Overall agreement in data items prioritised by both groups improved in round 2 (discordant items – 11% in round 1 and 4% in round 2).

Conclusion

Providing participants with additional feedback during the process of item prioritisation did not promote prioritisation of items during development of a core set. In the development of health core sets, where often many items are prioritised, further work to determine how to clearly and optimally communicatee feedback in a manner that promotes consensus effectively is required. Specifically, qualitative work with relevant stakeholders, exploring and clarifying the concepts of prioritisation and consensus, is warranted.

Introduction

Comparing and synthesising data across studies is fundamental to answering important clinical questions about health and clinical care. The ability to compare and combine data is, however, dependent on the compatibility of data. The development and use of a core data set, such as a core outcome set (COS), is one established approach to improve data synthesis [1]. COS are being increasingly developed for use in comparative effectiveness trials but are also important to standardise reporting in other contexts, such as disease- or treatment-specific registries. As such, there is growing interest in core set development methodology to optimise the efficiency and effectiveness of the development process and Delphi surveys are commonly used to prioritise core set items [2,3].

The Delphi survey method uses a simple framework in which researchers have flexibility in the methods used to seek consensus on a core set [4]. A key feature of this process is the provision of feedback to participants in subsequent survey rounds, which enables other participants’ opinions to be considered before re-prioritisation of items. Whilst ‘describe how participants receive feedback’ is an item on a checklist to standardise COS development [3], there is no agreed format of how it should be presented. Previous studies have shown the importance of feedback presentation, since patients and healthcare professionals prioritise different types of outcomes [5] and providing feedback from individual stakeholder groups may improve consensus between groups and agreement in items prioritised [6]. But how stakeholders interpret and utilise feedback, is unknown. Most Delphi studies use a Likert scale to score items [4,7] and therefore feedback is commonly presented as a summary statistic such as a mean or median value, with or without an indicator of range (e.g., standard deviation, inter-quartile range) [4]. Several published studies have noted that rating scales, such as Likert, lead to less variation in scores across items as repetition of scores is allowed (compared to ranking) [8] and that health Delphi studies tend to prioritise more items due to their perceived importance [8]. Together, in a health Delphi, this results in many items being scored highly and does not allow differentiation between the items considered the most important; a problem when the purpose of Delphi surveys is to reduce the number of items prioritised in each round. Given the increasing interest in the development of core sets in health service research, it seems pertinent to identify ways to improve the efficiency of prioritisation and consensus.

The aim of this nested randomised controlled trial (RCT) was to explore whether the provision of summarised feedback in addition to usual feedback (enhanced instructions), promoted the prioritisation of items compared to those receiving usual feedback only (standard feedback). This was explored within a Delphi survey for the prioritisation of items in a core data set for an obesity surgery registry.

Methods

Design

This was a prospective, parallel-group RCT nested within an ongoing study to develop core data sets for an international bariatric surgery registry [9]. The core data set development study comprised 3 steps: (i) identification of a long list of data items; (ii) the Delphi survey and (iii) the consensus meeting to finalise the core data sets. This RCT was nested within the Delphi survey (step ii). Methods for the original core data sets’ development study have been reported previously [9]. Further details of the Delphi survey step are briefly outlined below to provide relevant context to the methodological RCT.

Participants

All members of the International Federation of Surgery for Obesity and Metabolic Disorders (IFSO) were sent an email from the IFSO president, explaining the core data set project and inviting them to take part in the survey, if they wished. IFSO comprises a wide range of healthcare professionals including bariatric surgeons, psychologists, dietitians, bariatric physicians, and specialist nurses. Since this core data set did not focus on patient-reported outcomes, patients were not included as participants in the consensus process, however an international PPI group of people who had undergone bariatric surgery advised on the project.

Ethical approval for the study was granted by the Faculty of Health Sciences Research Ethics Committee (FREC) (ref. 116384) at the University of Bristol, UK. In line with other Delphi surveys completed by our group, consent was implied by completion of the survey; this approach was discussed with and approved by the FREC. Participants were not made aware that they were participating in a methodological study and were to be randomised to either standard or enhanced instructions in the round 2 survey since this may have biased the results. All procedures performed in studies involving human participants were in accordance with the ethical standards of the institutional and/or national research committee and with the 1964 Helsinki declaration and its later amendments or comparable ethical standards.

Delphi survey

Survey development.

A comprehensive list of potentially relevant items for the core sets was identified from earlier work, including a previous COS for bariatric surgery trials [5], a bariatric registry data dictionary project [10], systematic literature reviews and patient and public involvement group discussions. Items from each source were collated into a single long list where duplicates were removed and overlapping themes combined. The study team, comprising methodologists and health professionals, grouped the items into broader domains which reflected the different phases of data collection within bariatric registries. These items were included in a 97-item questionnaire divided into sections which would go on to form the four core data sets – 1. baseline information, 2. effectiveness outcomes, 3. surgical procedure information and 4. potential complications and side-effects of surgery. Free text boxes were included at the end of each section to enable stakeholders to propose new items. The electronic questionnaire was hosted on REDCap data capture software [11,12] and sent to all IFSO members via email.

Prioritisation of items.

In round 1 of the Delphi survey, all participants received identical surveys. They were asked to rate the importance of each item for inclusion in the final core data sets on a Likert scale of 1–9. A score of 7–9 indicated an item of critical importance, 4–6 an item important but not critical, and 1–3 an item of limited importance. Median scores for each item were calculated for all participants combined and for each stakeholder group (bariatric surgeons, dietitians, specialist nurses, psychologists, other). For round 2, in line with COMET guidelines [4], all data items were retained, regardless of their scoring in round 1. All responding participants received the same survey as in round 1 plus anonymised feedback for each item detailing their own score, their peer group’s median score, the other stakeholder groups (combined) median score and the overall median score. In addition to the feedback scores, patients were randomised to receive either ‘standard’ or ‘enhanced’ instructions to guide them in re-rating the items in round 2 (see ‘Randomisation’ below).

The percentage of participants rating each item 7–9 (critically important) after round 2 was calculated. In line with previous core sets [13,14] and COS guidance documents [3,1316], consensus to prioritise an item was defined, a priori, as being achieved if at least 70% of responding participants scored the item 7–9.

Items were categorised into three groups to facilitate discussion and voting on the final core set at the consensus meeting. These were: (1) items rated 7–9 by ≥95% participants to be ratified ‘in’ if no objections; (2) items rated 7–9 by 70–95% of all participants to be discussed in small groups and voted on by the whole group; and (3) items rated 7–9 by <70% of all participants to be ratified ‘out’ if no objections. Full details of the consensus meeting discussions and voting criteria have been described previously [9]. The final core data sets comprised 12 items [9].

Instructions.

In the round 1 survey, all participants received identical information about core sets and instructions on how to complete the survey (Fig 1). For the round 2 survey, two types of instructions (standard and enhanced) were developed (Fig 2). Standard instructions emphasised the provision of feedback from peer groups and other healthcare professionals and offered simple guidance on how to use the feedback presented when re-rating the data items for round 2 (Fig 2a). This is similar to that used in previous Delphi surveys to develop COS [13,17]. The aim of the enhanced instructions was to provide participants with summarised data of the ‘Top 5’ data items for each core set in order to facilitate re-rating data items in round 2. The Top 5 items were the five items prioritised (scored 7–9) by the highest percentage of participants in round 1. As with standard instructions, the introductory information emphasised the provision of feedback from peer groups and other healthcare professional groups but in addition, highlighted that the Top 5 data items for each core set were also provided (Fig 2b). For each of the four core sets, the Top 5 data items for each of the four sections were displayed (Fig 2b) with instructions on how to re-rate items.

thumbnail
Fig 1. Information and instructions presented to all participants to guide them in the completion of the questions in the round 1 survey.

https://doi.org/10.1371/journal.pone.0348136.g001

thumbnail
Fig 2. Information about the feedback (results) provided and instructions for how to re-rate data items in round 2.

Participants randomised to either standard (a) or enhanced (b) instructions in the round 2 survey.

https://doi.org/10.1371/journal.pone.0348136.g002

Randomisation

Participants who answered at least one question in round 1 of the Delphi survey, were randomised to receive either the standard or enhanced instructions in round 2. For those participants who had more than one identity number due to multiple openings of the survey, the most complete entry was included in the randomisation and the others, excluded. If completion rates were the same for all entries by one individual, the first entry was included. Participants with an odd numbered record ID were assigned standard instructions and those with an even numbered record ID, enhanced instructions. Randomisation was performed by the REDCap administrator using REDCap software [11,12]. Because of non-responders and participants with more than one ID number, randomisation to each instruction group was not strictly 1:1.

Sample size

This RCT was nested within an ongoing Delphi survey to develop a COS and, as such, the maximum possible sample size was restricted by the size of the Delphi survey. This approach reflects similar methodological studies nested within COS development studies where hypothesis testing is primarily exploratory.

Data analyses

The objective of this study was to explore the effects of instruction type (standard or enhanced) on the prioritisation of core items across two Delphi survey rounds. As such, only respondents who answered at least one question in both rounds 1 and 2 were included in the analyses. To establish whether respondents altered their responses in round 2 from round 1, data from round 1 respondents were grouped according to which set of instructions they received in round 2.

Response rates.

Response rates were calculated for each version of the survey (standard or enhanced) with the total number of participants per randomisation group as the denominator.

Prioritisation of items.

The number and percentage of items prioritised (rated 7–9 by ≥70% of participants) by each group was calculated for both rounds.

Agreement in items prioritised.

The number and percentage of concordant items (items that both groups did or did not prioritise) and discordant items (items that only one group prioritised) prioritised in each round was also calculated and the level of agreement between groups determined using kappa statistic (κ) [18]. The level of agreement was classified as poor (κ ≤ 0.0), slight (κ 0.01–0.20), fair (κ 0.21–0.40), moderate (κ 0.41–0.60), substantial (κ 0.61–0.80), or almost perfect or perfect (κ 0.81–1.0) [18]. All statistical analyses were performed in Stata version 17 [19].

Differences in prioritisation of ‘Top 5’ items.

The differences between groups in the percentage of participants prioritising round 1 ‘Top 5’ items (i.e., the 5 items for which the highest proportion of participants in each group scored the item as 7–9 in round 1) was calculated by subtracting the item scores in round 1 from round 2. A positive number signified an increase in the number of participants prioritising the item.

Results

Response rates and stakeholder characteristics of randomisation groups

All members of IFSO were invited by email to take part in the survey. The survey link was live for 6 months (April-September 2021) to allow a reasonable time for as many people as possible to complete the round 1 and round 2 surveys. A reminder was sent to all IFSO members in July 2021. In round 1, 272 participants answered at least one question; 260/272 (96%) participants completed section 1, 198 (73%) completed section 2 and 192 (71%) completed the whole survey. The least number of questions answered was three, by five people. All 272 round 1 respondents were randomised to receive either standard (n = 131, 48%) or enhanced (n = 141, 52%) instructions in round 2. In round 2, 123 participants (57 (44%) receiving standard instructions and 66 (47%) receiving enhanced instructions), answered at least one question (response rate of 45%). The majority (122/123 (99%)) completed section 1, 119 (97%) completed section 2 and 114 (93%) completed the whole survey. The least number of questions answered was 12, by one person. There was no difference in round 2 response rates between the two instruction groups. Details of participants’ characteristics, who answered both survey rounds, are presented in Table 1. Most respondents were surgeons, with a higher percentage of surgeons responding in the standard instruction group than the enhanced group (60% vs 52%). Dietitians were also well represented although similarly across both groups (21% standard vs 24% enhanced). The distribution of specialities in this sub-study, was broadly similar to all participants that completed round 1 [9]. Experience was similar between the groups with 60% of standard instructions and 59% enhanced group having >10 years’ experience. Respondents represented 35 countries with around one quarter from the United Kingdom.

thumbnail
Table 1. Details of participants completing both rounds 1 and 2 for each instruction group.

https://doi.org/10.1371/journal.pone.0348136.t001

Prioritisation of items

The number of items prioritised (scored 7–9 by ≥70%) by participants increased in both groups between rounds. The standard instructions group prioritised 58 (60%) items in round 1 and 70 (72%) in round 2. These additional 12 items were the result of thirteen further items being prioritised and one item that was prioritised in round 1, not being prioritised in round 2 (Table 2). The enhanced instructions group prioritised 53 items (55%) in round 1 and 66 (68%) in round 2–13 extra items. Differences in the items prioritised (scored 7–9 by ≥70% stakeholders) by at least one group in round 1, 2, or both are shown in Table 2. Items not prioritised by any group in either round are shown in Supplementary Table 1, and are also available in the original publication of the main study [9].

thumbnail
Table 2. Items prioritised (scored 7-9 by ≥70% stakeholders) by at least one group in at least one round.

https://doi.org/10.1371/journal.pone.0348136.t002

Agreement of items prioritised

In rounds 1 and 2 respectively, 86 (89%) and 93 (96%) items were concordant (prioritised or not prioritised by both groups) and 11 (11%) and 4 (4%) were discordant (prioritised or not prioritised by only one group) (Table 3).

thumbnail
Table 3. Agreement between groups in prioritised items.

https://doi.org/10.1371/journal.pone.0348136.t003

In round 1, agreement between groups in prioritised items was ‘substantial’ for core sets 1, 2 and 3b (κ = 0.61, 0.76 and 0.71, respectively) and perfect (κ = 1.0) for core set 3a (Table 3). The overall agreement for the four core sets was ‘substantial’ (κ = 0.77). In round 1, the standard instructions group prioritised 58 items and the enhanced instructions group 53, although there were 11 discordant items (three in core set 1, four in core set 2 and four in core set 3b) (Tables 2 and 3).

In round 2, agreement was found to be substantial for core set 1 (κ = 0.74), almost perfect for core sets 2 and 3a (κ = 0.93 and 0.82 respectively) and perfect for core set 3b (κ = 1.0). The overall agreement for the four core sets combined was κ = 0.90 indicating almost perfect agreement. In round 2, the standard instructions group prioritised 70 items and the enhanced instructions group 66 items. These 66 items were common to both instruction groups. The four discordant items prioritised by the standard instructions group were ‘details of the MDT’ (core set 1), ‘pre-surgery

weight loss’ (core set 1), ‘how well the pancreas produces insulin’ (core set 2) and ‘height of staples used’ (core set 3b) (Tables 2 and 3).

Overall, agreement between groups in prioritised items improved in round 2 for three of the four core sets. Overall agreement was higher in round 2 than round 1 (κ = 0.9 versus κ = 0.77) with only 4% discordant items, compared to 11% in round 1 (Table 3).

Differences in prioritisation of ‘Top 5’ items

The differences between groups in the percentage of participants prioritising round 1 ‘Top 5’ items are shown in Table 4. Differences between rounds 1 and 2 in the percentages of participants receiving standard instructions who prioritised the Top 5 items ranged from −2.8% (name of surgical procedure 96.6% to 93.8%) to +14.1% (hiatus hernia repair 82.8% to 96.9%). The mean change across the four core sets was + 2.3%. In the enhanced instructions group, differences ranged from −6.7% (hiatus hernia 92.0% to 85.3%) to +20.5% (other medical conditions not directly related to obesity 64.1% to 84.6%), with a mean change across the four core sets of +3.2%.

thumbnail
Table 4. Differences between groups in the percentage of participants prioritising round 1 ‘Top 5’ items.

https://doi.org/10.1371/journal.pone.0348136.t004

Discussion

This study examined the effects of providing summarised feedback on the prioritisation of items during a Delphi survey to reach consensus on the items to be included in a core data set. It was an exploratory randomised design, within the confines of an already planned core set development study, and therefore not powered sufficiently, so findings should be interpreted with this in mind. Results indicated that, overall, enhancing participant instructions in round 2 by including information about which five items were rated the highest in round 1, does not further promote consensus. Rather, both instruction groups prioritised a larger number of items in round 2 than in round 1 and there was little difference in the percentage of respondents prioritising the ‘Top 5’ items in round 2. Additional survey rounds due to lack of prioritisation and subsequent loss of participants due to survey fatigue may diminish the rigor of the core sets. Core sets are a useful tool in standardising research outputs, and gaining consensus is at the heart of this process.

Growing interest in core set methodology suggests that core set developers are trying to establish the best methodology to facilitate prioritisation and consensus. In the last five years (2019-current) [7], 64 studies were registered with the COMET database as investigating ‘COS methods research’. These studies explore numerous aspects of core set development including the ordering of survey items (e.g., patient-reported before/after clinical outcomes) [20], Likert rating scales [21,22] and presentation of feedback to participants in subsequent round(s) [6]. However, currently, there is no definitive method for developing COSs. Guidance documents created to standardise the development and reporting of a COS [3,7,15,16] navigate users through the process but do not prescribe methods for accomplishing COS generation and researchers continue to take different approaches [4]. The COS-STAP checklist states that the definition of consensus and how patients will receive feedback during the consensus process should be described [3], but does not provide guidance about developing survey and/or prioritisation instructions. This study has identified research questions about the role of instructions to achieve consensus during core set development but also highlights other areas of the Delphi process that warrant further exploration.

We hypothesised that the use of enhanced instructions, which provided summarised feedback in the form of the top five items prioritised by participants in round 1 for each core set, would efficiently guide participants through a sea of ‘critically important’ health items to those voted ‘critically important’ by the greatest proportion of participants. We reasoned that participants would be in a more informed position to re-rate the top five items higher, if in agreement, or lower, to indicate disagreement, thereby promoting fewer items being prioritised and consensus reached on a minimal core set. Instead, however, we observed that a greater number of items were prioritised, making the finalisation of a minimal core set, more challenging. The absence of prioritisation in both groups suggests that the presentation of median scores with or without a ‘Top 5’ did not facilitate prioritisation further and the reasons for this need to be considered. It is plausible that, in an effort to improve consensus (by modifying their opinions more in line with others’), stakeholders overlooked the aim of prioritising only ‘critically important’ items, by scoring more items as ‘important’. This finding highlights complexities in the mechanisms by which instructions and feedback influence participants’ behaviour in Delphi surveys, suggesting the need for in-depth qualitative work to explore participants’ interpretations and understanding during core set development. Indeed, it may have been a valuable addition to this work to have interviewed participants following completion of the surveys, however it was outside the scope of the study. Such work may focus on improving participants’ understanding of the concepts of ‘prioritisation’ and ‘consensus’ and how instructions may optimally distinguish between the two while simultaneously promoting both.

Previous qualitative work has explored participants’ understanding of the purpose of a COS and the Delphi process [23,24]. Through online feedback surveys [24] and in-person interviews [23], it was evident that participants’ understanding was variable [23,24] and was affected by previous experience with COS development and/or the Delphi survey process [23]. It was concluded that participants would benefit from repeated guidance on the principles for COS development during voting [24] and that further guidance and support needed to be accessible and salient [23]. Whilst the COMET Initiative website provides resources for lay audiences to explain the concepts of COSs and Delphi consensus process [25], these lack the detail required to differentiate between prioritisation and consensus. Understanding the relationship between these two concepts is important in finalising a minimum core set in a timely manner. Biggane’s recommendation [23] of considering the most appropriate medium(s) to communicate the core set study is pertinent. Whilst written instructions have historically been the norm and are undoubtable the simplest and quickest form of communication, they are perhaps not the most engaging and may even be ambiguous to readers. Use of more visual resources [23] such as demonstration videos [26] may help in engaging and educating participants. A novel ‘live’ approach in COS development was recently published [27]; researchers ran a Delphi ‘hackathon’ where all participants simultaneously completed the online Delphi surveys [27]. This live event involved an introductory session explaining the methodology and purpose of the study and provided breakout rooms for any questions regarding any aspect of the process [27]. Subsequent rounds emphasised that the goal of the Delphi survey was to reach consensus and participants were able to ask questions via a chat function [27]. The ability for participants to access timely guidance and support is undoubtedly valuable, and something which is not possible with current Delphi survey methods. It is vitally important that those taking part in a Delphi survey understand the significance of the role of feedback since utilising feedback to reconsider opinions and responses in order to gain consensus is fundamental to this process. If participants are not using the provided feedback effectively, then the process is not fit for purpose. Further qualitative work to explore how to clarify and communicate the concepts of prioritisation and consensus to all stakeholder groups in Delphi instructions, whether written, verbal or visual, is needed.

In addition to improving communication, ways by which to expedite prioritisation could be explored. The James Lind Alliance (JLA) specialises in prioritising research questions to be answered in specific healthcare fields [28,29]. The methodology involves asking relevant stakeholders what are the most important questions to research, collating this information and then asking stakeholders to rank or choose their top 10 research questions to be taken to a workshop for further discussion30 – a process very similar to core set development using Delphi techniques. As with all methodologies, there are advantages and disadvantages (as described in their guidebook [30]), but ultimately the outcome is a prioritised set of data items.

A strength of this study was that participants were not aware of their allocation to receive different instructions in the second survey round and responses were therefore free from performance bias. This study used a randomised design, in which participants who were assigned an identity number in the first survey round, were randomised by the REDCap software [11,12], to receive standard or enhanced instructions in the second survey round. Whilst randomisation ensures balanced groups, a potential limitation of using REDCap software to perform randomisation in this study was that all identity numbers were assigned to standard or enhanced feedback, regardless of whether participants answered questions or opened the survey numerous times. By including participants with multiple identity numbers and non-responders, the randomisation of participants to the two groups was not equal. This could have resulted in selection bias, however, in this study, baseline characteristics were similar between both groups and allocation did not affect response rates. Though desirable, not all nested COS methodology studies are randomised [21,31]. In addition, because this nested study was opportunistic in nature, it was not powered to be able to detect meaningful differences between the randomised groups and as such, results should be interpreted with caution and need further validation. A larger sample size would have increased the confidence of detecting a true effect of the different instruction types between the two randomisation groups and would have allowed more in-depth statistical analyses of the effects on individual stakeholder groups. When working within study limitations, efforts to engage and retain participants should be prioritised to increase the sample size and improve the strength and validity of the study. Response rates in round 2 were below half, lower than our previous core sets [13,17,32,33]. A participant was considered to have ‘responded’ if at least one core data item question had been answered. In round 1, this would likely have resulted in the inclusion of individuals curious about the survey content but then choosing not to proceed beyond the first few questions or first section (almost a quarter of participants did not continue to section 2). Ultimately, this would have led to inflated response rates in round 1 and exaggerated attrition rates in round 2. Comprehensive datasets strengthen the validity of study findings since missing data can introduce bias depending on the reason for omission. Omission, and therefore attrition, can be attributed to a number of factors including the length of the survey and the time elapsed between the first and final round [4], but in this instance could also be attributed to the inclusion of non-invested individuals. Ultimately, there is no guidance as to the number of completed questions required for inclusion as a ‘respondent’ or the length of the survey or the duration of the core set development process. But further work to explore completion inclusion thresholds and methods to promote completion and thereby reduce attrition would all be beneficial in promoting a comprehensive dataset. Finally, a key reflection in the writing up of this study is the recognition of the role that qualitative work would have played in progressing our understanding of how feedback was interpreted and applied.

Conclusion

Using a nested study approach within the development of a registry core data set to explore core set methodology, this study suggests that enhanced instructions emphasising the top items prioritised do not promote the prioritisation of items in a Delphi survey during core set development. Due to the complexities of the core set development process, it is likely that there are multiple factors at play. The absence of prioritisation in both groups perhaps suggests a lack of, or limited, understanding of the pathway to gaining consensus. Further work to explore how participants interpret and respond to feedback and instructions to simultaneously prioritise items and reach consensus with others is warranted. Educating and communicating these crucial elements will improve the efficiency and the value of the core set.

Supporting information

S1 Table. Proportion of stakeholders scoring items as ‘critically important’ (score 7–9).

https://doi.org/10.1371/journal.pone.0348136.s001

(DOCX)

References

  1. 1. Clarke M. Standardising outcomes for clinical trials and systematic reviews. Trials. 2007;8:39. pmid:18039365
  2. 2. Kearney A, Gargon E, Mitchell JW, Callaghan S, Yameen F, Williamson PR, et al. A systematic review of studies reporting the development of core outcome sets for use in routine care. J Clin Epidemiol. 2023;158:34–43. pmid:36948407
  3. 3. Kirkham JJ, Gorst S, Altman DG, Blazeby JM, Clarke M, Tunis S, et al. Core Outcome Set-STAndardised Protocol Items: the COS-STAP Statement. Trials. 2019;20(1):116. pmid:30744706
  4. 4. Williamson PR, Altman DG, Bagley H, Barnes KL, Blazeby JM, Brookes ST. The COMET Handbook: version 1.0. Trials. 2017;18(3):280.
  5. 5. Coulman KD, Howes N, Hopkins J, Whale K, Chalmers K, Brookes S, et al. A Comparison of Health Professionals’ and Patients’ Views of the Importance of Outcomes of Bariatric Surgery. Obes Surg. 2016;26(11):2738–46. pmid:27138600
  6. 6. Brookes ST, Macefield RC, Williamson PR, McNair AG, Potter S, Blencowe NS, et al. Three nested randomized controlled trials of peer-only or multiple stakeholder group feedback within Delphi surveys during core outcome and information set development. Trials. 2016;17(1):409. pmid:27534622
  7. 7. COMET Initiative. https://comet-initiative.org/Studies/SearchResults. Accessed 2024 January 30.
  8. 8. Del Grande C, Kaczorowski J. Rating versus ranking in a Delphi survey: a randomized controlled trial. Trials. 2023;24(1):543. pmid:37596699
  9. 9. Coulman KD, Chalmers K, Blazeby J, Dixon J, Kow L, Liem R, et al. Development of a Bariatric Surgery Core Data Set for an International Registry. Obes Surg. 2023;33(5):1463–75. pmid:36959437
  10. 10. Akpinar EO, Marang-van de Mheen PJ, Nienhuijs SW, Greve JWM, Liem RSL. National Bariatric Surgery Registries: an International Comparison. Obes Surg. 2021;31(7):3031–9. pmid:33786743
  11. 11. Harris PA, Taylor R, Minor BL, Elliott V, Fernandez M, O’Neal L, et al. The REDCap consortium: Building an international community of software platform partners. J Biomed Inform. 2019;95:103208. pmid:31078660
  12. 12. Harris PA, Taylor R, Thielke R, Payne J, Gonzalez N, Conde JG. Research electronic data capture (REDCap)--a metadata-driven methodology and workflow process for providing translational research informatics support. J Biomed Inform. 2009;42(2):377–81. pmid:18929686
  13. 13. Avery KNL, Chalmers KA, Brookes ST, Blencowe NS, Coulman K, Whale K, et al. Development of a Core Outcome Set for Clinical Effectiveness Trials in Esophageal Cancer Resection Surgery. Ann Surg. 2018;267(4):700–10. pmid:28288055
  14. 14. Avery KNL, Wilson N, Macefield R, McNair A, Hoffmann C, Blazeby JM, et al. A Core Outcome Set for Seamless, Standardized Evaluation of Innovative Surgical Procedures and Devices (COHESIVE): A Patient and Professional Stakeholder Consensus Study. Ann Surg. 2023;277(2):238–45. pmid:34102667
  15. 15. Kirkham JJ, Davis K, Altman DG, Blazeby JM, Clarke M, Tunis S, et al. Core Outcome Set-STAndards for Development: The COS-STAD recommendations. PLoS Med. 2017;14(11):e1002447. pmid:29145404
  16. 16. Kirkham JJ, Gorst S, Altman DG, Blazeby JM, Clarke M, Devane D, et al. Core Outcome Set-STAndards for Reporting: The COS-STAR Statement. PLoS Med. 2016;13(10):e1002148. pmid:27755541
  17. 17. Coulman KD, Hopkins J, Brookes ST, Chalmers K, Main B, Owen-Smith A, et al. A Core Outcome Set for the Benefits and Adverse Events of Bariatric and Metabolic Surgery: The BARIACT Project. PLoS Med. 2016;13(11):e1002187. pmid:27898680
  18. 18. Viera AJ, Garrett JM. Understanding interobserver agreement: the kappa statistic. Fam Med. 2005;37(5):360–3. pmid:15883903
  19. 19. StataCorp. Stata Statistical Software: Release 17. College Station, TX: StataCorp LLC. 2021.
  20. 20. Brookes ST, Chalmers KA, Avery KNL, Coulman K, Blazeby JM, ROMIO study group. Impact of question order on prioritisation of outcomes in the development of a core outcome set: a randomised controlled trial. Trials. 2018;19(1):66. pmid:29370827
  21. 21. Lange T, Kopkow C, Lützner J, Günther K-P, Gravius S, Scharf H-P, et al. Comparison of different rating scales for the use in Delphi studies: different scales lead to different consensus and show different test-retest reliability. BMC Med Res Methodol. 2020;20(1):28. pmid:32041541
  22. 22. Remus A, Smith V, Wuytack F. Methodology in core outcome set (COS) development: the impact of patient interviews and using a 5-point versus a 9-point Delphi rating scale on core outcome selection in a COS development study. BMC Med Res Methodol. 2021;21(1):10. pmid:33413129
  23. 23. Biggane AM, Williamson PR, Ravaud P, Young B. Participating in core outcome set development via Delphi surveys: qualitative interviews provide pointers to inform guidance. BMJ Open. 2019;9(11):e032338. pmid:31727660
  24. 24. Turnbull AE, Dinglas VD, Friedman LA, Chessare CM, Sepúlveda KA, Bingham CO 3rd, et al. A survey of Delphi panelists after core outcome set development revealed positive feedback and methods to facilitate panel member participation. J Clin Epidemiol. 2018;102:99–106. pmid:29966731
  25. 25. COMET Initiative. https://comet-initiative.org/Resources/PlainLanguage. Accessed 2024 February 15.
  26. 26. Hall DA, Smith H, Heffernan E, Fackrell K, Core Outcome Measures in Tinnitus International Delphi (COMiT’ID) Research Steering Group. Recruiting and retaining participants in e-Delphi surveys for core outcome set development: Evaluating the COMiT’ID study. PLoS One. 2018;13(7):e0201378. pmid:30059560
  27. 27. Lang KM, de Waal E, Bereczky T, Harrison K, Geissler J, Baolanos N. Delphi hackathon – a new approach to develop Core Outcome Sets for blood cancers. In: HARMONY Alliance. 2023;1–8.
  28. 28. James Lind Alliance. https://www.jla.nihr.ac.uk/. Accessed 2025 May 22.
  29. 29. Richards HS, Staruch RMT, Kinsella S, Savovic J, Qureshi R, Elliott D, et al. Top ten research priorities in global burns care: findings from the james lind alliance global burns research priority setting partnership. Lancet Glob Health. 2025;13(6):e1140–50. pmid:40286806
  30. 30. JLA Guidebook. https://www.jla.nihr.ac.uk/jla-guidebook 2025.
  31. 31. Webbe J, Allin B, Knight M, Modi N, Gale C. How to reach agreement: the impact of different analytical approaches to Delphi process results in core outcomes set development. Trials. 2023;24(1):345. pmid:37217933
  32. 32. McNair AGK, Whistance RN, Forsythe RO, Macefield R, Rees J, Pullyblank AM, et al. Core outcomes for colorectal cancer surgery: a consensus study. PLoS Med. 2016;13(8):e1002071. pmid:27505051
  33. 33. Potter S, Holcombe C, Ward JA, Blazeby JM, BRAVO Steering Group. Development of a core outcome set for research and audit studies in reconstructive breast surgery. Br J Surg. 2015;102(11):1360–71. pmid:26179938