Skip to main content
Advertisement
Browse Subject Areas
?

Click through the PLOS taxonomy to find articles in your field.

For more information about PLOS Subject Areas, click here.

  • Loading metrics

Identifying academic success and underperformance: The discriminative power of very short answer questions and multiple-choice questions

  • Elise V. van Wijk,

    Roles Conceptualization, Data curation, Formal analysis, Investigation, Methodology, Project administration, Validation, Writing – original draft, Writing – review & editing

    Affiliation Center for Innovation in Medical Education, Leiden University Medical Center, Leiden, The Netherlands

  • Floris M. van Blankenstein,

    Roles Conceptualization, Investigation, Methodology, Project administration, Supervision, Validation, Writing – review & editing

    Affiliation Center for Innovation in Medical Education, Leiden University Medical Center, Leiden, The Netherlands

  • Bastian N. Ruijter,

    Roles Data curation, Writing – review & editing

    Affiliation Department of Gastroenterology and Hepatology, Leiden University Medical Center, Leiden, The Netherlands

  • Jos H. T. Rohling,

    Roles Data curation, Writing – review & editing

    Affiliation Department of Cell & Chemical Biology, Leiden University Medical Center, Leiden, The Netherlands

  • Jolein van der Kraan,

    Roles Data curation, Writing – review & editing

    Affiliation Department of Gastroenterology and Hepatology, Leiden University Medical Center, Leiden, The Netherlands

  • Friedo W. Dekker,

    Roles Methodology, Validation, Writing – review & editing

    Affiliations Center for Innovation in Medical Education, Leiden University Medical Center, Leiden, The Netherlands, Department of Clinical Epidemiology, Leiden University Medical Center, Leiden, The Netherlands

  • Alexandra M. J. Langers

    Roles Conceptualization, Investigation, Methodology, Project administration, Supervision, Validation, Writing – review & editing

    a.m.j.langers@lumc.nl

    Affiliations Center for Innovation in Medical Education, Leiden University Medical Center, Leiden, The Netherlands, Department of Gastroenterology and Hepatology, Leiden University Medical Center, Leiden, The Netherlands

Abstract

Background

Multiple-choice questions (MCQs) are widely used in medical education, but are criticized for cueing and guessing. Very short answer questions (VSAQs), which require students to generate responses independently, may better assess knowledge. While VSAQs demonstrate higher item discrimination within individual exams, their effectiveness in distinguishing academic performance across multiple assessments remains unclear – representing a key gap in the validation of VSAQs under Messick’s framework, specifically the category of “relations to other variables”. This study examines whether VSAQs or MCQs more effectively distinguish students of varying performance levels across multiple summative examinations.

Methods

We analyzed retrospective data from six mixed-format examinations with VSAQs and MCQs of three cohorts of first- and second-year medical students. Academic performance was measured using grade point average (GPA) across assessments. Linear regression assessed the relationship of each question format with GPA, while ROC curves and C-statistics evaluated their ability to identify poor and excellent performing students (lowest and highest quintile of GPA).

Results

VSAQs showed higher item discrimination (Rir-values) than MCQs in all exams. VSAQs also had a stronger positive association with GPA compared to MCQs, and higher C-statistics, indicating superior discriminative ability.

Conclusion

VSAQs outperform MCQs in distinguishing academic performance levels across multiple assessments. Their integration into examinations enhances discriminative ability and may facilitate earlier identification of poor and excellent performing students, enabling targeted interventions and support of students.

Introduction

Assessment of undergraduate medical students predominantly relies on multiple-choice questions (MCQs), especially in the single best answer (SBA) format, due to their high reliability, rapid answering time, and the efficiency of machine marking [1,2]. However, this format also allows for guessing and cueing (i.e., answering questions by relying on cues from the question stem or answer options rather than on actual content knowledge) and encourages a recognition-based study approach, which may introduce noise into MCQ scores and diminish their ability to accurately reflect students’ true understanding [3]. In contrast, the very short answer question (VSAQ), an open-ended question requiring a concise answer, mitigates these issues by eliminating both cueing and guessing [4]. Consequently, VSAQs tend to exhibit higher item discrimination within individual examinations compared to MCQs [3,57], meaning that they can better differentiate between students with varying levels of performance on a given test. Higher discriminative ability indicates that an item more effectively distinguishes students based on their underlying level of knowledge or competence. Although students need more time to answer VSAQs compared to MCQs, which may reduce the number of questions per exam when testing time is fixed [3], the higher item discrimination of VSAQs implies that fewer questions may be needed per content domain to reliably differentiate between levels of student performance.

Messick conceptualizes validity as a unified construct supported by multiple interrelated resources of evidence, including content, response processes, internal structure, relations to other variables, and consequences of testing. Internal structure refers to the extent to which relationships among test items reflect the intended construct, whereas relations to other variables (external validity) concerns whether test scores relate as expected to other relevant measures, such as academic performance [8]. Establishing both internal and external validity is essential to justify the use of assessment formats in high-stakes educational contexts.

Previous research on VSAQs has primarily focused on internal structure, demonstrating consistently superior item discrimination compared to multiple-choice questions (MCQs) within individual examinations [3,57]. However, higher item discrimination within a single exam does not necessarily translate into a stronger ability to identify students as poor or excellent performers across multiple examinations. This distinction is illustrated by the findings of Eijsvogels et al. [9], who showed that extended matching questions (EMQs)– despite superior item discrimination compared to MCQs within examinations [10] – were less effective at identifying poor performing students across multiple assessments. Although EMQs (a form of multiple-choice question in which a single list of answer options is matched to several related clinical scenarios or questions) and VSAQs differ in format, this example underscores that internal structure alone is insufficient to support broader validity claims. Similarly, while VSAQs demonstrate superior item discrimination within individual examinations compared to MCQs [3,57], it remains uncertain whether performance on VSAQs meaningfully relates to broader academic outcomes, which is central to establishing their external validity. Addressing this uncertainty represents a key gap in the validation of VSAQs within Messick’s unified framework.

Beyond evaluating the item discrimination of individual assessment questions, it is therefore important to investigate whether certain question formats are better suited for identifying poor and excellent performing students across multiple examinations. This broader perspective can offer insights into the consistency and robustness of these formats in distinguishing students across different performance levels, and formats with higher discriminative power may facilitate the early identification of underperforming students, thereby enabling timely interventions. In this study, we aim to 1) examine the relationship between question format (VSAQs versus MCQs) and academic performance; 2) evaluate the ability of VSAQs and MCQs to identify poor and excellent performing students. To address these aims, we first assess the item discrimination of both question formats within each examination, thereby verifying the assumption that VSAQs have superior discriminative ability within examinations [3,57]. We use the VSAQ- and MCQ-scores from two summative mixed-format examinations administered during the first and second year of an undergraduate medical curriculum. Our analysis includes two student populations: 1) all students who participated in the first-year examination, including those who may later leave the program, and 2) nominal students who participated in both the first- and second-year examinations. The first population offers a broader performance range, particularly among lower performing students, while the second population offers more datapoints (i.e., questions) per student, enhancing the reliability of the analysis.

Methods

Setting

This retrospective cohort study was conducted in Leiden University Medical Center (LUMC), the Netherlands. The Dutch medical curriculum comprises a three-year bachelor’s program followed by a three-year master’s program. Although the bachelor courses are primarily assessed with MCQs, other formats such as extended matching, comprehensive integrated puzzle, open essay, and VSAQs are also included in the written assessments. Additionally, students participate in longitudinal training on various CanMEDS competencies [11] beyond the role of Medical Expert, such as communication skills, leadership, health promotion, and collaboration. To assess the discriminative ability of VSAQs and MCQs we analyzed summative examinations of two medical bachelor’s courses: “Regulation and Metabolism” (RM), a first-year fundamental course, and “Diseases of the Abdomen” (DA), a second-year clinical course. These courses (6 and 7 weeks, respectively) address metabolic and gastrointestinal topics. These courses were chosen because their summative assessment contained both VSAQs and MCQs within the same examination. In our prior study [5], we compared the psychometric properties of MCQs and VSAQs in formative assessments for both courses. Subsequently, the course coordinators added VSAQs to the previously MCQ-only summative exams, resulting in mixed-format examinations. This mixed format has now been used in both courses for three consecutive years (student cohorts 2020–2021; 2021–2022; 2022–2023), resulting in six summative mixed-format examinations which we selected for our analyses. This mixed-format design enabled direct within-examination comparisons between question formats under identical educational and assessment conditions. Teaching and exam preparation follows the curriculum program for all cohorts. Students receive scheduled instructional sessions, guided self-study materials, and practice assessments aligned with the exam formats. All students have equivalent access to these learning resources prior to the examination.

All assessments were administered digitally through RemindoToets (Paragin) system [12] and yielded study credits. All teachers involved in the question writing process were required to participate in a course about writing exam questions, where guidelines for writing high quality exam questions were addressed, based on existing literature [13]. Furthermore, all exam questions are reviewed by an Examination Review Board as part of a routine quality assurance system for assessments. VSAQs were marked accordingly to the standard institutional procedure. Predefined correct answers were automatically scored; all other responses were reviewed manually by the examiner. Correct but unanticipated answers (e.g., minor spelling variations) were added to the accepted answer list. Scoring was dichotomous (0/1), with a second examiner consulted when uncertainty arose.

Participants

The first population, RM participants, included all first-year bachelor medical students from the 2020–2021, 2021–2022, and 2022–2023 cohorts who completed the first sitting of the summative RM assessment (n = 902). The second population, RM & DA participants, comprised all second-year bachelor medical students from the same cohorts who completed the first sitting of both the RM and DA summative assessments in consecutive years (n = 797). Students who attempted less than 75% of all exams taken into consideration for the GPA calculation were excluded from both populations (n = 51).

Study design and data collection

We analyzed data from two summative mixed-format examinations from the RM and DA courses, administered between 2021 and 2024. For each assessment, we extracted individual questions, question formats, question scores (coded as 1 for correct and 0 for incorrect), residual item reliability-values (Rir) for each question, and student IDs from the Remindo assessment system. The Rir-value, representing the correlation between one test item and all other items of the test, measures item discrimination by indicating how well each question correlates with overall student performance on that test (after excluding that specific question) [14]. To assess the comparability of question formats within the examinations, we examined the distribution of each question format across the different themes. The data were accessed between 10/11/2024 and 17/12/2024 for our research. The data were anonymised using ID numbers, ensuring that the authors had no access to information that could identify individual participants.

We measured academic performance with the grade point average (GPA), a widely recognized and standardized indicator of academic success [15,16]. In our context, GPA is defined as a credit-weighted average of all numeric grades from the first and second year of study. For each student, GPA was calculated based on all courses that included a written examination. In addition to written exams, several courses also comprised other forms of examinations (e.g., oral exams, presentations, or practical assessments) which were scored by a pass or fail, thereby providing a broader representation of academic performance. Courses were weighted according to their European Credit Transfer and Accumulation System (ECTS) credits. An overview of all courses included in the GPA calculation and their corresponding ECTS weights is provided in S1 table. The RM and DA course awarded 7 and 8 ECTS, respectively. The GPA was computed by multiplying each course’s final grade by its corresponding ECTS credits and aggregating these values. All exam grades from the first and second-year courses were retrieved from the university’s administrative system. In our system (used in several European countries), grades range from 1 to 10, with 10 being the highest possible score. A grade of 5.5 or higher is considered a passing grade and is required to earn study credits. Failing grades (defined as <5.5) were included in the GPA calculation to provide a realistic representation of overall academic performance. We evaluated GPA based on two approaches: one using the most recent grade from the last sitting (including retake exams) and another using only the grade from the first sitting. For the first population (RM participants) we calculated the GPA based on first-year courses, while for the second population (RM & DA participants) the GPA included courses from both the first- and second-year. Based on these GPAs, students were categorized by quintiles into “poor performing” (first quintile), “average performing” (second through fourth quintile), or “excellent performing” (fifth quintile) students.

Data analysis

Descriptive statistics were calculated and presented for each examination, including total average scores, average VSAQ and MCQ scores, and the distribution of VSAQs and MCQs across the themes. To compare the different question formats, we calculated separate z-scores for VSAQs and MCQs based on their respective absolute scores. For the RM participants, z-scores were derived from their VSAQ and MCQ scores on the RM assessment. For the RM & DA participants, z-scores were calculated separately for VSAQs and MCQs using absolute scores from both the RM and DA assessments, ensuring that each student had one VSAQ z-score and one MCQ z-score reflecting their performance across both exams. Item discrimination was determined using the mean of the Rir-values for each question format [13]. The Rir-value for each question was calculated in relation to the entire exam.

Linear regression analysis was performed with GPA as linear outcome variable and the z-scores of the VSAQ and MCQ as covariates, separately for each cohort and across all cohorts. First, univariate regression analyses were conducted to assess the variance explained by each question format individually (reported as adjusted R2). Next, both formats were included in a multivariate model to account for shared variance and determine their relative predictive value. Student GPAs from the courses of the first year (RM participants) and the first two year (RM & DA participants) were used as outcome. Similarly, the same student GPAs were used to create new categorical variables: “poor performing” versus “other” (average and excellent) and “excellent performing” versus “other” (average and poor). Receiver Operating Characteristic (ROC) curves were used to evaluate the ability of MCQs and VSAQs to distinguish between poor and non-poor performing students, and also between excellent and non-excellent performing students. The Concordance-statistic (C-statistic or area under the ROC curve) was calculated as a measure of discriminatory power, reflecting how well these formats differentiate between performance levels. Significant differences between the C-statics of VSAQs and MCQs were assessed using paired bootstrapping with 95% confidence intervals. The difference was considered significant when the confidence interval (CI) of the score difference did not contain the zero. Bootstrapping was used to compare the paired C-statistics because it provides robust confidence intervals without relying on strict distributional assumptions. All statistical analyses were performed using R version 4.1.0 (R Foundation for Statistical Computing, Vienna, Austria).

Ethical approval

This study received ethical approval from the LUMC Educational Research Review Board (OEC/ERRB/20241008/1). Our ethical review board decided that informed consent was not required, as this study did not involve an intervention or experimental manipulation. Rather, VSAQs were introduced as part of an incremental improvement of the existing assessment practice, initially on a small scale to gain experience with this question format. The data analyzed were retrospective and already available from the assessment system as part of the students’ regular curriculum. We collected no additional demographic or personal data, and all data were aggregated and anonymized to ensure that individual students could not be identified.

Results

Descriptives and item discrimination

The students included in our study for each mixed-format examination, along with the average total, VSAQ, and MCQ scores are presented in Table 1. In all mixed-format examinations, the average Rir-values of the VSAQs were consistently higher compared to the MCQs (Table 1). The distribution of question formats across the themes was generally even, with no themes exclusively assessed by a single format (S2 Table).

thumbnail
Table 1. The scores and Rir-values of VSAQs and MCQs within the mixed-format examinations.

https://doi.org/10.1371/journal.pone.0349318.t001

Relationship between question format and academic performance

We conducted a linear regression analysis to examine the effects of VSAQ and MCQ z-scores on GPA. First, we assessed the variance explained by each question format separately using univariate models (R2 values). Among all RM participants, the R² values ranged from 0.61 to 0.67 for VSAQ z-score and from 0.44 to 0.58 for MCQ z-score. In all RM & DA participants, the R² values ranged from 0.68 to 0.71 for VSAQ z-score and from 0.59 to 0.67 for MCQ z-score. The R2 values represent the proportion of variance in GPA that can be explained by the VSAQ or MCQ z-score.

In the multiple regression analyses standardized beta coefficients (β) were used to compare the relative strength of associations between the VSAQ and MCQ z-scores and GPA; larger absolute β values indicate stronger relationships. The multiple regression analyses showed a significant positive association between both VSAQ and MCQ z-scores and GPA (Table 2). However, for all RM participants (n = 902), the VSAQ z-score had a stronger association with GPA when using first sitting grades (β = 0.63, t(903) = 21.28, p < .001, 95% CI [0.57, 0.69]) than the MCQ z-score (β = 0.36, t(903) = 12.17, p < .001, 95% CI [0.30, 0.42]), with the model explaining 68% of the variance in GPA (F(2, 899) = 945.79, p < .001, adjusted R2 = .68). Similarly, in all RM & DA participants (n = 797), the VSAQ z-score had a stronger positive association (β = 0.53, t(794) = 21.78, p < .001, 95% CI [0.48, 0.57]) compared to the MCQ z-score (β = 0.36, t(794) = 14.84, p < .001, 95% CI [0.31, 0.41]), with the model accounting for 77% of the variance (F(2, 794) = 1303.59, p < .001, adjusted R2 = .77). These results were similar when using last sitting grades and for the analyses of the separate cohorts (Table 2).

thumbnail
Table 2. Linear regression model parameters with grade point average as outcome based on the first and last sitting grades.

https://doi.org/10.1371/journal.pone.0349318.t002

Identification of poor and excellent performing students

We calculated the C-statistic to assess the ability of VSAQ and MCQ scores to identify poor performing (lowest GPA quintile) and excellent performing (highest GPA quintile) students. The C-statistic reflects the model’s ability to distinguish between individuals with and without the outcome of interest. Values closer to 1.0 indicate better discrimination, whereas a value of 0.5 indicates performance no better than chance. In all RM participants (n = 902), VSAQ z-scores showed higher C-statistics (i.e., greater discriminative ability) for both poor and excellent performing students compared to MCQ z-scores (Fig 1A). However, for poor performing students using first sitting grades, the difference between the C-statistics of VSAQ and MCQ was not significant. In all RM & DA participants (n = 797), the C-statistic of the VSAQ z-score remained significantly higher than that of the MCQ z-score for poor performing students using first sitting grades. For excellent performing students, there was no significant difference in discriminative ability (Fig 1B). The findings across separate cohort analyses followed the same pattern; either the C-statistic for the VSAQ z-score was significantly higher than that of the MCQ z-score, or there was no significant difference between them (Figs S1 and S2).

thumbnail
Fig 1. Receiver Operating Characteristic (ROC) curves for both poor and excellent performing students using the first sitting or last sitting grades from A) the total RM participants, and B) the total RM & DA participants.

Red = VSAQs; Blue = MCQs; CI = 95% confidence interval; TPR = true positive rate; FPR = false positive rate. NS = non-significant. C-statistics are shown above each graph, which were calculated from the area under the ROC curve.

https://doi.org/10.1371/journal.pone.0349318.g001

Discussion

In this study, we examined the relationship between question format (VSAQs versus MCQs) and academic performance based on GPA, and evaluated the ability of both formats to distinguish poor-performing students from all other students, and excellent-performing students from all other students. Our findings reveal that while both question formats are suitable for distinguishing students across performance levels, VSAQs consistently demonstrate superior effectiveness compared to MCQs in both the linear regression analysis and ROC curves. Notably, the difference in discriminative ability between VSAQs and MCQs was more pronounced among first-year students (RM participants) compared to the combined group of first- and second-year students (RM & DA participants). A possible explanation for this observation is the exclusion of mostly poor performing students who did not nominally progress to the second year, resulting in a more homogeneous group of students with less variability in academic performance. Consistent with previous research [3,57], VSAQs also demonstrated higher item discrimination than MCQs within the individual examinations. Our findings complement those of Eijsvogels et al. [9] by examining a different question format – VSAQs – that shares certain psychometric features with EMQs, but may function differently in practice. The observed differences in outcomes likely reflect format-specific characteristics. Although EMQs reduce reliance on guessing compared to traditional MCQs, they still provide a list of plausible options that could lead to guessing or cueing. In contrast, VSAQs require students to generate concise responses without external cues, inherently minimizing the likelihood of guessing and offering a more direct measure of their knowledge. Furthermore, Eijsvogels et al. [9] penalized incorrect answers on MCQs, which may have enhanced the discriminative performance of this question format. By normalizing performance scores within mixed-format examinations, we ensured a fair comparison, eliminating potential biases from variations in exam quality or instructional approach. Additionally, our analysis included multiple examinations and a larger sample size, enhancing the robustness of our findings.

The GPA used in this study to assess academic performance is primarily based on assessments that predominantly consist of MCQs, which could bias the results toward MCQs due to their alignment with the format used to calculate GPA. However, despite this potential bias, our findings indicate a clear advantage for VSAQs over MCQs in distinguishing poor and excellent performing students. The superior performance of VSAQs can be attributed to their format, which requires students to generate responses independently, eliminating opportunities for guessing and relying instead on their actual understanding [1720]. This characteristic makes VSAQs a more accurate reflection of student’s knowledge level. While guessing may inflate MCQ scores for poor performing students in a single examination, this effect will diminish when performance is averaged across multiple exams, thereby reducing noise introduced by guessing. Consequently, the improved discriminative performance of VSAQs may have practical educational value by enabling more accurate identification of students at different levels of academic achievement.

Strengths and limitations

This study is the first to evaluate the discriminative ability of VSAQs across multiple assessments using GPA as a measure of academic performance (evidence of relations to other variables), rather than focusing exclusively on item discrimination with individual exams (evidence of internal structure). By utilizing real-world summative assessment data from three distinct student cohorts, we minimized the biases associated with voluntary participation and ensured a robust, and representative sample size. This also allowed us to aggregate results and average out variance across examinations and populations, thereby enhancing the reliability of our findings. The inclusion of both first-year students with greater variation in academic performance, and nominal second-year students allowed for a more comprehensive analysis of performance across different populations, while mitigating selection bias. Moreover, mixed-format examinations enabled within-subjects comparisons of VSAQs and MCQs, under comparable instructional and assessment conditions. Standardizing scores further reduced variability and ensured reliable comparisons. However, this study also has limitations. While GPA is an objective and widely used measure of academic achievement, it simplifies complex learning outcomes and may overlook critical thinking and skill development [15]. Additionally, the variability in content and distribution between VSAQs and MCQs within real-world examinations could introduce bias. However, we mitigated this by aggregating data from multiple examinations and cohorts, standardizing scores, and conducting analyses across two student populations. Lastly, this study was conducted within a single medical school, which may limit the generalizability of the findings to other institutions and curricula. Although the inclusion of multiple cohorts and examinations strengthens the robustness of our findings, replication across different educational contexts is needed to determine whether the observed advantages of VSAQs are consistent across settings with varying practices and student populations.

Implications for practice and future research

Our findings indicate that VSAQs have a higher discriminative ability than MCQs, effectively distinguishing students both within individual examinations and across multiple examinations based on GPA. These findings suggest that VSAQs may represent a more robust and valid question format for evaluating academic performance and warrant further consideration within medical education assessment. Incorporation of VSAQs into assessments may enhance the ability to identify students who could benefit from early interventions and support, while also recognizing those who excel. However, given that this study was conducted within a single institution, multicenter studies and prospective cohort research are needed to confirm the generalizability of these findings and to evaluate the practical implications of implementing VSAQs across diverse educational settings before widespread adoption can be recommended.

Although the observed differences in discriminative ability between VSAQs and MCQs were modest, they were consistently observed across examinations and student cohorts. In educational practice, even relatively small improvements in discrimination may be meaningful, particularly in assessments that inform decisions regarding student progression or academic support. Improved discrimination may enhance the ability to distinguish between students with different levels of knowledge and understanding, thereby supporting more accurate assessment decisions. The potential educational benefits should therefore be considered alongside practical implementation requirements.

The practical implementation of VSAQs also warrants consideration. Adoption of VSAQs may require investments in examiner training, and adaptation of assessment procedures. However, training can be integrated into existing faculty development and item-writing workshops, allowing educators to develop VSAQs while simultaneously preparing examination content. Furthermore, previous research has shown that grading is feasible and efficient in large cohorts [5,6]. Therefore, although implementation requires investment of time and institutional support, the operational burden may be manageable and should be weighed against the potential gains in assessment validity and discrimination. Future studies could also examine grading efficiency, examiner experiences, and inter-rater reliability across different educational settings to further inform large-scale implementation of VSAQs.

Moreover, the stronger relationship between VSAQ performance and GPA suggests that VSAQs may warrant further investigation as a potential predictor of academic performance. Future prospective longitudinal studies are needed to determine whether performance on VSAQs is associated with subsequent academic outcomes and whether such assessments could have a role in selection procedures. Future studies could also expand academic performance measures beyond GPA, incorporating assessments of skill development and work-place learning. These additional metrics would provide a more holistic evaluation of the effectiveness of VSAQs and MCQs across different educational contexts. Additionally, exploring whether VSAQs are associated with long-term academic and professional success, including performance in real-world clinical settings, would provide deeper insights into their utility.

Conclusion

This study suggests that VSAQs may be more effective than MCQs in distinguishing undergraduate medical students across different academic performance levels. Our findings indicate that VSAQ scores are associated with broader academic performance beyond the exam itself, contributing to the “relations to other variables” dimension of Messick’s unified validity framework. Integrating VSAQs into assessments may enhance the discriminative power and robustness of assessments and could support early identification of both poor and excellent performing students, potentially enabling more targeted educational support. Future research is needed to explore the generalizability of these findings across diverse educational settings and to evaluate the extent to which VSAQs may predict long-term academic achievements.

Supporting information

S1 Fig. Receiver Operating Characteristic (ROC) curves for both poor and excellent performing students using the first sitting or last sitting grades from the RM participants cohort 2020–2021 (RM 2021), cohort 2021–2022 (RM 2022), and cohort 2022–2023 (RM 2023).

https://doi.org/10.1371/journal.pone.0349318.s001

(TIFF)

S2 Fig. Receiver Operating Characteristic (ROC) curves for both poor and excellent performing students using the first sitting or last sitting grades from the RM & DA participants cohort 2020–2021 (RM 2021 & DA 2022), cohort 2021–2022 (RM 2022 & DA 2023), and cohort 2022–2023 (RM 2023 & DA 2024).

https://doi.org/10.1371/journal.pone.0349318.s002

(TIFF)

S1 Table. Courses included in the calculation of the GPA and their corresponding ECTS credits.

https://doi.org/10.1371/journal.pone.0349318.s003

(PDF)

S2 Table. Distribution of VSAQs and MCQs across the themes within the examinations.

https://doi.org/10.1371/journal.pone.0349318.s004

(PDF)

Acknowledgments

The authors wish to thank all study coordinators and teachers for their critical appraisal of the research protocol and valuable contribution to the manuscript.

References

  1. 1. Al-Rukban MO. Guidelines for the construction of multiple choice questions tests. J Family Community Med. 2006;13(3):125–33. pmid:23012132
  2. 2. Schuwirth LWT, van der Vleuten CPM. Written assessment. ABC of learning and teaching in medicine. Oxford: Wiley-Blackwell. 2017. p. 65–9.
  3. 3. Mee J, Pandian R, Wolczynski J, Morales A, Paniagua M, Harik P, et al. An experimental comparison of multiple-choice and short-answer questions on a high-stakes test for medical students. Adv Health Sci Educ Theory Pract. 2024;29(3):783–801. pmid:37665413
  4. 4. Sam AH, Westacott R, Gurnell M, Wilson R, Meeran K, Brown C. Comparing single-best-answer and very-short-answer questions for the assessment of applied medical knowledge in 20 UK medical schools: Cross-sectional study. BMJ Open. 2019;9(9):e032550. pmid:31558462
  5. 5. van Wijk EV, Janse RJ, Ruijter BN, Rohling JHT, van der Kraan J, Crobach S, et al. Use of very short answer questions compared to multiple choice questions in undergraduate medical students: An external validation study. PLoS One. 2023;18(7):e0288558. pmid:37450485
  6. 6. Sam AH, Field SM, Collares CF, van der Vleuten CPM, Wass VJ, Melville C, et al. Very-short-answer questions: reliability, discrimination and acceptability. Med Educ. 2018;52(4):447–55. pmid:29388317
  7. 7. Sam AH, Peleva E, Fung CY, Cohen N, Benbow EW, Meeran K. Very Short Answer Questions: A Novel Approach To Summative Assessments In Pathology. Adv Med Educ Pract. 2019;10:943–8. pmid:31807109
  8. 8. Messick S. Validity of psychological assessment: validation of inferences from persons’ responses and performances as scientific inquiry into score meaning. Am Psychol. 1995;50:741–9.
  9. 9. Eijsvogels TMH, van den Brand TL, Hopman MTE. Multiple choice questions are superior to extended matching questions to identify medicine and biomedical sciences students who perform poorly. Perspect Med Educ. 2013;2(5–6):252–63. pmid:24203858
  10. 10. Fenderson BA, Damjanov I, Robeson MR, Veloski JJ, Rubin E. The virtues of extended matching and uncued tests as alternatives to multiple choice questions. Hum Pathol. 1997;28(5):526–32. pmid:9158699
  11. 11. Frank JR, Danoff D. The CanMEDS initiative: implementing an outcomes-based framework of physician competencies. Med Teach. 2007;29(7):642–7. pmid:18236250
  12. 12. Paragin. RemindoToets. https://www.paragin.nl/remindotoets/
  13. 13. Schuwirth L, Pearce J. Determining the quality of assessment items in collaborations: aspects to discuss to reach agreement developed by the Australian Medical Assessment Collaboration. https://research.acer.edu.au/higher_education/42 2014. 2026 March.
  14. 14. De Champlain AF. A primer on classical test theory and item response theory for assessments in medical education. Med Educ. 2010;44(1):109–17. pmid:20078762
  15. 15. Brookhart SM, Guskey TR, Bowers AJ, McMillan JH, Smith JK, Smith LF. A century of grading research. Rev Educ Res. 2016.
  16. 16. York TT, Gibson C, Rankin S. Defining and measuring academic success. Pract Assess Res Eval. 2015;20.
  17. 17. Sam AH, Hameed S, Harris J, Meeran K. Validity of very short answer versus single best answer questions for undergraduate assessment. BMC Med Educ. 2016;16(1):266. pmid:27737661
  18. 18. Schuwirth LW, van der Vleuten CP, Donkers HH. A closer look at cueing effects in multiple-choice questions. Med Educ. 1996;30(1):44–9. pmid:8736188
  19. 19. Desjardins I, Touchie C, Pugh D, Wood TJ, Humphrey-Murto S. The impact of cueing on written examinations of clinical decision making: a case study. Med Educ. 2014;48(3):255–61. pmid:24528460
  20. 20. Sam AH, Wilson R, Westacott R, Gurnell M, Melville C, Brown CA. Thinking differently - Students’ cognitive processes when answering two different formats of written question. Med Teach. 2021;43(11):1278–85. pmid:34126840