Skip to main content
Advertisement
Browse Subject Areas
?

Click through the PLOS taxonomy to find articles in your field.

For more information about PLOS Subject Areas, click here.

  • Loading metrics

Structured approaches to improving outcomes of oral exam in medical and paramedical education: A systematic review and evidence-based framework

  • Amir Torab-Miandoab,

    Roles Data curation, Formal analysis, Methodology, Software, Visualization, Writing – original draft, Writing – review & editing

    Affiliation Medical Education Research Center, Health Management and Safety Promotion Research Institute, Tabriz University of Medical Sciences, Tabriz, Iran

  • Mansour Ghafourifard,

    Roles Conceptualization, Methodology, Supervision, Writing – original draft, Writing – review & editing

    Affiliation Medical Education Research Center, Health Management and Safety Promotion Research Institute, Tabriz University of Medical Sciences, Tabriz, Iran

  • Sogand Habibi-Chenaran,

    Roles Data curation, Formal analysis, Validation, Writing – original draft

    Affiliation Department of Health Information Technology, School of Management and Medical Informatics, Tabriz University of Medical Sciences, Tabriz, Iran

  • Saeideh Ghaffarifar

    Roles Conceptualization, Formal analysis, Methodology, Project administration, Supervision, Validation, Writing – original draft

    sa.ghafarifar@gmail.com

    Affiliation Medical Education Research Center, Health Management and Safety Promotion Research Institute, Tabriz University of Medical Sciences, Tabriz, Iran

Abstract

Introduction

Oral examinations are central to medical and paramedical education for assessing reasoning, communication, and professional competencies, but remain challenged by issues of reliability, bias, and standardization. This systematic review synthesizes evidence on issues, requirements, and solutions to optimize oral exam outcomes in these fields.

Materials and methods

Following PRISMA guidelines, a comprehensive search was conducted across 14 databases (e.g., PubMed, Scopus, Web of Science) up to June 23, 2025, using PICO-framed queries focused on medical and paramedical students. Inclusion criteria encompassed English-language full text studies on oral assessments; exclusions included editorials and reviews. Three independent reviewers screened titles, abstracts, and full texts, extracting data on study characteristics, challenges, requirements, solutions, and outcomes. Quality appraisal utilized QUADAS, with thematic analysis and matrix patterns to identify correlations.

Results

From 25,594 records, 102 studies were included after deduplication and screening. The number of publications increased substantially after 2000, with the majority originating from the field of medicine (70.6%) and primarily addressing undergraduate (44.1%) and postgraduate (37.3%) education. Structured oral examinations (SOEs) dominated (56.9%), outperforming traditional formats in reliability (median Cronbach’s alpha 0.75; inter-rater interclass correlation coefficient (ICC) 0.47–0.82) and validity (modest correlations with objective structured clinical examinations (OSCEs), r=0.10–0.74). Key challenges included examiner variability (66.7%), student anxiety (41.2%), and logistical constraints (34.3%). The important requirements included assessment design, examiner calibration, technological integration, and equity. Solutions were emphasized on training (60.8%), standardization (56.9%), and technology (27.5%), yielding improved pass rates (50–100%) and satisfaction (72–96%). A comprehensive framework recommends multi-station objective structured viva examinations (OSVEs), hybrid delivery, combined rubrics, and artificial intelligence (AI)-assisted scoring for optimization. Limitations included generalizability (91.3%) and small samples (49.5%).

Conclusions

This review outlines key challenges and strategies to improve the quality of oral examinations. Standardized oral examinations, supported by technology, offer fair and reliable assessments. The findings provide a framework for educators and researchers to optimize their use, while future studies should validate predictive validity and explore AI-driven solutions to reduce bias and enhance scalability.

Introduction

An oral examination—also known as a viva or oral assessment—is a face-to-face evaluation method in which students verbally respond to examiners’ questions. It can be conducted one-on-one or in front of a panel and often involves discussion, case analysis, or problem-solving tasks [1]. The format allows examiners to probe the depth of understanding, clarify ambiguous responses, and assess how students organize and express their thoughts under pressure [2].

Oral exams are widely utilized across various disciplines and levels of education, particularly in high-stakes contexts [3]. In the medical and paramedical fields—including medicine, nursing, dentistry, physiotherapy, pharmacy, and public health—oral assessments play a critical role in evaluating students’ clinical decision-making, ethical reasoning, and professional judgment. They are also commonly used in interprofessional education settings, interdisciplinary case conferences, and competency-based assessments, reflecting their broad applicability and value in healthcare education [4].

The oral examination format offers several advantages over written and multiple-choice tests. It provides the opportunity for interactive assessment, immediate feedback, and tailored questioning based on student responses [5]. Oral exams closely mirror real-world clinical interactions, making them superior in assessing non-cognitive skills such as empathy, communication, professionalism, and the ability to think on one’s feet [6]. Its adaptability also allows examiners to explore complex reasoning processes, identify misconceptions, and evaluate critical thinking in ways that traditional written exams often fail to achieve [7].

Despite these strengths, oral exams have been persistently scrutinized for issues related to objectivity, reliability, inter-examiner variability, and susceptibility to bias. Concerns about consistency in questioning, examiner subjectivity, and limited standardization have raised questions about their fairness and validity, particularly when used in high-stakes settings [8]. Furthermore, the evolving complexity of medical education—with increased emphasis on competency-based assessment, digital innovation, and equity—has underscored the need for structured, evidence-informed improvements to oral examination methods [9].

Several educational frameworks and assessment models have been proposed to optimize oral exams, including structured oral examinations (SOEs), objective structured clinical examinations (OSCEs), and digital or video-recorded formats. However, the implementation of these strategies varies widely across institutions and regions, and the evidence regarding their effectiveness, feasibility, and educational impact remains scattered across the literature. A comprehensive understanding of the current challenges, best practices, and innovations is essential to inform educators, accreditation bodies, and policy-makers aiming to enhance the quality and outcomes of oral assessments [10].

In spite of their critical role in professional qualification and competency assurance, oral examinations in medical and paramedical sciences face persistent challenges related to standardization, validity, and educational effectiveness [11]. The lack of uniform criteria, examiner training, and robust quality assurance mechanisms has led to variability in outcomes and concerns about fairness and reliability. Moreover, emerging educational demands—such as competency-based learning, interprofessional collaboration, and technological integration—require updated and evidence-based approaches to oral assessment format [12].

While numerous interventions and modifications have been suggested in the literature to address these issues, there is a paucity of systematically synthesized evidence that comprehensively maps the existing challenges, stakeholder requirements, and viable solutions. Without such synthesis, efforts to reform oral examination practices risk being fragmented, inconsistent, or misaligned with educational goals. This systematic review aims to critically evaluate and integrate current research on the issues, requirements, and potential solutions to optimize oral examinations in the context of medical and paramedical education.

Materials and methods

This systematic review was conducted in accordance with the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines, which outline a standardized framework for reporting systematic reviews and meta-analyses [13]. This study was approved by the Ethics Committee of Tabriz University of Medical Sciences (Approval Code: IR.TBZMED.VCR.REC.1404.283).

Review question

This systematic review was guided by the following PICO question:

Among medical and paramedical students undergoing oral examinations (Population), what assessment strategies, requirements or interventions (Intervention), compared to traditional methods or no interventions (Comparison), are effective in improving oral examination outcomes (Outcome)?

Sources of information

A thorough literature search was performed across multiple electronic databases up to June 23 2025. These databases included PubMed, Web of Science, Scopus, CINAHL, ERIC, BIE, PsycINFO, IEEE, ProQuest, MEDLINE, Cochrane Library, Embase, SID, and ISC. In addition to peer-reviewed publications, we searched grey literature sources, including conference proceedings, books, dissertations, university repositories, institutional reports, and Google Scholar to reduce publication bias. These sources were identified through database searches, manual reference screening, and institutional repositories. Although grey literature was screened, only studies meeting all predefined eligibility criteria were included in the final review. To maintain awareness of new research, email alerts were configured in the databases to notify the authors of new studies matching the inclusion criteria, based on saved search queries as of June 23, 2025.

Inclusion and exclusion criteria

Studies were eligible for inclusion if they focused on oral examinations within the context of medical and paramedical education. Only English-language articles with full text availability were considered. The editorials, commentaries, review papers, and opinion pieces were excluded from the review. There were no limitations on the year of publication.

Search strategy

Search keywords were developed based on the study objectives, previously published literature, and input from subject-matter experts in medical informatics, health information management, and medical education. The detailed search syntax is outlined in Table 1. A medical librarian and an information specialist reviewed and approved the final search strategy. It is also noted that this search protocol was not pre-registered.

Initially, potentially relevant articles were identified based on their titles and imported into EndNote. Duplicate entries were then removed. A more detailed screening followed, which involved reviewing the abstracts and full texts to further refine the list of eligible studies.

Data collection procedure

Three researchers (A.T-M., S.Gh., and M.Gh.) independently evaluated the articles using a standardized spreadsheet. Each researcher assessed the titles for the inclusion and exclusion criteria, and articles that received two or more votes for exclusion were removed. In the next phase, the abstracts of the remaining studies were assessed independently using the same criteria, and those with at least two exclusion votes were discarded. Finally, the full texts were independently reviewed by all three authors, and studies flagged by at least two researchers were excluded.

A uniform data extraction form was employed to gather pertinent details from the selected articles. Information extracted included title, authors, year, country, study objective, study design, target population, educational setting, type of exam, technological component, challenges, requirements and needs, solutions, outcomes/key findings and limitation. All extracted data were consolidated and reviewed for accuracy by an independent investigator to ensure completeness and correctness. No data were missing from the final dataset.

Quality evaluation

To support the inclusion and exclusion decisions, a formal quality appraisal was performed using the Quality Assessment of Diagnostic Accuracy Studies (QUADAS) checklist. This tool is designed to assess the methodological rigor and trustworthiness of diagnostic studies, either as a whole or by evaluating their validity and reliability separately [13].

The checklist comprises 14 questions, each scored as “yes,” “no,” or “unclear.” Two researchers independently assessed each study for potential sources of bias. A “yes” was assigned when the required information was clearly reported, “no” if the information was absent, and “unclear” if the report lacked sufficient detail. Based on existing guidelines, studies scoring 7 or more out of 14, with a majority of “yes” responses, were considered high quality, while those scoring below 7 were classified as low quality. Any disagreements between the reviewers were resolved through mutual discussion and consensus.

To visualize and report the findings, various software tools were used, including Microsoft Excel, XMind, an online word cloud generator, and the yED Graph Editor.

Results

PRISMA-based study selection

The comprehensive literature search across all databases identified 25,594 records. Following the removal of duplicates (n = 14,842), 10,752 unique records remained for initial title screening. Of these, 9,973 were excluded due to irrelevance based on title review, leaving 779 records for abstract screening. After evaluating abstracts, 598 records were excluded, resulting in 181 articles subjected to full text assessment for eligibility. Among these, 79 articles were excluded for the following reasons: non-English language (n = 3), failure to meet inclusion criteria (n = 72), and unavailability of the full text (n = 4). Ultimately, 102 studies fulfilled the eligibility criteria and were included in the systematic review. The study selection process is illustrated in the PRISMA flow diagram (Fig 1).

Trends in scholarly output and global research distribution

The histogram provided in Fig 2 illustrates the temporal distribution of publications on oral examinations in medical and paramedical education, based on data retrieved from multiple databases spanning 1890 to 2025. The findings indicate that research on oral examinations in medical and paramedical education was virtually absent prior to the mid-20th century, after which a gradual increase in scholarly output emerged. This growth accelerated markedly from the late 20th century onward, culminating in a peak in recent years, particularly approaching 2025. Early contributions during this period often focused on problem-solving formats within specific disciplines, such as family medicine, before evolving toward more diverse approaches in the late 20th and early 21st centuries. These later developments encompassed structured group examinations, as well as the integration of technological tools into viva voce assessments. Collectively, reviewing the trend reflects a sustained and expanding academic emphasis on oral assessment methods, driven by shifts in educational paradigms, advancements in assessment methodologies, and the growing recognition of oral examinations as a critical element of competency-based evaluation. Moreover, the accompanying world map depicts the geographical distribution of publications, with notable concentrations in North America, Europe, Oceania and selected regions of Asia, underscoring the global reach and collaborative nature of research in this domain (Fig 2).

thumbnail
Fig 2. The trend of articles in the field of oral exam in medical and paramedical education in databases (The world map shown in Fig 2 was created by the authors using a public-domain base map (Natural Earth - http://www.naturalearthdata.com/)).

https://doi.org/10.1371/journal.pone.0355461.g002

Keyword mapping and thematic patterns

The word cloud, derived from a comprehensive analysis of literature on oral examinations in medical and paramedical education retrieved from multiple databases, visualizes the most frequently used terms in this research domain (Fig 3). Prominent keywords such as “Examine,” “Oral,” “Assess,” “Exam,” “Question,” “Student,” “Score,” “Valid,” and “Reliable” highlight core themes related to assessment methodologies, evaluation criteria, and the reliability of oral testing. Additional terms, including “Structure,” “Performance,” “Feedback,” and “Skill,” indicate an emphasis on structured assessment frameworks, performance appraisal, and the provision of constructive feedback. The presence of context-specific terms such as “Clinical,” “Medical,” “Case,” and “Viva” reflects the integration of oral examinations into clinical training and applied educational contexts. Collectively, this visualization underscores a sustained scholarly interest in the development and refinement of valid, reliable, and structured oral assessment tools tailored to the pedagogical requirements of medical and paramedical education.

thumbnail
Fig 3. Frequent words cloud based on findings.

https://doi.org/10.1371/journal.pone.0355461.g003

Overview of study designs

Analysis of the study designs reported in the 102 included articles revealed a diverse methodological landscape. Observational studies were the most common (n = 31), often incorporating inter-rater reliability assessments, generalizability analyses, and performance evaluations to examine real-world assessment practices. Experimental and quasi-experimental designs (n = 15) encompassed randomized controlled trials, crossover studies, and other interventional approaches, frequently aimed at comparing structured versus traditional examination formats and their effects on validity, reliability, and learner performance. Cross-sectional studies (n = 12) provided snapshot evaluations of perceptions, experiences, and scores, while qualitative studies (n = 10) explored thematic dimensions of stakeholder perspectives. Mixed-methods studies (n = 8) combined quantitative metrics with qualitative insights, offering a more comprehensive understanding of assessment processes. Additional designs included descriptive (n = 6), retrospective (n = 5), and comparative (n = 4) studies, which contributed foundational and historical perspectives. Reviews (n = 3), generalizability studies (n = 2), and other specialized approaches—such as pilot studies, survey-based investigations, and case studies (n = 6 collectively)—further underscored the methodological diversity supporting advancements in competency-based assessment within health professions education.

Study populations and educational contexts

Among the included studies, the majority (n = 72; 70.6%) were conducted in the field of medicine, frequently within subspecialties such as surgery (n = 12), anesthesiology (n = 5), psychiatry (n = 4), pharmacology (n = 4), physiology (n = 4), and emergency medicine (n = 4), with additional representation from vascular surgery, internal medicine, and ophthalmology. Nursing was the second most represented discipline (n = 9; 8.8%), followed by dentistry (n = 3; 2.9%), physiotherapy (n = 2; 2.0%), and osteopathy (n = 2; 2.0%). Less frequently represented specialties included pharmacy, occupational therapy, radiology, and clinical psychology (each n = 1–2). This distribution highlights a predominant focus on medical disciplines, particularly those involving clinical and procedural training, where oral examinations serve as a critical component of competency assessment.

Educational levels spanned the full continuum of health professions training. Undergraduate learners constituted the largest cohort (n = 45; 44.1%), encompassing students in medicine (bachelor of medicine, bachelor of surgery)MBBS(or equivalent), nursing, dentistry, and physiotherapy at early to mid-program stages (first- to fourth-year). Postgraduate trainees—including residents, fellows, and certification candidates—were represented in 38 studies (37.3%), with particular emphasis on residency training (post-graduate year (PGY) 1–5) in specialties such as surgery and anesthesiology. Faculty members and practicing professionals were the target population in 12 studies (11.8%), often in the context of certification, recertification, or examiner training. The remaining studies (n = 7; 6.9%) involved mixed or unspecified levels, including pre-registration interns and community health workers.

Sample sizes ranged from small pilot studies (9–20 participants) to large multi-year cohorts exceeding 1,000 candidates, with a median of approximately 80 participants per study (interquartile range: 30–150). Where reported, participant numbers were often stratified by subgroups—such as residents by year (e.g., 36 first-year vs. 32 second-year trainees) or examination format (e.g., 63 in structured oral examinations vs. 60 in conventional formats). Reported response rates averaged 70–80%, although some studies noted attrition (e.g., 42 dropouts from 150 enrolled). Gender distribution, inconsistently documented, showed a female predominance in nursing and allied health cohorts (e.g., 42 females vs. 13 males in one physiology study).

Studies were conducted across diverse institutional settings, most commonly university-affiliated medical schools and teaching hospitals (n = 78; 76.5%), including institutions such as the University of Michigan, Harvard Medical School affiliates, and the All-India Institute of Medical Sciences. Multi-center collaborations were frequent (n = 18; 17.6%), involving networks such as the Southern Association for Vascular Surgery (USA) and international partnerships (e.g., University Sains Malaysia with the University of Otago, New Zealand). Professional board and certification contexts—such as the American Board of Emergency Medicine and the Royal College of Physicians and Surgeons of Canada—featured in six studies (5.9%).

Oral examination formats and technological integration

Examinations were categorized by structural characteristics. Structured oral examinations were the most prevalent (n = 58; 56.9%), typically incorporating predefined questions, standardized scenarios, scoring rubrics, or checklists to enhance objectivity and reliability. Many of these formats were modeled on certification board practices (e.g., American Board of Surgery; Royal College of Physicians and Surgeons of Canada) and included variants such as mock oral examinations (n = 18; 17.6%) designed for preparatory purposes. Traditional, unstructured oral examinations accounted for 24 cases (23.5%) and were characterized by examiner-led, open-ended questioning, with greater subjectivity and variability in delivery. Virtual or online formats were reported in 12 cases (11.8%), reflecting adaptations to remote assessment—particularly in the post–COVID-19 period—while hybrid approaches appeared in 5 cases (4.9%). Less common formats included dialogic group orals (n = 1; 1.0%) and immersive simulations (n = 1; 1.0%). Three cases (2.9%) were not explicitly classified as examinations but evaluated curricula or surveys that incorporated oral assessment elements (Table 2).

thumbnail
Table 2. Distribution of oral exam types (N = 102) and technological components employed (N = 40 Cases with Technology).

https://doi.org/10.1371/journal.pone.0355461.t002

Exam characteristics varied in duration (7 minutes to 4 hours, where specified) and delivery method, with face-to-face formats predominating (n = 78; 76.5%), followed by virtual (n = 12; 11.8%) and hybrid (n = 5; 4.9%) modalities. Question types were primarily case-based or clinical (n = 85; 83.3%), with open-ended formats in 42 cases (41.2%). Scoring methods in structured examinations favored rubric-based or global rating scales (e.g., Likert-type or percentage-based), whereas traditional formats relied predominantly on subjective consensus grading.

Overall, technological integration was limited, with 62 cases (60.8%) reporting no technology use and relying exclusively on in-person delivery with paper-based tools (e.g., printed checklists or question cards). Among the 40 cases (39.2%) that incorporated technology, the most common application was video recording (n = 18; 17.6%), used for independent rating, bias mitigation, and post-examination review (e.g., inter-rater reliability analysis). Virtual platforms such as Zoom, Microsoft Teams, or Google Meet were employed in 12 cases (11.8%), primarily for remote mock orals. Digital tools for feedback and data collection (e.g., Google Forms, Microsoft Forms, REDCap) were used in 9 cases (8.8%). Additional innovations included simulation software (Second Life; n = 1), audio recording applications (Audacity or Soundnote; n = 2), and image or presentation software (e.g., PowerPoint, OsiriX) for case display (n = 5). In structured examinations, technology frequently supported objectivity through encrypted question banks or electronic checklists, whereas adoption in traditional formats was limited (4 of 24 cases; 16.7%) (Table 2).

Examiners were generally multidisciplinary faculty or specialists (e.g., surgeons, psychiatrists, nurses), with numbers per examination ranging from 1 to 140. Examiner training was reported in 28 cases (27.5%), typically involving orientation or standardization sessions; in contrast, many unstructured formats reported the use of untrained examiners, contributing to procedural variability. Inter-rater reliability, assessed using measures such as Cohen’s kappa or the intraclass correlation coefficient (ICC), was reported in 15 cases (14.7%), reflecting ongoing efforts to reduce subjectivity.

Challenges and proposed solutions in oral examinations

The identified challenges were systematically classified into six overarching domains through thematic coding: (1) examiner-related variability (e.g., subjectivity, bias, inconsistency), (2) student-related factors (e.g., stress, anxiety, preparedness), (3) logistical and resource constraints (e.g., time limitations, faculty availability), (4) reliability and validity concerns (e.g., low inter-rater agreement, limited generalizability), (5) format-specific limitations (e.g., subjectivity in traditional examinations, technical failures in virtual formats), and (6) equity-related issues (e.g., gender or racial bias, language barriers).

Examiner variability emerged as the most frequently cited challenge, reported in 68 studies (66.7%), reflecting persistent concerns over inconsistent scoring and halo effects. Student stress and anxiety were documented in 42 studies (41.2%), particularly in high-stakes or unstructured assessment settings. Logistical challenges were identified in 35 studies (34.3%), including time pressures and limited resources. Reliability and validity concerns were highlighted in 52 studies (51.0%), often attributed to small sample sizes or case-specific variance. Format-specific challenges, such as technical disruptions in virtual examinations, were reported in 18 studies (17.6%), while equity-related concerns were noted in 15 studies (14.7%).

These issues were more prevalent in traditional or unstructured examinations (28 studies; 27.5% of total) compared to structured formats, in which problems such as content overload or inflexible time limits were reported but were generally less severe. The shift toward virtual examinations—often accelerated by the COVID-19 pandemic (12 studies; 11.8%)—introduced challenges such as unstable internet connectivity but also improved accessibility for certain cohorts.

Proposed solutions were proactive and multifaceted, grouped into the following categories: examiner training and calibration, question and scoring standardization, use of multiple examiners, integration of technological tools, implementation of feedback mechanisms, and structural or format adjustments. Examiner training was the most frequently recommended intervention (62 studies; 60.8%), typically involving orientations, workshops, or peer review sessions to mitigate bias and improve scoring consistency. Standardization of questions and scoring—through structured checklists, rubrics, or predefined case scenarios—was suggested in 58 studies (56.9%), with the aim of enhancing reliability, often achieving inter-rater reliability coefficients exceeding 0.80 in multi-case assessments. The use of multiple examiners was advocated in 45 studies (44.1%), frequently employing consensus grading or independent scoring to reduce individual variability.

Technological interventions, such as video recording for post-hoc review or the use of secure online platforms, were reported in 28 studies (27.5%), particularly for addressing bias and enabling remote assessment. Feedback mechanisms, including post-examination debriefings and self-assessment opportunities, were noted in 32 studies (31.4%) to support learner development and promote iterative improvements in examination processes. Structural modifications—such as increasing the number of cases or adopting hybrid examination formats—were proposed in 40 studies (39.2%) to address logistical and equity-related challenges.

Solutions were frequently tailored to the specific challenge domain: for example, examiner training and scoring standardization addressed examiner variability in 85% of relevant studies, whereas technological interventions mitigated virtual examination challenges in 90% of such cases. In 15 studies (14.7%), pilot testing and validation procedures were implemented to ensure feasibility and optimize assessment design. Overall, structured and mock examination formats demonstrated superior effectiveness in addressing identified challenges, with reported improvements in inter-rater reliability (kappa coefficients >0.70) and enhanced learner satisfaction. The challenges and solutions are summarized in Fig 4.

thumbnail
Fig 4. Distribution of challenges and solution in oral exam.

https://doi.org/10.1371/journal.pone.0355461.g004

Requirements and needs for effective oral examinations

The analysis of extracted literature revealed a comprehensive set of requirements and needs essential for the effective design, implementation, and evaluation of oral examinations in medical and paramedical education. These requirements were categorized into several thematic domains (Table 3):

thumbnail
Table 3. Thematic categorization of requirements and needs for oral exam in medical and a paramedical education.

https://doi.org/10.1371/journal.pone.0355461.t003

  1. Assessment Design and Structure: A strong emphasis was placed on the use of SOEs to enhance fairness, reliability, and standardization. Studies consistently advocated for standardized question formats, blueprinted content, and validated scoring rubrics to ensure alignment with curricular objectives and clinical competencies. The inclusion of multiple cases and examiners was recommended to reduce variability and improve psychometric robustness.
  2. Examiner Preparation and Calibration: Effective oral assessment was found to depend heavily on examiner training, including calibration sessions, mentorship programs, and consensus-building activities. Mechanisms to mitigate examiner bias—such as double-blind procedures, standardized rating anchors, and demographic balancing—were frequently cited. The need for ongoing professional development and feedback loops for examiners was highlighted to maintain consistency and improve scoring accuracy.
  3. Candidate Experience and Support: Numerous studies underscored the importance of reducing student anxiety through clear communication of exam expectations, mock oral sessions, and supportive environments. Structured feedback mechanisms—both formative and summative—were deemed critical for enhancing learning outcomes and guiding remediation. Recommendations included adequate time allocation, transparent scoring criteria, and flexibility in question delivery to accommodate diverse learner needs.
  4. Technological Integration: While traditional face-to-face formats remained dominant, there was growing interest in virtual oral examinations, particularly in response to logistical and accessibility challenges. Requirements for stable platforms, secure proctoring, and user-friendly interfaces were frequently mentioned, alongside the need for digital tools to support feedback and scoring. App-based assessments and video-enhanced formats were explored as innovative approaches to simulate clinical scenarios and enhance realism.
  5. Curricular Alignment and Validity: Oral examinations were increasingly expected to assess higher-order cognitive skills, including clinical reasoning, decision-making, and communication. Alignment with competency-based frameworks (e.g., CanMEDS, ACGME) was emphasized to ensure relevance and educational impact. The need for content validity, construct validity, and predictive validity was recurrent, with several studies proposing pilot testing and psychometric analysis to support exam refinement.
  6. Operational and Logistical Considerations: Successful implementation required adequate faculty resources, coordinated scheduling, and appropriate examination settings. Recommendations included standardized documentation, recording protocols, and quality assurance processes to support transparency and reproducibility. Multi-institutional collaboration and centralized oversight were proposed to enhance consistency across programs.

Outcomes of oral examinations

Reliability emerged as a central concern, with structured formats consistently outperforming traditional ones. Cronbach’s alpha values ranged from 0.52 to 0.99 across studies, with a median of 0.75 indicating moderate to high internal consistency for SOEs. Inter-rater reliability, measured via ICCs or Cohen’s kappa, was generally moderate (kappa = 0.47–0.82), improving post-standardization or examiner training (from 0.49 to 0.82). Generalizability coefficients suggested that 6–10 cases or examiners were needed for reliability ≥0.80. Variability in examiner severity contributed significantly to score variance (84.1%), underscoring the need for calibration. Practice effects enhanced performance over repeated assessments (pass rates increasing from 75% to 100%), but halo effects were noted in candidate-specific rating.

Content and construct validity were supported in most studies, with SOEs demonstrating strong alignment with clinical competencies. Correlations between oral scores and written exams or OSCEs were modest (r = 0.10–0.74), indicating oral exams assess distinct skills like clinical reasoning and communication. Criterion validity was evident in associations with postgraduate performance (r = 0.37–0.45 with prior certification scores) and reduced racial grading disparities post-standardization. However, some studies reported low predictive validity for board certification pass rates (no significant correlation).

Pass rates varied widely (50–100%), with structured formats yielding higher consistency (77–100%). Mean scores ranged from 56.53% to 95.3%, often higher in hybrid or virtual models (91.98% in NCLEX/oral groups). Significant improvements were observed with training levels (PGY-3: 50% vs. PGY-5: 92%, p = 0.006) and interventions like mock orals (89% vs. national 86%). Gender and ethnic differences were noted, with females occasionally underperforming due to bias, though mitigated in structured formats.

Student satisfaction was high for structured and virtual formats (72–96% agreement on fairness and reduced stress), with preferences for SOEs over traditional viva (64–90%). Qualitative themes included enhanced reflection, reduced anxiety, and better alignment with real-world practice. Examiners valued SOEs for objectivity (83–100% agreement) but noted workload burdens (NASA TLX score: 59.6 ± 14.1). Feasibility was affirmed in resource-constrained settings, though time requirements (25–37.5 hours for 150 students) posed challenges.

Compared to written exams, orals better assessed higher-order skills (problem-solving, r = 0.80) but showed lower reliability without structure. Virtual formats were comparable to in-person (no significant differences in scores or pass rates, p > 0.05), with benefits in accessibility and cost (76–97% agreement). Feasibility metrics highlighted minimal costs ($34.60 per resident) but emphasized examiner training needs.

Limitations of included studies

The most dominant limitation was restricted generalizability, appearing in 94 studies (91.3%). This was frequently attributed to single-institution or single-center designs, discipline-specific foci (e.g., psychiatry, pharmacology, or vascular surgery), or contextual factors such as regional or cultural settings, which limit extrapolation to broader educational or clinical environments. Small sample sizes emerged as the second most prevalent issue, cited in 51 studies (49.5%), often involving cohorts of fewer than 50 participants (e.g., 20–30 students or residents), thereby compromising statistical power, representativeness, and the ability to detect meaningful effects.

Potential biases were reported in 61 studies (59.2%), encompassing various forms such as selection bias (e.g., convenience or volunteer sampling), response bias in surveys, examiner bias due to familiarity or subjectivity, and reporting bias from self-administered questionnaires. This theme underscores inherent challenges in oral examination research, where interpersonal dynamics and non-randomized designs amplify subjectivity. Lack of quantitative data or rigorous statistical analysis was evident in 22 studies (21.4%), with many relying on qualitative insights, descriptive statistics, or unvalidated instruments (moderate Cronbach’s alpha values), which hinders objective evaluation of outcomes like reliability or validity.

Additional recurring limitations included a narrow scope or focus on specific topics/subjects (21.4%), such as single disciplines or exam formats, potentially overlooking interdisciplinary applications; insufficient details on examiner training or calibration (19.4%), which could affect scoring consistency and inter-rater reliability; and dependence on subjective measures (e.g., self-reported perceptions or surveys; 15.5%), introducing variability and potential inaccuracies. Absence of long-term follow-up or prospective data was noted in 13 studies (12.6%), limiting insights into sustained impacts on learning or clinical performance. Less frequent but notable issues encompassed incomplete or truncated data (8.7%), low survey response rates (6.8%), technical challenges (e.g., connectivity in virtual formats; 5.8%), retrospective designs (4.9%), and quasi-experimental approaches (1.9%), all of which pose risks to causal inference and reproducibility (see Table 4 for a comprehensive summary of these themes). All details of the articles included in the study are provided in S1 File.

thumbnail
Table 4. Summary of recurrent limitations Identified in the included studies.

https://doi.org/10.1371/journal.pone.0355461.t004

Patterns and associations across study variables

Based on the study dataset, we conducted a qualitative thematic analysis to identify patterns, correlations, and potential associations among the extracted variables. Table 5 presents the identified patterns across variable pairs.

thumbnail
Table 5. Matrix of observed patterns between study variables.

https://doi.org/10.1371/journal.pone.0355461.t005

Comprehensive framework for optimized oral examinations

Following a comprehensive analysis of all extracted data from the included studies, a unified framework for the implementation of oral examinations in medical and paramedical education was developed. Adherence to this framework is expected to optimize the outcomes of these assessments, enhancing their effectiveness and reliability in educational settings. Table 6 and Fig 5 illustrate this framework in detail.

thumbnail
Table 6. Comprehensive framework of optimized oral exam execution in medical and paramedical education.

https://doi.org/10.1371/journal.pone.0355461.t006

thumbnail
Fig 5. Comprehensive framework of optimized oral exam execution in medical and paramedical education.

https://doi.org/10.1371/journal.pone.0355461.g005

Discussion

This systematic review synthesizes evidence from 102 studies on the issues, requirements, and solutions for optimizing oral examinations in medical and paramedical education, highlighting a predominant focus on structured formats, persistent challenges such as examiner variability, and emerging technological integrations. The findings reveal that SOEs were the most common format, demonstrating superior reliability and validity compared to traditional unstructured viva voce, with improvements in inter-rater agreement post-standardization. These outcomes align with recent literature emphasizing the psychometric advantages of SOEs. For instance, a meta-analysis by Rahman et al. reported high validity and reliability for structured viva in health professions education, with overall acceptability rates of 79.8% among learners (p < 0.001), mirroring our observation of enhanced learner satisfaction in structured formats [14]. Similarly, a study by Pernar et al. on SOEs in undergraduate medical education found them effective in assessing clinical reasoning and providing real-time feedback, which corroborates our data on SOEs’ alignment with higher-order skills like decision-making and communication [15].

The geographical and temporal distribution of studies, with a surge in publications post-2000 and concentrations in North America and Europe, reflects evolving educational paradigms driven by competency-based frameworks such as CanMEDS and ACGME. This trend may explain the shift toward structured assessments, as global accreditation bodies increasingly demand evidence-based evaluation methods to ensure fairness and equity [16]. In comparison, a scoping review by Janke et al. on high-stakes examinations in higher education identified similar growth in oral assessment research, attributing it to the need for authentic evaluations that simulate clinical interactions, particularly in resource-constrained settings like those in Asia and Oceania represented in our review [17]. The predominance of medicine over paramedical fields like nursing in our included studies underscores a disciplinary bias, potentially due to the high-stakes nature of medical certification exams, where oral formats are integral for assessing procedural and ethical competencies. This is consistent with findings from a systematic review by Wang et al., which highlighted SOEs’ effectiveness in lab-based physiology sessions for medical students, improving performance metrics by addressing misconceptions through interactive probing [18].

Challenges identified in our review, such as examiner variability and student anxiety, persist despite interventions, likely attributable to inherent subjectivity in human-led assessments and the psychological pressures of high-stakes environments. These issues are exacerbated in traditional formats, where inter-rater reliability is lower, as evidenced by our data showing unstructured exams contributing to 27.5% of format-specific problems. Comparative studies support this: a randomized trial by Akkaraju demonstrated that oral exams, when unstructured, amplify bias and stress, reducing validity, whereas structured variants mitigate these through rubrics and calibration [19]. The post-COVID-19 emergence of virtual formats introduced new challenges like technical disruptions, but also solutions such as video recording for bias mitigation. This evolution is echoed in a narrative review by Khalaf et al. on online exams in dental education during the pandemic, which reported comparable outcomes to in-person assessments (p > 0.05) but highlighted connectivity issues as barriers, similar to our findings on equity concerns including language barriers and access disparities [20]. The limited technological integration in our review may stem from institutional inertia and resource limitations, particularly in low- and middle-income countries, where pre-pandemic reliance on face-to-face methods predominated. However, post-pandemic analyses, such as those by Hytönen et al., indicate accelerated adoption of digital tools like Zoom for OSCE-integrated orals, improving accessibility and reducing costs ($34.60 per resident in one study), which aligns with our framework’s recommendation for hybrid synchronous delivery [21].

Solutions proposed in our synthesis, including examiner training and standardization, were effective in enhancing reliability, likely because they address root causes like halo effects and rater drift through calibration and rubrics. This is substantiated by a quasi-experimental study by Lepp et al., which integrated oral exams across pharmacy courses, yielding higher satisfaction and performance due to structured feedback mechanisms [22]. The thematic requirements in Table 3, emphasizing curricular alignment and technological needs, explain improved outcomes in structured formats by ensuring validity through blueprinting and psychometric validation. For example, correlations between oral scores and OSCEs in our review indicate that orals assess unique non-cognitive domains, a finding replicated in a validity study by Pernar et al., where SOEs reduced racial grading disparities post-standardization [23]. The observed patterns in Table 5, such as the shift from traditional to virtual formats post-2010, can be interpreted as responses to global disruptions like COVID-19, fostering innovation in assessment to maintain educational continuity while enhancing equity [24].

Limitations prevalent in our included studies, including restricted generalizability and small sample sizes, mirror broader challenges in assessment research, where single-institution designs limit external validity. These align with critiques in a systematic review by Memon et al. on oral exams in postgraduate medical education, which noted biases from non-randomized methodologies [25]. Despite these, our comprehensive framework (Table 6) offers a blueprint for optimized execution, integrating evidence-based choices like multi-station OSVEs and AI-assisted scoring, which could mitigate biases and improve feasibility in diverse settings.

In terms of implications, our findings underscore the need to extend innovations beyond medicine into paramedical education, where oral examinations remain underutilized despite their potential to assess essential competencies. Targeted interventions should also address inequities in technological adoption across low- and middle-income countries. Future research must prioritize multicenter and longitudinal studies to evaluate the predictive validity of virtual oral exams for clinical performance, while exploring artificial intelligence and natural language processing to enhance objectivity and minimize bias [26]. By implementing these strategies, oral examinations can evolve into robust, equitable, and competency-driven tools that better prepare health professionals for real-world practice.

Conclusion

This systematic review of 102 studies highlights that SOEs provide superior reliability, validity, fairness, and learner satisfaction compared to traditional viva formats. Key strategies for improvement include examiner training, question and scoring standardization, and the use of multiple examiners, all of which reduce subjectivity and enhance psychometric robustness. Technological tools such as video recording, virtual platforms, and digital scoring systems further support transparency, accessibility, and equity. Nevertheless, challenges persist, including examiner variability, student anxiety, logistical constraints, and equity-related issues such as language barriers and unequal access to technology. The underrepresentation of paramedical fields underscores the need for broader application and evaluation beyond medicine. Future research should focus on multicenter and longitudinal designs to strengthen generalizability and examine the predictive validity of oral exams for clinical performance. Emerging innovations—such as AI- and NLP-assisted scoring, VR/AR-enabled environments, and programmatic assessment frameworks—offer promising avenues for enhancing objectivity and scalability. In conclusion, oral examinations remain indispensable for assessing competencies such as clinical reasoning, ethical judgment, and communication. Through structured formats, examiner calibration, technological integration, and alignment with competency-based frameworks, oral exams can evolve into reliable, equitable, and learner-centered assessments that better prepare health professionals for real-world practice.

Supporting information

S1 File. Additional file 1.

Details of the included studies.

https://doi.org/10.1371/journal.pone.0355461.s001

(XLSX)

Acknowledgments

The authors sincerely thank the Medical Education Research Center, Tabriz University of Medical Sciences, for its support of this study. We also acknowledge all researchers whose published work contributed to the evidence synthesized in this systematic review.

References

  1. 1. Theobold AS. Oral exams: a more meaningful assessment of students’ understanding. J Stat Data Sci Educ. 2021;29(2):156–9.
  2. 2. Haque M, Ibtisam RS, Mustafa T, Qayyum S, Tahir QU, Melsing SB, et al. Oral examinations: What medical students and examiners think! Comparison of opinions on oral examination. Int J Pathol. 2018:66–73.
  3. 3. Erol H, Akın U. Opinions of school administrators about the oral exam carried out to select school principals. Int Online J Educ Sci. 2015;7(4).
  4. 4. Goins SM, French RJ, Martin JG. The use of structured oral exams for the assessment of medical students in their radiology clerkship. Curr Probl Diagn Radiol. 2023;52(5):330–3. pmid:37032291
  5. 5. Hazen H. Use of oral examinations to assess student learning in the social sciences. J Geogr Higher Educ. 2020;44(4):592–607.
  6. 6. Oliven A, Nave R, Baruch A. Long experience with a web-based, interactive, conversational virtual patient case simulation for medical students’ evaluation: comparison with oral examination. Med Educ Online. 2021;26(1):1946896. pmid:34180780
  7. 7. Cogil C, Gutierrez B. Integrate a brief oral examination for improved patient outcomes. J Nurse Practition. 2024;20(8):105125.
  8. 8. Abbasi Kasani H, Shams Mourkani G, Seraji F, Abedi H. Identifying the weaknesses of formative assessment in the e-learning management system. J Med Edu. 2020;19(2).
  9. 9. Pfiffner S, Albazi E, Musa A, Altinok G, Johnson SC, Harb A. The new diagnostic radiology oral exam: challenges, opportunities, and future directions. Acad Radiol. 2024;31(5):2190–1. pmid:38184415
  10. 10. Hosseini FA, Hemati M, Shaygan M, Gheysari S, Jaberi A, Ghobadi M, et al. Exploring the challenges and needs of nursing students in relation to OSCE exam stress: a qualitative study. PLoS One. 2025;20(7):e0327898. pmid:40658703
  11. 11. Lahri S, Meyer R. “We’re getting there”: Registrar and examiner perspectives on structured oral examinations in emergency medicine. J Coll Med South Afr. 2025;3(1):206. pmid:40951605
  12. 12. Burch V, McGuire J, Buch E, Sathekge M, M’bouaffou F, Senkubuge F, et al. Feasibility and acceptability of web-based structured oral examinations for postgraduate certification: mixed methods preliminary evaluation. JMIR Form Res. 2024;8:e40868. pmid:38064633
  13. 13. Torab-Miandoab A, Samad-Soltani T, Jodati A, Akbarzadeh F, Rezaei-Hachesu P. A unified component-based data-driven framework to support interoperability in the healthcare systems. Heliyon. 2024;10(15):e35036. pmid:39161828
  14. 14. Abuzied AIH, Nabag WOM. Structured viva validity, reliability, and acceptability as an assessment tool in health professions education: a systematic review and meta-analysis. BMC Med Educ. 2023;23(1):531. pmid:37491301
  15. 15. Pernar LIM, Askari R, Breen EM. Oral examinations in undergraduate medical education - what is the “value added” to evaluation? Am J Surg. 2020;220(2):328–33. pmid:31918844
  16. 16. Frank JR, Snell L, Sherbino J. CanMEDS 2015 Physician Competency Framework. Ottawa: Royal College of Physicians and Surgeons of Canada; 2015.
  17. 17. Janke KK, Westberg SM, Lee J. A review of the benefits and drawbacks of high-stakes final examinations in higher education. High Educ. 2023;86(6):1–24.
  18. 18. Wang L, Khalaf AT, Lei D, Gale M, Li J, Jiang P, et al. Structured oral examination as an effective assessment tool in lab-based physiology learning sessions. Adv Physiol Educ. 2020;44(3):453–8. pmid:32795125
  19. 19. Akkaraju S. The oral exam--learning for mastery and appreciating it. J Effect Teach High Educ. 2023;6(1):66–80.
  20. 20. Khalaf K, El-Kishawi M, Moufti MA, Al Kawas S. Introducing a comprehensive high-stake online exam to final-year dental students during the COVID-19 pandemic and evaluation of its effectiveness. Med Educ Online. 2020;25(1):1826861. pmid:33000704
  21. 21. Hytönen H, Näpänkangas R, Karaharju-Suvanto T, Eväsoja T, Kallio A, Kokkari A, et al. Modification of national OSCE due to COVID-19 - implementation and students’ feedback. Eur J Dent Educ. 2021;25(4):679–88. pmid:33369812
  22. 22. Lepp GA, Westberg SM, Lee J, Janke KK. Exploring the unanticipated value of an oral exam integrating content across courses. Curr Pharm Teach Learn. 2025;17(5):102302. pmid:39987592
  23. 23. Saab SS, Pollack S, Lerner V, Banks E, Salva CR, Colbert-Getz J. Validity study of an end-of-clerkship oral examination in obstetrics and gynecology. J Surg Educ. 2023;80(2):294–301. pmid:36266228
  24. 24. Akimov A, Malin M. When old becomes new: a case study of oral examination as an online assessment tool. Assess Eval High Educ. 2020;45(8):1205–21.
  25. 25. Memon MA, Joughin GR, Memon B. Oral assessment and postgraduate medical examinations: establishing conditions for validity, reliability and fairness. Adv Health Sci Educ Theory Pract. 2010;15(2):277–89. pmid:18386152
  26. 26. Sabin M, Jin KH, Smith A. Oral exams in shift to remote learning. In: Proceedings of the 52nd ACM Technical Symposium on Computer Science Education; 2021 Mar 3-7; Virtual Event, USA. New York: ACM; 2021. pp. 666–72.