Oral examinations (vivas) are face-to-face assessments used widely in medical and paramedical education to evaluate clinical reasoning, communication, professionalism, and decision-making. Their interactive nature allows examiners to probe thinking, clarify responses, and assess students’ performance under pressure. Despite these advantages, oral exams have long been criticized for limited objectivity, inter-examiner variability, and challenges in standardization—issues that are particularly concerning when exams are high stakes.
This systematic review followed PRISMA guidance. The authors searched 14 electronic databases (including PubMed, Scopus, Web of Science, CINAHL, MEDLINE, Embase, and IEEE) and screened grey literature up to June 23, 2025. A PICO-framed question focused on medical and paramedical students and interventions or strategies designed to improve oral examination outcomes. Three independent reviewers performed screening and data extraction. Quality appraisal used QUADAS and thematic analysis with matrix methods to identify patterns and correlations.
From 25,594 retrieved records, 102 studies met inclusion criteria. The number of publications rose markedly after 2000. Most studies (70.6%) were from medicine, and the bulk addressed undergraduate (44.1%) or postgraduate (37.3%) education. Study designs, settings, and sample sizes varied; many studies were limited in generalizability and often had small samples.
Across included studies the most frequent problems were:
Additional recurrent concerns were lack of standardized rubrics, potential bias (including differential item functioning), and limited quality assurance mechanisms.
The review identified core requirements that underpin improved oral examinations:
Interventions most frequently reported as beneficial included:
Other innovations described across studies included multi-station formats (often aligned with OSCE principles), hybrid in-person/digital delivery models, and early explorations of AI and natural language processing for scoring assistance.
Reported psychometric and outcome data across studies included:
Based on synthesis of the literature, the authors recommend a comprehensive framework emphasizing:
The framework highlights examiner training, robust quality assurance, data governance, and piloting before high-stakes implementation.
Limitations reported across the body of evidence included threats to generalizability (noted in 91.3% of studies) and frequent small sample sizes (49.5%). Many studies lacked longitudinal follow-up or validation of predictive validity. Early explorations of AI and automated scoring were promising but remain under-validated; the review calls for robust studies to evaluate predictive performance, fairness (differential item functioning), and regulatory compliance.
The review concludes that standardized oral examinations, supported by examiner training, structured rubrics, multi-station designs, and appropriate technology, can improve reliability, validity, and perceived fairness compared with traditional viva formats. The authors recommend adoption of an evidence-based framework—centered on OSVEs, hybrid delivery, combined rubrics, and validated AI-assisted tools—while emphasizing the need for further research to confirm predictive validity, scale interventions, and evaluate bias mitigation strategies. Educators and institutions should pilot structured approaches with rigorous quality assurance before broad high-stakes deployment.