Delirium is a frequent and clinically important complication in pediatric intensive care units. Reliable, validated assessment instruments are necessary to identify delirium, monitor clinical course, and evaluate interventions. The review aimed to identify pediatric delirium assessment tools used in acute care settings and to evaluate their measurement properties and the certainty of available evidence.
The authors conducted a systematic search of multiple databases without language or date restrictions, including MEDLINE, EMBASE, PsycINFO, CINAHL, the Cochrane Library, and Web of Science. From 8,378 retrieved records, studies were screened for eligibility. Eligible reports were original observational, cross-sectional, or validation studies that evaluated at least one measurement property of a pediatric delirium tool in acute care.
Data extraction, study selection, and risk-of-bias assessment were performed independently by two reviewers. Risk of bias and evidence appraisal used adapted COSMIN (COnsensus-based Standards for the selection of health Measurement INstruments) methods and the Quality Assessment of Diagnostic Accuracy Studies-2 (QUADAS-2) for diagnostic accuracy. Overall certainty of evidence was rated using a COSMIN-adapted GRADE approach.
Forty-one studies met inclusion criteria. The included literature covered 15 language versions of six delirium assessment instruments used in pediatric critical care:
Among these, CAPD had the largest evidence base overall. Several tools were evaluated across multiple language versions, whereas some instruments were represented by single-study evaluations.
The review focused on commonly reported measurement properties such as criterion validity, reliability (including interrater reliability), and convergent validity. Major findings reported by the authors include:
High-certainty evidence supported the criterion validity of CAPD and the Confusion Assessment Method tools (preschool- and pediatric-CAM-ICU). This indicates a strong relationship with a reference standard or diagnostic comparator as assessed in the included studies.
Reliability for CAPD and the CAM-ICU tools was supported with moderate-certainty evidence, reflecting acceptable interrater consistency reported across multiple investigations.
SOS-PD demonstrated a favorable profile despite being evaluated in fewer studies: the review found moderate-certainty evidence for its criterion validity and high-certainty evidence for convergent validity and reliability. This suggests SOS-PD performed consistently on measures of agreement with related constructs and had stable interrater performance where tested.
Evidence for CDAS and PEDS was limited and derived from single studies, preventing firm conclusions about their broader measurement properties or generalizability.
The authors highlight that evidence across instruments was uneven. Most primary studies focused on criterion validity and interrater reliability, while other important domains received little attention. Specific gaps included:
Rare assessment of measurement error, which is important to understand score precision and minimal detectable change.
Limited evaluation of cross-cultural validity despite multiple language versions being available.
Sparse data on subgroup performance (for example by age groups or clinical subpopulations), which limits conclusions about differential performance in distinct pediatric cohorts.
An overarching methodological challenge is that criterion validity in delirium research is difficult to define because no perfect diagnostic gold standard exists for pediatric delirium; this complicates interpretation of diagnostic accuracy and validity metrics.
Across the 41 included studies, SOS-PD emerged as having the most favorable certainty profile across evaluated measurement properties, although it was assessed in fewer studies compared with CAPD, which had the largest overall evidence base. The review authors recommended that future research should expand beyond criterion validity and interrater reliability to address measurement error, cross-cultural validity, subgroup performance, and real-world implementation considerations.
These findings summarize the current state of evidence for commonly used pediatric ICU delirium instruments and identify priorities for further validation work to improve screening, diagnosis, and evaluation of delirium in critically ill children.