Childhood mental health conditions such as ADHD, anxiety, and depression affect an estimated 13–20% of children, yet many cases go undetected and untreated (reported ranges 25–62% undetected). Objective, scalable tools could augment traditional screening and clinical interviews. Prior work in pediatric digital phenotyping has largely focused on single modalities, leaving open which physiological or behavioral signals are most informative and whether combining signals improves detection while remaining practical to implement.
This study enrolled 103 children aged 4–8 who completed a roughly seven-minute structured behavioral assessment while wearing multiple sensors and providing speech samples. Machine-learning models were trained to predict clinician-assigned diagnoses derived from a gold-standard clinical interview. The analysis aimed to compare modalities, device locations, and tasks to balance predictive performance against feasibility and implementation burden.
Collected signals included electrodermal activity, cardiovascular measures, skin temperature, movement, and speech features separated into acoustic and linguistic components. Data came from multiple body locations and devices to test whether particular placements or sensors contributed disproportionately to diagnostic discrimination. The report evaluates each modality's contribution to model performance and how combinations affected overall accuracy.
Children completed a short, structured behavioral assessment lasting approximately seven minutes. The protocol was designed to elicit physiological and vocal responses relevant to emotional and attentional states. Specific task-level details, such as task names or transcripts, are presented in the source but are summarized here as components of a brief, standardized assessment used to capture multimodal biobehavioral signals.
Supervised machine-learning models were trained to discriminate clinical diagnoses of ADHD, anxiety, and depression using the multimodal sensor and speech data. The reference standard for model training and evaluation was a gold-standard clinical interview. Models were evaluated primarily by their discrimination ability, with area under the receiver operating characteristic curve (AUC) reported as the principal performance metric.
Across diagnostic targets, machine-learning discrimination performance achieved AUC values in the range of 0.74–0.92. The manuscript compares single-modality models and multimodal combinations, as well as the influence of device location and task selection, to identify configurations that yielded the best trade-off between performance and implementation complexity.
A key finding was that combining wearable-derived model predictions with caregiver report substantially increased case detection sensitivity. Specifically, sensitivity rose by 35–54 percentage points compared with caregiver report alone, while maintaining moderate-to-high specificity. In this combined approach, the method detected two to three times more clinician-confirmed cases than caregiver report alone, indicating that brief multimodal assessment can meaningfully augment standard screening approaches.
The authors developed an implementation-burden score to quantify feasibility concerns associated with different sensor combinations, locations, and tasks. Using this score, they demonstrated that near-best predictive performance could be achieved at relatively low burden for some diagnostic targets. This finding highlights potential pathways to design brief, low-burden assessment protocols that still provide clinically useful objective signals.
The study was conducted under University of Vermont Institutional Review Board approval (CHRBSS00001218). Several authors (B.C.L., E.W.M., R.S.M., and N.C.) are co-founders and equity holders of Biobe, Inc.; B.C.L. serves as CEO. The company did not fund the study and had no role in its design, conduct, analysis, or publication decision. Individual-level data cannot be shared publicly under the IRB; de-identified summary statistics and aggregated outputs supporting figures and tables are available from the corresponding author upon reasonable request and execution of a data-use agreement.
Declared funding sources include a National Science Foundation Graduate Research Fellowship, U.S. National Science Foundation award 2046440, and NIH K23MH123031. The manuscript is a preprint and has not been peer reviewed; it should not guide clinical practice without further validation.
In this cohort of 103 children aged 4–8, a brief, multimodal wearable and speech-based assessment produced machine-learning models that discriminated ADHD, anxiety, and depression with AUC values between 0.74 and 0.92. Combining model outputs with caregiver report substantially increased sensitivity and detected substantially more clinician-confirmed cases than caregiver report alone. An accompanying implementation-burden score indicated that high performance may be achievable with low-burden setups for some targets, supporting the feasibility of brief wearable assessments as an objective complement to caregiver-reported screening. The work is presented as preliminary (preprint) and emphasizes that additional validation and peer review are required before clinical adoption.