Investigators developed a sparse, fully disclosed point-of-care ultrasound (POCUS) risk equation intended to predict difficult videolaryngoscopy. The equation's analytic form was recovered from data rather than fixed a priori, and its structural properties were machine-checked by formal proof. The project follows TRIPOD+AI 2024 reporting principles and is reported as, to the authors' knowledge, the first formally verified clinical risk predictor.
The derivation cohort comprised 259 adults undergoing elective videolaryngoscopy at a single centre (Clínica Universidad de Navarra) in a prospective, single-operator design. The primary outcome was a clinically relevant difficult videolaryngoscopy state labelled no-Easy airway, observed in 68 of 259 patients (26.3%). Ethical approval and trial registration details are reported in the source.
A pre-defined library of 71 candidate terms constructed from nine anterior-airway POCUS features was screened using Sequentially Thresholded Least Squares (STLSQ) with bootstrap stability selection (B = 300). Screening retained a seven-term logistic equation. Two interactions met the prespecified stability criterion (|c|/sigma_c > 2): skin-to-epiglottis × skin-to-hyoid-bone distance and tongue volume × sagittal tongue area. A two-term bootstrap-stable model was pre-specified for robustness analysis.
Internal validation consisted of repeated cross-validation (5 × 10 repeats) supplemented by temporal and device hold-outs. Pre-specified assessments quantified overfitting and optimism. Model behaviour under these internal validation strategies was reported alongside calibration metrics and recalibration procedures.
The seven-term equation achieved a cross-validated C-statistic of 0.966 with an optimism-corrected value reported as 0.968. Performance persisted across temporal and device hold-outs, with C-statistics in the range 0.94–0.97.
Calibration-in-the-large matched the observed prevalence. The cross-validated calibration slope was 0.90, attenuating to 0.625 in out-of-time evaluation; applying a standard recalibration restored the slope to 0.92 without reducing discrimination. These metrics quantify modest overfitting and show how simple recalibration can address temporal attenuation.
A pre-specified two-term bootstrap-stable model was evaluated as a robustness check. This simpler model reproduced discrimination similar to the seven-term model (reported C-statistic range 0.964–0.968). Reported events-per-parameter for the models was 34 and shrinkage was 0.99, supporting that the predictive performance is not an artefact of the screening stage.
Decision-curve analysis indicated positive net benefit compared with a clinical baseline across decision thresholds from 10% to 50%. The source reports this as evidence the equation would provide clinical utility within a range of plausible decision thresholds in practice.
Five behavioural properties of the deployed equation were formalised and machine-checked in the Lean 4 proof assistant. All five Lean 4 theorem scripts compiled without error, and the authors disclose that the theorem statements and standardisation constants are provided in full in the manuscript and repository.
The deployable equation, standardisation constants, and the five Lean 4 theorem statements are disclosed in the manuscript. The Lean 4 proof project and the analysis, validation, and figure-generation scripts are publicly available at a Zenodo repository under a PolyForm Noncommercial License; the repository includes a script that reproduces and verifies the deployed equation without protected screening hyperparameters. The de-identified analysis dataset is available under controlled access governed by the institutional data-access committee at Universidad de Navarra; requests are processed within 30 days for non-commercial research. Two STLSQ screening hyperparameters used for model discovery remain available under a research collaboration agreement while a patent application covering the screening methodology is pending.
The study was conducted in a single-operator cohort and the POCUS inputs are operator-dependent. The authors explicitly state that external validation requires prior harmonisation of the measurement protocol and operator credentialing. They also disclose a pending patent application covering the described methods and name two authors on that application as competing interests. No other competing interests or third-party payments for the work were declared in the source.
Overall, the reported result is a sparse, interpretable, and formally verified POCUS-based logistic risk equation with high internally validated discrimination for predicting difficult videolaryngoscopy, accompanied by transparent code, a machine-checked proof, and controlled-access data for reproducibility. External generalisability and operational deployment will require standardised measurement procedures and independent validation.