Untreated single-vision control arms in paediatric myopia efficacy trials are increasingly difficult to justify and retain. The authors examine whether validated virtual control arms created from published models of untreated childhood axial elongation can provide population-level estimates of untreated growth and thus reduce reliance on untreated allocations in trials.
Five published axial-elongation models were implemented unchanged in an open-source software tool provided as supplementary material. Each model predicts untreated axial elongation using baseline variables: age, cycloplegic spherical equivalent, sex, and ethnicity. Predictions were anchored at the cohort's baseline axial length (AL) and evaluated at the actual follow-up times for each participant.
Validation used a multi-ethnic untreated cohort of 242 myopic children drawn from China, Vietnam and India. Axial length was measured at approximately 6 and 12 months. Ethnic strata included East Asian participants (reported as Chinese and Vietnamese) and South Asian participants (Indian). The authors report subgroup sizes where relevant (for example, n = 71 for the Vietnamese 12-month subgroup).
Model predictions were compared with observed AL changes using bias, root-mean-square error (RMSE), and prediction-interval coverage. The authors defined pre-specified acceptability thresholds: absolute bias less than 0.03 mm and prediction-interval coverage greater than or equal to 0.90. Equivalence testing and p-values were used to assess whether observed bias fell within the pre-specified equivalence bounds.
At the group level, two models—the regional generalised estimating equation (GEE) and a meta-regression model—reproduced mean East Asian axial elongation at 6 months with negligible bias. The GEE model bias at 6 months was −0.013 mm and met the equivalence criterion (equivalence ±0.03 mm, p = 0.014). At 12 months the GEE bias was −0.004 mm, although equivalence was not established in an underpowered 12-month Vietnamese subgroup (n = 71; p = 0.068).
Older age-only models under-predicted axial elongation by approximately 0.07 to 0.12 mm, indicating these simpler models were not well matched to the validation cohort.
A key finding was that published individual prediction intervals were too narrow: observed coverage across participants was 0.77, below the pre-specified target of 0.90. Thus while mean predictions were generally accurate for some strata, the models underestimated individual-level uncertainty.
The validated models matched mean untreated growth for East Asian participants at 6 months (and closely at 12 months in pooled analysis), conditional on cohort independence. By contrast, Indian (South Asian) axial elongation did not align with any existing model strata; Indian growth fell between strata and was not matched by an available model. This indicates a current gap: existing published models do not adequately represent South Asian paediatric axial elongation.
Older models that used age only as a predictor performed worse across ethnic groups in this cohort, systematically under-predicting growth.
The authors emphasise that the tool functions as a group-level instrument for estimating mean untreated axial elongation and therefore may be useful as a virtual control arm in trials at the population level. It is not a reliable tool for predicting individual patient axial length change because published prediction intervals were too narrow and individual uncertainty was underestimated.
Key limitations reported by the authors include cohort dependence of the validated models and limited follow-up duration in the validation cohort (approximately 6 and 12 months). The 12-month subgroup analyses were underpowered in some ethnic strata (for example, the Vietnamese subgroup cited at n = 71), which limited statistical conclusions about equivalence at 12 months.
Another important limitation is the lack of an existing model stratum that matches Indian (South Asian) axial elongation patterns. The authors note that the tool's value for estimating treatment effect requires back-testing against a trial with a known untreated arm and longer follow-up—ideally 24 to 36 months—to confirm performance over the timeframes common in myopia control trials.
The authors state that all data produced in the study are available upon reasonable request and that the relevant software tool is provided in the supplementary files. Ethical approvals were obtained from multiple regional ethics committees covering participating sites in India, China and Vietnam. The authors declare no competing interests and report that relevant participant consent and institutional approvals were secured.
From this external validation, the principal implication is that some published models can reproduce mean untreated axial elongation for East Asian populations at short-term follow-up and could inform group-level virtual control arms in paediatric myopia trials. However, models currently underestimate individual-level uncertainty and do not cover South Asian (Indian) growth patterns. The authors recommend further validation by back-testing the tool against randomized trials with an untreated arm and longer follow-up (24–36 months) before deployment in confirmatory trial settings. Development or calibration of models that include South Asian strata is also needed to extend applicability.
Validated models implemented in an open-source tool can reproduce mean untreated East Asian axial elongation at 6 months, but current prediction intervals are too narrow for individual forecasts and South Asian children remain unserved by existing model strata. Further longer-term back-testing and model expansion are necessary before virtual control arms could replace untreated allocations in paediatric myopia trials.