Interruptions in antiretroviral therapy (ART) supply compromise individual treatment and raise the risk of virologic failure and drug resistance, undermining population-level viral suppression. India’s National AIDS Control Organization (NACO) runs one of the world’s largest public ART programmes, where frequent regimen transitions, evolving formulations, changing treatment guidelines, and procurement-driven fluctuations make reliable demand forecasting difficult. The authors developed an end-to-end, regimen-specific forecasting workflow designed to support procurement planning during periods of unstable demand.
The analysis used monthly national ART consumption data from January 2013 through December 2024. To protect confidentiality during method development, a validated privacy-preserving synthetic dataset was used for pipeline construction, and the final evaluation was performed on the real national consumption time series. The real aggregate data were provided by NACO under a data-use agreement and are not publicly available.
The study compared three broad model classes comprising five models in total:
The goal was regimen-specific forecasting rather than a one-size-fits-all model, acknowledging that different formulations and volumes may behave differently.
A fixed forecast horizon of 18 months was used for all models. Training and test windows differed between synthetic and real-data experiments because real data were available only through February 2024. For synthetic data, models were trained through June 2023 and tested on July 2023–December 2024. For real data, models were trained through August 2022 with the test window covering September 2022–February 2024.
Forecasting performance was assessed using signed percentage deviation (SPD), a metric that preserves the direction of error: positive SPD indicates under-prediction (forecast lower than actual), negative SPD indicates over-prediction. Models were evaluated and selected by the smallest absolute deviation, while retaining SPD to inform directional tendencies.
To account for the operational consequences of under-prediction, the authors derived a regimen-specific model-error buffer based on observed model performance. This buffer was applied only to held-out under-prediction, thereby padding forecasts where models historically tended to under-predict. The full forecasting workflow, including buffering and directional error reporting, was deployed through a no-code dashboard to support procurement planners.
On the synthetic benchmark dataset the best-performing models achieved small absolute deviations for several regimens. Reported examples include an absolute deviation as low as 0.46% for adult ABC+3TC and as high as 11.92% for adult AZT+3TC. These synthetic benchmarks provided a controlled environment to compare classical, transformer, and hybrid approaches.
When evaluated on real consumption data, no single modeling approach dominated across all regimen series. Classical methods like Holt–Winters or ARIMA remained competitive for some formulations, while transformer and hybrid models produced better predictive outcomes for others. Specific real-data examples reported by the authors:
These regimen-level differences motivated a portfolio approach to model selection rather than a single universal model.
Several formulations, especially low-volume and transition regimens, remained difficult to forecast accurately, revealing persistent operational uncertainty. Pediatric regimens showed systematic directional errors across all models, a more concerning pattern than isolated magnitude errors. For example:
These consistent directional biases imply factors (for instance, regimen transitions, shifting pediatric dosing practices, or small counts) that are not well captured by the evaluated time-series models.
The forecasting workflow, including selection rules, regimen-specific buffering, and directional error reporting, was operationalized through a no-code dashboard to aid procurement planning. The authors emphasise that the approach is intended to strengthen decision support during unstable demand periods rather than to replace existing public-health procurement systems.
The results support a portfolio approach to ART demand forecasting in national HIV programmes: select models at the regimen level, report directional errors transparently, and apply cautious model-error buffering where under-prediction risk is present. Performance differences across regimens indicate that hybrid and transformer methods can outperform classical techniques for some series, but classical approaches remain adequate for others. The persistent forecasting challenges for pediatric and low-volume regimens underscore ongoing operational uncertainty.
Limitations noted in the source include reliance on aggregate national consumption data and the preprint status of the work. The article is a preprint that has not been peer reviewed and therefore should not be used alone to guide clinical or procurement practice.
The real national aggregate ART consumption records used for final evaluation are not publicly available and were provided under a data-use agreement with NACO. A validated synthetic dataset, the full analysis code, and the dashboard source code are available on Zenodo at DOI 10.5281/zenodo.21839042. The authors report that Johns Hopkins issued a Not Human Subjects Research determination for the secondary analysis of de-identified data, and the authors declared no competing interests.