This preprint describes a collaborative real-time multi-model forecasting system deployed in Germany for the 2024/25 autumn–winter respiratory season. The system was designed to provide short-term forecasts of multiple public health surveillance indicators and to build on methodological lessons from forecasting efforts during the COVID-19 pandemic. The authors emphasize practical, operational forecasting in a setting where multiple routine surveillance streams are available and subject to revisions.
Forecasts were produced for four primary surveillance indicators: general practitioner consultations for acute respiratory infections (ARI), hospitalizations for severe acute respiratory infections (SARI), and laboratory-confirmed cases of seasonal influenza and RSV. The study used openly available routine surveillance data published by the Robert Koch Institute. The authors note that all indicators experienced retrospective revisions, which affected real-time forecasting and required integration of correction methods.
Because the surveillance indicators were revised after initial publication, the forecasting pipeline incorporated a nowcasting step to adjust for reporting delays and retrospective data changes. Nowcasts aimed to provide corrected current estimates on which forecasts could be based. Overall, nowcasting performance was described as convincing, though some models were influenced by specific calendar effects: for example, Christmas break-related reporting dynamics produced an upward bias in early January for a subset of approaches.
A total of nine models were run in real time. When multiple models targeted the same indicator, their outputs were combined into an ensemble forecast. The ensemble was evaluated alongside individual models and simple benchmark models to assess added value. The authors report that most forecasting models outperformed simple benchmarks, and that the ensemble was among the better-performing approaches. However, unlike some previous collaborative forecasting projects, the ensemble did not consistently outperform the best individual models across all targets.
The forecasts produced during the season were, on the whole, well calibrated. Performance gains relative to baseline benchmarks were more substantial for age-stratified targets than for pooled targets, and were concentrated at short forecast horizons — specifically at lead times of two to three weeks. These findings indicate the system provided useful short-term situational awareness, particularly when stratification by age was available and when policy- or clinical-relevant horizons were short.
A persistent challenge identified by the authors was accurate anticipation of peak timing and peak magnitude of disease activity. Several models tended to predict epidemic curves that were too flat and signaled a turnaround earlier than observed. The preprint gives a specific example for SARI, where many models predicted a transition to decline in late January whereas observed data showed a peak timing closer to mid-February. These mismatches highlight limits of current model formulations and data inputs for predicting inflection points in seasonal respiratory epidemics.
The authors reflect on lessons for organizing collaborative forecasting efforts in the post-COVID-19 context. They discuss practical aspects of running multi-team, multi-target forecasting in routine public health settings and note the potential of AI-supported modelling to augment forecasting workflows. The preprint does not provide implementational details or empirical evaluation of AI approaches; the mention of AI is framed as a future potential rather than a validated component in this deployment.
All analyses used publicly available surveillance data from the Robert Koch Institute; specific data links for ARI, SARI, RSV, and influenza series are cited in the source. The authors declared no competing interests and confirmed compliance with ethical guidelines. Multiple funders supported the work, including German federal and research foundations; funding details are provided in the preprint. The authors also note that this is a preprint and has not been peer reviewed, and therefore the findings should not be used to directly guide clinical practice.