Operational decisions—staffing, capacity, inventory and replenishment—are commonly informed by count forecasts. Transparent starting points include independent Poisson or fixed-share multinomial models. However, persistent users, shared shocks and latent rate heterogeneity can produce extra-Poisson variation and persistent composition changes that increase tail exposure and shortage risk. The relevant criterion for a model is whether it alters the chosen action and thereby affects realized loss (shortages or holding costs), not merely whether it fits observed variance.
Reinforced-count simulation (RCS) is introduced as a compact diagnostic that stresses mean-based decisions under plausible composition persistence. It asks whether an action calibrated under independence remains adequate when type shares persist, rather than asserting the reinforced generator as the true data-generating process.
The RCS workflow separates three tasks: estimate the predictable mean, represent unresolved variation around that mean, and evaluate the induced decision. The computational sequence is:
A calibration warning is reported when the independence-selected action increases expected cost by at least 10% or reduces fill rate by at least five percentage points relative to an over-dispersion-aware alternative. These thresholds are reporting conventions and can be replaced by application-specific tolerances.
Two public datasets provide cross-sectional and temporal validation:
RAND Health Insurance Experiment (RAND-HIE): 20,190 annual counts of outpatient physician visits, grouped into four self-rated health categories (excellent, good, fair, poor). This provides 200 deterministic, stratified 50/50 training/test splits to evaluate held-out decision transfer within the sampled population.
NYC 311 Service Requests (erm2-nwe9): daily counts for five pre-specified categories (Illegal Parking, Noise – Residential, Blocked Driveway, Street Condition, Water System) from 1 January 2022 through 31 December 2024. The panel comprises 1,096 days and 5,480 category-day counts with no missing cells. Rolling origins use the preceding 365 days for calibration and the next 28 days for evaluation, producing 26 non-overlapping test windows.
All candidate generators share the same period mean forecast and differ in their treatment of unresolved variation:
The objective across models is to select the smallest integer quantile corresponding to q/(q + h) for the scalar newsvendor setup or the robust counterpart for Scarf’s model. For multi-type coordinated policies, a can-order (s, c, S) rule is evaluated with holding, lost-event and a fixed setup charge per replenishment occasion.
One-period loss for demand d, stock s, holding cost h and lost-event cost q is the standard newsvendor loss. Experiments set h = 1 and evaluate multiple q values (including 2, 5, 10, 20) in sensitivity analyses. Fill rate is defined as fulfilled demand divided by demand. Regret is reported as percent cost increase relative to the best candidate action in the stated test environment. For rolling-origin averages, paired percentage improvements are computed within each origin and averaged; 95% confidence intervals use the standard error across the 26 origins.
When multi-type count samples are available, the reinforcement concentration is estimated by maximum likelihood under a Dirichlet–multinomial composition model after conditioning on each period’s total count. Optimization is performed over a bounded interval [0.05, 5000].
A simulation study evaluates estimation stability as a function of the number of observed periods (14, 28, 56, 112, 224) with multiple replications. When only mean demand is available, concentration is not identifiable; the workflow then reports a break-even grid indicating the largest reinforcement concentration at which the recommended integer decision first differs from the independent decision, serving as a sensitivity range.
Four robustness experiments probe model and parameter boundaries:
Summarize 28-day NYC windows by median variance-to-mean ratio, mean absolute cross-category correlation and median 95th-percentile-to-mean ratio; evaluate associations with Spearman correlation to show dispersion and correlation can move independently.
Evaluate a finite-memory matrix urn across a grid of cross-share and memory parameters using long simulations to retain dispersion, lag-one autocorrelation, pairwise correlation, cost and coordination effects.
Conduct policy transfer across three multi-type environments (independent negative-binomial margins, shared-gamma negative-binomial, finite-memory reinforced) by calibrating policies in each environment and scoring them across all test environments to distinguish marginal over-dispersion from dependence that affects coordination.
Cross assumed design q/h and true evaluation q/h over {2,5,10,20} to produce a regret matrix showing whether model choice remains important under economic input misspecification.
Simulations use a master random seed and deterministic sub-seeds. Classical moment validation used 12,000 explicit urn paths; scalar demand experiments used 10,000–150,000 draws depending on calculation. Monte Carlo confidence intervals are based on paired replications or rolling origins. All code, fixed aggregates, environment versions and generated tables are provided in the reproducibility package (S1 Code) with exact query scripts and an MIT license.
RAND-HIE person-year visit counts exhibited over-dispersion in every self-rated health group. Variance-to-mean ratios ranged from 6.43 (excellent health) to 9.66 (poor health). Approximately one third of respondents in the first three groups reported zero visits; 95th percentiles were 9, 10 and 13 visits, with the poor-health group 95th percentile at 21. Group maxima ranged from 69 to 77 visits.
NYC daily category means ranged from 166.6 to 1,266.9 and every raw variance-to-mean ratio exceeded five. Residential noise was highly episodic (maximum daily count 6,559 and a raw ratio of 423.6). The 28-day rolling windows show that windows in the upper dispersion quartile had median variance-to-mean 29.69 and median 95th-percentile-to-mean ratio 1.44, while mean absolute cross-category correlation did not systematically increase with dispersion; the Spearman association between dispersion and absolute correlation was not strong (p = 0.287), illustrating that large dispersion changes need not be accompanied by larger pairwise correlation.
Using 200 stratified 50/50 splits, scalar stock decisions were compared across independent Poisson, negative-binomial, reinforced count and empirical quantile rules. At q/h = 5 differences were modest: RCS reduced cost by 1.11% and increased fill rate by about eight percentage points compared to Poisson. At q/h = 20 the gap widened: Poisson held-out cost was 17.379 with fill rate 76.48%, while the reinforced decision had cost 14.757 and fill rate 91.14%, a paired cost improvement of 15.01%. Negative-binomial and empirical decisions performed similarly or slightly better in this scalar setting, consistent with marginal over-dispersion being the primary driver for scalar newsvendor actions.
Temporal validation across 26 rolling origins used a Poisson GLM with day-of-week and month indicators and a training-window linear trend to remove predictable calendar structure before calibrating residual generators. At q/h = 10, the RCS mean paired cost improvement was 18.7% with 95% confidence interval 12.4%–25.0%, and fill rate increased from 93.3% to 96.8%. These results illustrate that RCS can materially improve held-out performance in temporal settings with pronounced episodic demand.
A transfer experiment across multi-type environments showed that each calibrated model was best in its matching environment; using a mismatched model produced regret in the range of approximately 6.0%–20.0%. Joint sensitivity analyses, concentration-learning curves and cost-ratio perturbations identify regimes where the diagnostic is useful and when parameter estimation error can dominate model choice.
RCS is presented as a reproducible check on mean-based decisions when composition dependence is uncertain. It is a diagnostic stress test rather than a claim about the true multivariate demand process. The workflow provides explicit steps for estimating means, inspecting residual dispersion, calibrating candidate generators and scoring candidate actions on held-out or simulated environments. Thresholds for calibration warnings are conventions and should be adapted to context.
All code, exact aggregation queries and data processing scripts used for the RAND-HIE and NYC 311 analyses are provided in the Supporting Information (S1 Code). RAND-HIE and NYC 311 source datasets are publicly accessible through the cited statsmode ls module and the NYC Open Data portal, respectively.