Determining appropriate sample size and statistical power for studies that estimate spillover effects in networks is essential to obtain reliable inference. This study focuses on sociometric, network-based designs where interventions were not randomized across individuals or components. The authors evaluated how several design parameters influence power: the number of network components, the total number of nodes, node degree, transitivity, and effect size. The analyses target the setting and estimator specified in the work rather than offering universally applicable formulas across all possible estimators or interference assumptions.
The authors conducted a comprehensive simulation study to assess how the above design parameters affect statistical power when estimating spillover. Simulated networks were constructed to vary systematically in component count, node counts, node degree, and transitivity. In addition to synthetic networks, the authors used a real-world sociometric network from the Transmission Reduction Intervention Project (TRIP) to demonstrate the behavior of power under empirically observed network structure. The analytic focus and reported results apply to the particular inverse probability weighting estimator employed, and the estimator's required assumptions were maintained throughout the simulations.
Across simulation scenarios the authors observed several consistent patterns. When holding the number of components constant, increasing the total number of nodes led to higher statistical power to detect spillover effects. Similarly, larger effect sizes produced increased power. By contrast, simply increasing the number of components while keeping the total number of nodes fixed did not necessarily improve power; in some scenarios power remained the same or decreased slightly as components increased when node count was held constant.
These findings underscore that total sample size (nodes) and effect magnitude are primary drivers of power in the evaluated settings, while component count interacts with node allocation and does not uniformly translate into power gains.
Network topology had a measurable impact on power estimates. Specifically, scenarios with higher average node degree—that is, where nodes had more connections—were associated with reduced power for estimating spillover effects. Likewise, greater transitivity (the tendency for connected nodes to form closed triangles) also corresponded to lower power.
The authors further examined the effect of component size distribution. Highly unbalanced networks, such as when most nodes belong to a single component and few nodes populate other components, produced drastic reductions in power. This result highlights the importance of considering not only total sample size but also how nodes are distributed across components when designing studies intended to estimate spillover.
Beyond simulations, the work presents a closed-form expression for power under the inverse probability weighting estimator used in the study. When applied to the same design settings, the closed-form calculations yielded behavior consistent with the simulation results: power increased with more nodes and larger effect sizes, and power did not necessarily increase with more components when the total number of nodes was fixed. In some cases the closed-form solution indicated that power could remain unchanged or slightly decrease as component count rose while nodes were held constant, reinforcing the simulation findings.
All reported results are conditional on the specific inverse probability weighting estimator and the interference assumptions required for its validity. The authors explicitly note that alternative estimators or different assumptions about interference among units could produce different power patterns. They also emphasize that these power and sample size insights apply to the analyzed estimator and network models; study designers should not generalize the numeric results to settings that deviate from the assumptions used here without further assessment.
The authors state that all simulation data produced in the study are available upon reasonable request. The real-world TRIP network data used are accessible from the National Addiction and Health Data Archive Program (NAHDAP) via the cited repository, with some files subject to restrictions. The preprint declares that ethical approval was provided by the IRB of the University of Rhode Island and that the authors have no competing interests. Readers are reminded that this manuscript is a preprint on medRxiv and has not undergone peer review; it should not be used to guide clinical practice without further validation.