Organization for Economic Cooperation and Development (OECD) estimates suggest that about 20% of health care spending may be wasteful or harmful. Efforts to discourage low-value care have produced mixed results. This randomized controlled trial evaluated whether personalized letters that combined peer comparison feedback with appeals to professional norms could change two specific low-value behaviors among primary care physicians (PCPs): ordering of vitamin D testing and prescribing nongeneric medications instead of generic medications.
The study was a nationwide randomized controlled trial conducted in Switzerland between November 2020 and December 2021. PCPs were randomly assigned to intervention groups addressing one of three targets (vitamin D testing, generic prescribing, or a cost intervention) or to a common control group. This report focuses on the arms for vitamin D testing and generic prescribing compared with the common control group.
Physicians in the intervention groups received a personalized information letter. Each letter combined two behavioral elements: feedback comparing the physician to peers (peer comparison feedback) and messaging that emphasized professional norms regarding the low-value service in question. The letters targeted either ordering of vitamin D blood tests or the prescribing of nongeneric medications (to encourage generic substitution).
There were two primary endpoints in the comparison reported here: (1) the number of vitamin D tests per 100 patients, and (2) the share (percentage) of medications prescribed as generics. The investigators estimated average treatment effects using linear regression. To assess effect heterogeneity across physicians, they used a causal forest approach.
A total of 1,816 PCPs were included in the comparison reported: 618 were randomly assigned to the vitamin D intervention, 597 to the generic prescribing intervention, and 601 to the common control group. The analysis compared each intervention arm with the shared control group.
The intervention targeting vitamin D testing produced a statistically significant reduction in test ordering. On average, the personalized peer-comparison letters reduced vitamin D testing by 3.66 tests per 100 patients compared with the control group (95% confidence interval [CI], −5.42 to −1.89; P<0.001). This finding indicates a measurable decrease in this low-value diagnostic activity after a single-letter behavioral intervention that emphasized peer norms.
The intervention that targeted prescribing behavior did not produce a statistically significant change in the average share of generic medications prescribed. The mean difference compared with control was +0.57 percentage points (95% CI, −0.68 to +1.81 percentage points; P=0.37). In other words, the peer-comparison letter did not reliably increase generic substitution at the population average level within this trial.
Heterogeneity analysis using causal forests revealed variation in physician responses. For vitamin D testing, estimated reductions among physician subgroups ranged approximately from one to seven fewer tests per 100 patients, indicating that some physicians responded more strongly than others to the intervention. For generic prescribing, higher baseline rates of generic prescribing were associated with increases in generic substitution following the intervention, while physicians with low baseline levels did not show increases. Importantly, investigators report no evidence that the interventions induced increases in low-value care among physicians who began with low baseline rates.
The authors conclude that personalized letters combining peer comparison feedback and professional-norm messaging reduced vitamin D testing but did not increase generic medication prescribing on average. The trial was funded by the Swiss National Science Foundation and registered in the AEA Randomized Controlled Trials Registry (AEARCTR-0004747). The report appears in NEJM Evidence (2026 Sep;5[9]:EVIDoa2400308) with DOI 10.1056/EVIDoa2400308 and PubMed ID 42640164.
The abstract does not describe several implementation details and limitations in depth. For example, the source does not report the precise content wording of the letters, the timing and frequency of the mailings, follow-up duration for outcome measurement beyond the study period, or potential concurrent initiatives that may have influenced prescribing or testing behavior. Details on clinician demographics beyond counts, patient-level outcomes, or cost implications are not provided in the abstract. These items may be described in the full article but are not reported in the source abstract.