Perioperative observational studies are commonly used to evaluate anesthesia and perioperative practices that are difficult to test in randomized trials. However, treatment selection in these studies is often driven by surgical procedure type, which introduces procedure-level confounding across heterogeneous cohorts. Scalable approaches that represent surgical context are needed to improve confounding adjustment in large perioperative datasets.
The authors developed a method to convert free-text surgical procedure names into continuous representations and evaluated whether including these representations as covariates improves agreement between observational analyses and randomized trial benchmarks.
The study used free-text procedure names from 627,624 adult perioperative records. Procedure names were embedded with open-source sentence-embedding models to produce vector representations for each procedural label. These high-dimensional embeddings were reduced using principal components analysis and then incorporated into entropy-balanced observational analyses as additional covariates for confounding adjustment.
Adjustment strategies compared in the analysis were: unweighted (crude), clinical covariate adjustment (standard covariates), and clinical covariate plus surgical-name embedding adjustment. The embedding approach was positioned as an additional layer of confounding control intended to capture procedure-specific context not fully described by conventional covariates.
To test the practical impact of surgical-name embeddings, the authors replicated three recent perioperative randomized trials using observational data and three repeated analyses per trial scenario. The trials used as benchmarks were:
For each replication, the authors compared results across the three adjustment strategies to determine whether adding procedure-name embeddings produced estimates more consistent with the randomized trial findings.
In the GA-CARES replication, the unweighted analysis and the clinical covariate–adjusted analysis suggested lower two-year mortality associated with total intravenous anesthesia. When surgical-name embeddings were added to the adjustment set, the mortality estimate was attenuated to a nonsignificant association that aligned with the randomized GA-CARES trial result (reported OR 0.84 [95% CI, 0.62–1.14]; p=0.274). This change supports the role of procedure-level confounding that was captured by the embeddings but incompletely adjusted for by clinical covariates alone.
For the PADDI replication, adding surgical-name adjustment reproduced the trial’s overall non-harm conclusion for 30-day surgical site infection. The reported estimate was OR 0.72 [95% CI 0.69–0.76], p<0.001 in the adjusted analysis incorporating embeddings. Importantly, the embedding adjustment preserved an expected protective association with postoperative nausea and/or vomiting, indicating that embeddings can correct bias for one outcome while retaining plausible associations for other outcomes.
In the GAP replication for perioperative gabapentin and postoperative length of stay, the full cohort was null across adjustment strategies. However, recovering trial-consistent null length-of-stay estimates within specific surgical subgroups (cardiac, thoracic, and abdominal) required the inclusion of surgical-name embeddings. This finding suggests that subgroup-level confounding by procedure type can be substantial and that embeddings help control such heterogeneity.
The authors tested whether simple categorical indicators (e.g., department) or randomly generated covariates could substitute for the embeddings. Neither department indicators nor random covariates reproduced the effect of the surgical-name embeddings, supporting the interpretation that embeddings capture meaningful, procedure-specific information beyond coarse administrative labels or spurious variables.
Free-text surgical-name embeddings produced from open-source sentence-embedding models provide a scalable approach to represent surgical context in large perioperative datasets. Incorporating these embeddings into entropy-balanced observational analyses improved concordance with randomized trial benchmarks in several replication exercises while preserving expected treatment associations for other outcomes.
The authors conclude that surgical-name embeddings are a useful additional layer of confounding adjustment in perioperative observational studies and may help mitigate bias introduced by heterogeneous procedure mixes.
A software package implementing the method is publicly available at https://pypi.org/project/surgical-embeddings/ and can be used to reproduce intermediate data from the work. The Institutional Review Board of Stanford University approved the research. The authors declared no competing interests.