Variant effect predictors are increasingly benchmarked against multiplexed assays of variant effect (MAVEs) rather than clinical labels to avoid label circularity. However, benchmarking against an assay introduces a fundamental constraint: the maximum correlation a predictor can show against assay readouts is limited by the assay's own reproducibility. Precision and reproducibility vary across genomic and experimental territories, so raw benchmark comparisons can misrepresent where predictors genuinely fail versus where the assay is insufficiently reliable.
The analysis scored nineteen predictors across sixteen strata derived from a frozen atlas of 64,178 saturation genome editing variants covering seven cancer-susceptibility genes. Published replicate scores and reported standard errors from these assays serve as the basis for estimating each territory's measurement reliability. The study uses these empirical reproducibility estimates to compute a per-territory reliability ceiling — the upper bound on observable correlation between predictor output and assay measurements imposed by assay noise.
From published replicate measurements and standard errors, the study estimates a reliability ceiling for each territory. These ceilings quantify how much of the predictor–assay correlation could be attributable to true signal rather than assay noise. The author emphasizes that ceilings vary substantially across territories: variability between strata in reliability is larger than variability between predictors. Where a territory lacks sufficient replicate information, the source reports that a reliability estimate cannot be provided.
Simulations show that correcting observed predictor–assay correlations for the assay's reliability ceiling changes apparent performance in a non-linear way. Above a ceiling of approximately 0.45, the correction reduces error (improves the adjusted interpretation of predictor performance); below this ceiling the correction amplifies error estimates. This relation implies that reliability correction is essential to interpret predictor rankings and absolute performance when assay reproducibility is limited or heterogeneous across territories.
Applying territory-resolved reliability correction materially alters the map of where predictors appear to fail. The study finds that much of the observed collapse in predictive performance at canonical splice sites is an assay property. After correction, the median shortfall in performance at splice sites relative to coding regions narrows from 1.7-fold to 1.4-fold. This convergence persists when either BARD1 or PALB2 is dropped from the atlas, but inverts when BRCA1 is omitted. Because the finding depends on these leave-one-gene-out folds and on a single deposit influencing the frontier parity, the author reports all three folds rather than asserting gene independence.
Importantly, the corrected map relocates genuine predictor failure: rather than canonical splice sites, true failure concentrates 11–50 base pairs into the intron, a region that the uncorrected analysis presents only modestly as problematic.
Considering MaveDB deposits more broadly, the analysis finds that 2,452 of 2,803 score sets carry, at the upper bound, the information required to compute a reliability estimate, although a conventional column-name search would find only about a tenth of them. Focusing on 674 human deposits with a computable ceiling, between 29.9% and 51.8% fall below a ceiling of 0.90, indicating that a substantial fraction of assays lack high reproducibility and thus limit benchmark interpretability.
When predictors are scored as classifiers against the assays' own functional calls in three genes, they separate damaging from tolerated variants better than their raw correlations with quantitative scores would suggest. Nevertheless, none of the evaluated predictors reaches the strongest evidence band at the 95%-specificity operating point. This indicates that while corrected performance can improve interpretation, current predictors still fall short of the most stringent clinical-evidence thresholds when evaluated against MAVE-derived functional calls.
The study concludes that benchmarks resolved by experimental territory should report a per-stratum reliability estimate, or explicitly state that the assay does not permit such an estimate. Providing per-territory ceilings requires nothing additional from data depositors when replicate statistics are published, and at the upper-bound assessment the approach applies to most (about 87%) of MaveDB entries. Reporting these ceilings makes clear where predictor performance is limited by assay noise versus where predictors genuinely fail, and it materially changes interpretations at splice-site extremes and in intronic regions.
All findings and numerical values above are drawn from the preprint and its abstract. Where the source does not report more detailed methods, exact simulation parameters, or deposit identifiers, those details were not reported in the source.