Multi-task models that predict compound bioactivity transfer information across assays but are only useful for experimental prioritization when accompanied by reliable uncertainty estimates. The work reported here benchmarks both prediction accuracy and uncertainty calibration for multiple multi-task predictors, and evaluates how well different uncertainty approaches can rank predictions by error to support selective experimental follow-up.
Evaluation used two collections of test assays drawn from ChEMBL: 100 test assays from ChEMBL27 and 100 test assays from ChEMBL37. Each panel was evaluated under two partitioning strategies: a random split and a clustering split. These splits were used to probe model behavior in settings with differing chemical similarity between training and test queries.
Five multi-task predictors were benchmarked; the article compares prediction accuracy and uncertainty calibration across these models. For uncertainty calibration, three sources of uncertainty were considered: the predictors' native uncertainty outputs, Deep Ensemble approaches, and MC Dropout. Risk-ranking comparisons involved two Gaussian process (GP) backbones used to translate model outputs into risk scores for queries.
The benchmark examined both calibration — how well reported uncertainties match observed errors — and error ranking — how effectively uncertainties order queries by expected error. The authors report mean calibration results across the panels and splits. For risk-ranking, they evaluated metrics that quantify how much mean absolute error (MAE) can be reduced by selecting low-risk subsets of queries; one reported metric is R50, the percentage reduction in MAE after retaining the lowest-risk half of queries.
Among the evaluated methods, adaptive deep kernel fitting (ADKF) achieved the best mean calibration results across the benchmark. Conversely, deep kernel transfer (DKT) was more effective at ranking prediction errors than ADKF when considering the predictors' native uncertainty. These findings separate two desirable properties: calibration (ADKF strength) and ranking ability (DKT strength under native uncertainty).
The authors introduce Influence Calibrated Support Reconstruction (ICSR) as a new risk score tailored for kernel-based predictors. ICSR combines measured support reconstruction errors with query-specific influence, and then stabilizes the weighted average of these quantities by shrinking toward the full-support mean. The method is intended to adjust Gaussian process uncertainty estimates without changing the underlying predicted activities, thereby producing a more reliable risk ordering for selective decision-making.
Applying ICSR with both DKT and ADKF backbones improved all three reported mean risk-ranking metrics compared with native uncertainty and an adapted neighborhood comparator across every assay panel and split. ICSR produced a mean half-query MAE reduction (R50) in the range of 14.2–19.6%, indicating that retaining the lowest-risk half of queries as ranked by ICSR reduced MAE by that percentage on average. The benchmark underscores the value of jointly assessing calibration and error ranking when judging uncertainty strategies for bioactivity prediction.
A retrospective case focusing on the SARS-CoV-2 main protease was included to demonstrate practical utility: ICSR was used to select more reliable predictions in that example. The source text reports this illustrative case but does not provide full experimental or numerical detail in the summary presented here.
The study supports the view that uncertainty estimation should be evaluated both for calibration and for its ability to prioritize low-error predictions. For kernel-based predictors, ICSR is presented as a post-processing risk-scoring strategy that can improve selective use of predictions by better ordering queries for experimental follow-up while leaving predicted activities unchanged.
The abstract and source summary report overall comparative outcomes, the set of datasets, splits, uncertainty methods, and the new ICSR approach with R50 improvements. However, full methodological details, model hyperparameters, exact implementations of the five multi-task predictors, precise definitions of the three mean risk-ranking metrics beyond R50, and per-assay numerical breakdowns were not reported in the source content provided here.
This benchmark study contrasts calibration and ranking performance of multi-task bioactivity predictors and uncertainty estimates on ChEMBL-derived assay panels. It finds a trade-off between calibration (best for ADKF) and ranking (stronger for DKT with native uncertainty), and proposes ICSR to improve risk ranking for kernel-based predictors. Reported improvements include a mean R50 gain of 14.2–19.6%, and a retrospective SARS-CoV-2 example illustrates potential application for prioritizing more reliable predictions.