Virtual screening is routinely used in early-stage drug discovery to prioritize candidate molecules from large libraries. Its practical value is primarily in reliably placing likely actives at the top of ranked lists when only a limited number of compounds can be experimentally tested. The authors frame method comparison as a calibration-aware top-k decision problem, arguing that standard aggregate metrics may obscure target-level heterogeneity and may not reflect decision quality under fixed experimental budgets.
This study asks whether a complete screening package, TASC-VS, which integrates test-time adaptation (TTA), calibration, and uncertainty-aware ranking, yields a reproducible top-k decision advantage within a controlled benchmark comparison.
The evaluation was designed as a frozen-benchmark, deployment-oriented comparison. All methods were run against the same curated subset of LIT-PCBA targets using an auditable workflow that links headline benchmark performance to fixed-budget decision endpoints. The study explicitly defines multi-layered evaluation goals to capture headline summaries, paired target-level effects, and concrete top-k decision value.
The analysis used a curated subset of 15 LIT-PCBA targets selected because required inputs were available for TASC-VS and comparators (target metadata, fold assignments, receptor preparations, ligand labels, Vina poses, RTMScore outputs, and TASC-VS input features). The curated subset contains 23,250 evaluated molecule-target records and 181 positives. By contrast, the full LIT-PCBA benchmark comprises 415,225 molecules and 7,844 actives; therefore the subset represents about 5.6% of molecules and 2.3% of actives from the published benchmark. Positives are present in seven of the 15 targets.
The 15 evaluated targets were ALDH1A1, MTORC1, VDR, ESR1_ANTAGONIST, FEN1, OPRK1, GBA, MAPK1, PPARG, ADRB2, KAT2A, PKM2, ESR1_AGONIST, IDH1, and TP53. Target-specific counts for total/positive were provided in the source and guided subsequent analyses.
The authors established a multi-layer evaluation framework that includes:
Top-k oriented endpoints reported include normalized enrichment factor at 1% (NEF1), Boltzmann-enhanced discrimination of ROC (BEDROC), precision among the top 20 candidates (P20), and false positives among the top 20 (FP20).
TASC-VS full is a pipeline combining three elements: test-time adaptation (TTA), post-hoc calibration, and uncertainty-aware ranking/penalization. Comparators included classical docking (Vina), rescoring frameworks such as RTMScore applied to Vina poses (RTMScore-on-Vina-poses), and other baseline methods used for controlled comparison. Component analyses separately assessed the contribution of calibration, uncertainty penalization, and TTA to observed effects.
On the seven targets that contained positives, TASC-VS full achieved NEF1 = 0.052 and BEDROC = 0.062. By comparison, RTMScore-on-Vina-poses achieved NEF1 = 0.0057 and BEDROC = 0.0279 on the same positive-target subset. Across all 15 targets, TASC-VS full had P20 = 0.0433 and FP20 = 19.13, whereas RTMScore-on-Vina-poses had P20 = 0.0067 and FP20 = 19.87.
These reported values indicate modest, directionally consistent benefits for TASC-VS full on the curated subset, particularly on top-k focused endpoints.
Component-level experiments showed that post-hoc calibration improved reliability metrics in the pipeline. Uncertainty penalization contributed a small benefit to top-20 selection. Test-time adaptation (TTA) altered reliability measures slightly but was not the principal driver of the main ranking gains achieved by the full package. The source reports these component findings qualitatively and via the comparative endpoint values above.
A paired P20 analysis comparing TASC-VS full against RTMScore-on-Vina-poses produced a mean benefit of 0.0367 with a 95% bootstrap confidence interval of 0.0067–0.0733. The raw (unadjusted) Wilcoxon test p-value was 0.026. However, multiplicity-adjusted tests were not significant: Holm-adjusted p = 1.0 and Benjamini–Hochberg q = 0.322. These results indicate that target-paired gains were present but not robust to multiple-testing correction in the evaluated comparison set.
An external stress test using Directory of Useful Decoys, Enhanced (DUD-E) data favored Vina over TASC-VS full. The authors interpret this as evidence that the observed advantage of TASC-VS full is sensitive to dataset composition and that performance on one curated benchmark subset does not guarantee broader superiority.
The study reframes virtual-screening method comparison around decision value under fixed experimental budgets rather than only global ranking statistics. On the curated 15-target LIT-PCBA subset, TASC-VS full showed modest, consistent top-k advantages and improved reliability after calibration. Component analyses suggest calibration and uncertainty handling contributed most to decision-oriented gains, while TTA had a smaller role.
However, the limited coverage of the full LIT-PCBA benchmark in the curated subset, the presence of targets with zero positives, and the unfavorable DUD-E stress-test result highlight dataset- and benchmark-specific sensitivity. The authors therefore caution that broader deployment of TASC-VS full requires separate validation on other targets and representative datasets.
TASC-VS full combines calibration, uncertainty-aware ranking, and TTA to provide calibration-aware top-k decision support within the evaluated curated LIT-PCBA subset. The package delivered modest top-k benefits versus selected comparators on this frozen benchmark, but external stress testing and multiple-testing corrections reduce confidence in generalizability. The authors recommend additional validation before broader deployment.
The analysis used the LIT-PCBA benchmark dataset available from the published LIT-PCBA resource (DOI provided in the source). The authors note that the curated subset evaluated here is not the full LIT-PCBA benchmark and explicitly state the limited subset coverage when interpreting results. Further validation on larger or different benchmarks is necessary to establish general applicability.