Enzyme turnover numbers (kcat) are essential quantitative parameters for kinetic models and enzyme-constrained genome-scale metabolic models (ecGEMs). Because measured kcat values are sparse in public databases, researchers increasingly turn to machine learning (ML) methods to estimate turnover numbers. Standard assessments of these predictors typically report global regression metrics on benchmark datasets, but the practical value of predicted kcat values depends on how prediction errors propagate through downstream mechanistic models such as ecGEMs.
The study summarized here benchmarks multiple contemporary ML-based kcat predictors against curated datasets, assesses training-set proximity, and evaluates how predicted kcat values affect ecGEM-derived growth predictions for Saccharomyces cerevisiae across multiple conditions. The central claim is that benchmark accuracy alone is insufficient to guarantee downstream usefulness; system-level, application-driven validation is required.
The authors benchmarked six current kcat predictors on a curated BRENDA-derived dataset and five predictors on an independent EnzyExtract dataset. On the BRENDA-derived benchmark the predictors achieved only moderate regression accuracy. Performance declined sharply on the EnzyExtract dataset: all predictors achieved R2 values of 0.20 or lower on EnzyExtract.
This marked drop in measured accuracy between benchmark datasets shows that predictor performance evaluated by conventional global metrics can be sensitive to the choice of benchmark and to differences between benchmark and training data.
To understand generalization, the authors compared each benchmark dataset with the available training data used by each predictor. They quantified exact sequence match overlap between benchmark entries and predictor training sets. For the BRENDA-derived dataset, overlap ranged from 24% to 78% exact sequence matches across predictors. For EnzyExtract the overlap was substantially lower, ranging from 9% to 26%.
While lower overlap correlated with poorer benchmark performance on EnzyExtract, overlap alone did not fully explain differences in generalization among predictors. Some predictors with similar training-set proximity still differed in how well their predictions extended to the independent dataset, indicating other determinants of out-of-sample behavior.
Predicted kcat values from the ML tools were used to parameterize ecGEMs of Saccharomyces cerevisiae. The authors evaluated predicted growth across 19 conditions and compared model outputs to experimentally observed growth variation. Across these conditions, none of the tool-specific ecGEMs consistently reproduced the experimental patterns in growth.
Importantly, the downstream performance of ecGEMs did not align with benchmark ranking. In glucose minimal medium the predictor that performed worst on standard benchmarks produced the most accurate growth prediction in the ecGEM context, whereas predictors with higher benchmark ranks yielded larger deviations in predicted growth.
This discordance demonstrates that global regression metrics and benchmark ranks are not reliable proxies for downstream utility in mechanistic models.
The authors traced the mismatch between benchmark accuracy and downstream performance to localized, high-leverage errors in predicted kcat values. Specifically, underpredicted turnover numbers for the mitochondrial ADP/ATP carrier limited adenine nucleotide exchange in yeast models. This underprediction imposed an apparent constraint on cytosolic ATP supply, which in turn altered growth predictions and the model-inferred phenotype.
When the researchers relaxed the constraint on the carrier's kcat within the ecGEMs, predicted growth shifted toward the experimental reference. This case illustrates that ML-derived kcat errors can affect not just quantitative outputs (growth rates) but also which phenotype or limiting process a mechanistic model appears to identify.
The findings argue that ML predictors of biological parameters like kcat should be validated within the downstream systems they are intended to support. Standard global benchmarking can mask localized prediction errors at network positions with outsized influence on model behavior. Application-driven validation — testing predicted parameters in situ in representative mechanistic models and conditions — is necessary to detect such high-leverage mispredictions.
Predictor developers should consider reporting not only global regression metrics but also assessments of training-set overlap and robustness tests that reflect downstream use cases. Modelers using ML-derived parameters should perform sensitivity analyses and targeted checks for high-leverage reactions (for example carriers or transporters) that can disproportionately affect system-level outputs.
The study used curated datasets derived from BRENDA and an independent EnzyExtract dataset for benchmarking. Predicted kcat values were integrated into ecGEMs of Saccharomyces cerevisiae and growth evaluated across 19 experimental conditions. The authors provide supporting data and code in linked repositories. Specific methodological details, code, and data are available at the project repositories cited by the authors.
The authors declare no competing interests.
Benchmark accuracy of ML-based kcat predictors does not guarantee accurate or reliable downstream behavior when those predictions parameterize mechanistic models. Generalization declines on independent datasets with lower training-set overlap, but training-set proximity does not fully explain differences among predictors. Localized, high-leverage prediction errors — exemplified by the mitochondrial ADP/ATP carrier in yeast — can distort growth predictions and the mechanistic interpretation of phenotype. The authors recommend application-driven validation of biological parameter predictors in the systems they are meant to support, supplementing conventional benchmark metrics.