In-beam positron emission tomography (PET) is a promising modality for monitoring dose delivery during carbon ion radiotherapy (CIRT), but converting measured positron-emitter activity to physical dose is challenging because the relationship is complex and nonlinear. This proof-of-concept study aimed to improve activity-to-dose mapping by developing decomposition-based deep learning frameworks that incorporate auxiliary physical supervision to guide intermediate representations toward physically meaningful signals.
The authors used idealized Monte Carlo (MC) simulations to generate training and validation data. Simulations were performed on computed tomography (CT) phantoms derived from 18 non-small cell lung cancer patients. The MC output provided ground-truth depth-dose distributions and positron-emitter activity distributions used to evaluate model performance.
Models were designed to predict laterally integrated one-dimensional depth-dose distributions for individual pencil-beam spots. Input features included the 5-minute cumulative positron-emitter activity profile following irradiation and CT Hounsfield unit (HU) profiles. The laterally integrated depth-dose was the primary prediction target compared against MC ground truth.
Three neural-network approaches were developed and compared:
DirectNet: a baseline model that maps the input activity and CT HU profiles directly to depth-dose without a decomposition pathway.
TemcoNet: a decomposition-based model that integrates a Transformer-based decomposition module. This model was supervised by cumulative post-irradiation activity at later time points (10, 15, and 20 minutes) as auxiliary signals to shape intermediate representations.
NucoNet: another decomposition-based model with a Transformer decomposition module supervised by nuclide-specific yields for key emitters (11C, 15O, and 10C). The nuclide-yield supervision provided a physically motivated constraint on the decomposition outputs.
Both TemcoNet and NucoNet used decomposition pathways intended to separate intermediate components that relate more directly to the temporal and nuclide-dependent behavior of post-irradiation activity.
The decomposition-based models received additional supervisory signals beyond the final dose target. TemcoNet used cumulative activity at later post-irradiation times (10, 15, 20 minutes) to inform temporal aspects of the activity-to-dose mapping. NucoNet used nuclide-yield information for 11C, 15O, and 10C to encourage the network to learn decomposed components corresponding to distinct nuclide contributions. DirectNet did not include these auxiliary physics-informed supervision pathways and served as a baseline for comparison.
When evaluated against MC ground truth, all models produced similar median range accuracy for the predicted depth-dose distributions. However, the two decomposition-based models produced substantially improved overall dose predictions compared with the baseline. Specifically, the mean relative error of dose prediction decreased from 2.36% for DirectNet to below 0.4% for both TemcoNet and NucoNet. In addition, the mean gamma passing rate at the 2 mm/2% criterion increased from 45.31% with the baseline model to approximately 96% for each decomposition-based model. These improvements indicate a large increase in agreement with MC-derived dose distributions when decomposition and physics-informed supervision are included.
Ablation experiments were performed to probe the role of the decomposition pathway and the auxiliary supervision. Results indicated that the decomposition pathway learned intermediate representations that were physically meaningful, rather than arbitrary latent features. Furthermore, inclusion of nuclide-yield supervision provided an additional benefit in dose prediction over temporal-activity supervision alone, supporting the hypothesis that supervising components tied to known physical emitters improves mapping fidelity.
This proof-of-concept work demonstrates that physics-informed, decomposition-based deep learning can substantially improve Monte Carlo–derived activity-to-dose mapping for in-beam PET monitoring in carbon ion radiotherapy. By combining Transformer-based decomposition modules with auxiliary supervision drawn from post-irradiation activity time points and nuclide yields, the models markedly reduced mean relative dose error and greatly increased gamma passing rates versus a direct-mapping baseline.
Limitations reported in the source include that the study used idealized MC simulations on CT phantoms from 18 patients; details on experimental validation, real-world measurement noise, clinical implementation, and broader generalizability were not reported in the abstract. As a proof-of-concept, the findings support further investigation and experimental validation to determine translation to clinical in-beam PET systems and treatment verification workflows.
Physics-informed decomposition networks, exemplified by TemcoNet and NucoNet, improved activity-to-dose mapping quality in MC-simulated CIRT scenarios compared with a direct-mapping baseline. Auxiliary supervision tied to temporal activity and nuclide-specific yields enabled the networks to learn physically interpretable intermediate representations and substantially reduce dose-prediction errors. Further work is needed to validate these methods under experimental conditions and to assess clinical applicability.