Comparative evaluations of trajectory inference (TI) methods typically assess full pipelines, obscuring how much geometric distortion arises specifically from dimensionality reduction (DR). The authors note that no prior study has systematically quantified how well DR methods alone preserve a known reference path when projecting high-dimensional single-cell data to two dimensions. They also identify an absence of dedicated metrics to measure path-preservation quality after DR. Because DR is a universal preprocessing step that shapes downstream TI results, its independent geometric effects merit direct evaluation.
Two single-cell datasets with known reference trajectories were used. The primary linear dataset is a CD4+ T‑cell surface-protein dataset comprising 3,096 cells and 51 proteins. A ground-truth linear reference path was constructed by selecting cells lying close to the first principal component (PC1) within a single cluster, yielding a known linear trajectory in the original high-dimensional protein-expression space. A second dataset contained a closed B‑cell cell‑cycle loop detected by persistent homology in a separate CyTOF dataset, providing a topologically distinct cyclic reference.
Sixteen DR methods were applied to both datasets. The manuscript reports method names and relative performance but does not invent additional method details. A panel of twelve geometric path-preservation metrics was computed for each DR result. These metrics spanned several categories: log-ratio distortions of length, curvature, and spatial similarity; Spearman rank correlations of pairwise distances and of segment lengths; and structural complexity measures such as self-intersection frequency and coiling. One metric, named SpatDistSpear, was identified within the panel as the single most discriminating metric in separating method fidelity groups.
To evaluate sensitivity to the number of points used to define the reference path, the authors varied the fraction of cells included in the path from 1% to 10% of the dataset (31–310 path points in the primary dataset) and recomputed the metrics and method rankings at each threshold. The same eleven density levels were tested on the cyclic reference to assess whether conclusions about metric and method stability generalize to a closed-loop topology. The goal was to determine whether changing the path-density threshold alters metric values, method rankings, or the overall conclusions of the benchmark.
On the linear PC1 reference, absolute values of all twelve metrics changed smoothly as the path-density threshold increased from 1% to 10%, consistent with the expected broadening of the reference band when more points are included. Despite these shifts in absolute metric values, composite method rankings aggregated across all twelve metrics remained stable across every density threshold tested: no method moved between performance tiers as the hyperparameter varied. A composite rank identified a consistent set of high-performing methods across densities (UMAP, MDS, CNPE, TSNE, SPE, LPMIP) and a consistent set of low-performing methods (SPMDS, LPP, DVE, LAPEIG, PHATE).
When considered alone, the SpatDistSpear metric discriminated methods into fidelity groups that did not fully match the composite ranking. SpatDistSpear separated a high-fidelity group (LPMIP, DM, MDS, SPMDS, DVE, CISOMAP, CNPE; all r > 0.80) from a mid-range group (LAPEIG, SPE, PHATE, UMAP, TSNE, FOSMOD) and a low-fidelity group (PFA, NNP, LPP). This pattern demonstrates that global distance preservation metrics and composite performance do not always agree on the same top tier of methods. Crucially, this disagreement in which metric identifies the best methods was density-independent rather than an artifact of the chosen path-density threshold.
The closed loop reproduced the same density-independence observed for the linear reference: although absolute metric values drifted with per-segment band width as the number of path points changed, composite method ranks again held constant across all eleven density levels. The identity of best and worst performers was largely, though not entirely, conserved between the two topologies. CNPE, LPMIP, SPE, and MDS performed well on both the linear path and the closed loop, while LAPEIG, DVE, and SPMDS performed poorly on both. By contrast, UMAP and TSNE were top performers on the linear path but dropped to the middle of the sixteen-method panel on the cyclic loop rather than to the worst tier. The authors interpret this as topology-dependence of DR benchmark rankings: a method's rank can depend on the shape of the reference trajectory rather than on the path-density threshold.
This work introduces a pipeline-independent evaluation focused directly on how DR methods distort trajectory geometry, filling an absent benchmarking dimension in existing TI comparisons. The main empirical finding is within-dataset rank stability: when a fixed reference-path threshold is chosen, composite DR method rankings remain robust to changes in the number of points used to define the path, for both linear and cyclic trajectories. Practically, this validates using a fixed path-density threshold as a reliable operating point in large-scale DR benchmarking. However, the partial reordering of top performers across different trajectory topologies indicates that DR benchmark rankings are trajectory-shape-dependent and should not be assumed to transfer from one geometry (e.g., linear) to another (e.g., cyclic). The authors declared no competing interests. The paper provides no additional procedural details beyond those summarized here.