Comparative genomics and pangenomics study related biological processes at different scales and with different data types: comparative genomics focuses on evolution between species (phylogenetics), while population genetics and pangenomics address evolution within species. Although they share fundamental mechanisms, their joint analysis has been limited by methodological and representational differences. The authors frame this gap around two active subfields: genome rearrangement studies in comparative genomics and graphical pangenomes in pangenomics. They argue that despite long-standing theoretical results in rearrangement models, these results have had limited impact on pangenomic analyses because of mismatches in problem formulations and scalability.
Classical rearrangement formulations inherit assumptions tailored to small sets of genomes or to phylogenetic trees. The authors note that several common assumptions in rearrangement problems — for example, the presence of a single underlying tree relating all genomes — can be inadequate when applied to many pangenomes. Graphical pangenome representations, which may encode complex variant structures across many individuals, do not always fit neatly into these traditional models. As a result, directly applying theoretical rearrangement results to pangenomic data can produce misleading or inapplicable outcomes.
On the practical side, modern pangenomes often contain very large numbers of individual genomes. The authors highlight that this scale makes many classical computational problems infeasible: NP-hard parsimony problems and exhaustive all-vs-all rearrangement-distance computations become impractical as pangenome size increases. These scaling barriers limit the use of exact rearrangement-based measures and motivate the need for new problem definitions and algorithms that account for the data volume and the graph-like structures used in pangenomics.
To address both theoretical and practical shortcomings, the authors propose the Complete Ancestral Reconstruction for Pangenomes (CARP) problem. CARP is intended as a problem formulation tailored to pangenomic settings while maintaining intuitive relationships to classical rearrangement problems. The formulation is described as overcoming limitations of prior models and as aligning more closely with the data structures used in graphical pangenomes.
The source text does not provide the formal definition of CARP, algorithmic approaches, complexity results, or experimental evaluation within the accessible abstract and front matter. Therefore, specific technical details of the CARP formulation, such as input specification, objective functions, constraints, or solution strategies, were not reported in the available source and cannot be summarized here.
The authors emphasize that CARP preserves intuitive connections to both established rearrangement problems and to the representation of variants in pangenome graphs. This suggests CARP is positioned as a bridge: it aims to make theoretical insights from rearrangement studies applicable to pangenomic analyses, and conversely to adapt pangenome representations so they can be analyzed with rearrangement-oriented concepts. The work highlights the conceptual similarity of central data structures used in both fields and seeks a unified setting in which within- and between-species evolutionary questions can be studied together.
The proposal of CARP addresses an important methodological gap by offering a problem formulation designed for the complexity and representation of modern pangenomes. If adopted and further developed, CARP could enable more meaningful measures of rearrangement complexity across large pangenomic datasets and foster cross-fertilization between comparative genomics theory and pangenomics practice.
Because the available source material is limited to the preprint title, author information, abstract, and front matter, the article’s detailed methods, proofs, algorithmic formulations, computational experiments, and empirical results are not present in the provided text. The authors’ formal definition of CARP, analyses of computational complexity, algorithmic solutions, benchmarks on pangenome data, and case studies were not reported in the excerpt; readers should consult the full preprint for those technical details.
Competing interests
The authors have declared no competing interest.
Note on the source
This summary is based solely on the bioRxiv preprint metadata and abstract provided. The preprint has not been peer reviewed, and the full technical content is available in the preprint PDF and full text for readers requiring complete definitions and results.