The authors sequenced seven Indonesian rice cultivars using PacBio HiFi technology, generating a total of 95.40 Gb of sequence data. Cultivar-specific consensus genomes were produced using a reference-guided approach anchored to the telomere-to-telomere Nipponbare assembly AGIS1.0. Per-cultivar sequencing coverage ranged from 27.92x to 41.58x. The resulting consensus assemblies spanned between 387.93 Mb and 390.54 Mb.
This reference-guided strategy standardized the genomic coordinates across cultivars by aligning and constructing consensus sequences relative to AGIS1.0, enabling direct, locus-level comparisons while preserving the reference coordinate framework.
Consensus genomes derived from the HiFi data exhibited high gene-space completeness as measured by BUSCO, with completeness values of approximately 98.3% to 98.5%. These BUSCO scores indicate that the assemblies capture the vast majority of conserved single-copy orthologs expected in rice, supporting downstream comparative analyses focused on predicted coding sequences and promoters.
Assembly sizes clustered within a narrow range (387.93–390.54 Mb), consistent with a reference-guided consensus approach that constrains overall genome span relative to the AGIS1.0 template.
Predicted proteins from the seven cultivar consensus genomes were analyzed with OrthoFinder, which assigned 99.1% of predicted proteins to 40,737 orthogroups. Of these orthogroups, 27,514 were identified as core orthogroups represented across all seven cultivars. This result indicates a highly conserved predicted gene content within the reference-guided assemblies, emphasizing that most protein-coding sequences are shared across the sampled Indonesian cultivars when evaluated in the AGIS1.0 coordinate system.
The high proportion of proteins assigned to orthogroups facilitates cross-cultivar comparisons at the gene-family and orthologous-gene level and supports identification of both conserved and variable loci relevant to agronomic traits.
The study performed a targeted locus-level assessment of genes and gene-family entries previously associated with key rice traits. Across 280 expected cultivar-by-locus combinations (derived from 40 genes or gene-family entries related to grain pigmentation, nitrogen and amino-acid metabolism, and starch properties), the analysis recovered 278 combinations.
This near-complete recovery demonstrates that the reference-guided consensus genomes reliably represent known trait-associated loci in these cultivars under the AGIS1.0 framework. The recovered loci cover traits central to grain quality and nutrient-related metabolism.
Comparative analysis of predicted proteins across the seven consensus genomes prioritized several candidate genes for further functional investigation. Noted candidates included ANS1, SBE2b, SSIIa/ALK, Wx/GBSSI, OsAAP6/qPC1, and SSI. These genes are implicated in processes such as pigment biosynthesis and starch metabolism, which are directly relevant to grain pigmentation and cooking/texture properties.
The study frames these gene candidates as testable priorities for downstream validation studies and for potential incorporation into genomics-assisted breeding strategies targeting grain quality and related traits.
Promoter sequences were compared in an AGIS1.0-anchored manner for loci with suitable anchoring. Of 269 completed promoter comparisons anchored to AGIS1.0, 159 passed predefined quality-control criteria. The remaining 110 comparisons were flagged due to concerns including gene-model discrepancies, boundary inconsistencies, synteny breaks, or structural issues in the alignments or assemblies.
Importantly, the set of flagged promoter comparisons accounted for more than 90% of the alignment-derived sequence variation observed across promoter comparisons. This concentration of variation within comparisons that failed quality control highlights the potential for assembly, annotation, or structural artifacts to inflate apparent regulatory divergence if rigorous QC is not applied.
The authors emphasize that careful evaluation of gene models, synteny, and structural integrity is necessary before interpreting promoter differences as functionally meaningful regulatory variation.
Collectively, these reference-guided genomic resources create a standardized framework to assess sequence variation in Indonesian rice germplasm relative to the AGIS1.0 reference. The high BUSCO completeness, extensive orthogroup representation, and near-complete recovery of targeted trait loci indicate that the consensus genomes are suitable for comparative analyses focused on coding sequences and promoter neighborhoods.
By prioritizing coding and regulatory candidate loci—while documenting promoter comparisons that require caution due to quality concerns—the work provides a prioritized set of targets for functional validation and for integration into genomics-assisted breeding pipelines.
Limitations and reported constraints: the source reports the sequencing coverage, assembly spans, BUSCO completeness, orthogroup counts, recovery of cultivar-by-locus combinations, prioritized candidate genes, and the promoter QC outcomes. No additional experimental results, functional validations, or downstream phenotypic correlations were reported in the source document.
These reference-guided consensus genomes and the associated locus-level comparisons are positioned as community resources to facilitate future functional genomics, validation experiments, and crop-improvement efforts focused on Indonesian rice germplasm.