This study is a retrospective longitudinal investigation conducted in Quebec over the period 1997–2027 to evaluate how well existing medico-administrative algorithms classify diabetes phenotypes. The principal aim in Phase 1 is to assess the diagnostic performance of the Corsenac et al. (2022) algorithms for four diabetes categories: type 1 diabetes (T1D), type 2 diabetes (T2D), latent autoimmune diabetes in adults (LADA) and other specific diabetes types. Performance metrics will be estimated separately for each phenotype using independent reference subsamples drawn from a first cohort (A1–A2–A3; n = 5,200).
The reference cohort is partitioned into three subsamples designed to capture diversity and realistic phenotype proportions while providing sufficient statistical power for subgroup and population analyses. The subsamples are described as:
Each subsample will act as an independent reference against which algorithm classifications are compared, enabling calculation of sensitivity, specificity, positive and negative predictive values, and other diagnostic performance metrics as appropriate. The subsample structure intends to reflect both the relative rarity of some phenotypes (notably LADA and certain specific types) and the commonness of T2D.
All records from subsamples A1, A2 and A3 will be probabilistically linked with medico-administrative and pharmaceutical claims held by the Régie de l'assurance maladie du Québec (RAMQ). The probabilistic linkage will be performed by the Institut de la statistique du Québec (ISQ). Linked data will provide administrative diagnostic codes, health service use, and prescription drug claims that are necessary for algorithm evaluation and for developing refined classification rules.
Following evaluation of the Corsenac et al. algorithms against the reference subsamples, the study will apply machine learning methods to refine algorithmic definitions for the four phenotypes. The plan is to use linked reference data to train and validate algorithmic refinements that better discriminate among T1D, T2D, LADA and other specific types in medico-administrative datasets. The source describes the intention to employ machine learning but does not specify the exact models, hyperparameters, feature lists or training procedures; those details were not reported in the source.
In Phase 2, the refined algorithms produced in Phase 1 will be applied to a second medico-administrative population-based cohort (cohort C; n = 50,000). The objective is to generate the first simultaneous population-level estimates of prevalence and incidence for the four diabetes phenotypes in Quebec using administrative data. These frequency estimates will be derived from the algorithmic classifications applied to cohort C and will be adjusted to represent the general population using reweighting techniques described below.
Analyses are restricted to individuals continuously covered by RAMQ's public drug insurance because pharmaceutical claims are integral to algorithm inputs and phenotype differentiation. RAMQ's public drug insurance covers approximately 46% of the Quebec population; therefore, direct analyses will use this subset of records. The study protocol describes planned reweighting to account for the partial coverage and to make results more representative of the wider Quebec population.
To address the fact that the analytic sample (continuous RAMQ public drug insurance enrollees) does not encompass the full Quebec population, the investigators will apply two reweighting strategies:
These approaches are described at a high level in the protocol; specific variable sets used for calibration or the models for inverse probability weighting were not reported in the source.
Feasibility of the study was approved by the various ethics review boards of partner institutions. The protocol is registered on ClinicalTrials.gov under identifier NCT06573905. The source states that results will be disseminated through scientific and public health channels. No specific timelines, publication targets, or dissemination platforms beyond these general plans were reported in the source.
Limitations noted or implicit in the protocol include reliance on RAMQ public drug insurance coverage (covering about 46% of the province), use of self-report and physician-confirmed diagnoses as reference sources (each with inherent limitations), and the absence in the source of detailed machine learning model specifications and reweighting variable lists. These details were not reported in the source article.