The authors developed a Graph Representation Specification to make extraction of computational graphs from clinical guideline decision algorithms reproducible. The specification includes an ontology, a motif catalogue, disambiguation conventions, decomposition rules, a deterministic validator, and a scoring engine. These components were designed to constrain representational choices that otherwise produce divergent graphs when coding guideline algorithms, so that any complexity metric derived from those graphs is not dependent on idiosyncratic coding.
Translating guideline decision procedures into computational graphs requires judgment. When multiple coders convert the same guideline into a graph without shared constraints, the resulting graphs can differ and any complexity measure derived from those graphs will inherit that variation. The objective was to develop and prospectively test an empirical method to make graph extraction reproducible, using the Clinical Guideline Complexity Index (CGCI) and four guideline algorithms as a case study. The authors emphasize establishing measurement reliability rather than validating the CGCI as a construct.
The team used an iterative, error-driven grammar induction approach. Steps were: measure inter-coder disagreement on extracted graphs; localize the dominant class of disagreement; induce a single grammar rule intended to resolve that class of disagreement; and prospectively test whether applying that rule improves coder agreement in the targeted class on fresh coding pairs.
To quantify reproducibility they pre-specified a topology-based endpoint named Decision Topology Agreement rather than relying on edge agreement. The authors argued that edge agreement is oversensitive to representational choices that do not affect the score, while a topology-based measure better captures agreement relevant to CGCI computation.
Two trained coders independently coded four guideline algorithms: diabetes, dyslipidemia, heart failure, and hypertension. The specification and the induced rules were refined based on observed disagreements and then tested prospectively.
The four guideline algorithms served as the test corpus for grammar induction and prospective validation. Coding focused on motifs and assessment topologies encountered in the algorithms; one specific source of induced rules arose from the diabetes comorbidity panel. Coders applied the specification and, after induction of a rule, independently coded fresh instances to test predicted improvements in agreement.
A grammar rule induced from the diabetes comorbidity panel (an assessment topology) produced a pre-specified prediction that figures in the heart-failure guideline, sharing the same motif, would converge when the rule was applied. On a newly and independently coded pair of heart-failure figures, the prediction was confirmed: the resultant absolute CGCI difference was approximately one.
Decision topology reproduced closely overall. The authors report decision-order agreement at or near 1.00 for three of the four guidelines, indicating high concordance in the decision sequence representation under the specification and applied rules.
The study identified that breadth counting—the way nodes or modifiers that expand decision breadth are tallied—was sensitive to representational choices. Introducing an explicit modifier-counting rule reduced the largest observed disagreement from 27 tokens to 4 tokens. Despite these improvements, some residual disagreement remained, but it was bounded and could be localized to specific and nameable representational choices. This localization makes further iterative grammar refinement feasible and targeted.
The authors conclude that graph-extraction reproducibility can be systematically improved through iterative grammar refinement, and that a prospectively derived rule can be confirmed to improve agreement in anticipated classes of representational motifs. These results establish the measurement foundation—specifically reliability—for a companion study that will interpret the Clinical Guideline Complexity Index (CGCI) as reflective of cognitive load. The authors note that they have established reliability rather than construct validity.
Finally, the authors suggest the method may apply to other contexts where graphs are extracted from structured source artifacts. Remaining disagreements are bounded and attributable to definable representational choices, which supports continued refinement of the specification. The article is a preprint and has not undergone peer review; details beyond those reported in the preprint were not provided in the source.