Data-extraction forms are routinely piloted in systematic reviews, but extant guidance and error-studies address primarily numerical data. When the unit of interest is a verbatim term transcribed from source text—as required by evidence maps that inventory terminology—established agreement measures and extraction conventions are lacking. This study piloted a form intended to extract operative-step terms from the surgical literature on deep endometriosis as a preparatory exercise for Phase 1 of the ATLAS-ONTO evidence-mapping project.
The pilot aimed to determine whether the extraction form produced acceptable inter-reviewer agreement using a set-overlap metric, and to identify the mechanisms of discordance that would require revision before larger-scale use.
Design and reporting: A blinded inter-reviewer agreement study was conducted and reported according to GRRAS. Two reviewers independently extracted operative-step terms from a purposive sample of 11 heterogeneous articles in English and French. Eight articles were ultimately extracted and seven contained a defined index for the target operative step.
Primary and secondary measures: The primary outcome was the mean per-article Jaccard index of normalized term sets, with an explicit rule defined for empty sets. Secondary metrics were Cohen’s kappa calculated for categorical attributes among exactly matching terms and simple agreement for anatomical-structure coding.
Pre-specified thresholds: Thresholds for acceptable agreement were fixed before extraction: Jaccard ≥ 0.70; kappa ≥ 0.60; simple agreement ≥ 0.80.
Classification and harmonization: Unmatched terms were classified retrospectively by mechanism. Following an analysis that identified discordance mechanisms and record-integrity defects, the same articles were reassessed after a harmonization session in which extraction conventions were agreed and a field (start-definition) was split.
Additional procedural notes: Two record-integrity defects discovered during analysis were documented together with their effect on the dataset and estimates. The pilot’s extraction forms (initial and revised), extraction manual with change log, datasets for every round, analysis scripts with R version and locale guard, and a Python cross-check were deposited in the Open Science Framework project associated with the study.
Sample and extractions: Eight articles were extracted; seven had a defined index term for the operative step under study. Across the dataset, 92 terms were recorded by the two reviewers.
Primary outcome: The extraction form did not meet the pre-specified primary threshold. The mean per-article Jaccard index was 0.581, and the pooled match proportion was 55/92 = 0.598.
Exact matches and discordances: Of the 92 terms, 55 matched exactly between reviewers and 37 did not. The unmatched terms were retrospectively attributed to four mechanisms: limited source coverage (21 instances), disagreements in span extent of extracted text (8 instances), comprehension of source language (7 instances), and a single residual discrepancy.
Secondary outcomes: Conditional on the 55 exactly matching terms, Cohen’s kappa for some categorical attributes was high (for example, end definition kappa = 0.930; laterality kappa = 0.781). Anatomical-structure coding simple agreement was 0.873. These conditional high kappas indicate that measures limited to exactly matching items did not detect the primary failure of set overlap.
Corrections and harmonization: Two record-integrity defects identified during the analysis were described and their effects reported; details of these defects were included in the project’s public materials. The team implemented four written extraction conventions and split the start-definition field to reduce ambiguities. After a harmonization session and reassessment on the same articles, the Jaccard index rose to 0.913. The authors note that this post-harmonization increase represents convergence driven by harmonization rather than independent reproducibility.
The pilot demonstrates that conventional agreement measures applied only to matched items (for example, kappa) can give a misleadingly optimistic impression of reproducibility when the extraction task is to capture sets of verbatim terms. A set-overlap measure such as the Jaccard index is more appropriate for initial piloting of forms that target terminology or free-text phrases because it reflects both presence/absence and overlap of extracted items across reviewers.
The primary mechanisms of discordance identified—source coverage differences, span extent disagreement, and language comprehension—are intrinsic to text-extraction tasks and are not specific to surgery or to deep endometriosis literature. Addressing these mechanisms requires clear articulation of an admissible source for extraction, explicit rules for text-span boundaries, and conventions for handling multilingual sources.
Harmonization can markedly increase agreement on the same material, but such convergence should not be conflated with independent reproducibility. The authors therefore recommend piloting with prespecified thresholds and metrics and then testing the revised form on new material to assess reproducibility independently.
For extraction forms intended to capture verbatim terminology or free-text items, the study recommends:
The revised form developed after this pilot will be tested on new material in Phase 1 of the ATLAS-ONTO project to assess reproducibility beyond harmonization-driven convergence.
All data and materials referenced in the manuscript are publicly available. The project deposited: the blank extraction forms (initial and revised), the extraction manual with its dated change log, the datasets from every round of the pilot (blinded round, two exploratory exercises, post-harmonization round, corrected dataset and directed audit), audited pre-extraction and appraisal files, the data dictionary, analysis scripts with R version and locale guard, the Python cross-check, the consolidated results file, and complete analytical outputs including post-registration corrections. The OSF registrations and folders cited in the source contain the corrected analysis script (version 1.1) and its outputs. No individual patient or participant data were used.
The authors declared no competing interests. The manuscript states that ethical guidelines were followed, necessary IRB or ethics approvals obtained, and that participant consent procedures and trial registration guidance were addressed where relevant. The authors confirm adherence to reporting guidelines and have provided supporting materials in the public OSF repository.