This cross-sectional pilot evaluated whether documented corrective actions in national medical safety incident reports in Japan emphasized individual vigilance (often termed Safety-I) or structural, system-level interventions (Safety-II). The analysis aimed to test the feasibility of automated, large-scale classification of free-text corrective-action entries and to produce summary indices describing corrective-action maturity across a national dataset.
The dataset comprised all corrective-action free-text fields from the 2010 release of the Japan Council for Quality Health Care (JCQHC) Medical Accident Information Collection Project. In total 11,507 entries were analysed, consisting of 8,804 near-miss (Hiyari-Hatto) reports and 2,703 accident (Jiko) reports. The analysis used only publicly available de-identified data from this national reporting repository.
Records were classified by a five-stage hybrid pipeline combining rule-based and machine-learning methods. The stages were: (1) an expert-developed rule dictionary; (2) TF-IDF vectorization with k-nearest-neighbor matching; (3) cosine-similarity matching; (4) a two-tier large-language-model (LLM) classifier; and (5) a conservative priority-cascade fallback to ensure a definitive label. The pipeline's objective was to assign each corrective-action free-text entry to a maturity level on a pre-defined 7-level scale (L0–L6).
Each corrective-action entry received an L0–L6 label representing ascending levels of corrective-action maturity. Two summary indices were specified:
Pre-specified minimal clinically important difference (MCID) criteria for between-group comparisons were a risk difference (RD) ≥ 2 percentage points and Cramer's V ≥ 0.10.
The hybrid system assigned a definitive L0–L6 label to every record (0% unresolved). Overall, 16.6% of records were classified as L0 (unclassifiable). Therefore 9,599 records were classifiable (non-L0).
Among the entire dataset, the most common label was L1, representing individual-vigilance actions; L1 accounted for 54.6% of all records. Across the classifiable subset, a minority of corrective actions met the study's definition of system-based interventions: the overall SSMR was reported as 11.82% (95% CI, 11.19–12.49%).
SSMR was higher for accident (Jiko) reports than for near-miss (Hiyari-Hatto) reports. The reported SSMR for accident reports was 18.12% (95% CI, 16.69–19.64%). The source text for the corresponding near-miss SSMR contained a corruption and the exact percentage for near-miss reports is not clearly reported in the source; the difference between groups was reported as RD = 8.66 percentage points. Cramer's V was reported as 0.120, and chi-square(1) = 136.97001.
These reported measures indicate that accident reports contained a substantively higher proportion of system-based corrective actions than near-miss reports, exceeding the pre-specified MCID thresholds.
Between-group comparisons used a chi-square test with Wilson 95% confidence intervals, and effect-size reporting via Cramer's V and the absolute risk difference. The analysis compared observed differences against MCID thresholds established before analysis: RD ≥ 2 percentage points and Cramer's V ≥ 0.10. The reported RD of 8.66 percentage points and Cramer's V = 0.120 both exceed these thresholds, indicating that the difference was not only statistically significant but also met the study's a priori definition of clinical significance.
The pilot demonstrates feasibility of automated, large-scale classification of corrective-action text in a national incident reporting system, with full coverage (no unresolved records) achieved by the hybrid pipeline. Key limitations reported in the source include a single-year interim analysis (2010) and ongoing development and validation of the classification dictionary and code. The classification dictionary, analysis code, and derived datasets were not publicly available at the time of the pilot report because they remain under active development; a planned release is stated for after completion of the full study. The author indicates these interim results support proceeding to a planned 16-year longitudinal analysis.
Ethical approval for the work was provided by the Tokyo-Kita Medical Center Clinical Research Ethics Review Committee (approval no. 558). The requirement for informed consent was waived because the study used publicly available de-identified records. The author declared no competing interests. Source data are publicly available from the JCQHC Medical Accident Information Collection Project; aggregate results are included in the manuscript and supplementary materials.