---
title: "Labeling strategy shapes machine learning detection of visual field progression in glaucoma"
id: "plos-one-8-labeling-matters-a-multicenter-machine-learning-study-on-visual-field"
canonical_url: "https://medichelpline.com/clinical-feed/plos-one-8-labeling-matters-a-multicenter-machine-learning-study-on-visual-field"
content_type: "clinical_feed_article"
specialty: "General"
source_name: "PLOS ONE (Medicine)"
source_url: "https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0357179"
published_at: "2026-08-28T14:00:00.000Z"
evidence_level: "Journal Feed"
license: "CC-BY-NC-4.0 / Informational Use"
---
# Labeling strategy shapes machine learning detection of visual field progression in glaucoma
## Provenance & Clinical Metadata
- **Canonical URL:** https://medichelpline.com/clinical-feed/plos-one-8-labeling-matters-a-multicenter-machine-learning-study-on-visual-field
- **Specialty:** [General](https://medichelpline.com/clinical-feed/general.md)
- **Primary Source:** PLOS ONE (Medicine)
- **Source URL:** [Original Journal Publication](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0357179)
- **Published At:** 2026-08-28T14:00:00.000Z
- **Evidence Rating:** Journal Feed
## Executive GIST (TL;DR)
- This multicenter retrospective study compared how two algorithm-derived labeling strategies affect machine learning (ML) performance for detecting **visual field** (VF) progression in **glaucoma** using data from five tertiary hospitals. - Two labeling approaches were evaluated without an independent clinical reference: an inclusive **Consensus label** (progression if at least one of five conventional algorithms flagged it) and a conservative **Wiggs’ label** (region-based event–threshold rule). - Five conventional progression algorithms contributed to the Consensus definition: MD slope, VFI slope, AGIS, CIGTS, and pointwise linear regression (PLR). - Four ML classifiers were trained separately with each label: support vector machine (SVM), random forest (RF), logistic regression (LR), and extreme gradient boosting (XGBoost). - Performance metrics included area under the ROC curve (AUC), sensitivity, specificity, precision–recall analysis summarized by average precision (AP). - Models trained with the Consensus label achieved excellent discrimination (AUC 0.92–0.95), high sensitivity (0.82–0.85), near-perfect specificity (0.99–1.00), and high AP (0.93–0.94). - Models trained with the Wiggs’ label showed lower discrimination (AUC 0.88–0.89), reduced sensitivity (0.63–0.72), moderate-to-high specificity (0.87–0.92), and lower AP (0.84–0.85), reflecting a stricter spatial requirement for progression. - Ablation analysis suggested Consensus-based performance relied on complementary information across heterogeneous algorithms rather than on any single criterion. - The authors conclude that labeling strategy is a major determinant of ML performance for VF progression detection and that model performance largely reflects compatibility with the chosen ground truth rather than validating one clinical standard over another. - The study underscores the need for careful ground-truth definition when developing and interpreting ML models for glaucoma progression research.
## Clinical Analysis & Structured Key Points
[ Skip to main content ](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0357179#main-content) Advertisement * [plos.org](https://plos.org/) * [Create account](https://community.plos.org/registration/new) * [Sign in](https://journals.plos.org/user/secure/login?page=%2Fplosone%2Farticle%3Fid%3D10.1371%2Fjournal.pone.0357179) * * About * Browse * Publish * [](https://journals.plos.org/plosone/ "PLOS One") * Search [advanced search](https://journals.plos.org/plosone/search) * [Browse Topics](https://journals.plos.org/plosone/subjectAreaBrowse) Browse Subject Areas ? Click through the PLOS taxonomy to find articles in your field. For more information about PLOS Subject Areas, click [here](https://github.com/PLOS/plos-thesaurus/blob/master/README.md "Link opens in new window"). [](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0357179) [](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0357179) * 0 [Save](https://journals.plos.org/plosone/article/metrics?id=10.1371/journal.pone.0357179#savedHeader) [Total Mendeley and Citeulike bookmarks.](https://journals.plos.org/plosone/article/metrics?id=10.1371/journal.pone.0357179#savedHeader) * 0 [Citation](https://journals.plos.org/plosone/article/metrics?id=10.1371/journal.pone.0357179#citedHeader) [Paper's citation count computed by Dimensions.](https://journals.plos.org/plosone/article/metrics?id=10.1371/journal.pone.0357179#citedHeader) * 22 [View](https://journals.plos.org/plosone/article/metrics?id=10.1371/journal.pone.0357179#viewedHeader) [PLOS views and downloads.](https://journals.plos.org/plosone/article/metrics?id=10.1371/journal.pone.0357179#viewedHeader) * 0 [Share](https://journals.plos.org/plosone/article/metrics?id=10.1371/journal.pone.0357179#discussedHeader) [Sum of Facebook, Twitter, Reddit and Wikipedia activity.](https://journals.plos.org/plosone/article/metrics?id=10.1371/journal.pone.0357179#discussedHeader) Open Access Peer-reviewed Research Article # Labeling matters: A multicenter machine learning study on visual field progression in Glaucoma * Hyobeen Kim, Roles Conceptualization, Data curation, Formal analysis, Funding acquisition, Investigation, Methodology, Writing – original draft Affiliation Department of Mathematics, Chonnam National University, Gwangju, Korea ⨯ * EunAh Kim, Roles Conceptualization, Data curation, Formal analysis, Funding acquisition, Investigation, Methodology, Writing – review & editing Affiliation Department of Ophthalmology, Samsung Changwon Hospital, Sungkyunkwan University School of Medicine, Changwon, Korea ⨯ * Sangwoo Moon, Roles Conceptualization, Data curation, Formal analysis, Investigation, Methodology, Writing – review & editing Affiliation Department of Ophthalmology, Pusan National University Yangsan Hospital, Pusan National University College of Medicine, Yangsan, Korea ⨯ * Sang Wook Jin, Roles Conceptualization, Data curation, Formal analysis, Writing – review & editing Affiliation Department of Ophthalmology, Dong-A University College of Medicine, Busan, Korea ⨯ * Jung Lim Kim, Roles Conceptualization, Data curation, Formal analysis, Writing – review & editing Affiliation Department of Ophthalmology, Busan Paik Hospital, Inje University College of Medicine, Busan, Korea ⨯ * Seung Uk Lee, Roles Conceptualization, Data curation, Formal analysis, Writing – review & editing Affiliation Department of Ophthalmology, Kosin University College of Medicine, Busan, Korea ⨯ * Jeong Rye Park , Roles Conceptualization, Data curation, Formal analysis, Investigation, Methodology, Supervision, Writing – review & editing * E-mail: parkjr@gist.ac.kr (RP); glaucoma@pusan.ac.kr (JL) Affiliation Department of Mathematical Sciences, Gwangju Institute of Science and Technology, Gwangju, Korea ⨯ * Jiwoong Lee Roles Conceptualization, Data curation, Formal analysis, Funding acquisition, Investigation, Methodology, Writing – original draft, Writing – review & editing * E-mail: parkjr@gist.ac.kr (RP); glaucoma@pusan.ac.kr (JL) Affiliations Department of Ophthalmology, Pusan National University College of Medicine, Busan, Korea, Biomedical Research Institute, Pusan National University Hospital, Busan, Korea [ ![ORCID logo](https://journals.plos.org/resource/img/orcid_16x16.png) https://orcid.org/0000-0002-1053-612X ](https://orcid.org/0000-0002-1053-612X "ORCID Registry") ⨯ # Labeling matters: A multicenter machine learning study on visual field progression in Glaucoma * Hyobeen Kim, * EunAh Kim, * Sangwoo Moon, * Sang Wook Jin, * Jung Lim Kim, * Seung Uk Lee, * Jeong Rye Park, * Jiwoong Lee ![PLOS](https://journals.plos.org/resource/img/logo-plos-full-color.svg) x * Published: August 28, 2026 * * [Article](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0357179) * [Authors](https://journals.plos.org/plosone/article/authors?id=10.1371/journal.pone.0357179) * [Metrics](https://journals.plos.org/plosone/article/metrics?id=10.1371/journal.pone.0357179) * [Comments](https://journals.plos.org/plosone/article/comments?id=10.1371/journal.pone.0357179) * [Media Coverage](http://plos.altmetric.com/details/doi/10.1371/journal.pone.0357179) * [Peer Review](https://journals.plos.org/plosone/article/peerReview?id=10.1371/journal.pone.0357179) * [Abstract](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0357179#abstract0) * [Introduction](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0357179#sec005) * [Materials and Methods](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0357179#sec006) * [Results](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0357179#sec007) * [Discussion](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0357179#sec008) * [Supporting information](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0357179#sec009) * [References](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0357179#references) * [Reader Comments](https://journals.plos.org/plosone/article/comments?id=10.1371/journal.pone.0357179) * [Figures](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0357179) ![Accessible Data Icon](https://journals.plos.org/resource/img/accessible_data.svg)Accessible Data [ See the data ![Link Icon](https://journals.plos.org/resource/img/data_link_icon.svg) ](https://github.com/kimhb1029/Visual_Field_Progression) This article includes the Accessible Data icon, an experimental feature to encourage data sharing and reuse. [Find out how research articles qualify for this feature.](https://theplosblog.plos.org/2023/07/accessible-data/) ## Abstract ### Background To compare machine learning (ML) performance for detecting visual field (VF) progression across different labeling strategies using a large multicenter dataset. ### Methods In this multicenter retrospective study, VF data were collected from five tertiary referral hospitals. Two algorithm-derived labeling approaches were evaluated without an independent clinical reference standard: an inclusive Consensus label, defined as progression detected by at least one of five conventional algorithms (mean deviation slope, Visual Field Index slope, Advanced Glaucoma Intervention Study, Collaborative Initial Glaucoma Treatment Study, and pointwise linear regression), and a conservative Wiggs’ label, based on a region-based event–threshold rule. Four ML classifiers, support vector machine, random forest, logistic regression, and extreme gradient boosting, were trained using each labeling strategy. Model performance was assessed using the area under the receiver operating characteristic curve (AUC), sensitivity, specificity, and precision–recall analysis summarized by average precision (AP). ### Results Using the Consensus label, all models demonstrated excellent discrimination (AUC, 0.92–0.95), with high sensitivity (0.82–0.85) and near-perfect specificity (0.99–1.00). Precision–recall analysis showed consistently high reliability of progression detection, with AP values ranging from 0.93 to 0.94. In contrast, models trained with the Wiggs’ label exhibited lower AUCs (0.88–0.89) and reduced sensitivity (0.63–0.72), while maintaining moderate-to-high specificity (0.87–0.92) and lower AP values (0.84–0.85), reflecting a stricter, region-based progression definition. Ablation analysis showed that Consensus-based performance was not driven by any single criterion, but rather by complementary information across heterogeneous progression algorithms. ### Conclusion In this multicenter study, the labeling strategy was a major determinant of ML performance in VF progression detection. The Consensus label enabled sensitive and reliable identification of progression with high specificity with respect to the Consensus label definition across heterogeneous clinical settings, whereas the Wiggs’ label provided conservative, spatially consistent confirmation. The observed performance differences primarily reflect model-label compatibility rather than the clinical validity of either detection system, underscoring that careful definition of ground truth is critical for interpreting ML-based glaucoma progression research. ## Figures ![Fig 4](https://journals.plos.org/plosone/article/figure/image?size=inline&id=10.1371/journal.pone.0357179.g004) ![Fig 1](https://journals.plos.org/plosone/article/figure/image?size=inline&id=10.1371/journal.pone.0357179.g001) ![Fig 2](https://journals.plos.org/plosone/article/figure/image?size=inline&id=10.1371/journal.pone.0357179.g002) ![Table 1](https://journals.plos.org/plosone/article/figure/image?size=inline&id=10.1371/journal.pone.0357179.t001) ![Table 2](https://journals.plos.org/plosone/article/figure/image?size=inline&id=10.1371/journal.pone.0357179.t002) ![Table 3](https://journals.plos.org/plosone/article/figure/image?size=inline&id=10.1371/journal.pone.0357179.t003) ![Table 4](https://journals.plos.org/plosone/article/figure/image?size=inline&id=10.1371/journal.pone.0357179.t004) ![Fig 3](https://journals.plos.org/plosone/article/figure/image?size=inline&id=10.1371/journal.pone.0357179.g003) ![Table 5](https://journals.plos.org/plosone/article/figure/image?size=inline&id=10.1371/journal.pone.0357179.t005) ![Table 6](https://journals.plos.org/plosone/article/figure/image?size=inline&id=10.1371/journal.pone.0357179.t006) ![Fig 4](https://journals.plos.org/plosone/article/figure/image?size=inline&id=10.1371/journal.pone.0357179.g004) ![Fig 1](https://journals.plos.org/plosone/article/figure/image?size=inline&id=10.1371/journal.pone.0357179.g001) ![Fig 2](https://journals.plos.org/plosone/article/figure/image?size=inline&id=10.1371/journal.pone.0357179.g002) ![Table 1](https://journals.plos.org/plosone/article/figure/image?size=inline&id=10.1371/journal.pone.0357179.t001) **Citation:** Kim H, Kim E, Moon S, Jin SW, Kim JL, Lee SU, et al. (2026) Labeling matters: A multicenter machine learning study on visual field progression in Glaucoma. PLoS One 21(8): e0357179. https://doi.org/10.1371/journal.pone.0357179 **Editor:** Suho Lim, Daegu Veterans Health Service Medical Center, KOREA, REPUBLIC OF **Received:** April 2, 2026; **Accepted:** August 13, 2026; **Published:** August 28, 2026 **Copyright:** © 2026 Kim et al. This is an open access article distributed under the terms of the [Creative Commons Attribution License](http://creativecommons.org/licenses/by/4.0/), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited. **Data Availability:** The raw multicenter clinical data are not publicly available because of institutional and ethical restrictions protecting patient privacy. Processed train/test datasets, sample data, and analysis code used in this study are publicly available through the following GitHub repository: . **Funding:** This research was supported by Global-Learning & Academic research institution for Master’s·PhD students, and Postdocs (LAMP) Program of the National Research Foundation of Korea (NRF) grant funded by the Ministry of Education (No. RS-2024-00442775) and by a grant of the Korea Health Technology R&D Project through the Korea Health Industry Development Institute (KHIDI), funded by the Ministry of Health & Welfare, Republic of Korea (grant number: HR20C0026), and a clinical research grant from Pusan National University Hospital in 2026. The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript. **Competing interests:** The authors have declared that no competing interests exist. ## Introduction Accurate and consistent determination of progressive visual field (VF) loss is a key challenge in glaucoma management. VF testing is inherently subject to test–retest variability, which has prompted the development of various statistical approaches to detect progression [[1](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0357179#pone.0357179.ref001)–[3](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0357179#pone.0357179.ref003)]. Among these, event-based methods (guided progression analysis, Advanced Glaucoma Intervention Study [AGIS], and Collaborative Initial Glaucoma Treatment Study [CIGTS] criteria) and trend-based methods (mean deviation [MD] slope, Visual Field Index [VFI] slope, pointwise linear regression [PLR], and permutation of PLR [PoPLR]) are widely used [[4](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0357179#pone.0357179.ref004)–[6](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0357179#pone.0357179.ref006)]. However, agreement among these progression algorithms is low (κ = 0.12–0.52), leading to inconsistent classification of progression and uncertainty in clinical decision-making [[5](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0357179#pone.0357179.ref005),[7](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0357179#pone.0357179.ref007)–[9](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0357179#pone.0357179.ref009)]. Furthermore, relying solely on any single algorithm carries an inherent risk of false-negative classification, as each method may miss progression signals that others detect [[10](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0357179#pone.0357179.ref010)]. This variability poses a fundamental challenge when constructing labeled datasets for supervised machine learning, as the choice of labeling strategy directly determines what the model learns to detect. To address this challenge, we proposed a “Consensus label” defined as positive when at least one of the five algorithms detected progression. This “any-positive” strategy was adopted based on both clinical and methodological rationale. Glaucoma is the leading cause of irreversible blindness worldwide, and because visual loss once incurred cannot be restored, early detection and timely intervention are paramount [[11](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0357179#pone.0357179.ref011)]. Given this irreversibility, the clinical consequence of a missed progression event (false negative) far outweighs the cost of a false-positive alert. Since each algorithm may detect progression signals that others miss, particularly given that different methods are sensitive to different patterns of VF loss, requiring agreement among multiple algorithms before assigning a positive label risks systematically discarding genuine progression signals present in the training data. The any-positive Consensus label was therefore designed to maximize sensitivity in label assignment, ensuring that eyes with any algorithm-detectable signal of deterioration are retained as positive cases, consistent with the clinical principle that any indicator of progression in an irreversible disease warrants heightened vigilance. Previous studies have proposed adopting a consensus-based ground-truth to address inconsistency among multiple progression algorithms [[4](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0357179#pone.0357179.ref004),[6](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0357179#pone.0357179.ref006)]. Sabharwal et al. [[6](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0357179#pone.0357179.ref006)] defined progression when at least four out of six criteria agreed and developed a deep learning model that integrated spatial and temporal information based on this consensus labeling. Recent studies across medical AI applications have shown that outcome definition and labeling strategy can substantially influence model performance and generalizability, often exceeding the impact of model architecture itself [[12](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0357179#pone.0357179.ref012)–[14](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0357179#pone.0357179.ref014)]. We therefore hypothesized that, in glaucoma progression modeling, the definition of VF progression used for labeling would substantially influence machine learning performance, independent of the underlying model. However, to our knowledge, prior studies have not directly and systematically compared how the choice of labeling strategy, independent of model architecture or dataset, affects machine learning performance in VF progression detection, and the relative impact of labeling definition on model learning remains poorly characterized. The application of machine learning to glaucoma-related tasks has grown substantially in recent years. A comprehensive survey of AI methods in glaucoma diagnosis reported that deep learning accounts for the largest proportion of published approaches (31.5%), yet conventional machine learning methods, particularly support vector machines (SVM), remain widely adopted, representing 25.9% of reported methods, underscoring the continued clinical and methodological relevance of traditional classifiers [[15](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0357179#pone.0357179.ref015)]. Beyond model architecture, the quality and definition of training labels have emerged as a critical determinant of model performance across ophthalmic AI applications. For instance, a recent multicenter study applying ensemble machine learning models, including logistic regression (LR), random forest (RF), and support vector classifier, to electronic health records for glaucoma identification demonstrated that severe class imbalance and heterogeneous data definitions substantially impaired model performance, and that resampling strategies designed to address label distribution directly improved classification outcomes [[16](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0357179#pone.0357179.ref016)]. These findings, alongside evidence from broader medical AI literature [[12](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0357179#pone.0357179.ref012)–[14](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0357179#pone.0357179.ref014)], indicate that the challenge of ground truth definition is not unique to glaucoma but represents a fundamental issue across AI-based clinical applications. The present study addresses this challenge directly by systematically comparing the effect of labeling strategy on machine learning performance in VF progression detection. In this study, we newly proposed an inclusive “Consensus label,” in which a case was considered progressive if any one of five major progression algorithms classified it as “True.” This label was directly compared with the spatiotempor
## Related Clinical Research

- [Nursing undergraduates’ willingness to work in long-term care: qualitative insights from a geriatr](https://medichelpline.com/clinical-feed/plos-one-16-willingness-of-nursing-undergraduates-to-work-in-long-term-care-facilities-a.md)
- [Cropland type determines arbuscular mycorrhizal fungi (AMF) density and diversity in Northwest Eth](https://medichelpline.com/clinical-feed/plos-one-23-cropland-type-shapes-arbuscular-mycorrhizal-fungi-amf-population-density-and.md)
- [Patient and Public Engagement in Clinical Trials: CTO’s Decade of Co‑production Lessons](https://medichelpline.com/clinical-feed/bmj-open-0-partnering-with-patients-and-the-public-in-the-clinical-trials-ecosystem-a.md)
- [Variable-rate feedback versus fixed-rate basal PCA: effect on 48-hour cumulative opioid consumption](https://medichelpline.com/clinical-feed/bmj-open-7-comparison-of-cumulative-opioid-consumption-between-variable-rate-feedback.md)
- [FRACTURE-ML: Nationwide machine-learning tool for population hip fracture prediction](https://medichelpline.com/clinical-feed/plos-medicine-2-a-clinical-decision-support-tool-for-accurate-hip-fracture-prediction-a.md)

## Navigation
- [← Back to General Feed](https://medichelpline.com/clinical-feed/general.md)
- [← All Clinical Specialties](https://medichelpline.com/clinical-feed.md)
## Medical & Regulatory Disclaimer

> [!CAUTION]
> MedicHelpline content is structured for research, educational, and professional discovery purposes. It does not constitute individual medical advice, clinical diagnosis, or treatment recommendations.
> Always verify dosing, contraindications, and regulatory alerts against official product labeling and primary regulatory sources before clinical decision-making.