Risk-of-bias (ROB) assessment is a core step in systematic reviews but is often repetitive, time-consuming, and subject to inter-rater variability. The authors evaluated whether a large language model (LLM) can serve as a virtual reviewer to perform ROB judgments using the QUIPS (Quality in Prognosis Studies) framework. Focusing on prognostic research in a single clinical discipline—neurology—they aimed to assess feasibility and compare agreement between the automated pipeline and human reviewers.
The study used a zero-shot prompting strategy to construct an LLM-based pipeline that emulates a human reviewer applying the QUIPS tool. The pipeline was engineered through targeted prompts to produce domain-specific ROB judgments. The authors applied this automated pipeline to articles that had been included in previously published systematic reviews of prognosis studies in neurology.
Agreement analyses compared two pairs: (1) the LLM-generated ROB assessments versus the original human ROB assessments documented in the source systematic reviews, and (2) the original human ROB assessments against each other. Statistical tests included Cohen’s weighted kappa to quantify agreement and Wilcoxon signed-rank tests to compare paired score distributions. Rank-biserial correlations were used to describe directionality of differences between human and LLM ratings.
The evaluation dataset comprised 298 individual articles drawn from 15 systematic reviews. These reviews spanned three neurological prognosis domains: epilepsy, traumatic brain injury, and stroke. Additional dataset and pipeline materials were made publicly available by the authors in a GitHub repository.
Agreement between raters was quantified using Cohen’s weighted kappa with 95% confidence intervals. The authors report an LLM-human overall agreement of Cohen’s weighted kappa = 0.22 (95% CI, 0.12–0.33). For human-human agreement in a small sample (n = 5), Cohen’s weighted kappa was reported as −0.25 (95% CI, −1.04–0.54). The Wilcoxon signed-rank test was applied across four bias domains and for overall risk scores; these tests were statistically significant (p < 0.05). Rank-biserial correlations indicated a tendency for human raters to assign higher risk scores than the LLM.
The study demonstrates feasibility of an LLM-based approach to automate ROB assessment with the QUIPS tool in neurology prognosis literature. Numerical findings include an LLM-human Cohen’s weighted kappa of 0.22 for overall risk scores, interpreted by the authors as limited agreement. A comparison limited to a small sample (n = 5) suggested that LLM-human agreement may not be inferior to human-human agreement; however, the human-human kappa in that subset was negative (−0.25) with a wide confidence interval, reflecting high uncertainty due to the small sample.
Statistical comparisons across four QUIPS bias domains and the overall scores reached significance in Wilcoxon signed-rank tests (p < 0.05), and rank-biserial correlations showed that human raters tended to score studies at higher risk than the LLM pipeline did.
The authors interpret their results as proof-of-concept evidence that a prompt-engineered large language model pipeline can automate QUIPS-based risk-of-bias assessments for prognosis studies in neurology. Agreement with human raters was limited but measurable, and in a very small sample the automated approach was not demonstrably inferior to human-human agreement.
Important limitations are inherent in the reported data: the human-human comparison is based on a small sample (n = 5), leading to wide confidence intervals and uncertain estimates. The authors emphasize the need for targeted methodological refinements, notably standardizing how QUIPS is implemented in automated workflows and validating the pipeline against expert reference ratings. They also note that while automation may reduce time and cost in systematic reviews, further work is required before deployment in routine evidence synthesis.
The authors declare no competing interests. They state that all relevant ethical guidelines were followed. The study materials—data, code, and the LLM-based pipeline—are available in a publicly accessible GitHub repository: https://github.com/SemKampman/QUIPS_Automated_ROB_Assessment. The preprint was posted to medRxiv and includes supplementary material and links for download.
A tailored, prompt-engineered LLM pipeline can feasibly generate QUIPS ROB assessments for prognosis studies in neurology. Reported LLM-human agreement was limited (Cohen’s weighted kappa = 0.22, 95% CI 0.12–0.33), and a very small human-human comparison (n = 5) yielded an uncertain negative kappa. Wilcoxon tests showed significant differences across bias domains and overall scores (p < 0.05), with human raters tending to assign higher risk scores than the LLM. The authors recommend methodological refinements—standardizing QUIPS implementation and validating automated outputs against expert ratings—before broader adoption, and they provide their code and data to support replication and further development.