---
title: "Zero-shot LLMs Outperform Rule-based and NER Methods for Colorectal Cancer Symptom Extraction"
id: "medrxiv-22-leveraging-large-language-models-for-colorectal-cancer-symptom-extraction-from"
canonical_url: "https://medichelpline.com/clinical-feed/medrxiv-22-leveraging-large-language-models-for-colorectal-cancer-symptom-extraction-from"
content_type: "clinical_feed_article"
specialty: "Oncology"
source_name: "medRxiv (Clinical Preprints)"
source_url: "https://www.medrxiv.org/content/10.64898/2026.09.15.26362961v1?rss=1"
published_at: "2026-09-20T12:00:00.000Z"
evidence_level: "Verified Feed"
license: "CC-BY-NC-4.0 / Informational Use"
---
# Zero-shot LLMs Outperform Rule-based and NER Methods for Colorectal Cancer Symptom Extraction
## Provenance & Clinical Metadata
- **Canonical URL:** https://medichelpline.com/clinical-feed/medrxiv-22-leveraging-large-language-models-for-colorectal-cancer-symptom-extraction-from
- **Specialty:** [Oncology](https://medichelpline.com/clinical-feed/oncology.md)
- **Primary Source:** medRxiv (Clinical Preprints)
- **Source URL:** [Original Journal Publication](https://www.medrxiv.org/content/10.64898/2026.09.15.26362961v1?rss=1)
- **Published At:** 2026-09-20T12:00:00.000Z
- **Evidence Rating:** Verified Feed
## Executive GIST (TL;DR)
- Study evaluated extraction of 46 colorectal cancer (CRC) symptoms from unstructured discharge notes in **MIMIC-IV** using rule-based, pretrained NER, and zero-shot large language models (LLMs). - Dataset: 2,704 CRC discharge notes; a 200-note adjudicated gold standard (two raters, pooled kappa=0.71, macro kappa=0.49) used for benchmarking. - Symptom target list derived from the Memorial Symptom Assessment Scale and EORTC QLQ-CR29, covering 46 cancer-related symptoms. - Methods compared: dictionary-based rule matching, pretrained clinical NER, zero-shot **Claude Haiku**, zero-shot **Gemini 3.5 Flash**, and two hybrid approaches applying post-hoc rule-based negation filtering to LLM outputs. - Primary evaluation metrics: Macro and Micro F1, precision, and recall against adjudicated gold standard. - Key results: **Gemini 3.5 Flash** achieved highest performance (Macro F1=0.70, Micro F1=0.86, Macro Precision=0.74); **Claude Haiku** followed (Macro F1=0.63, Macro Recall=0.71). - Rule-based and NER methods performed substantially worse (rule-based Macro F1=0.44; NER Macro F1=0.38). - Applying rigid post-hoc negation filtering to LLM outputs reduced performance (Gemini+Hybrid Macro F1=0.58; Claude+Hybrid Macro F1=0.54) by overturning correct LLM predictions due to fixed-window matching. - Conclusion: **zero-shot LLMs** offer a scalable, accurate alternative to manual chart review and traditional NLP pipelines for oncology symptom surveillance without institution-specific rule development or model training. - Cautions: Preprint not peer reviewed; negation correction should be validated for syntactic scope before applying to LLM outputs. - Data access: MIMIC-IV notes available to credentialed PhysioNet users; code repository link provided in source.
## Clinical Analysis & Structured Key Points
Leveraging Large Language Models for Colorectal Cancer Symptom Extraction from MIMIC-IV Clinical Notes | medRxiv Skip to main content Leveraging Large Language Models for Colorectal Cancer Symptom Extraction from MIMIC-IV Clinical Notes View ORCID Profile Youran Lee , Ivo Dinov , Xiaosu Hu , Yun Jiang doi: https://doi.org/10.64898/2026.09.15.26362961 Youran Lee University of Michigan Find this author on Google Scholar Find this author on PubMed Search for this author on this site ORCID record for Youran Lee For correspondence: youranrl{at}umich.edu Ivo Dinov University of Michigan Find this author on Google Scholar Find this author on PubMed Search for this author on this site Xiaosu Hu University of Michigan Find this author on Google Scholar Find this author on PubMed Search for this author on this site Yun Jiang University of Michigan Find this author on Google Scholar Find this author on PubMed Search for this author on this site Abstract Info/History Metrics Supplementary material Data/Code Preview PDF Abstract Background: Much of the symptom burden in colorectal cancer (CRC) patients is documented in unstructured discharge-note narrative, and manual extraction is not scalable. Whether large language models (LLMs) outperform rule-based and named entity recognition (NER) methods has not been rigorously benchmarked. Objective: To benchmark rule-based, NER, and zero-shot LLM methods for extracting 46 cancer-related symptoms from CRC discharge notes against an adjudicated ground truth. Methods: We analyzed 2,704 discharge notes from CRC patients in MIMIC-IV. A 46-symptom target list was built from the Memorial Symptom Assessment Scale and the EORTC QLQ-CR29. Four approaches -- dictionary-based rule matching, pretrained clinical NER, and zero-shot Claude Haiku and Gemini 3.5 Flash -- plus two hybrid variants (LLM output with post-hoc rule-based negation filtering) were evaluated against a 200-note gold standard adjudicated by two raters (pooled kappa=0.71, macro kappa=0.49), using Macro/Micro F1, precision, and recall. Results: Gemini 3.5 Flash performed best (Macro F1=0.70, Micro F1=0.86, Macro Precision=0.74), followed by Claude Haiku (Macro F1=0.63, Macro Recall=0.71); both substantially outperformed rule-based (Macro F1=0.44) and NER (Macro F1=0.38) methods. Post-hoc negation filtering paradoxically degraded LLM performance (Gemini+Hybrid Macro F1=0.58; Claude+Hybrid Macro F1=0.54) by overriding correct predictions through rigid, fixed-window matching. Conclusions: Zero-shot LLMs substantially outperform rule-based and NER approaches for CRC symptom extraction; post-hoc negation correction should not be applied to LLM outputs without syntactic scope validation. Implications for Practice: Zero-shot LLM extraction offers a scalable, accurate alternative to manual chart review and traditional NLP pipelines for oncology symptom surveillance, without institution-specific rule development or model training. Competing Interest Statement The authors have declared no competing interest. Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: This study utilized the MIMIC-IV Clinical Notes dataset (version 2.2), a large-scale de-identified clinical note database from Beth Israel Deaconess Medical Center, and get approved under a valid data use agreement (DUA) with PhysioNet. https://physionet.org/content/mimiciv/3.1/ I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes Data Availability The data used in this study are available from the MIMIC-IV database to credentialed users who complete the required training from PhysioNet. and obtain appropriate data access. https://github.com/youranrl-dot/crc_symptom_llm Copyright The copyright holder for this preprint is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. It is made available under a CC-BY-NC-ND 4.0 International license . Back to top Previous Next Posted September 20, 2026. Download PDF Supplementary Material Data/Code Email Thank you for your interest in spreading the word about medRxiv. NOTE: Your email address is requested solely to identify you as the sender of this article. Your Email * Your Name * Send To * Enter multiple addresses on separate lines or separate them with commas. You are going to email the following Leveraging Large Language Models for Colorectal Cancer Symptom Extraction from MIMIC-IV Clinical Notes Message Subject (Your Name) has forwarded a page to you from medRxiv Message Body (Your Name) thought you would like to see this page from the medRxiv website. Your Personal Message CAPTCHA This question is for testing whether or not you are a human visitor and to prevent automated spam submissions. Share Leveraging Large Language Models for Colorectal Cancer Symptom Extraction from MIMIC-IV Clinical Notes Youran Lee , Ivo Dinov , Xiaosu Hu , Yun Jiang medRxiv 2026.09.15.26362961; doi: https://doi.org/10.64898/2026.09.15.26362961 Share This Article: Copy Citation Tools Leveraging Large Language Models for Colorectal Cancer Symptom Extraction from MIMIC-IV Clinical Notes Youran Lee , Ivo Dinov , Xiaosu Hu , Yun Jiang medRxiv 2026.09.15.26362961; doi: https://doi.org/10.64898/2026.09.15.26362961 Citation Manager Formats BibTeX Bookends EasyBib EndNote (tagged) EndNote 8 (xml) Medlars Mendeley Papers RefWorks Tagged Ref Manager RIS Zotero Tweet Widget Facebook Like Google Plus One Subject Area Health Informatics Subject Areas All Articles Addiction Medicine (619) Allergy and Immunology (903) Anesthesia (334) Cardiovascular Medicine (4828) Dentistry and Oral Medicine (479) Dermatology (419) Emergency Medicine (651) Endocrinology (including Diabetes Mellitus and Metabolic Disease) (1633) Epidemiology (15974) Forensic Medicine (33) Gastroenterology (1206) Genetic and Genomic Medicine (7065) Geriatric Medicine (734) Health Economics (1066) Health Informatics (5060) Health Policy (1437) Health Systems and Quality Improvement (1775) Hematology (587) HIV/AIDS (1347) Infectious Diseases (except HIV/AIDS) (16307) Intensive Care and Critical Care Medicine (1174) Medical Education (671) Medical Ethics (155) Nephrology (727) Neurology (7291) Nursing (368) Nutrition (1084) Obstetrics and Gynecology (1243) Occupational and Environmental Health (1005) Oncology (3614) Ophthalmology (1055) Orthopedics (398) Otolaryngology (457) Pain Medicine (472) Palliative Medicine (139) Pathology (709) Pediatrics (1805) Pharmacology and Therapeutics (739) Primary Care Research (768) Psychiatry and Clinical Psychology (5919) Public and Global Health (9770) Radiology and Imaging (2432) Rehabilitation Medicine and Physical Therapy (1456) Respiratory Medicine (1250) Rheumatology (642) Sexual and Reproductive Health (773) Sports Medicine (579) Surgery (782) Toxicology (107) Transplantation (305) Urology (293)
## Related Clinical Research

- [Cancer type–specific chromatin regulator mutations linked to better survival after immune checkpoi](https://medichelpline.com/clinical-feed/medrxiv-17-cancer-type-specific-profiling-of-chromatin-regulator-mutations-identifies.md)
- [Healthy lifestyle and pan-cancer risk: UK Biobank prospective analysis](https://medichelpline.com/clinical-feed/british-journal-of-cancer-2-healthy-lifestyle-and-cancer-risk-a-pan-cancer-analysis-in-the-uk-biobank.md)
- [Human-like sialome (CMAH loss) links hyperglycemia to accelerated colorectal cancer progression](https://medichelpline.com/clinical-feed/biorxiv-0-human-like-sialome-remodeling-links-hyperglycemia-to-colorectal-cancer.md)
- [KaroLiver cohort: population-based dataset of curative liver-directed treatment for colorectal liv](https://medichelpline.com/clinical-feed/medrxiv-9-cohort-profile-karoliver-a-population-based-cohort-of-patients-receiving.md)
- [Mendelian Randomisation Study Linking the Oral Microbiome to Oral, Oropharyngeal, and Tongue Cance](https://medichelpline.com/clinical-feed/pubmed-42583788.md) (DOI: 10.3290/j.ohpd.c_2778)

## Navigation
- [← Back to Oncology Feed](https://medichelpline.com/clinical-feed/oncology.md)
- [← All Clinical Specialties](https://medichelpline.com/clinical-feed.md)
## Medical & Regulatory Disclaimer

> [!CAUTION]
> MedicHelpline content is structured for research, educational, and professional discovery purposes. It does not constitute individual medical advice, clinical diagnosis, or treatment recommendations.
> Always verify dosing, contraindications, and regulatory alerts against official product labeling and primary regulatory sources before clinical decision-making.