Vollständiger Abstract
Worum geht es in dieser Arbeit?
Abstract Motivation Geographic biases in cancer genomics research limit precision oncology’s global applicability. Major consortia represent only 16 countries and sample less than 1% of 1.4 million publications worldwide. While automated literature mining could address these biases, the lack of benchmark datasets prevents rigorous development and validation of extraction models. We present a FAIR-compliant gold-standard benchmark of 129 manually curated cancer genomics studies spanning 46 countries, tripling geographic representation versus existing consortia. This dataset enables systematic evaluation of automated geographic information and molecular data extraction approaches and establishes a reference task for document-level geographic information extraction models. Results Our benchmark highlights critical challenges for AI-powered literature mining. GeoBoost2 achieves 55% macro-F1 and 46% micro-F1 for patient-origin extraction across 129 articles, exposing precision–recall tradeoffs requiring domain adaptation. Seventy percent of studies include author affiliations from countries unrelated to patient cohorts, showing that metadata-based approaches systematically misattribute geographic provenance. Of 129 studies, 96 provide molecular data for 263,472 cancer patients, revealing substantial untapped literature-derived resources absent from major consortia and supporting training and comparison of domain-specific NLP models. Availability and implementation The complete benchmark dataset, including manual annotations, automated extraction outputs, evaluation scripts, and supplementary documentation, is available at Zenodo with DOI 10.5281/zenodo.18259159.
Bibliografischer Nachweis
Publikationsdaten
- Autor:innen
- Maricel G Kann, Karen O’Connor, Petra Tembei, Graciela Gonzalez Hernandez
- Quelle
- Bioinformatics Advances
- Publikation
- 2026-01-01
- Band / Ausgabe
- Nicht angegeben
- Seiten
- Nicht angegeben
- ISSN / ISBN
- 2635-0041
- Zitationen
- 0 laut Crossref
- Referenzen
- 0 hinterlegt
Zitieren
Zitierfähiger Nachweis
Maricel G Kann, Karen O’Connor, Petra Tembei, Graciela Gonzalez Hernandez (2026). A FAIR gold-standard benchmark for geographic entity recognition in cancer genomics literature for biomedical NLP. Bioinformatics Advances. https://doi.org/10.1093/bioadv/vbag225
Kontext
Themen, Förderung und Nutzung
Lizenzhinweise: Lizenz 1