The gold standard dataset consists of full clinical case-summary pairs manually selected. These clinical cases were primarily sourced from the PubMed database, using filters based on publication type (case reports) and publication languages.
Language(s)
Spanish
Dataset description link
Year
2025
Domain
Health
Annotations
Summaries
Format
txt
Data access
Public
Data link
Publication
Rodríguez-Ortega M, Rodríguez-Lopez E, Lima-López S, Escolano C, Melero M, Pratesi L, Vigil-Giménez L, Fernandez L, Farré-Maduell E, Krallinger M. Overview of MultiClinSum task at BioASQ 2025: evaluation of clinical case summarization strategies for multiple languages: data, evaluation, resources and results. In Proceedings of CLEF 2025, CEUR Workshop Proceedings.
Publication link
License
CC-BY-4.0
NLP Topic
Number of units
998
Tokens
599698
Sentences
26827
Documents
998
Size
998.00MB

