The dataset is composed of news articles from different municipalities in the province of Alicante, related to areas such as sports, culture, leisure, and festivities, and has been developed and validated by experts in text adaptation. For each article, two adapted versions are provided: one in Plain Language (PL) format and another in Easy-to-Read (E2R) format, following the corresponding accessibility guidelines.
Language(s)
Spanish
Dataset description link
Year
2025
Domain
News
Annotations
E2R and PL formats adaptations
Format
txt
Data access
Registration
Publication
Botella-Gil, B. et al. 2025. Overview of CLEARS at IberLEF 2025: Challenge for Plain Language and Easy-to-Read Adaptation for Spanish texts. Procesamiento del Lenguaje Natural, 75, pp. 393-400.
NLP Topic
Number of units
3000
Documents
3000
Size
3000.00MB
Training set size
2100
Test set size
900

