The dataset is composed of Spanish tweets published between October 2020 and October 2021 by 18 media outlets from different Spanish-speaking countries, each linked to its corresponding news article (including URL, cleaned HTML, headline, subtitle, body text, and images). Each tweet is labeled as “clickbait” or “non-clickbait.” In the case of clickbait, it is accompanied by manually created spoilers.
Language(s)
Spanish
Dataset description link
Year
2025
Domain
News
Text types
Tweets
Annotations
Each tweet is labeled as “clickbait” or “non-clickbait.” In the case of clickbait, it is accompanied by manually created spoilers.
Format
txt
Data access
Public
Publication
Mordecki, G. et al. 2025. Overview of TA1C at IberLEF 2025: Detecting and Spoiling Clickbait in Spanish-Language News. Procesamiento del Lenguaje Natural, 75, pp. 523-535.
NLP Topic
Number of units
3500
Type of units
Tweets
Documents
3500
Size
3500.00MB
Training set size
2800
Test set size
700

