SatirA 2025

The SatirA 2025 dataset consists of a collection of audio segments (approximately 25 hours) and their transcriptions, extracted from satirical programs and news broadcasts in Spanish and annotated as satire or non-satire.

Language(s)
Spanish
Year
2025
Domain
Diverse
Annotations
Each instance is labeled with a binary category indicating whether the content is satirical or non-satirical, based on the communicative intent of the segment.
Data access
Registration

Publication
Pan R., et al. 2025. Overview of SatiSPeech at IberLEF 2025: Multimodal Audio-Text Satire Classification in Spanish. Procesamiento del Lenguaje Natural, 75, pp. 513-522.
NLP Topic
Number of units
8000
Documents
8000
Size
8000.00MB
Training set size
6000
Test set size
2000

If you have published a result better than those on the list, send a message to odesia-comunicacion@lsi.uned.es indicating the result and the DOI of the article, along with a copy of it if it is not published openly.