The SatirA 2025 dataset consists of a collection of audio segments (approximately 25 hours) and their transcriptions, extracted from satirical programs and news broadcasts in Spanish and annotated as satire or non-satire.
Language(s)
Spanish
Dataset description link
Year
2025
Domain
Diverse
Annotations
Each instance is labeled with a binary category indicating whether the content is satirical or non-satirical, based on the communicative intent of the segment.
Data access
Registration
Publication
Pan R., et al. 2025. Overview of SatiSPeech at IberLEF 2025: Multimodal Audio-Text Satire Classification in Spanish. Procesamiento del Lenguaje Natural, 75, pp. 513-522.
NLP Topic
Number of units
8000
Documents
8000
Size
8000.00MB
Training set size
6000
Test set size
2000

