Datasets

Below is information about Spanish textual data sets created with the goal of solving NLP tasks. In this case, these are collections of texts, generally enriched with annotations.

Filter by

Search by keyword

Domain

NLP topic

Language

Year

HOMO-LAT-2025

Social
Spanish (Argentina) , Spanish (Bolivia) , Spanish (Chile) , Spanish (Colombia) , Spanish (Dominican Republic) , Spanish (Mexico) , Spanish (Peru) , Spanish (Uruguay)
Published in 2025
7,100
7100.00MB
hate detection
DIMEMEX-2025

Social
Spanish (Mexico)
Published in 2025
3,000
3000.00MB
hate detection
HOMO-MEX 2023

Social
Spanish (Mexico)
Published in 2023
11,000
Tweets
hate detection
DA-VINCIS 2023

Social
Spanish (Mexico)
Published in 2023
4,731
Tweets
processing events
DA-VINCIS

Social
Spanish (Mexico)
Published in 2022
5,453
Tweets
processing events

If you have published a result better than those on the list, send a message to odesia-comunicacion@lsi.uned.es indicating the result and the DOI of the article, along with a copy of it if it is not published openly.

Datasets

Filter by

HOMO-LAT-2025

DIMEMEX-2025

HOMO-MEX 2023

DA-VINCIS 2023

DA-VINCIS