Datasets

Below is information about Spanish textual data sets created with the goal of solving NLP tasks. In this case, these are collections of texts, generally enriched with annotations.

Filter by

Search by keyword

Domain

NLP topic

Language

Year

HOMO-LAT-2025

Social
Spanish (Argentina) , Spanish (Bolivia) , Spanish (Chile) , Spanish (Colombia) , Spanish (Dominican Republic) , Spanish (Mexico) , Spanish (Peru) , Spanish (Uruguay)
Published in 2025
7,100
7100.00MB
hate detection
DIMEMEX-2025

Social
Spanish (Mexico)
Published in 2025
3,000
3000.00MB
hate detection
Homo-MEX

Spanish (Mexico)
Published in 2024
18,200
hate detection
HOMO-MEX 2023

Social
Spanish (Mexico)
Published in 2023
11,000
Tweets
hate detection
Mexican Aggressiveness Corpus

Spanish (Mexico)
Published in 2020
10,475
Tweets
hate detection
MEX-A3T-profiling

Spanish (Mexico)
Published in 2018
5,000
Tweets
hate detection

If you have published a result better than those on the list, send a message to odesia-comunicacion@lsi.uned.es indicating the result and the DOI of the article, along with a copy of it if it is not published openly.

Datasets

Filter by

HOMO-LAT-2025

DIMEMEX-2025

Homo-MEX

HOMO-MEX 2023

Mexican Aggressiveness Corpus

MEX-A3T-profiling