MentalRiskES-2025

The dataset contains Spanish messages about gambling collected from digital platforms such as Telegram, Twitch, Reddit, and Ludopatía.org. Each user is labeled according to their risk level (high or low) and type of addiction.

Language(s)
Spanish
Year
2025
Domain
Health
Annotations
Each instance in the dataset is labeled with a risk level (high or low) and a specific type of addiction (betting, online gaming, trading/crypto, or loot boxes), enabling both binary and multiclass classification.
Format
json
Data access
Registration

Publication
Mármol-Romero, A. et al. (2025). Overview of MentalRiskES at IberLEF 2025: Early Detection of Addiction Risk in Spanish. Procesamiento del Lenguaje Natural, Revista, 75: 425-440.
NLP Topic
Number of units
32342
Documents
32342
Size
32342.00MB
Training set size
22,491
Test set size
9,407
Development set size
444

If you have published a result better than those on the list, send a message to odesia-comunicacion@lsi.uned.es indicating the result and the DOI of the article, along with a copy of it if it is not published openly.