MultiParaDetox-2025-es

MultiParaDetox is a dataset for text detoxification, a style transfer task that converts toxic expressions into a neutral register. It extends the ParaDetox approach to multiple languages, enabling the automatic creation of parallel corpora for detoxification.

Language(s)
Spanish
French
Hindi
Italian
Ukrainian
Year
2024
Domain
Diverse
Annotations
corpus paralelos
Data access
Public

Publication
Daryna Dementieva, Nikolay Babakov, Amit Ronen, Abinew Ali Ayele, Naquee Rizwan, Florian Schneider, Xintong Wang, Seid Muhie Yimam, Daniil Moskovskiy, Elisei Stakovskii, Eran Kaufman, Ashraf Elnagar, Animesh Mukherjee, and Alexander Panchenko. 2025. Multilingual and Explainable Text Detoxification with Parallel Corpora. In Proceedings of the 31st International Conference on Computational Linguistics, pages 7998–8025, Abu Dhabi, UAE. Association for Computational Linguistics.
Number of units
1000
Type of units
Tweets
Training set size
400
Test set size
600

If you have published a result better than those on the list, send a message to odesia-comunicacion@lsi.uned.es indicating the result and the DOI of the article, along with a copy of it if it is not published openly.