MultiParaDetox is a dataset for text detoxification, a style transfer task that converts toxic expressions into a neutral register. It extends the ParaDetox approach to multiple languages, enabling the automatic creation of parallel corpora for detoxification.
Language(s)
Spanish
French
Hindi
Italian
Ukrainian
Dataset description link
Year
2024
Domain
Diverse
Annotations
corpus paralelos
Annotation guide link
Data access
Public
Publication
Daryna Dementieva, Nikolay Babakov, Amit Ronen, Abinew Ali Ayele, Naquee Rizwan, Florian Schneider, Xintong Wang, Seid Muhie Yimam, Daniil Moskovskiy, Elisei Stakovskii, Eran Kaufman, Ashraf Elnagar, Animesh Mukherjee, and Alexander Panchenko. 2025. Multilingual and Explainable Text Detoxification with Parallel Corpora. In Proceedings of the 31st International Conference on Computational Linguistics, pages 7998–8025, Abu Dhabi, UAE. Association for Computational Linguistics.
Publication link
NLP Topic
Number of units
1000
Type of units
Tweets
Training set size
400
Test set size
600

