The EA-MT dataset

The dataset used in this task is composed of source-language (English) sentences that contain named entities that are potentially complex from a machine translation perspective. These entities may be rare, ambiguous, or unknown to translation systems, posing an additional challenge beyond conventional lexical translation. The goal of the dataset is to evaluate the ability of machine translation systems to correctly handle such elements, ensuring their accurate transfer into the target language without loss of meaning or disambiguation errors. Consequently, the dataset is specifically designed to test the robustness of models in handling difficult named entity cases, which constitute a critical aspect of real-world translation applications.

Language(s)
Spanish
English
Arabic
Deuch
French
Italian
Korean
Chinese
Year
2025
Annotations
Each sentence is labeled with its translations in different languages.
Data access
Public

Publication
Simone Conia, Min Li, Roberto Navigli, and Saloni Potdar. 2025. SemEval-2025 Task 2: Entity-Aware Machine Translation. In Proceedings of the 19th International Workshop on Semantic Evaluation (SemEval-2025), pages 2535–2557, Vienna, Austria. Association for Computational Linguistics.
Documents
57611
Size
57611.00MB
Training set size
7278
Test set size
49606

If you have published a result better than those on the list, send a message to odesia-comunicacion@lsi.uned.es indicating the result and the DOI of the article, along with a copy of it if it is not published openly.