machine translation

The EA-MT dataset

The dataset used in this task is composed of source-language (English) sentences that contain named entities that are potentially complex from a machine translation perspective. These entities may be rare, ambiguous, or unknown to translation systems, posing an additional challenge beyond conventional lexical translation. The goal of the dataset is to evaluate the ability of machine translation systems to correctly handle such elements, ensuring their accurate transfer into the target language without loss of meaning or disambiguation errors.

XC-Translate-2025-en-es

The dataset is a gold benchmark for evaluating the performance of MT systems in translating text containing entities with names that are ignificantly different across a set of 10 diverse languages. First the entities of interest for the task have been selected, and then sentences containing these entities are generated with GPT-4. Each sentence is translated into 10 target languages by at least three native translators.