EmoSPeech-2024: Multimodal Speech-text Emotion Recognition in Spanish - Multimodal AER

This task introduces a new modality in the AER challenge by incorporating audio. The objective is to integrate text and voice signals to determine the emotion conveyed in each segment, based on five of Ekman's six basic emotions: anger, disgust, fear, joy, and sadness, as well as a neutral emotion.

Publication
Pan et al. (2024). Overview of EmoSPeech at IberLEF 2024: Multimodal Speech-text Emotion Recognition in Spanish. Procesamiento del Lenguaje Natural, Revista, 73: 359-368.

Task results

System MacroF1 Sort ascending
BSC–UPC 0.8669
THAU–UPM 0.8248
CogniCIC 0.7123
ITST 0.6876
UNED–UNIOVI 0.6709

If you have published a result better than those on the list, send a message to odesia-comunicacion@lsi.uned.es indicating the result and the DOI of the article, along with a copy of it if it is not published openly.