Datasets

Below is information about Spanish textual data sets created with the goal of solving NLP tasks. In this case, these are collections of texts, generally enriched with annotations.

  • DrugTEMIST -2024

    Health
    Spanish , English , Italian
    Published in 2024
    1,000
    Clinical case reports
    (named) entity recognition

  • CardioCCC-2024

    Health
    Published in 2024
    508
    Clinical case reports
    (named) entity recognition

  • HOPE

    Spanish
    Published in 2024
    9,205
    sentiment analysis

  • DIPROMATS

    Spanish , English
    Published in 2024
    9,591
    text classification

  • EmoSPeech

    Diverse
    Spanish
    Published in 2024
    3,750
    sentiment analysis

  • GenoVarDis

    Biology
    Spanish
    Published in 2024
    633
    (named) entity recognition

  • MedProcNER/ProcTEMIST corpus 2023

    Health
    Spanish
    Published in 2023
    1,000
    Clinical notes
    (named) entity recognition

  • Rest-Mex 2023 Sentiment

    Tourism
    Spanish (Mexico)
    Published in 2023
    359,565
    Tripadvisor opinions
    sentiment analysis

  • FinancES 2023

    Finance
    Spanish
    Published in 2023
    7,980
    News headlines
    sentiment analysis

  • DIPROMATS-ES 2023

    Politics
    Spanish , English
    Published in 2023
    9,591
    Tweets
    text classification

  • MEDDOPLACE Corpus: Gold Standard annotations for Medical Documents Place-related Content Extraction

    Health
    Spanish
    Published in 2023
    1,000
    Clinical notes
    (named) entity recognition

  • ClinAIS 2023

    Health
    Spanish
    Published in 2023
    1,038
    Clinical notes
    text classification

  • MultiCoNER v2 ES

    General
    Spanish
    Published in 2023
    264,207
    Wiki sentences Questions Search queries
    (named) entity recognition

  • MINT ES

    General
    Spanish , English
    Published in 2023
    1,991
    Tweets
    sentiment analysis

  • MultiCoNER-ES

    Diverse
    Spanish
    Published in 2022
    233,987
    Wikipedia Questions Search queries
    (named) entity recognition

If you have published a result better than those on the list, send a message to odesia-comunicacion@lsi.uned.es indicating the result and the DOI of the article, along with a copy of it if it is not published openly.