The dataset contains multigenre content in Arabic, English, and Spanish. The dataset is collected from various fact-checking domains through Google Fact-check Explorer API2, complete with detailed metadata and an evidence corpus sourced from the web.
Language(s)
Spanish
English
Arabic
Dataset description link
Year
2025
Domain
News
Text types
Afirmaciones
Annotations
label indicating whether a numerical claim is true, false or contradictory
Format
json
Data access
Public
Publication
Venktesh, V., et al. Overview of the CLEF-2025 CheckThat! lab task 3 on fact-checking numerical claims. Working Notes of CLEF, 2025.
Publication link
NLP Topic
Number of units
3689
Training set size
1506
Test set size
482
Development set size
587

