PastReader-2025

The dataset is composed of historical press publications in the public domain digitized by the National Library of Spain (BNE). It covers content from the 17th to the 20th century and is available in PDF format along with corresponding OCR files downloadable as plain text. The quality of the OCR varies due to factors such as the digitization date and the preservation state of the originals. Part of the corpus includes collaborative manual corrections, which constitute highly valuable resources as reference data (“ground truth”) for the training and evaluation of automatic systems.

Language(s)
Spanish
Year
2025
Annotations
The dataset includes bibliographic and structural tagging information associated with each publication and issue, such as the title, date, issue number, and collection metadata provided by the BNE. In addition, the OCR-derived full-text files preserve basic structural cues (e.g., page and issue segmentation), enabling alignment between scanned images, PDFs, and their corresponding textual content for further processing and annotation.
Format
txt
Data access
Public

Publication
Montejo-Ráez, A. et al. 2025. Overview of PastReader at IberLEF 2025: Transcribing Texts From the Past. Procesamiento del Lenguaje Natural, 75, pp. 453-460.
Number of units
12195
Documents
12195
Size
12195.00MB
Training set size
8959
Test set size
2736
Development set size
500

If you have published a result better than those on the list, send a message to odesia-comunicacion@lsi.uned.es indicating the result and the DOI of the article, along with a copy of it if it is not published openly.