automatic transcription
PastReader-2025
- Read more about PastReader-2025
- Log in or register to post comments
The dataset is composed of historical press publications in the public domain digitized by the National Library of Spain (BNE). It covers content from the 17th to the 20th century and is available in PDF format along with corresponding OCR files downloadable as plain text. The quality of the OCR varies due to factors such as the digitization date and the preservation state of the originals. Part of the corpus includes collaborative manual corrections, which constitute highly valuable resources as reference data (“ground truth”) for the training and evaluation of automatic systems.

