Info: OCR-Based Text Mining of 18th-Century Manuscripts for Digital Literary Corpus Construction Topics I phdassistance.com
Published: 5th August in OCR-Based Text Mining of 18th-Century Manuscripts for Digital Literary Corpus Construction Topics I phdassistance.com
Our academic writing and marking services can help you
Digitisation of historical archives has tremendously contributed to the preservation of cultural heritage and analysis of historical documents through computation. With the increasing digitisation of numerous 18th-century research papers by different institutions such as libraries, the need for text extraction and construction of historical document corpora has greatly increased. Recent technological advancements have been made in artificial intelligence, optical character recognition, transformer-based recognition of handwritten text and natural language processing. These have led to greater ease of use and analysis of historical texts and documents. Some of the problems that need to be addressed in relation to this topic include improvement in the accuracy of optical character recognition, preservation of textual authenticity, analysis of damaged historical texts and creation of digital corpora of documents. Current literature on the subject indicates that there is a need for incorporation of artificial intelligence, historical document processing and OCR-Based Text Mining in digital corpus construction.
The fast-paced development of the Historical Digitisation project has greatly improved accessibility to historical textual corpora, providing researchers with an opportunity to analyse a vast amount of 18th-Century studies with the help of computer-assisted methods. Optical Character Recognition (OCR) is a critical component in the process of transforming scanned historical documents into readable text; nonetheless, OCR mistakes still impact the quality and reliability of digital archives. Hill and Hengchen (2019) found that OCR errors affect different digital humanities tasks, such as topic modelling, collocation, authorship attribution, and vector space modelling, differently. The results obtained by Hill and Hengchen (2019) show that OCR quality should be measured with respect to particular research objectives and not through conventional measures of accuracy only. At the same time, there are no ways to measure the accuracy of OCR in the absence of manually compiled ground-truth data, which is usually not available in the case of historical textual corpora.
The research analyses the impact of OCR errors on different historical text analysis approaches and the possibility of predicting the quality of OCR in cases where no data can be used as a reference. Nonetheless, there is no suggestion of an automated system that can predict the quality of OCR on a per-task basis for large historical corpora. The development of such intelligent systems would provide valuable support for OCR Research Services as well as improve historical corpus creation.
Hill, M. J., & Hengchen, S. (2019). Quantifying the Impact of Dirty OCR on Historical Text Analysis: Eighteenth Century Collections Online as a Case Study. Digital Scholarship in the Humanities, 34(4), 825–843.
Due to the increasing availability of large datasets of historical texts, corpus linguistics has become more popular as well as digital scholarship. Examples of available corpora include resources such as Eighteenth Century Collections Online (ECCO), which allows us to access 18th-century manuscripts and conduct research on language development, literary trends, and history in general. Nevertheless, according to Tolonen, Mäkelä, Ijaz, and Lahti (2021), the quality of OCR does not mean the corpus quality itself. The authors have listed other aspects that should be considered during the research on corpus linguistics, including reprint duplication, publication bias, metadata issues, and unbalanced corpus representativeness. As it was mentioned, no evaluation method considers those aspects besides OCR quality itself. Thus, it becomes necessary to find an integrated approach that will take into account both OCR quality and corpus representativeness.
This study examines the quality of OCR, corpus biases, reproductions, inconsistencies in metadata, and publications in ECCO. Nonetheless, these quality factors are assessed separately without formulating a reliability assessment framework that can combine these quality factors. There is great potential in the development of a corpus reliability index by combining several quality factors to facilitate OCR Services and historical text mining.
Tolonen, M., Mäkelä, E., Ijaz, A., & Lahti, L. (2021). Corpus Linguistics and Eighteenth Century Collections Online (ECCO). Research in Corpus Linguistics, 9(1).
The rise of digitised historical archives has provided researchers with new possibilities of exploring extensive literary archives through computational means. Optical Character Recognition (OCR) has proved to be a necessary tool in the process of making texts from 18th-Century studies machine-readable; nonetheless, OCR mistakes affect the analysis of complex linguistic and literary features. Liimatta (2024) examined how OCR quality impacts the automated detection of register features in Eighteenth Century Collections Online (ECCO). The findings reveal the resilience of some linguistic features as well as distortion of others depending on declining OCR quality. Nevertheless, despite the revealed dependence between OCR mistakes and analysis of registers, the research is limited by the examination of linguistic features and does not include higher literary structures such as narrative structure, style, interactions between characters, and thematic progressions. Advanced frameworks for analysing how OCR mistakes affect the higher literary structures could enrich Digital Humanities Research.
The research focuses on assessing the robustness of linguistic register features based on different OCR quality levels, yet the impact of OCR errors on other literary elements and narration is out of scope of this research. The problem of incorporating OCR quality assessment into literary element detection has attracted little attention so far, which makes room for intelligent analytical frameworks creation for better DDC building.
Liimatta, A. (2024). OCR Quality and the Resilience of Algorithmic Identification of Linguistic Register Features in Eighteenth Century Collections Online. Journal of Data Mining and Digital Humanities.
Fast-growing technologies of Large Language Models (LLMs) have brought significant changes to Optical Character Recognition (OCR) post-correction in terms of increasing legibility of digitised historical documents. Thus, they open up numerous prospects for developing the technology of Historical Digitisation, especially in cases of digitisation of large collections of 18th-Century studies, for which OCR correction requires substantial amounts of time and effort. Kanerva et al. (2025) studied the effectiveness of using LLMs for OCR post-correction of historical English and Finnish documents, proving that LLMs can be very effective in reducing errors in English text while pointing out certain limitations of LLM use in OCR correction. It is clear from their results that LLMs may change historical wording in their attempt to increase OCR legibility, produce fictional information or even distort text authenticity during the correction process. Post-correction practices have so far been limited to increasing OCR efficiency regardless of validation of the historical authenticity of the edited text.
The paper measures the effectiveness of the LLM-based OCR post-correction process based on decreases in character errors rate in historical datasets. Nevertheless, it fails to establish the auditing methodology that would be able to detect fake corrections, measure semantic accuracy, and maintain historical integrity while performing automatic post-correction. This is a significant potential area for improvement in OCR Research.
Kanerva, J., Ledins, C., Käpyaho, S., & Ginter, F. (2025). OCR Error Post-Correction with LLMs in Historical Documents: No Free Lunches. arXiv:2502.01205..
The digital preservation of historical archives provides unique opportunities for the preservation of cultural heritage as well as the ability for computational analysis of historical documents. With the growth of handwritten document collections, Handwritten Text Recognition (HTR) is required to convert the manuscripts to machine-readable form for further scholarly research. In the review paper by AlKendi, Gechter, Heyberger, and Guyeux (2024), recent developments in handwritten text recognition are discussed with emphasis on improvements in deep learning and transformer recognition models. Although much has been achieved in this area, there are still some challenges in the use of HTR systems associated with the diversity of handwriting, lack of appropriate benchmark datasets, linguistic diversity and inadequate evaluation metrics for historical manuscripts. Currently available HTR models and their benchmarks are developed using small language-specific datasets, making it impossible for the models to generalize on a large scale of historical archives.
The research gives an overview of the various methods employed in handwritten text recognition (HTR), datasets used, and future work needed in the area but fails to provide any benchmark datasets to evaluate transformer-based HTR models. This creates room for improving OCR Research through the creation of benchmark datasets to aid in the evaluation of models and generalisation to different datasets.
AlKendi, W., Gechter, F., Heyberger, L., & Guyeux, C. (2024). Advancements and Challenges in Handwritten Text Recognition: A Comprehensive Survey. Journal of Imaging, 10(18).
Need assistance finalising your dissertation topic? Selecting a strong, researchable topic can be challenging — but you don’t have to do it alone.
Our research consultants can help refine your ideas, identify literature gaps, and guide you toward a topic that aligns with current academic trends and your programme requirements.
Contact us to begin one-on-one topic development and refinement with PhdAssistance.com Research Lab.
PhDAssistance. (n.d.). Cybersecurity in business Dissertation Topics Retrieved January 28th, from https://phdassistance.com/topic/cybersecurity-business/
Jalolova, M., and Musawwir, M. “Cybersecurity in business Dissertation Topics for PhD Scholars.” PhDAssistance, https://phdassistance.com/topic/cybersecurity-business/ Accessed 28th January 2026.
Jalolova, M., and Musawwir, M. “Cybersecurity in business Dissertation Topics for PhD Scholars.” PhDAssistance, PhDAssistance, Web. 28th January 2026.
Jalolova, M., and Musawwir, M., n.d. Cybersecurity in business Dissertation Topics for PhD scholars. [online] Available at: https://phdassistance.com/topic/cybersecurity-business/ [Accessed 28th January 2026].
Jalolova M., Musawwir M. Cybersecurity in business Dissertation Topics for PhD scholars [Internet]. PhDAssistance; [cited 2026 28th January]. Available from: https://phdassistance.com/topic/cybersecurity-business/
Jalolova, M., and Musawwir, M. (n.d.). Cybersecurity in business Dissertation Topics for PhD scholars. Retrieved 28th January 2026, from https://phdassistance.com/topic/cybersecurity-business/
Jalolova, M., and Musawwir, M., Cybersecurity in business Dissertation Topics (n.d.) https://phdassistance.com/topic/cybersecurity-business/ accessed 28th January 2026.
Free resources to assist you with your university studies!