phdassistance

OCR-Based Text Mining of 18th-Century Manuscripts for Digital Literary Corpus Construction

 Info: OCR-Based Text Mining of 18th-Century Manuscripts for Digital Literary Corpus Construction Topics I phdassistance.com

Published: 5th August in OCR-Based Text Mining of 18th-Century Manuscripts for Digital Literary Corpus Construction Topics I phdassistance.com

Share this:

Introduction

Digitisation of historical archives has tremendously contributed to the preservation of cultural heritage and analysis of historical documents through computation. With the increasing digitisation of numerous 18th-century research papers by different institutions such as libraries, the need for text extraction and construction of historical document corpora has greatly increased. Recent technological advancements have been made in artificial intelligence, optical character recognition, transformer-based recognition of handwritten text and natural language processing. These have led to greater ease of use and analysis of historical texts and documents. Some of the problems that need to be addressed in relation to this topic include improvement in the accuracy of optical character recognition, preservation of textual authenticity, analysis of damaged historical texts and creation of digital corpora of documents. Current literature on the subject indicates that there is a need for incorporation of artificial intelligence, historical document processing and OCR-Based Text Mining in digital corpus construction.

Proposed PhD Topic 1: Predicting OCR Reliability Without Ground Truth: A Task-Specific Error-Tolerance Framework for 18th-Century Digital Corpora
Background Context:

The fast-paced development of the Historical Digitisation project has greatly improved accessibility to historical textual corpora, providing researchers with an opportunity to analyse a vast amount of 18th-Century studies with the help of computer-assisted methods. Optical Character Recognition (OCR) is a critical component in the process of transforming scanned historical documents into readable text; nonetheless, OCR mistakes still impact the quality and reliability of digital archives. Hill and Hengchen (2019) found that OCR errors affect different digital humanities tasks, such as topic modelling, collocation, authorship attribution, and vector space modelling, differently. The results obtained by Hill and Hengchen (2019) show that OCR quality should be measured with respect to particular research objectives and not through conventional measures of accuracy only. At the same time, there are no ways to measure the accuracy of OCR in the absence of manually compiled ground-truth data, which is usually not available in the case of historical textual corpora.

PhD-Level Verification:

The research analyses the impact of OCR errors on different historical text analysis approaches and the possibility of predicting the quality of OCR in cases where no data can be used as a reference. Nonetheless, there is no suggestion of an automated system that can predict the quality of OCR on a per-task basis for large historical corpora. The development of such intelligent systems would provide valuable support for OCR Research Services as well as improve historical corpus creation.

Research Questions:
  • How to predict task-specific OCR reliability without ground truth data sets?
  • What are the appropriate quality measures that best facilitate historical text mining?
  • Is there an improvement in predicting OCR quality that helps in constructing a Digital Literary Corpus?
  • Contributions at the PhD-Level:
  • Task-specific OCR reliability prediction.
  • Integration of OCR quality estimation into historical corpus construction processes.
  • Decision support system to enhance reliability in digital humanities text mining.
  • Suggested Readings:

    Hill, M. J., & Hengchen, S. (2019). Quantifying the Impact of Dirty OCR on Historical Text Analysis: Eighteenth Century Collections Online as a Case Study. Digital Scholarship in the Humanities, 34(4), 825–843.

    Historical Manuscript Digitisation
    Proposed PhD Topic 2: Beyond Precision and Recall: A Composite Reliability Index Combining Reprint Bias and OCR Quality for 18th-Century Text Corpora
    Background Context:

    Due to the increasing availability of large datasets of historical texts, corpus linguistics has become more popular as well as digital scholarship. Examples of available corpora include resources such as Eighteenth Century Collections Online (ECCO), which allows us to access 18th-century manuscripts and conduct research on language development, literary trends, and history in general. Nevertheless, according to Tolonen, Mäkelä, Ijaz, and Lahti (2021), the quality of OCR does not mean the corpus quality itself. The authors have listed other aspects that should be considered during the research on corpus linguistics, including reprint duplication, publication bias, metadata issues, and unbalanced corpus representativeness. As it was mentioned, no evaluation method considers those aspects besides OCR quality itself. Thus, it becomes necessary to find an integrated approach that will take into account both OCR quality and corpus representativeness.

    PhD-Level Verification:

    This study examines the quality of OCR, corpus biases, reproductions, inconsistencies in metadata, and publications in ECCO. Nonetheless, these quality factors are assessed separately without formulating a reliability assessment framework that can combine these quality factors. There is great potential in the development of a corpus reliability index by combining several quality factors to facilitate OCR Services and historical text mining.

    Research Questions:
  • How can OCR accuracy and corpus bias be combined in a coherent reliability framework?
  • What are the quality measures that have the greatest impact on corpus reliability in a Historical context?
  • Could a composite measure of reliability contribute to the field of Digital Humanities study on historical corpora?
  • PhD-Level Contributions:
  • A composite reliability index for historical corpora.
  • Combination of OCR accuracy, metadata consistency, and corpus representation bias.
  • Framework enabling reliable Historical Manuscript Digitisation and corpus analysis.
  • Suggested Readings:

    Tolonen, M., Mäkelä, E., Ijaz, A., & Lahti, L. (2021). Corpus Linguistics and Eighteenth Century Collections Online (ECCO). Research in Corpus Linguistics, 9(1).

    Proposed Dissertation topic 3: From Register to Narrative: Extending OCR-Error Impact Analysis to Literary-Structural Features in 18th-Century Fiction Corpora
    Background Context:

    The rise of digitised historical archives has provided researchers with new possibilities of exploring extensive literary archives through computational means. Optical Character Recognition (OCR) has proved to be a necessary tool in the process of making texts from 18th-Century studies machine-readable; nonetheless, OCR mistakes affect the analysis of complex linguistic and literary features. Liimatta (2024) examined how OCR quality impacts the automated detection of register features in Eighteenth Century Collections Online (ECCO). The findings reveal the resilience of some linguistic features as well as distortion of others depending on declining OCR quality. Nevertheless, despite the revealed dependence between OCR mistakes and analysis of registers, the research is limited by the examination of linguistic features and does not include higher literary structures such as narrative structure, style, interactions between characters, and thematic progressions. Advanced frameworks for analysing how OCR mistakes affect the higher literary structures could enrich Digital Humanities Research.

    PhD Level Verification:

    The research focuses on assessing the robustness of linguistic register features based on different OCR quality levels, yet the impact of OCR errors on other literary elements and narration is out of scope of this research. The problem of incorporating OCR quality assessment into literary element detection has attracted little attention so far, which makes room for intelligent analytical frameworks creation for better DDC building.

    Research Questions:
  • What is the effect of OCR errors on extracting literary and narrative structures?
  • Which features of literature are more resistant to OCR errors?
  • Can we build OCR-aware models to enhance HMD for literary studies?
  • PhD-Level Contributions:
  • Development of an OCR-aware model for literary structure analysis.
  • Integration of features of narration, style, and language into historical text mining.
  • Enhancement of the reliability of computational literary analysis based on digitised historical documents.
  • Suggested Readings:

    Liimatta, A. (2024). OCR Quality and the Resilience of Algorithmic Identification of Linguistic Register Features in Eighteenth Century Collections Online. Journal of Data Mining and Digital Humanities.

    Proposed Dissertation Topic 4: Correction or Fabrication? Auditing LLM-Based OCR Post-Correction for Textual Authenticity in 18th-Century Literary Editing.
    Background Context:

    Fast-growing technologies of Large Language Models (LLMs) have brought significant changes to Optical Character Recognition (OCR) post-correction in terms of increasing legibility of digitised historical documents. Thus, they open up numerous prospects for developing the technology of Historical Digitisation, especially in cases of digitisation of large collections of 18th-Century studies, for which OCR correction requires substantial amounts of time and effort. Kanerva et al. (2025) studied the effectiveness of using LLMs for OCR post-correction of historical English and Finnish documents, proving that LLMs can be very effective in reducing errors in English text while pointing out certain limitations of LLM use in OCR correction. It is clear from their results that LLMs may change historical wording in their attempt to increase OCR legibility, produce fictional information or even distort text authenticity during the correction process. Post-correction practices have so far been limited to increasing OCR efficiency regardless of validation of the historical authenticity of the edited text.

    PhD-Level Verification:

    The paper measures the effectiveness of the LLM-based OCR post-correction process based on decreases in character errors rate in historical datasets. Nevertheless, it fails to establish the auditing methodology that would be able to detect fake corrections, measure semantic accuracy, and maintain historical integrity while performing automatic post-correction. This is a significant potential area for improvement in OCR Research.

    Research Questions:
  • How can LLM-generated OCR corrections preserve historical authenticity?
  • Which validation techniques can identify fabricated or altered historical content?
  • Can explainable auditing improve Digital Corpus construction using LLM-assisted OCR correction?
  • Contributions at the PhD-Level:
  • Development of an explainable auditing framework for LLM-based OCR post-correction.
  • Integration of semantic fidelity assessment with automated historical text correction.
  • Guidelines for trustworthy AI-assisted editing of digitised historical manuscripts.
  • Suggested Readings:

    Kanerva, J., Ledins, C., Käpyaho, S., & Ginter, F. (2025). OCR Error Post-Correction with LLMs in Historical Documents: No Free Lunches. arXiv:2502.01205..

    Proposed Dissertation Topic 5: Building a Large-Scale Multi-Hand Benchmark for Transformer-Based Handwritten Text Recognition of 18th-Century Manuscripts
    Background Context:

    The digital preservation of historical archives provides unique opportunities for the preservation of cultural heritage as well as the ability for computational analysis of historical documents. With the growth of handwritten document collections, Handwritten Text Recognition (HTR) is required to convert the manuscripts to machine-readable form for further scholarly research. In the review paper by AlKendi, Gechter, Heyberger, and Guyeux (2024), recent developments in handwritten text recognition are discussed with emphasis on improvements in deep learning and transformer recognition models. Although much has been achieved in this area, there are still some challenges in the use of HTR systems associated with the diversity of handwriting, lack of appropriate benchmark datasets, linguistic diversity and inadequate evaluation metrics for historical manuscripts. Currently available HTR models and their benchmarks are developed using small language-specific datasets, making it impossible for the models to generalize on a large scale of historical archives.

    PhD-Level Verification:

    The research gives an overview of the various methods employed in handwritten text recognition (HTR), datasets used, and future work needed in the area but fails to provide any benchmark datasets to evaluate transformer-based HTR models. This creates room for improving OCR Research through the creation of benchmark datasets to aid in the evaluation of models and generalisation to different datasets.

    Research Questions:
  • How can transformer-based HTR models improve recognition across diverse historical handwriting styles?
  • Which benchmark characteristics best support reliable historical text recognition?
  • Can large-scale benchmark datasets improve Digital Corpus construction from handwritten historical collections?
  • PhD-Level Contributions:
  • Development of a large-scale benchmark dataset for historical handwritten text recognition.
  • Integration of transformer-based evaluation methods for diverse manuscript collections.
  • Framework supporting scalable transcription and preservation of historical archives for digital scholarship.
  • Suggested Readings:

    AlKendi, W., Gechter, F., Heyberger, L., & Guyeux, C. (2024). Advancements and Challenges in Handwritten Text Recognition: A Comprehensive Survey. Journal of Imaging, 10(18).

    Need assistance finalising your dissertation topic? Selecting a strong, researchable topic can be challenging — but you don’t have to do it alone.
    Our research consultants can help refine your ideas, identify literature gaps, and guide you toward a topic that aligns with current academic trends and your programme requirements.
    Contact us to begin one-on-one topic development and refinement with PhdAssistance.com Research Lab.

    Share this:

    Cite this work

    Study Resources

    Free resources to assist you with your university studies!

    Research Questions