Info: Natural Language Processing of Clinical Notes for Automated ICD-10 Coding Accuracy Topics I phdassistance.com
Published: 17th August in Natural Language Processing of Clinical Notes for Automated ICD-10 Coding Accuracy Topics I phdassistance.com
Our academic writing and marking services can help you
The increasing number of EHRs demands effective and precise coding methods for clinical cases. Manual ICD-10 coding is usually associated with low speed and the inability to cope with long and complicated clinical texts. Modern advancements in artificial intelligence and NLP for Automated ICD-10 Coding offer possibilities to identify useful information and code diseases more effectively. Yet, the problems of large coding spaces, data imbalance, ambiguous terminology, long clinical documents, and the lack of model interpretability remain open to investigation.
The growing use of electronic health records has offered the chance of automation in the process of extracting diagnoses from unstructured medical texts. The use of Natural Language Processing in Medical Coding will enable the translation of clinical notes into coded diseases, thereby alleviating the burden of manual work and ensuring consistency. But coding is quite difficult when the clinical condition is uncommon. Lee & Im (2025) studied the automatic classification of ICD based on clinical text and found promising results in various stages of cause-of-death classification. Nevertheless, the researchers’ results have revealed that there is greater difficulty in classifying rare illnesses and underrepresented categories owing to scarcity of training samples. Another finding of the researchers was the existence of ambiguity and confusion in clinical language, as well as similarity between diseases. Such findings give rise to the possibility of undertaking doctoral research and developing a special framework that will enhance classification accuracy of rare and clinically important diseases while keeping performance on common ones at the same high level.
This research shows the variability of model performance on rare compared to common diagnostic categories. This is a clear research gap when it comes to developing classification methods that are optimised for rare clinical conditions. A PhD research project can focus on developing strategies for data augmentation, transfer learning, and class-specific learning for rare ICD-10 categories.
Lee, S., & Im, G. (2025). Machine learning for automated cause-of-death classification from 2021 to 2022 in Korea: development and validation of an ICD-10 prediction model.
The incorporation of AI into documentation in the healthcare industry has brought about the use of Automated Medical Coding, especially in situations where there are huge numbers of electronic documents that need to be processed in an efficient manner. However, accurate predictions might not be enough in practical settings within the medical settings. Medical coders, physicians, auditors, and health administrators should understand the process behind an automatically coded document. The systematic review on the use of artificial intelligence in ICD coding done by Mousavi Baigi et al. (2025) revealed poor interpretability and transparency as the key problems. They pointed out that even if the predictions produced using deep learning models are highly accurate, very little information is generated regarding the clinical evidence behind such predictions. The need for transparency, interdisciplinary collaboration, and auditable coding was stressed. Such gaps in the existing literature present the scope for developing an architecture where, along with predicting the proper classification, it is possible to determine the clinical evidence behind it.
This research demonstrates the variability in the performance of models when dealing with rare classes versus common ones. This represents a research gap in developing classification models optimised for rare diseases. The following research ideas could be proposed in a PhD study: data augmentation for rare ICD-10 categories; transfer learning for rare ICD-10 categories; class-specific learning for rare ICD-10 categories.
Mousavi Baigi, S. F., Sarbaz, M., Darroudi, A., et al. (2025). Artificial Intelligence-based Automated International Classification of Diseases Coding: A Systematic Review.
Clinical documentation is full of facts that do not always get understood by applying the keyword approach to the documentation. It could mean that the fact stated within the note is about a condition present now, something diagnosed in the past, a disease that is suspected, negative findings, or something that happened before being admitted to the hospital. Clinical Notes NLP can help identify and interpret these different forms of clinical information by considering the context in which medical terms are used. Gaspar et al. (2025) evaluated the effectiveness of NLP in detecting instances of bleeding in hospital discharge notes against traditional ICD-10-based methods. The results showed that narrative processing provided clinically relevant information that would be challenging to detect using structured codes only. The NLP method also managed to discern clinically relevant and severe bleeding cases, as well as historical ones, while the rule-based one failed to represent temporal information effectively. Therefore, the results indicate the importance of contextual analysis of clinical narratives as an area for further doctoral study due to the ambiguous nature of the meaning of some medical terms depending on temporal or clinical context.
According to Gaspar et al., NLP can capture clinically significant data in discharge summaries that are not captured by conventional ICD-10 methods. This study was also able to prove the importance of differentiating between historical events and severity during identification of the need for further temporal reasoning and validation. As PhD level, this gives the chance of exploring the contextual understanding of clinical narratives which include temporal information, negations, and historical versus current diseases. The suggested research can help in developing a context-aware classification model as well as identifying the types of errors that occur in context.
Gaspar, F., Zayene, M., Coumau, C., et al. (2025). Natural Language Processing and ICD-10 Coding for Detecting Bleeding Events in Discharge Summaries: Comparative Cross-Sectional Study.
Large-scale clinical classification requires a lot of labelled medical documents, but making high-quality labels is expensive since trained health care specialists must analyse the documents and assign corresponding disease codes, especially since some disease groups can be quite frequent and others rare. Kaur and Ginige have analysed different algorithms to perform automated coding and have emphasised the role of machine learning and deep learning methods to assign ICD codes from clinical data. This reanalysis can serve as a starting point for further study in the development of better learning strategies that can take advantage of both labelled and unlabeled medical data. Active learning is one example where the model is allowed to pick the most valuable data to annotate by the expert. This approach has the potential to decrease the need for annotation of all data and to improve the process of learning difficult and rare diagnostic classes.
Kaur & Ginige have shown how useful the algorithms can be when applied to ICD coding tasks, providing an initial background for future work on more effective and efficient machine learning-based coding. ICD-10 Coding Automation reduce the manual effort involved in assigning diagnostic codes. One important issue that requires further research is the problem of generating large enough labeled clinical datasets. PhD research may involve the application of active learning for identifying the most useful clinical documents to label manually, and evaluating the improvements in annotation and classification of ICD-10 codes achieved in this way.
Kaur, R., & Ginige, J. A. (2018). Comparative analysis of algorithmic approaches for auto-coding with ICD-10-AM and ACHI.
The large number of categories in ICD-10 makes classification a complex problem for AI. There is a possibility that medical documents must be classified from general diseases to very specific diagnostic codes, where similar terms can occur for related codes. This means that considering each code as a separate class will make the problem even more complex. Moons et al. have explored deep learning algorithms that could classify clinical data into ICD codes and emphasised the importance of hierarchical objectives when it comes to handling big coding space and a small sample size. AI Medical Coding therefore requires approaches that can understand relationships between different levels of the ICD-10 classification system. Moons et al.’s study indicated that hierarchical methods may assist in helping the model learn about relationships between classification levels. But noisy clinical documentation and increasing coding hierarchy complexity continue to be challenging issues for automatic large-scale classification. That opens the way for building a hierarchical structure in which the model learns about the disease group, category, subcategory, and finally ICD code.
Moons et al. showed the possibilities of deep neural networks and hierarchical learning in automatic classification according to ICD codes. This shows that one may use hierarchical relations for solving the problems of large-scale classification. PhD level research might include an analysis of hierarchical approaches on noisy data sets. The experiment would be conducted to analyze the performance of the flat and hierarchical approaches through performance on a code level and hierarchical errors to evaluate the benefit from using ICD-10 hierarchy.
Moons et al. (2020). Deep neural networks for automatically classifying clinical data into ICD codes.
Need assistance finalising your dissertation topic in natural language processing in clinical notes? Selecting a strong, researchable topic can be challenging — but you don’t have to do it alone.
Our research consultants can help refine your ideas, identify literature gaps, and guide you toward a topic that aligns with current academic trends and your programme requirements.
Contact us to begin one-on-one topic development and refinement with PhdAssistance.com Research Lab.
PhDAssistance. (n.d.). Cybersecurity in business Dissertation Topics Retrieved January 28th, from https://phdassistance.com/topic/cybersecurity-business/
Jalolova, M., and Musawwir, M. “Cybersecurity in business Dissertation Topics for PhD Scholars.” PhDAssistance, https://phdassistance.com/topic/cybersecurity-business/ Accessed 28th January 2026.
Jalolova, M., and Musawwir, M. “Cybersecurity in business Dissertation Topics for PhD Scholars.” PhDAssistance, PhDAssistance, Web. 28th January 2026.
Jalolova, M., and Musawwir, M., n.d. Cybersecurity in business Dissertation Topics for PhD scholars. [online] Available at: https://phdassistance.com/topic/cybersecurity-business/ [Accessed 28th January 2026].
Jalolova M., Musawwir M. Cybersecurity in business Dissertation Topics for PhD scholars [Internet]. PhDAssistance; [cited 2026 28th January]. Available from: https://phdassistance.com/topic/cybersecurity-business/
Jalolova, M., and Musawwir, M. (n.d.). Cybersecurity in business Dissertation Topics for PhD scholars. Retrieved 28th January 2026, from https://phdassistance.com/topic/cybersecurity-business/
Jalolova, M., and Musawwir, M., Cybersecurity in business Dissertation Topics (n.d.) https://phdassistance.com/topic/cybersecurity-business/ accessed 28th January 2026.
Free resources to assist you with your university studies!