NLP in healthcare is a broad research area. A more specific research question is how transformer-based NER models can identify clinical entities in Arabic–English code-switched notes. Clinical notes in UAE healthcare settings may contain Arabic-script text, English clinical terminology and transliterated Arabic; their prevalence and distribution should be established through corpus analysis rather than assumed. Within that setting, a scholar can build, annotate and evaluate a named entity recognition (NER) model that extracts diagnoses, symptoms and medications from such text.
Who this is for: PhD scholars in computational linguistics, computer science, health informatics and data science; clinician-informaticians; and supervisors scoping NLP projects in the UAE and GCC.
Scope: This article separates published evidence, cited to source, from research hypotheses, labelled as such, and design guidance from PhD Assistance. This is only research methodology advice and not medical advice; none of the extraction systems outlined here should ever be used to determine any patient’s course of treatment. Please confirm all access, ethics and governance policies with your institution and health authority before submission.
Prepared by PhD Assistance’s research-methodology team. The technical content should undergo subject-matter review by qualified specialists in clinical NLP and UAE research governance before publication.
The problem is not that Arabic NLP is hard in general; it is that the clinical variety of it is largely under-resourced. A scoping review of freely available Arabic corpora identified 48 sources, categorised as multipurpose, dialectal, sentiment, speech and image-based, and concluded that Arabic is underrepresented relative to English; no clinical or medical corpus appeared among them [1]. This reported scarcity identifies a potential research opportunity, subject to verification through an updated and reproducible literature search, and it is also the project’s largest risk, because a scholar who cannot obtain text cannot run the study.
The research gap and the search behind it. A preliminary literature review suggests a potential gap in publicly available, annotated Arabic–English code-switched clinical NER datasets from the UAE and wider Gulf region. This gap requires confirmation through a reproducible search of relevant bibliographic databases, conference proceedings and research repositories — including ACL Anthology, PubMed, Scopus, IEEE Xplore, arXiv, LREC and WANLP. The final review should document the search date, Boolean search strings, inclusion and exclusion criteria, and screening results.
Recent work narrows, but does not close, this gap, and should not be read as leaving Arabic clinical NER entirely unbuilt. Human-verified, distant-supervision benchmarking has compared recurrent and transformer architectures for Arabic clinical mental-health NER [17]; large language models have separately been applied to Arabic mental-health entity extraction [19]; and a comparable span-based annotation and agreement framework has been demonstrated for Korean–English mixed-language clinical notes [18]. What this recent work does not supply is specific: (1) a mapped inventory of existing Arabic clinical NER datasets and tasks; (2) a treatment of Arabic–English code-switching — as distinct from monolingual Arabic clinical text — as its own language and annotation problem; and (3) a publicly available, annotated UAE-specific Arabic–English code-switched clinical benchmark, whose absence requires independent verification rather than assertion.
This approach makes the proposed research gap and methodological contribution open to systematic academic examination. In a thesis, one can identify a particular text genre, construct a corpus using a specific methodology, analyse the corpus with an acceptable level of inter-rater reliability, test particular model types on the corpus, and evaluate entity extraction performance at the language level. These tasks become separate chapters, each delivering an object for a reviewer’s consideration.

Annotated synthetic clinical note titled “One clinical note, three languages”, labelled as a synthetic example that is not a real patient note. Three lines in a progress note include words from Arabic script, English clinical terms, and Arabic words written in the Latin alphabet, each highlighted in a coloured box labelled Diagnosis, Symptom, Medication, Dosage, and Negation, respectively, and labelled AR, EN, and TRANSLIT. Below the lines, there are explanations that script switching breaks the tokeniser, transliterated Arabic cannot be recognised by either English or Arabic models, negation may be encoded in either language and clinical abbreviations in English occur everywhere.
Three properties of this text defeat assumptions built into most clinical NLP pipelines. Script switching happens mid-sentence, so a tokeniser tuned to one script fragments the other. Transliterated Arabic — Arabic words written in Latin characters may not align consistently with the pretrained vocabularies and language representations of English and Arabic models, potentially affecting tokenisation and recognition accuracy.
This is merely a hypothesis with regard to your context, as the frequency and nature of code-switching in the UAE’s clinical records have not been studied in the existing literature. Quantifying it — what proportion of notes code-switch, in which fields, and with what entity density — is a legitimate first study, and it de-risks everything that follows.
English-only clinical encoders may inadequately represent Arabic-script segments because their tokenizers and pre-training data are not designed for Arabic linguistic structures. Subword fragmentation and vocabulary mismatch can reduce representation quality, although the actual impact should be measured empirically rather than assumed to be total failure. A second, subtler limitation is that clinical models are pre-trained on English clinical registers, so the very advantage they carry — exposure to abbreviations, dosage patterns and note structure — does not extend to the Arabic segments.
General Arabic models have the mirror problem. AraBERT [4], CAMeLBERT [5] and ARBERT/MARBERT [6] are pre-trained on news, web and social text; MARBERT in particular targets dialectal Arabic. None was pre-trained on clinical registers, and the CAMeLBERT study found that proximity of pre-training variant to the fine-tuning data mattered more than pre-training size [5] — which is precisely the argument for treating clinical Arabic as its own variant rather than assuming a general Arabic model transfers.
Multilingual models such as mBERT [2] and XLM-R [3] cover both scripts in one vocabulary, which is a genuine advantage here. Arabic transliteration and Arabizi introduce spelling variation and non-standard character sequences that may be inadequately represented in existing Arabic and multilingual model training data, even where some patterns overlap with multilingual pre-training text. Their effect should be evaluated separately using transliteration-specific test subsets, with normalisation rules, alternative transliteration forms, digit-to-character mapping and handling of inconsistent spellings specified in advance, and original-versus-normalised performance compared directly. The effect of transliteration on model performance therefore represents a testable research hypothesis rather than an established finding.
The schema represents the study’s critical decision, the one which most studies tend to leave unspecified. Make decisions regarding this first, before any text can be classified, such as determining what kinds of entities will be there (diagnoses, symptoms, medications, dosages, procedures); whether dosages will be classified as entities or as attributes of medications; whether negations are entities, attributes of other entities, or separate layers for annotation; whether tokens that are transliterations of their Arabic counterparts are labelled in the same category; and the annotation boundaries of tokens that span multiple writing systems.
Every decision influences the reported F1 score. These decisions should therefore be prespecified in the annotation protocol before model evaluation begins. A smaller schema with clearly defined categories and higher annotator agreement may be more defensible than an elaborate schema with inconsistent annotation.
Federal Law No. 2 of 2019 concerning the use of information and communications technology in health fields establishes restrictions on the handling of health data associated with healthcare services provided within the UAE. Article 13 states that such health data and information may not be stored, processed, generated or transformed outside the UAE unless a resolution is issued by the Health Authority in coordination with the Ministry. For this proposed study, the researcher must therefore identify the location of data storage and processing, including hosting, backups and any cross-border transfers, and establish the applicable approval requirements before accessing or processing clinical records. The applicability of other provisions of the Law must be assessed separately against the official operative text and the proposed research activities [11].
The executive regulation for Federal Law No. 2 of 2019 is set out in Cabinet Resolution No. 32 of 2020 [12]. Ministerial Resolution No. 51 of 2021 separately specifies the circumstances and conditions under which health information may be stored, processed, generated or transferred outside the State; its exceptions and attached safeguards should be read from, and cited to, the operative text itself rather than summarised as a general permission [13]. A proposal should also assess the UAE Personal Data Protection Law, Federal Decree-Law No. 45 of 2021, for applicability and exemptions relevant to the specific processing activity, without assuming it applies identically to every health-data use.
Beyond federal legislation, Abu Dhabi’s Department of Health and the Dubai Health Authority set their own requirements, and each institution has its own research ethics committee and data custodian. A defensible methodology chapter names the specific health authority, institutional ethics committee and data custodian the researcher will approach, rather than treating governance as a single generic approval step.
Corpus sampling should also be specified before data collection begins: whether one or several hospital sites will contribute notes; which note types are in scope (for example, progress notes, discharge summaries, triage notes); whether inpatient and outpatient records are pooled or analysed separately; how repeated records from the same patient are handled across the sampling frame; and how the target corpus size and annotation workload are estimated from expected entity density. These decisions determine representativeness and should be stated as inclusion and exclusion criteria, not left implicit.
De-identification is a research task, not a preprocessing footnote. The i2b2/UTHealth corpus work shows what a documented protocol looks like: an explicit list of identifier categories, annotated at scale, with the annotation process itself reported [8]. In code-switched text, the task is harder because names and places appear in two scripts and in transliteration, so a de-identification system validated on English will under-detect. Report residual re-identification risk rather than claiming the text is anonymous, and have the approach approved rather than assumed.
Name the fallback. If corpus access is delayed, a study built on synthetic code-switched clinical text — generated to a documented specification and validated by clinicians for plausibility — can establish the schema, the annotation protocol and the model comparison. It is publishable with the narrower claim stated: it evaluates the method, not performance on real UAE records.
Annotation quality bounds everything downstream, so it is measured, not asserted. Double-annotate a defined proportion of the corpus — all of it where resources allow, otherwise a prespecified random sample large enough to estimate agreement with usable precision. Report agreement by entity type, not pooled, and separately for Arabic-script, English and transliterated segments, because agreement typically differs across them and the differences are a finding.
Choose the agreement statistic deliberately. Cohen’s kappa assumes a fixed set of classifiable items, which span-based NER does not provide, since annotators disagree about where entities begin and end as well as what they are. For span-based NER, inter-annotator agreement can be evaluated using pairwise entity-level precision, recall and F1 under prespecified exact or partial matching rules; the relationship between such measures and chance-corrected agreement statistics is well characterised [9]. A complementary chance-corrected measure, such as an appropriately defined Krippendorff’s alpha, may also be considered — recent clinical NER work has reported span-level F1 alongside relaxed Krippendorff’s alpha [18]. The selected statistics should match the annotation unit and boundary protocol, and the adjudication decision should state which measure governed it.
Prespecify the adjudication process, the guideline-revision cycle, and when the guideline is frozen. A guideline that keeps changing during annotation produces a corpus annotated to several different standards — a defect that surfaces at viva.
An annotation guideline should record, for each span, its source language or script, its entity category and the specific annotation decision applied. The following is an illustrative example only, not real patient data:
Text segment | Language | Entity | Annotation decision |
Arabic diagnosis | Arabic script | Diagnosis | Diagnosis label |
English medical term | English | Symptom | Symptom label |
Arabizi expression | Transliteration | Symptom | Agreed canonical category |
Medication name | English / Arabic | Medication | Medication label |
Negation expression | Arabic / English | Negation | Contextual negation annotation |
These are illustrative categories, not real patient data.
A defensible thesis compares named model families, on fixed checkpoints, on one fixed benchmark with one fixed split, not a single model against published numbers from elsewhere.
Reproducing this comparison requires naming the actual checkpoint used for every model family, not only the architecture family, and explaining how model families will be compared fairly when their tokenisers and input representations differ:
Model/checkpoint | Architecture and size | Tokenizer / maximum length | Pre-training domain | Implementation and reporting notes |
mBERT (bert-base-multilingual-cased) | Transformer encoder; ~177M | WordPiece; 512 | Multilingual Wikipedia | Fine-tune with task-specific head; report seed and checkpoint version |
XLM-R (xlm-roberta-base / large) | Transformer encoder; ~270M / ~550M | SentencePiece; 512 | Multilingual CommonCrawl | Specify whether base or large is used |
AraBERT v2 (aubmindlab/bert-base-arabertv2) | Transformer encoder; ~136M | WordPiece; 512 | Arabic news, Wikipedia and web | Specify variant and Farasa pre-segmentation |
CAMeLBERT-Mix (CAMeL-Lab/bert-base-arabic-camelbert-mix) | Transformer encoder; ~110M | WordPiece; 512 | Mixed MSA, dialectal and classical Arabic | Identify the exact CAMeLBERT variant |
MARBERT (UBC-NLP/MARBERT) | Transformer encoder; ~163M | WordPiece; 512 | Dialectal Arabic and Twitter | Report separately from ARBERT |
ARBERT (UBC-NLP/ARBERT) | Transformer encoder; ~163M | WordPiece; 512 | Modern Standard Arabic | Report separately from MARBERT |
Bio_ClinicalBERT | Transformer encoder; approximately 110M, subject to checkpoint verification | WordPiece; 512, subject to checkpoint verification | English clinical notes | Candidate baseline; exact checkpoint pending selection |
LLM prompting baseline | Decoder-only LLM | Vendor tokenizer; model-dependent | Vendor-dependent | Record provider, model version, access date and prompt configuration |
Also record the hardware used for fine-tuning and, where feasible, repeat key runs under more than one random seed to report variance alongside the point estimate.
Model family and examples | What it brings | What to watch |
Multilingual: mBERT, XLM-R | Shared multilingual representations across Arabic and English scripts | Clinical-domain mismatch and potentially limited representation of transliterated forms |
Arabic-specific: AraBERT, CAMeLBERT, ARBERT, MARBERT | Arabic morphology, dialect and language-specific representation | Pre-training data may not reflect clinical language |
Clinical (English): Bio_ClinicalBERT | Exposure to English clinical abbreviations, dosage patterns and note structures | Arabic-script segments may experience vocabulary and representation mismatch |
Transliteration normalisation: applicable model plus normalisation | Maps selected transliterated forms towards predefined canonical forms | Normalisation can introduce errors; compare original and normalised inputs |
LLM prompting baseline | Enables an initial extraction baseline without task-specific labelled training data | Prompt sensitivity, reproducibility and data-governance requirements |
Report the comparison as a finding, including a null one. The hypothesis that an Arabic-specific model beats a multilingual one on this text is reasonable and untested; so is the reverse. Fix the split by patient rather than by note, so that notes from the same patient cannot appear in both training and test data — a leak that inflates every number reported.
Three low-resource strategies can be considered, each with distinct methodological advantages and limitations. Transfer learning — fine-tuning from a general Arabic or multilingual checkpoint, optionally with an intermediate step on unlabelled in-domain text — is the default, and its benefit depends on how close that checkpoint’s variant is to clinical text [5]. Data augmentation by means of artificial code-switched sentences created based on known rules will expand the coverage without extra costs but will train the model in code-switching patterns which might be unnatural for humans; use natural texts only for evaluation. Active learning will save annotation costs by choosing only informative instances, but the resulting labelled sample will be non-random.
Code-switching benchmarks such as LinCE [7] offer a template for how multiple code-switched corpora and tasks are packaged for comparable evaluation, even though its language pairs are not Arabic–English clinical.
Provide results for entity-level precision, recall, and F1 score rather than token-level accuracy, as the latter is misleading since there are many non-entity tokens in the text. Calculate these scores for two types of matching: strict matching, where the exact span and class should match, and relaxed or partial matching, where at least some overlap is allowed. Both are reported for biomedical entity recognition tasks [10]. State which is primary before results are seen.
The comparison across model families should be interpreted statistically, not read off a single reported number:
Evaluation component | Required methodology |
Precision | Correctly predicted entities relative to all predicted entities |
Recall | Correctly predicted entities relative to all reference entities |
Strict F1 | Exact span and entity type agreement |
Relaxed F1 | Prespecified partial-overlap matching |
Macro F1 | Average performance across entity classes |
Micro F1 | Aggregate entity-level performance |
Confidence intervals | Prespecified bootstrap or other justified estimation procedure |
Model comparison | Paired evaluation on the same held-out test cases |
Error analysis | Boundary, type, transliteration, negation and script errors |
The statistical comparison should account for the paired nature of model predictions, since every model is evaluated on the same held-out notes. State in advance how uncertainty (confidence intervals) and multiple model comparisons will be handled, and report macro and micro averages separately given the expected class imbalance across entity types; patient-level splitting should be preserved throughout, including for any bootstrap resampling.
Three further requirements separate a strong evaluation from a weak one:
Extraction is only half of the value: unlinked strings are not interoperable. Mapping entities to SNOMED CT for clinical findings and procedures [14], ICD-11 for diagnoses [15] and RxNorm for medications [16] turns a model output into something reusable. Treat the mapping as a measured step with its own accuracy, not an assumed one, and report the proportion of mentions that map, that map ambiguously, and that have no target concept — the last being especially likely for Arabic-script and transliterated mentions.
State the limit. This produces structured research data. It does not produce a clinical decision-support tool, and a thesis that claims the latter will be asked for prospective evaluation it has not done.
Dimension | Evidence to hold before submission | Risk if unresolved | Evidence status (Available / Pending / Not established) |
Data source | A named site, a custodian and a written access pathway | The commonest cause of abandoned clinical-NLP PhDs | To be confirmed |
De-identification | An identifier category list, a documented method and an approved residual-risk statement | Ethics approval refused, or corpus unusable | To be confirmed |
Entity schema | A written schema with every boundary rule decided | Results not comparable across annotators or with the literature | To be confirmed |
Annotation protocol | A frozen guideline, double-annotated sample, adjudication rule | A corpus annotated to several standards | To be confirmed |
Agreement target | A prespecified threshold, measured per entity type and per language segment [9] | Agreement reported only when favourable | To be confirmed |
Evaluation plan | Primary matching criterion, per-entity and per-segment reporting, patient-level split | Reviewers unable to judge the contribution | To be confirmed |
Any dimension without evidence behind it is a decision still to be taken or a limitation to declare — neither is acceptable left unstated.
Existing evidence and research gaps
Existing evidence | Main contribution | Remaining gap |
Ahmed et al. 2022 [1] | 48 freely available Arabic corpora catalogued; none clinical | Does not establish or rule out a UAE-specific clinical resource |
Arabic encoders [4, 5, 6] | Strong Arabic models; variant proximity is relevant | Clinical-register transfer remains unmeasured |
Multilingual encoders [2, 3] | Shared vocabulary across scripts and cross-lingual transfer | Clinical-register and transliteration handling remain untested |
LinCE [7] | Packaged benchmark for code-switched tasks | Not Arabic–English and not clinical |
i2b2/UTHealth [8] | Documented de-identification annotation protocol | English-only; does not address code-switched or transliterated identifiers |
Arabic mental-health NER [17] | Human-verified benchmark comparing recurrent and transformer architectures | Monolingual Arabic mental-health data, not UAE code-switched clinical notes |
Korean–English clinical NER [18] | Span-based annotation and agreement framework | Different language pair; methodological lessons only |
Arabic LLM-based NER [19] | LLM-based extraction from Arabic mental-health text | Not UAE-specific or code-switched; not benchmarked against fine-tuned transformers on the same split |
Existing evidence | Research implication | How the proposed UAE study addresses the gap |
Ahmed et al. 2022 [1] | Establishes the need to investigate clinical Arabic corpus availability | Targets a UAE/Gulf clinical setting with a documented access and governance route |
Arabic encoders [4, 5, 6] | Arabic models may be adaptable, but clinical transfer needs measurement | Benchmarks Arabic-specific encoders against clinical text |
Multilingual encoders [2, 3] | Cross-script transfer does not establish transliteration performance | Adds a transliteration-specific evaluation subset |
LinCE [7] | Demonstrates benchmark packaging for code-switched tasks | Adapts the benchmark approach to Arabic–English clinical text |
i2b2/UTHealth [8] | De-identification needs a documented protocol | Includes Arabic-script and transliterated identifier categories with residual-risk reporting |
Arabic mental-health NER [17] | Arabic clinical-adjacent NER is feasible, but code-switching remains untested | Extends the comparison to Arabic–English code-switching |
Korean–English clinical NER [18] | Span-level annotation and agreement methods can inform mixed-language NER | Adapts the annotation and agreement design for Arabic–English clinical text |
Arabic LLM-based NER [19] | LLM prompting is a possible baseline for Arabic clinical-adjacent extraction | Includes an LLM prompting baseline alongside fine-tuned transformers under one evaluation protocol |
Research questions, with the inferential stance stated. RQ1 (primary): what entity-level F1 do multilingual, Arabic-specific and clinical transformers achieve on code-switched clinical notes, under one fixed split and schema? The null is that the families do not differ; the alternative is two-sided. RQ2: How does performance differ across Arabic-script, English and transliterated segments — estimation, not testing. RQ3: Does transliteration normalisation improve entity-level F1, and at what cost in introduced errors? RQ4 (exploratory): what proportion of extracted entities map to SNOMED CT, ICD-11 or RxNorm, and where does mapping fail?
Contribution. A completed thesis would deliver an annotated code-switched clinical corpus with a published schema and measured agreement; a benchmark comparison of model families on a fixed split; per-language-segment error analysis; and a documented de-identification and governance route reusable for other Gulf clinical-NLP work. Corpus size will bound generalisation, single-site text limits spectrum, annotation reflects the guideline rather than ground truth, and no clinical-use claim is supportable within a doctorate.
Common methodological weaknesses that may require proposal revision include: no named data source; no annotation plan or agreement measurement; no baseline comparison; token-level rather than entity-level metrics; splits made by note rather than by patient; de-identification assumed rather than validated; and a decision-support claim the design cannot support.
General Arabic models such as AraBERT, CAMeLBERT and ARBERT/MARBERT are pre-trained on news, web and social text, not clinical registers, so vocabulary and domain mismatch can reduce representation quality for diagnoses, symptoms and medication mentions. The actual effect should be measured against a clinical corpus rather than assumed.
No. Entity-level strict and relaxed F1 should be reported per entity type and separately for Arabic-script, English and transliterated segments, together with confidence intervals and a structured error analysis, because a single pooled figure can hide where the model actually fails.
Cohen’s kappa assumes a fixed, enumerable set of items to classify, which span-based NER does not provide because annotators can disagree about where an entity begins and ends. Pairwise entity-level precision, recall and F1, optionally alongside a chance-corrected statistic such as Krippendorff’s alpha, is better suited to span annotation.
No. Entity extraction, even when mapped to SNOMED CT, ICD-11 or RxNorm, produces structured research data. It does not constitute a validated clinical decision-support tool, and a thesis should not claim readiness for clinical deployment without prospective evaluation that a doctoral timeline rarely supports.
Start a clinical-NLP methodology review with PhD Assistance’s research-methodology team: your proposed topic, corpus plan and evaluation design checked against this framework, delivered as a written recommendations summary you can take into a supervisor meeting.
To start, share your draft topic, your target programme, and what you know about clinical text access and ethics approval. The review supports your own proposal writing and does not replace it, and it does not provide clinical guidance.
Related support: research proposal development · machine learning and statistical analysis · systematic review