Skip to main content

phdassistance

Transformer-Based Named Entity Recognition for Arabic–English Code-Switched Clinical Notes in UAE Electronic Health Records

NLP in healthcare is a broad research area. A more specific research question is how transformer-based NER models can identify clinical entities in Arabic–English code-switched notes. Clinical notes in UAE healthcare settings may contain Arabic-script text, English clinical terminology and transliterated Arabic; their prevalence and distribution should be established through corpus analysis rather than assumed. Within that setting, a scholar can build, annotate and evaluate a named entity recognition (NER) model that extracts diagnoses, symptoms and medications from such text.

Who this is for: PhD scholars in computational linguistics, computer science, health informatics and data science; clinician-informaticians; and supervisors scoping NLP projects in the UAE and GCC.

Scope: This article separates published evidence, cited to source, from research hypotheses, labelled as such, and design guidance from PhD Assistance. This is only research methodology advice and not medical advice; none of the extraction systems outlined here should ever be used to determine any patient’s course of treatment. Please confirm all access, ethics and governance policies with your institution and health authority before submission.

Prepared by PhD Assistance’s research-methodology team. The technical content should undergo subject-matter review by qualified specialists in clinical NLP and UAE research governance before publication.

1. Why clinical text in the UAE is a distinct NLP research problem

The problem is not that Arabic NLP is hard in general; it is that the clinical variety of it is largely under-resourced. A scoping review of freely available Arabic corpora identified 48 sources, categorised as multipurpose, dialectal, sentiment, speech and image-based, and concluded that Arabic is underrepresented relative to English; no clinical or medical corpus appeared among them [1]. This reported scarcity identifies a potential research opportunity, subject to verification through an updated and reproducible literature search, and it is also the project’s largest risk, because a scholar who cannot obtain text cannot run the study.

The research gap and the search behind it. A preliminary literature review suggests a potential gap in publicly available, annotated Arabic–English code-switched clinical NER datasets from the UAE and wider Gulf region. This gap requires confirmation through a reproducible search of relevant bibliographic databases, conference proceedings and research repositories — including ACL Anthology, PubMed, Scopus, IEEE Xplore, arXiv, LREC and WANLP. The final review should document the search date, Boolean search strings, inclusion and exclusion criteria, and screening results.

Recent work narrows, but does not close, this gap, and should not be read as leaving Arabic clinical NER entirely unbuilt. Human-verified, distant-supervision benchmarking has compared recurrent and transformer architectures for Arabic clinical mental-health NER [17]; large language models have separately been applied to Arabic mental-health entity extraction [19]; and a comparable span-based annotation and agreement framework has been demonstrated for Korean–English mixed-language clinical notes [18]. What this recent work does not supply is specific: (1) a mapped inventory of existing Arabic clinical NER datasets and tasks; (2) a treatment of Arabic–English code-switching — as distinct from monolingual Arabic clinical text — as its own language and annotation problem; and (3) a publicly available, annotated UAE-specific Arabic–English code-switched clinical benchmark, whose absence requires independent verification rather than assertion.

This approach makes the proposed research gap and methodological contribution open to systematic academic examination. In a thesis, one can identify a particular text genre, construct a corpus using a specific methodology, analyse the corpus with an acceptable level of inter-rater reliability, test particular model types on the corpus, and evaluate entity extraction performance at the language level. These tasks become separate chapters, each delivering an object for a reviewer’s consideration.

2. What code-switched and transliterated clinical notes actually look like

Article - Recreation Image 1 - PA

Annotated synthetic clinical note titled “One clinical note, three languages”, labelled as a synthetic example that is not a real patient note. Three lines in a progress note include words from Arabic script, English clinical terms, and Arabic words written in the Latin alphabet, each highlighted in a coloured box labelled Diagnosis, Symptom, Medication, Dosage, and Negation, respectively, and labelled AR, EN, and TRANSLIT. Below the lines, there are explanations that script switching breaks the tokeniser, transliterated Arabic cannot be recognised by either English or Arabic models, negation may be encoded in either language and clinical abbreviations in English occur everywhere.

Three properties of this text defeat assumptions built into most clinical NLP pipelines. Script switching happens mid-sentence, so a tokeniser tuned to one script fragments the other. Transliterated Arabic — Arabic words written in Latin characters may not align consistently with the pretrained vocabularies and language representations of English and Arabic models, potentially affecting tokenisation and recognition accuracy.

This is merely a hypothesis with regard to your context, as the frequency and nature of code-switching in the UAE’s clinical records have not been studied in the existing literature. Quantifying it — what proportion of notes code-switch, in which fields, and with what entity density — is a legitimate first study, and it de-risks everything that follows.

3. Why English clinical NLP and general Arabic models both fall short

English-only clinical encoders may inadequately represent Arabic-script segments because their tokenizers and pre-training data are not designed for Arabic linguistic structures. Subword fragmentation and vocabulary mismatch can reduce representation quality, although the actual impact should be measured empirically rather than assumed to be total failure. A second, subtler limitation is that clinical models are pre-trained on English clinical registers, so the very advantage they carry — exposure to abbreviations, dosage patterns and note structure — does not extend to the Arabic segments.

General Arabic models have the mirror problem. AraBERT [4], CAMeLBERT [5] and ARBERT/MARBERT [6] are pre-trained on news, web and social text; MARBERT in particular targets dialectal Arabic. None was pre-trained on clinical registers, and the CAMeLBERT study found that proximity of pre-training variant to the fine-tuning data mattered more than pre-training size [5] — which is precisely the argument for treating clinical Arabic as its own variant rather than assuming a general Arabic model transfers.

Multilingual models such as mBERT [2] and XLM-R [3] cover both scripts in one vocabulary, which is a genuine advantage here. Arabic transliteration and Arabizi introduce spelling variation and non-standard character sequences that may be inadequately represented in existing Arabic and multilingual model training data, even where some patterns overlap with multilingual pre-training text. Their effect should be evaluated separately using transliteration-specific test subsets, with normalisation rules, alternative transliteration forms, digit-to-character mapping and handling of inconsistent spellings specified in advance, and original-versus-normalised performance compared directly. The effect of transliteration on model performance therefore represents a testable research hypothesis rather than an established finding.

4. Defining the entity schema: diagnoses, symptoms, medications and negation

The schema represents the study’s critical decision, the one which most studies tend to leave unspecified. Make decisions regarding this first, before any text can be classified, such as determining what kinds of entities will be there (diagnoses, symptoms, medications, dosages, procedures); whether dosages will be classified as entities or as attributes of medications; whether negations are entities, attributes of other entities, or separate layers for annotation; whether tokens that are transliterations of their Arabic counterparts are labelled in the same category; and the annotation boundaries of tokens that span multiple writing systems.

Every decision influences the reported F1 score. These decisions should therefore be prespecified in the annotation protocol before model evaluation begins. A smaller schema with clearly defined categories and higher annotator agreement may be more defensible than an elaborate schema with inconsistent annotation.

5. Building a clinical corpus ethically: access, de-identification and approval

Federal Law No. 2 of 2019 concerning the use of information and communications technology in health fields establishes restrictions on the handling of health data associated with healthcare services provided within the UAE. Article 13 states that such health data and information may not be stored, processed, generated or transformed outside the UAE unless a resolution is issued by the Health Authority in coordination with the Ministry. For this proposed study, the researcher must therefore identify the location of data storage and processing, including hosting, backups and any cross-border transfers, and establish the applicable approval requirements before accessing or processing clinical records. The applicability of other provisions of the Law must be assessed separately against the official operative text and the proposed research activities [11].

The executive regulation for Federal Law No. 2 of 2019 is set out in Cabinet Resolution No. 32 of 2020 [12]. Ministerial Resolution No. 51 of 2021 separately specifies the circumstances and conditions under which health information may be stored, processed, generated or transferred outside the State; its exceptions and attached safeguards should be read from, and cited to, the operative text itself rather than summarised as a general permission [13]. A proposal should also assess the UAE Personal Data Protection Law, Federal Decree-Law No. 45 of 2021, for applicability and exemptions relevant to the specific processing activity, without assuming it applies identically to every health-data use.

Beyond federal legislation, Abu Dhabi’s Department of Health and the Dubai Health Authority set their own requirements, and each institution has its own research ethics committee and data custodian. A defensible methodology chapter names the specific health authority, institutional ethics committee and data custodian the researcher will approach, rather than treating governance as a single generic approval step.

Corpus sampling should also be specified before data collection begins: whether one or several hospital sites will contribute notes; which note types are in scope (for example, progress notes, discharge summaries, triage notes); whether inpatient and outpatient records are pooled or analysed separately; how repeated records from the same patient are handled across the sampling frame; and how the target corpus size and annotation workload are estimated from expected entity density. These decisions determine representativeness and should be stated as inclusion and exclusion criteria, not left implicit.

De-identification is a research task, not a preprocessing footnote. The i2b2/UTHealth corpus work shows what a documented protocol looks like: an explicit list of identifier categories, annotated at scale, with the annotation process itself reported [8]. In code-switched text, the task is harder because names and places appear in two scripts and in transliteration, so a de-identification system validated on English will under-detect. Report residual re-identification risk rather than claiming the text is anonymous, and have the approach approved rather than assumed.

Name the fallback. If corpus access is delayed, a study built on synthetic code-switched clinical text — generated to a documented specification and validated by clinicians for plausibility — can establish the schema, the annotation protocol and the model comparison. It is publishable with the narrower claim stated: it evaluates the method, not performance on real UAE records.

6. Annotation design and inter-annotator agreement

Annotation quality bounds everything downstream, so it is measured, not asserted. Double-annotate a defined proportion of the corpus — all of it where resources allow, otherwise a prespecified random sample large enough to estimate agreement with usable precision. Report agreement by entity type, not pooled, and separately for Arabic-script, English and transliterated segments, because agreement typically differs across them and the differences are a finding.

Choose the agreement statistic deliberately. Cohen’s kappa assumes a fixed set of classifiable items, which span-based NER does not provide, since annotators disagree about where entities begin and end as well as what they are. For span-based NER, inter-annotator agreement can be evaluated using pairwise entity-level precision, recall and F1 under prespecified exact or partial matching rules; the relationship between such measures and chance-corrected agreement statistics is well characterised [9]. A complementary chance-corrected measure, such as an appropriately defined Krippendorff’s alpha, may also be considered — recent clinical NER work has reported span-level F1 alongside relaxed Krippendorff’s alpha [18]. The selected statistics should match the annotation unit and boundary protocol, and the adjudication decision should state which measure governed it.

Prespecify the adjudication process, the guideline-revision cycle, and when the guideline is frozen. A guideline that keeps changing during annotation produces a corpus annotated to several different standards — a defect that surfaces at viva.

An annotation guideline should record, for each span, its source language or script, its entity category and the specific annotation decision applied. The following is an illustrative example only, not real patient data:

Text segment

Language

Entity

Annotation decision

Arabic diagnosis

Arabic script

Diagnosis

Diagnosis label

English medical term

English

Symptom

Symptom label

Arabizi expression

Transliteration

Symptom

Agreed canonical category

Medication name

English / Arabic

Medication

Medication label

Negation expression

Arabic / English

Negation

Contextual negation annotation

These are illustrative categories, not real patient data.

7. Transformer models compared: multilingual, Arabic-specific and clinical

A defensible thesis compares named model families, on fixed checkpoints, on one fixed benchmark with one fixed split, not a single model against published numbers from elsewhere.

Reproducing this comparison requires naming the actual checkpoint used for every model family, not only the architecture family, and explaining how model families will be compared fairly when their tokenisers and input representations differ:

Model/checkpoint

Architecture and size

Tokenizer / maximum length

Pre-training domain

Implementation and reporting notes

mBERT (bert-base-multilingual-cased)

Transformer encoder; ~177M

WordPiece; 512

Multilingual Wikipedia

Fine-tune with task-specific head; report seed and checkpoint version

XLM-R (xlm-roberta-base / large)

Transformer encoder; ~270M / ~550M

SentencePiece; 512

Multilingual CommonCrawl

Specify whether base or large is used

AraBERT v2 (aubmindlab/bert-base-arabertv2)

Transformer encoder; ~136M

WordPiece; 512

Arabic news, Wikipedia and web

Specify variant and Farasa pre-segmentation

CAMeLBERT-Mix (CAMeL-Lab/bert-base-arabic-camelbert-mix)

Transformer encoder; ~110M

WordPiece; 512

Mixed MSA, dialectal and classical Arabic

Identify the exact CAMeLBERT variant

MARBERT (UBC-NLP/MARBERT)

Transformer encoder; ~163M

WordPiece; 512

Dialectal Arabic and Twitter

Report separately from ARBERT

ARBERT (UBC-NLP/ARBERT)

Transformer encoder; ~163M

WordPiece; 512

Modern Standard Arabic

Report separately from MARBERT

Bio_ClinicalBERT

Transformer encoder; approximately 110M, subject to checkpoint verification

WordPiece; 512, subject to checkpoint verification

English clinical notes

Candidate baseline; exact checkpoint pending selection

LLM prompting baseline

Decoder-only LLM

Vendor tokenizer; model-dependent

Vendor-dependent

Record provider, model version, access date and prompt configuration

Also record the hardware used for fine-tuning and, where feasible, repeat key runs under more than one random seed to report variance alongside the point estimate.

Model family and examples

What it brings

What to watch

Multilingual: mBERT, XLM-R

Shared multilingual representations across Arabic and English scripts

Clinical-domain mismatch and potentially limited representation of transliterated forms

Arabic-specific: AraBERT, CAMeLBERT, ARBERT, MARBERT

Arabic morphology, dialect and language-specific representation

Pre-training data may not reflect clinical language

Clinical (English): Bio_ClinicalBERT

Exposure to English clinical abbreviations, dosage patterns and note structures

Arabic-script segments may experience vocabulary and representation mismatch

Transliteration normalisation: applicable model plus normalisation

Maps selected transliterated forms towards predefined canonical forms

Normalisation can introduce errors; compare original and normalised inputs

LLM prompting baseline

Enables an initial extraction baseline without task-specific labelled training data

Prompt sensitivity, reproducibility and data-governance requirements

Report the comparison as a finding, including a null one. The hypothesis that an Arabic-specific model beats a multilingual one on this text is reasonable and untested; so is the reverse. Fix the split by patient rather than by note, so that notes from the same patient cannot appear in both training and test data — a leak that inflates every number reported.

8. Low-resource strategies: transfer learning, augmentation and active learning

Three low-resource strategies can be considered, each with distinct methodological advantages and limitations. Transfer learning — fine-tuning from a general Arabic or multilingual checkpoint, optionally with an intermediate step on unlabelled in-domain text — is the default, and its benefit depends on how close that checkpoint’s variant is to clinical text [5]. Data augmentation by means of artificial code-switched sentences created based on known rules will expand the coverage without extra costs but will train the model in code-switching patterns which might be unnatural for humans; use natural texts only for evaluation. Active learning will save annotation costs by choosing only informative instances, but the resulting labelled sample will be non-random.

Code-switching benchmarks such as LinCE [7] offer a template for how multiple code-switched corpora and tasks are packaged for comparable evaluation, even though its language pairs are not Arabic–English clinical.

9. Evaluation examiners expect: entity-level F1 and per-language error analysis

Provide results for entity-level precision, recall, and F1 score rather than token-level accuracy, as the latter is misleading since there are many non-entity tokens in the text. Calculate these scores for two types of matching: strict matching, where the exact span and class should match, and relaxed or partial matching, where at least some overlap is allowed. Both are reported for biomedical entity recognition tasks [10]. State which is primary before results are seen.

The comparison across model families should be interpreted statistically, not read off a single reported number:

Evaluation component

Required methodology

Precision

Correctly predicted entities relative to all predicted entities

Recall

Correctly predicted entities relative to all reference entities

Strict F1

Exact span and entity type agreement

Relaxed F1

Prespecified partial-overlap matching

Macro F1

Average performance across entity classes

Micro F1

Aggregate entity-level performance

Confidence intervals

Prespecified bootstrap or other justified estimation procedure

Model comparison

Paired evaluation on the same held-out test cases

Error analysis

Boundary, type, transliteration, negation and script errors

The statistical comparison should account for the paired nature of model predictions, since every model is evaluated on the same held-out notes. State in advance how uncertainty (confidence intervals) and multiple model comparisons will be handled, and report macro and micro averages separately given the expected class imbalance across entity types; patient-level splitting should be preserved throughout, including for any bootstrap resampling.

Three further requirements separate a strong evaluation from a weak one:

  • Per-entity reporting. Performance may vary substantially between medication and symptom entities, depending on their frequency, annotation complexity and linguistic variation; a pooled figure hides that.
  • Per-language-segment reporting. Report Arabic-script, English and transliterated segments separately. This is the central claim of the thesis, and pooling it away forfeits the contribution.
  • Confidence intervals and a stated error analysis. Sample misclassified spans systematically, categorise the errors — boundary, type, script-boundary, transliteration, abbreviation, negation scope — and report the distribution.

10. From extracted entities to standard terminologies and research use

Extraction is only half of the value: unlinked strings are not interoperable. Mapping entities to SNOMED CT for clinical findings and procedures [14], ICD-11 for diagnoses [15] and RxNorm for medications [16] turns a model output into something reusable. Treat the mapping as a measured step with its own accuracy, not an assumed one, and report the proportion of mentions that map, that map ambiguously, and that have no target concept — the last being especially likely for Arabic-script and transliterated mentions.

State the limit. This produces structured research data. It does not produce a clinical decision-support tool, and a thesis that claims the latter will be asked for prospective evaluation it has not done.

11. PhD Assistance clinical-NLP corpus and annotation design checklist

Dimension

Evidence to hold before submission

Risk if unresolved

Evidence status (Available / Pending / Not established)

Data source

A named site, a custodian and a written access pathway

The commonest cause of abandoned clinical-NLP PhDs

To be confirmed

De-identification

An identifier category list, a documented method and an approved residual-risk statement

Ethics approval refused, or corpus unusable

To be confirmed

Entity schema

A written schema with every boundary rule decided

Results not comparable across annotators or with the literature

To be confirmed

Annotation protocol

A frozen guideline, double-annotated sample, adjudication rule

A corpus annotated to several standards

To be confirmed

Agreement target

A prespecified threshold, measured per entity type and per language segment [9]

Agreement reported only when favourable

To be confirmed

Evaluation plan

Primary matching criterion, per-entity and per-segment reporting, patient-level split

Reviewers unable to judge the contribution

To be confirmed

Any dimension without evidence behind it is a decision still to be taken or a limitation to declare — neither is acceptable left unstated.

12. How to turn this blueprint into an approvable PhD proposal

Existing evidence and research gaps

Existing evidence

Main contribution

Remaining gap

Ahmed et al. 2022 [1]

48 freely available Arabic corpora catalogued; none clinical

Does not establish or rule out a UAE-specific clinical resource

Arabic encoders [4, 5, 6]

Strong Arabic models; variant proximity is relevant

Clinical-register transfer remains unmeasured

Multilingual encoders [2, 3]

Shared vocabulary across scripts and cross-lingual transfer

Clinical-register and transliteration handling remain untested

LinCE [7]

Packaged benchmark for code-switched tasks

Not Arabic–English and not clinical

i2b2/UTHealth [8]

Documented de-identification annotation protocol

English-only; does not address code-switched or transliterated identifiers

Arabic mental-health NER [17]

Human-verified benchmark comparing recurrent and transformer architectures

Monolingual Arabic mental-health data, not UAE code-switched clinical notes

Korean–English clinical NER [18]

Span-based annotation and agreement framework

Different language pair; methodological lessons only

Arabic LLM-based NER [19]

LLM-based extraction from Arabic mental-health text

Not UAE-specific or code-switched; not benchmarked against fine-tuned transformers on the same split

12. How to turn this blueprint into an approvable PhD proposal

Existing evidence

Research implication

How the proposed UAE study addresses the gap

Ahmed et al. 2022 [1]

Establishes the need to investigate clinical Arabic corpus availability

Targets a UAE/Gulf clinical setting with a documented access and governance route

Arabic encoders [4, 5, 6]

Arabic models may be adaptable, but clinical transfer needs measurement

Benchmarks Arabic-specific encoders against clinical text

Multilingual encoders [2, 3]

Cross-script transfer does not establish transliteration performance

Adds a transliteration-specific evaluation subset

LinCE [7]

Demonstrates benchmark packaging for code-switched tasks

Adapts the benchmark approach to Arabic–English clinical text

i2b2/UTHealth [8]

De-identification needs a documented protocol

Includes Arabic-script and transliterated identifier categories with residual-risk reporting

Arabic mental-health NER [17]

Arabic clinical-adjacent NER is feasible, but code-switching remains untested

Extends the comparison to Arabic–English code-switching

Korean–English clinical NER [18]

Span-level annotation and agreement methods can inform mixed-language NER

Adapts the annotation and agreement design for Arabic–English clinical text

Arabic LLM-based NER [19]

LLM prompting is a possible baseline for Arabic clinical-adjacent extraction

Includes an LLM prompting baseline alongside fine-tuned transformers under one evaluation protocol

 

Research questions, with the inferential stance stated. RQ1 (primary): what entity-level F1 do multilingual, Arabic-specific and clinical transformers achieve on code-switched clinical notes, under one fixed split and schema? The null is that the families do not differ; the alternative is two-sided. RQ2: How does performance differ across Arabic-script, English and transliterated segments — estimation, not testing. RQ3: Does transliteration normalisation improve entity-level F1, and at what cost in introduced errors? RQ4 (exploratory): what proportion of extracted entities map to SNOMED CT, ICD-11 or RxNorm, and where does mapping fail?

Contribution. A completed thesis would deliver an annotated code-switched clinical corpus with a published schema and measured agreement; a benchmark comparison of model families on a fixed split; per-language-segment error analysis; and a documented de-identification and governance route reusable for other Gulf clinical-NLP work. Corpus size will bound generalisation, single-site text limits spectrum, annotation reflects the guideline rather than ground truth, and no clinical-use claim is supportable within a doctorate.

Common methodological weaknesses that may require proposal revision include: no named data source; no annotation plan or agreement measurement; no baseline comparison; token-level rather than entity-level metrics; splits made by note rather than by patient; de-identification assumed rather than validated; and a decision-support claim the design cannot support.

Frequently asked questions

General Arabic models such as AraBERT, CAMeLBERT and ARBERT/MARBERT are pre-trained on news, web and social text, not clinical registers, so vocabulary and domain mismatch can reduce representation quality for diagnoses, symptoms and medication mentions. The actual effect should be measured against a clinical corpus rather than assumed.

No. Entity-level strict and relaxed F1 should be reported per entity type and separately for Arabic-script, English and transliterated segments, together with confidence intervals and a structured error analysis, because a single pooled figure can hide where the model actually fails.

Cohen’s kappa assumes a fixed, enumerable set of items to classify, which span-based NER does not provide because annotators can disagree about where an entity begins and ends. Pairwise entity-level precision, recall and F1, optionally alongside a chance-corrected statistic such as Krippendorff’s alpha, is better suited to span annotation.

No. Entity extraction, even when mapped to SNOMED CT, ICD-11 or RxNorm, produces structured research data. It does not constitute a validated clinical decision-support tool, and a thesis should not claim readiness for clinical deployment without prospective evaluation that a doctoral timeline rarely supports.

Request a clinical-NLP topic and methodology review

Start a clinical-NLP methodology review with PhD Assistance’s research-methodology team: your proposed topic, corpus plan and evaluation design checked against this framework, delivered as a written recommendations summary you can take into a supervisor meeting.

  1. Initial topic assessment — your question mapped against published Arabic, code-switched and clinical NLP work, and where the novelty claim rests.
  2. Corpus and annotation review — data source, de-identification, entity schema, guideline and agreement plan.
  3. Modelling and evaluation check — model families to compare, split design, matching criterion and per-segment reporting.
  4. Recommendations and action points — a written summary of what is defensible, what needs evidence and what to change, to take into a supervisor meeting.

To start, share your draft topic, your target programme, and what you know about clinical text access and ethics approval. The review supports your own proposal writing and does not replace it, and it does not provide clinical guidance.

Related support: research proposal development · machine learning and statistical analysis · systematic review

Reference

  1. Ahmed A, Ali N, Alzubaidi M, Zaghouani W, Abd-alrazaq AA, Househ M. Freely available Arabic corpora: a scoping review. Comput Methods Programs Biomed Update. 2022;2:100049. doi:10.1016/j.cmpbup.2022.100049
  2. Devlin J, Chang M-W, Lee K, Toutanova K. BERT: pre-training of deep bidirectional transformers for language understanding. Proc 2019 Conf North American Chapter ACL: Human Language Technologies. 2019:4171–4186.
  3. Conneau A, Khandelwal K, Goyal N, et al. Unsupervised cross-lingual representation learning at scale. Proc 58th Annual Meeting ACL. 2020:8440–8451.
  4. Antoun W, Baly F, Hajj H. AraBERT: transformer-based model for Arabic language understanding. Proc 4th Workshop on Open-Source Arabic Corpora and Processing Tools (OSACT4). 2020:9–15.
  5. Inoue G, Alhafni B, Baimukan N, Bouamor H, Habash N. The interplay of variant, size, and task type in Arabic pre-trained language models. Proc 6th Arabic Natural Language Processing Workshop (WANLP). 2021:92–104.
  6. Abdul-Mageed M, Elmadany A, Nagoudi EMB. ARBERT & MARBERT: deep bidirectional transformers for Arabic. Proc 59th Annual Meeting ACL and 11th IJCNLP (Vol 1: Long Papers). 2021:7088–7105.
  7. Aguilar G, Kar S, Solorio T. LinCE: a centralized benchmark for linguistic code-switching evaluation. Proc 12th Language Resources and Evaluation Conference (LREC). 2020:1803–1813.
  8. Stubbs A, Uzuner Ö. Annotating longitudinal clinical narratives for de-identification: the 2014 i2b2/UTHealth corpus. J Biomed Inform. 2015;58:S20–S29. doi:10.1016/j.jbi.2015.07.020
  9. Hripcsak G, Rothschild AS. Agreement, the F-measure, and reliability in information retrieval. J Am Med Inform Assoc. 2005;12(3):296–298. doi:10.1197/jamia.M1733
  10. Segura-Bedmar I, Martínez P, Herrero-Zazo M. SemEval-2013 Task 9: extraction of drug-drug interactions from biomedical texts (DDIExtraction 2013). Proc 7th International Workshop on Semantic Evaluation (SemEval 2013). 2013:341–350.
  11. United Arab Emirates. Federal Law No. 2 of 2019 Concerning the Use of Information and Communications Technology in Health Fields, Arts 13, 16 and 20. Issued 6 February 2019; Official Gazette No. 647; in force 14 May 2019. https://uaelegislation.gov.ae/en/legislations/1209
  12. United Arab Emirates Cabinet. Cabinet Resolution No. 32 of 2020 Concerning the Executive Regulation of Federal Law No. 2 of 2019. Issued 22 April 2020; Official Gazette No. 677; in force 30 October 2020. https://uaelegislation.gov.ae/en/legislations/1444
  13. United Arab Emirates, Ministry of Health and Prevention. Ministerial Decree No. 51 of 2021 Regarding Cases in Which Health Data and Information May Be Stored or Transferred Outside the Country. 2021. Available at: https://mohap.gov.ae/en/w/ministerial-decree-no.-51-of-the-year-2021-regarding-cases-in-which-health-data-and-information-may-be-stored-or-transferred-outside-the-country
  14. SNOMED International. SNOMED CT. https://www.snomed.org (accessed 3 October 2026).
  15. World Health Organization. International Classification of Diseases, 11th Revision (ICD-11). https://icd.who.int (accessed 3 October 2026).
  16. US National Library of Medicine. RxNorm. https://www.nlm.nih.gov/research/umls/rxnorm (accessed 3 October 2026).
  17. Jalil AA, Sbera R. Arabic Clinical Mental Health Named Entity Recognition via Distant Supervision: Recurrent vs. Transformer-Based Architectures. Polytechnic Journal. 2026;16(2) 3. doi:10.59341/2707-7799.1882.
  18. Jang EH, Aguirre J, Lee S, Moon H, Cha WC. Span-based annotation framework for LLM-based clinical named entity recognition: development and validation using Korean emergency department notes. JAMIA Open. 2025;8(6) ooaf157. doi:10.1093/jamiaopen/ooaf157.
  19. Alfaifi A, Alhuzali H, Alasmari A. Named Entity Recognition in Arabic Mental Health Using Large Language Models. ACM Transactions on Asian and Low-Resource Language Information Processing. 2026;25(7) 57:1–27. doi:10.1145/3811876.