“Bias in healthcare AI” is a discussion topic. This is a measurable one: does a 30-day readmission model achieve the same discrimination, calibration and error rates for Emirati and expatriate patients — and if not, can a doctoral study show whether the gap comes from the data, the label or the model, and test whether it can be reduced without losing overall accuracy?
Who this is for: PhD scholars in health informatics, data science, public health and health services research; clinician-data scientists; and supervisors scoping AI-ethics projects in the UAE and GCC.
Scope. This article separates published evidence, cited to source, from research hypotheses, labelled as such, and design guidance. It is research-methodology guidance only. It does not direct discharge planning, resource allocation, or any decision about an individual patient or group; clinicians and the institution under their own governance make those decisions.
Readmission models are consequential because they are used to flag patients for discharge planning, to allocate follow-up calls and home visits, and as inputs to quality reporting. A model that under-ranks one group does not merely mis-score it — it withholds the attention the score is meant to trigger.
The research gap, and the search behind it. A UAE readmission prediction model already exists: a 2026 study at Al Amal Psychiatric Hospital in Dubai developed and validated a LASSO logistic model on 14,994 admissions from 9,263 patients (2018–2025), reaching AUC 0.649 (95% CI 0.623–0.678) in validation and outperforming LACE (AUC 0.557) [1]. Notably, that study reports differential missingness by nationality — non-UAE nationals had different odds of missing marital-status data — but does not report model performance stratified by nationality [1]. Our documented search did not identify a published fairness audit of a readmission model in a UAE or GCC tertiary setting reporting subgroup discrimination, calibration and error rates with confidence intervals. A claim of this kind is defensible only as the output of a reproducible search: record the databases queried, the Boolean strings, the date run, the limits and the screening criteria, and re-run before submission.
The gap is specific: the model exists, the audit does not. That is a stronger starting position than an unbuilt field, because a scholar can audit an existing model rather than build one first.
Define the prediction task first: the index admission, the discharge event that starts the clock, what counts as a readmission (all-cause or condition-specific, planned or unplanned), the 30-day window and its censoring rules, and the time origin at which the model scores. State whether the model ranks patients for a finite intervention budget or classifies them at a fixed threshold, because the fairness metric that matters differs between the two.
Use determines the harm. A model allocating a limited number of follow-up calls is a ranking problem, where what matters is who appears in the top decile. A model triggering a protocol at a cut-off is a classification problem, where error rates at that threshold matter. Audit the model as it is used, not in the abstract.
Nationality is not a mechanism; it is a label that bundles several. Treat each candidate attribute as a proxy for something specific and say which:
Attribute | Mechanism it plausibly proxies | What to watch |
Nationality (Emirati / expatriate) | Entitlement, continuity of care, likelihood of remaining in-country | Coarse; bundles several mechanisms at once |
Insurance tier | Access to follow-up, which facilities are in network | Varies by emirate and employer; may change mid-episode |
Primary language | Documentation completeness, history-taking depth | Often recorded inconsistently or not at all |
Length of residence | Record depth; prior admissions visible to the model | Frequently missing; may need derivation |
Prespecify the groups, a minimum cell size defined on events rather than patients, and the rule for missing attribute data — excluded, grouped as unknown, or imputed. Subgroup analyses may be underpowered, particularly when event counts are small; therefore, estimate expected precision in advance and prespecify which subgroup estimates will be considered informative, so compute expected confidence-interval widths in advance and state which subgroups can support an informative estimate [10]. Overlapping confidence intervals should not be used alone to determine whether subgroup performance differs. Compare the disparity directly using an appropriate statistical or resampling method and report the effect estimate with its confidence interval.
An ethics point, not an afterthought. Analysing a sensitive attribute to audit for disparity is a different purpose from using it as a model feature. State which you are doing, and have the distinction approved.
This is the most important section for a UAE study, and the one most often skipped. The label — “readmitted within 30 days” — is not observed directly; it is observed through the data system. Three mechanisms can make it differ systematically by group:
The canonical demonstration is Obermeyer and colleagues’ analysis of a widely used population-health algorithm that predicted health-care cost as a proxy for health need: because less was spent on Black patients at equal illness, patients at the same risk score were sicker, and the proportion identified for extra care rose from 17.7% to 46.5% when the label was corrected [4]. The mechanism transfers even though the setting does not — a label encoding access rather than the outcome of interest will encode the access disparity.
This is a hypothesis for your setting, not a finding. Whether label capture differs by nationality in UAE records is exactly what a first study should quantify — by linking to a wider data source where possible, by modelling censoring explicitly, and by reporting capture rates per subgroup as a study characteristic.
Audit something specific. The LACE index — length of stay, acuity, comorbidity and emergency visits — was derived and validated for early death or unplanned readmission after discharge [2], and the HOSPITAL score was derived for potentially avoidable 30-day readmissions in medical patients [3]. Both are transparent, cheap to compute and widely used, which makes them legitimate comparators rather than straw men; the UAE study above found LACE performed poorly in its setting (AUC 0.557) [1], which is itself the kind of local finding worth replicating.
Against those, train a modern model — regularised regression or gradient boosting — on local data. Split by patient, not by admission, so repeat admissions cannot appear in both training and test data; that leak inflates every number, including the subgroup numbers. Freeze the model specification, preprocessing rules and decision thresholds before subgroup performance is evaluated, and evaluate it on an independent dataset not used for model development or tuning.

The fairness audit pipeline: seven stages, and the decision each one forces. A research-design schematic showing seven stacked stages — data and cohort, subgroup definition, model under audit, fairness metrics, disparity diagnosis, mitigation test, and reporting — each with the decision to prespecify and what it produces. Colour bands mark whether a stage produces measurement, diagnosis, intervention or reporting. A note states that stage 5 separates an audit from a comparison, that metrics must be fixed before results are seen, and that the output is evidence about a model rather than a clinical or policy decision.
Metric | What it equalises | When it is the right question |
Demographic parity | Flag rate across groups | Rarely in clinical use; ignores real differences in risk |
Equalised odds | True-positive and false-positive rates [5] | A fixed-threshold model where both error types carry harm |
Equal opportunity | True-positive rate only [5] | When missing a high-risk patient is the dominant harm |
Calibration within groups | Predicted risk matches observed risk, per group | When the score is read as a probability by a clinician |
Predictive parity | Positive predictive value across groups | When the flag triggers a finite, costly intervention |
They conflict, and that is a mathematical result rather than a matter of effort. Where base rates differ between groups, a model cannot in general satisfy calibration within groups and equal false-positive and false-negative rates at the same time [6]. The practical consequence for a thesis: choose the metric set before seeing results, justify it from how the model is used, and report the others anyway so the trade-off is visible. Choosing a metric after seeing which one looks favourable is choosing the answer.
An observed subgroup gap is not yet a finding. Attributing it is what separates an audit from a comparison, and it is the thesis’s analytical core.
Report the decomposition, including the residual you cannot attribute. A thesis reporting how much of the gap each source explains, and how much remains unexplained, is stronger than one claiming a clean answer.
Four families, each with a stated cost. Reweighting training data to balance subgroup representation is simple and leaves the model class unchanged, but can degrade overall calibration. Threshold adjustment — different operating points per group — equalises chosen error rates transparently, but means the same score carries a different action depending on group membership, which needs explicit ethical and governance justification. Group-aware recalibration fixes calibration within groups while leaving ranking untouched, so it helps when the score is read as a probability, not when it is used to rank. Fairness-constrained training optimises accuracy subject to a fairness constraint and is implemented in open toolkits such as Fairlearn [7]; it can cost more accuracy than simpler methods for the same gain.
Report the cost, not just the gain. For each method, give the fairness metric, overall discrimination and calibration before and after, and the subgroup-level effect — mitigation that improves one group by harming another is a result to report, not to hide. State in advance what accuracy loss would make a mitigation unacceptable.
Report per subgroup, never pooled only: AUROC and AUPRC with confidence intervals; calibration slope and calibration-in-the-large, with a flexible curve estimated by a smoother rather than risk deciles; sensitivity, specificity and positive predictive value at the prespecified threshold; and the flag rate. State the minimum event count below which a subgroup estimate is not reported.
Map the write-up against TRIPOD+AI at protocol stage rather than at write-up [8], and appraise the model under audit using PROBAST+AI, which assesses quality, risk of bias and applicability for prediction models developed with regression or AI methods [9]. Using both signals to an examiner that the thesis knows the difference between reporting its own work and appraising someone else’s model.
Governance should be mapped across applicable federal requirements, emirate-level health-authority requirements and institution-specific ethics and data-custodian requirements. Federal Law No. 2 of 2019 — in force 14 May 2019 — provides that health data relating to health services provided in the State may not be stored, processed, generated or transferred outside the UAE without health-authority or ministerial approval (Art. 13), restricts use for non-health purposes without written consent or a stated exception (Art. 16), and sets a minimum retention period (Art. 20) [11]. Cabinet Resolution No. 32 of 2020 issues the executive regulation [12], and Ministerial Resolution No. 51 of 2021 permits transfer outside the State in stated circumstances, including scientific research, subject to conditions [13]. These are separate legislative records, to be cited separately; read the conditions for the exception you rely on from the instrument itself, as this article has not verified that Resolution’s operative text. Abu Dhabi’s Department of Health and the Dubai Health Authority set their own requirements, and the relevant institution may require ethics and/or data-custodian approval.
Two points specific to fairness work. Analysing nationality to detect disparity is a legitimate purpose that must still be declared and approved; and a subgroup result can identify a small group as well as describe it, so agree reporting granularity and minimum cell sizes with the ethics committee before analysis.
Dimension | Evidence to hold before submission | Risk if unresolved |
Subgroup definition | Prespecified groups, mechanisms named, minimum event counts | Post-hoc grouping; findings not defensible |
Fairness metric | A metric set fixed in advance, justified from model use | Choosing the metric after seeing the result |
Disparity source | A decomposition plan across data, label and model | A gap observed but unattributed |
Mitigation | A protocol stating acceptable accuracy loss, tested per subgroup | Mitigation that helps one group by harming another |
Reporting | A TRIPOD+AI-mapped plan [8] and a PROBAST+AI appraisal [9] | Reviewers unable to judge the audit |
Governance | Applicable federal, emirate-level and institution-specific approvals identified; sensitive-attribute purpose declared | Analysis complete but unpublishable |
Any dimension without evidence behind it is a decision still to be taken or a limitation to declare — neither is acceptable left unstated.
Evidence gap map
Existing evidence | Main contribution | Remaining gap |
Bokhari et al. 2026 [1] | A validated UAE 30-day readmission model; LACE performed poorly locally; differential missingness by nationality noted | No performance stratified by nationality; no fairness audit |
LACE [2]; HOSPITAL [3] | Transparent, widely used baselines | Derived elsewhere; no subgroup fairness evidence for UAE populations |
Obermeyer et al. 2019 [4] | Label choice produced measured racial bias at scale | US cost data; mechanism transfers, the numbers do not |
Hardt et al. [5]; Chouldechova [6] | Formal fairness criteria and their incompatibility | Methodological; not applied to a Gulf clinical cohort |
TRIPOD+AI [8]; PROBAST+AI [9] | Reporting and appraisal standards | Standards, not evidence about any model |
Research questions, with the inferential stance stated. RQ1 (primary): do subgroup AUROC, calibration and error rates differ between Emirati and expatriate patients for a frozen readmission model? The null is no difference, and the alternative is two-sided. RQ2: how much of any observed gap is attributable to data completeness, label capture and model behaviour, respectively — decomposition, reported with its unexplained residual. RQ3: do the chosen mitigation methods reduce the gap, and at what cost to overall discrimination and calibration — estimation, not testing. RQ4 (exploratory): does the pattern hold across insurance tier, language and length of residence, recognising that these estimates will be imprecise.
Contribution. A completed thesis would deliver subgroup performance and calibration estimates with confidence intervals for a UAE readmission model; an attributed decomposition of any disparity across data, label and model; a measured comparison of mitigation methods with their performance costs; and a reusable audit protocol mapped to TRIPOD+AI and PROBAST+AI. Single-site data bounds generalisation, subgroup estimates will be imprecise, label capture cannot be fully observed without linkage, and no clinical or policy claim is supportable within a doctorate.
Proposals get sent back for: no measurable outcome; no metric chosen in advance; a gap reported without attribution; no mitigation tested; splits made by admission rather than by patient; subgroup results without confidence intervals or minimum cell sizes; sensitive attributes analysed without a declared purpose; and a policy recommendation the design cannot support.
A fairness audit can form a viable PhD study when it addresses a substantive research gap and goes beyond measurement—for example, by investigating the sources of disparity, evaluating mitigation strategies and producing a defensible methodological contribution. An audit reporting a gap without diagnosing its source is a paper, not a thesis.
The one that matches how the model is used — equalised odds for a fixed-threshold alert, predictive parity for a finite intervention budget, calibration within groups when clinicians read the score as a probability. Fix it before results and report the others anyway.
Yes. Where base rates differ, calibration within groups and equal error rates generally cannot both hold [6], so a model can satisfy one criterion and fail another on the same data.
Then the audit is not feasible as specified, and that is worth knowing before you start. Insurance tier or language may be available where nationality is not; each proxies something different, so the research question changes with the variable.
No. Auditing an existing deployed or published model can be a legitimate and methodologically focused design, provided the model specification can be frozen and an independent evaluation dataset not used for model development or tuning can be obtained.
No. It can report audit findings and discuss what they would imply if confirmed. Decisions about care pathways rest with clinicians and the institution through their own governance.
PhD Assistance’s research-methodology team can review your proposed topic, audit design and feasibility against this framework.
To start, share your draft topic, your target programme, and what you know about data access and ethics approval. The review supports your own proposal writing and does not replace it, and it does not provide clinical guidance.
Related support: research proposal development · machine learning and biostatistics · systematic review