Skip to main content

phdassistance

Algorithmic Fairness Auditing of a Machine-Learning 30-Day Readmission Model Across Emirati and Expatriate Subgroups in a UAE Tertiary Hospital: A PhD Research Framework

“Bias in healthcare AI” is a discussion topic. This is a measurable one: does a 30-day readmission model achieve the same discrimination, calibration and error rates for Emirati and expatriate patients — and if not, can a doctoral study show whether the gap comes from the data, the label or the model, and test whether it can be reduced without losing overall accuracy?

Who this is for: PhD scholars in health informatics, data science, public health and health services research; clinician-data scientists; and supervisors scoping AI-ethics projects in the UAE and GCC.

Scope. This article separates published evidence, cited to source, from research hypotheses, labelled as such, and design guidance. It is research-methodology guidance only. It does not direct discharge planning, resource allocation, or any decision about an individual patient or group; clinicians and the institution under their own governance make those decisions.

1. Why readmission models are a high-stakes fairness problem in UAE hospitals

Readmission models are consequential because they are used to flag patients for discharge planning, to allocate follow-up calls and home visits, and as inputs to quality reporting. A model that under-ranks one group does not merely mis-score it — it withholds the attention the score is meant to trigger.

The research gap, and the search behind it. A UAE readmission prediction model already exists: a 2026 study at Al Amal Psychiatric Hospital in Dubai developed and validated a LASSO logistic model on 14,994 admissions from 9,263 patients (2018–2025), reaching AUC 0.649 (95% CI 0.623–0.678) in validation and outperforming LACE (AUC 0.557) [1]. Notably, that study reports differential missingness by nationality — non-UAE nationals had different odds of missing marital-status data — but does not report model performance stratified by nationality [1]. Our documented search did not identify a published fairness audit of a readmission model in a UAE or GCC tertiary setting reporting subgroup discrimination, calibration and error rates with confidence intervals. A claim of this kind is defensible only as the output of a reproducible search: record the databases queried, the Boolean strings, the date run, the limits and the screening criteria, and re-run before submission.

The gap is specific: the model exists, the audit does not. That is a stronger starting position than an unbuilt field, because a scholar can audit an existing model rather than build one first.

2. What a 30-day readmission model predicts — and how hospitals use it

Define the prediction task first: the index admission, the discharge event that starts the clock, what counts as a readmission (all-cause or condition-specific, planned or unplanned), the 30-day window and its censoring rules, and the time origin at which the model scores. State whether the model ranks patients for a finite intervention budget or classifies them at a fixed threshold, because the fairness metric that matters differs between the two.

Use determines the harm. A model allocating a limited number of follow-up calls is a ranking problem, where what matters is who appears in the top decile. A model triggering a protocol at a cut-off is a classification problem, where error rates at that threshold matter. Audit the model as it is used, not in the abstract.

3. Defining subgroups responsibly: nationality, insurance, language and residence

Nationality is not a mechanism; it is a label that bundles several. Treat each candidate attribute as a proxy for something specific and say which:

Attribute

Mechanism it plausibly proxies

What to watch

Nationality (Emirati / expatriate)

Entitlement, continuity of care, likelihood of remaining in-country

Coarse; bundles several mechanisms at once

Insurance tier

Access to follow-up, which facilities are in network

Varies by emirate and employer; may change mid-episode

Primary language

Documentation completeness, history-taking depth

Often recorded inconsistently or not at all

Length of residence

Record depth; prior admissions visible to the model

Frequently missing; may need derivation

Prespecify the groups, a minimum cell size defined on events rather than patients, and the rule for missing attribute data — excluded, grouped as unknown, or imputed. Subgroup analyses may be underpowered, particularly when event counts are small; therefore, estimate expected precision in advance and prespecify which subgroup estimates will be considered informative, so compute expected confidence-interval widths in advance and state which subgroups can support an informative estimate [10]. Overlapping confidence intervals should not be used alone to determine whether subgroup performance differs. Compare the disparity directly using an appropriate statistical or resampling method and report the effect estimate with its confidence interval.

An ethics point, not an afterthought. Analysing a sensitive attribute to audit for disparity is a different purpose from using it as a model feature. State which you are doing, and have the distinction approved.

4. Label bias: when the outcome itself is measured differently across groups

This is the most important section for a UAE study, and the one most often skipped. The label — “readmitted within 30 days” — is not observed directly; it is observed through the data system. Three mechanisms can make it differ systematically by group:

  • Readmission to a different facility that the data source does not see, recording a true readmission as a non-event.
  • Patients leaving the country after discharge, which removes them from the risk set in a way that correlates with nationality and visa status.
  • Differential loss to follow-up, where a group with less continuous care appears to have fewer events because fewer are captured.

The canonical demonstration is Obermeyer and colleagues’ analysis of a widely used population-health algorithm that predicted health-care cost as a proxy for health need: because less was spent on Black patients at equal illness, patients at the same risk score were sicker, and the proportion identified for extra care rose from 17.7% to 46.5% when the label was corrected [4]. The mechanism transfers even though the setting does not — a label encoding access rather than the outcome of interest will encode the access disparity.

This is a hypothesis for your setting, not a finding. Whether label capture differs by nationality in UAE records is exactly what a first study should quantify — by linking to a wider data source where possible, by modelling censoring explicitly, and by reporting capture rates per subgroup as a study characteristic.

5. Baseline and machine-learning models: LACE, HOSPITAL score and gradient boosting

Audit something specific. The LACE index — length of stay, acuity, comorbidity and emergency visits — was derived and validated for early death or unplanned readmission after discharge [2], and the HOSPITAL score was derived for potentially avoidable 30-day readmissions in medical patients [3]. Both are transparent, cheap to compute and widely used, which makes them legitimate comparators rather than straw men; the UAE study above found LACE performed poorly in its setting (AUC 0.557) [1], which is itself the kind of local finding worth replicating.

Against those, train a modern model — regularised regression or gradient boosting — on local data. Split by patient, not by admission, so repeat admissions cannot appear in both training and test data; that leak inflates every number, including the subgroup numbers. Freeze the model specification, preprocessing rules and decision thresholds before subgroup performance is evaluated, and evaluate it on an independent dataset not used for model development or tuning.

6. Fairness metrics explained — and why they cannot all be satisfied at once

30-day readmission prediction; Emirati expatriate subgroup fairness; equalised odds calibration within groups; predictive parity trade-off; label bias clinical prediction; LACE HOSPITAL score; TRIPOD+AI PROBAST+AI reporting; Fairlearn fairness toolkit; health equity PhD topics UAE; clinical AI research proposal

The fairness audit pipeline: seven stages, and the decision each one forces. A research-design schematic showing seven stacked stages — data and cohort, subgroup definition, model under audit, fairness metrics, disparity diagnosis, mitigation test, and reporting — each with the decision to prespecify and what it produces. Colour bands mark whether a stage produces measurement, diagnosis, intervention or reporting. A note states that stage 5 separates an audit from a comparison, that metrics must be fixed before results are seen, and that the output is evidence about a model rather than a clinical or policy decision.

Metric

What it equalises

When it is the right question

Demographic parity

Flag rate across groups

Rarely in clinical use; ignores real differences in risk

Equalised odds

True-positive and false-positive rates [5]

A fixed-threshold model where both error types carry harm

Equal opportunity

True-positive rate only [5]

When missing a high-risk patient is the dominant harm

Calibration within groups

Predicted risk matches observed risk, per group

When the score is read as a probability by a clinician

Predictive parity

Positive predictive value across groups

When the flag triggers a finite, costly intervention

They conflict, and that is a mathematical result rather than a matter of effort. Where base rates differ between groups, a model cannot in general satisfy calibration within groups and equal false-positive and false-negative rates at the same time [6]. The practical consequence for a thesis: choose the metric set before seeing results, justify it from how the model is used, and report the others anyway so the trade-off is visible. Choosing a metric after seeing which one looks favourable is choosing the answer.

7. Diagnosing where disparity comes from: data, labels or model

An observed subgroup gap is not yet a finding. Attributing it is what separates an audit from a comparison, and it is the thesis’s analytical core.

  • Compare feature availability and missingness rates by subgroup; refit on the complete-case intersection and see whether the gap persists. If the disparity narrows materially, this would provide evidence that incomplete data contribute to the observed subgroup difference.
  • Estimate capture rates per subgroup (Section 4). If the gap shrinks under a censoring-aware or linked-data label, it is partly label bias.
  • If the gap survives matched features and a corrected label, examine the model: per-subgroup calibration curves, error analysis on misranked cases, and whether a simpler model shows the same pattern.

Report the decomposition, including the residual you cannot attribute. A thesis reporting how much of the gap each source explains, and how much remains unexplained, is stronger than one claiming a clean answer.

8. Mitigation methods and their performance trade-offs

Four families, each with a stated cost. Reweighting training data to balance subgroup representation is simple and leaves the model class unchanged, but can degrade overall calibration. Threshold adjustment — different operating points per group — equalises chosen error rates transparently, but means the same score carries a different action depending on group membership, which needs explicit ethical and governance justification. Group-aware recalibration fixes calibration within groups while leaving ranking untouched, so it helps when the score is read as a probability, not when it is used to rank. Fairness-constrained training optimises accuracy subject to a fairness constraint and is implemented in open toolkits such as Fairlearn [7]; it can cost more accuracy than simpler methods for the same gain.

Report the cost, not just the gain. For each method, give the fairness metric, overall discrimination and calibration before and after, and the subgroup-level effect — mitigation that improves one group by harming another is a result to report, not to hide. State in advance what accuracy loss would make a mitigation unacceptable.

9. Reporting a fairness audit: subgroup metrics, calibration and TRIPOD+AI

Report per subgroup, never pooled only: AUROC and AUPRC with confidence intervals; calibration slope and calibration-in-the-large, with a flexible curve estimated by a smoother rather than risk deciles; sensitivity, specificity and positive predictive value at the prespecified threshold; and the flag rate. State the minimum event count below which a subgroup estimate is not reported.

Map the write-up against TRIPOD+AI at protocol stage rather than at write-up [8], and appraise the model under audit using PROBAST+AI, which assesses quality, risk of bias and applicability for prediction models developed with regression or AI methods [9]. Using both signals to an examiner that the thesis knows the difference between reporting its own work and appraising someone else’s model.

10. Ethics and governance of sensitive attributes in UAE health data

Governance should be mapped across applicable federal requirements, emirate-level health-authority requirements and institution-specific ethics and data-custodian requirements. Federal Law No. 2 of 2019 — in force 14 May 2019 — provides that health data relating to health services provided in the State may not be stored, processed, generated or transferred outside the UAE without health-authority or ministerial approval (Art. 13), restricts use for non-health purposes without written consent or a stated exception (Art. 16), and sets a minimum retention period (Art. 20) [11]. Cabinet Resolution No. 32 of 2020 issues the executive regulation [12], and Ministerial Resolution No. 51 of 2021 permits transfer outside the State in stated circumstances, including scientific research, subject to conditions [13]. These are separate legislative records, to be cited separately; read the conditions for the exception you rely on from the instrument itself, as this article has not verified that Resolution’s operative text. Abu Dhabi’s Department of Health and the Dubai Health Authority set their own requirements, and the relevant institution may require ethics and/or data-custodian approval.

Two points specific to fairness work. Analysing nationality to detect disparity is a legitimate purpose that must still be declared and approved; and a subgroup result can identify a small group as well as describe it, so agree reporting granularity and minimum cell sizes with the ethics committee before analysis.

11. PhD Assistance clinical-AI fairness audit framework

Dimension

Evidence to hold before submission

Risk if unresolved

Subgroup definition

Prespecified groups, mechanisms named, minimum event counts

Post-hoc grouping; findings not defensible

Fairness metric

A metric set fixed in advance, justified from model use

Choosing the metric after seeing the result

Disparity source

A decomposition plan across data, label and model

A gap observed but unattributed

Mitigation

A protocol stating acceptable accuracy loss, tested per subgroup

Mitigation that helps one group by harming another

Reporting

A TRIPOD+AI-mapped plan [8] and a PROBAST+AI appraisal [9]

Reviewers unable to judge the audit

Governance

Applicable federal, emirate-level and institution-specific approvals identified; sensitive-attribute purpose declared

Analysis complete but unpublishable

Any dimension without evidence behind it is a decision still to be taken or a limitation to declare — neither is acceptable left unstated.

12. How to turn this blueprint into an approvable PhD proposal

Evidence gap map

Existing evidence

Main contribution

Remaining gap

Bokhari et al. 2026 [1]

A validated UAE 30-day readmission model; LACE performed poorly locally; differential missingness by nationality noted

No performance stratified by nationality; no fairness audit

LACE [2]; HOSPITAL [3]

Transparent, widely used baselines

Derived elsewhere; no subgroup fairness evidence for UAE populations

Obermeyer et al. 2019 [4]

Label choice produced measured racial bias at scale

US cost data; mechanism transfers, the numbers do not

Hardt et al. [5]; Chouldechova [6]

Formal fairness criteria and their incompatibility

Methodological; not applied to a Gulf clinical cohort

TRIPOD+AI [8]; PROBAST+AI [9]

Reporting and appraisal standards

Standards, not evidence about any model

Research questions, with the inferential stance stated. RQ1 (primary): do subgroup AUROC, calibration and error rates differ between Emirati and expatriate patients for a frozen readmission model? The null is no difference, and the alternative is two-sided. RQ2: how much of any observed gap is attributable to data completeness, label capture and model behaviour, respectively — decomposition, reported with its unexplained residual. RQ3: do the chosen mitigation methods reduce the gap, and at what cost to overall discrimination and calibration — estimation, not testing. RQ4 (exploratory): does the pattern hold across insurance tier, language and length of residence, recognising that these estimates will be imprecise.

Contribution. A completed thesis would deliver subgroup performance and calibration estimates with confidence intervals for a UAE readmission model; an attributed decomposition of any disparity across data, label and model; a measured comparison of mitigation methods with their performance costs; and a reusable audit protocol mapped to TRIPOD+AI and PROBAST+AI. Single-site data bounds generalisation, subgroup estimates will be imprecise, label capture cannot be fully observed without linkage, and no clinical or policy claim is supportable within a doctorate.

Proposals get sent back for: no measurable outcome; no metric chosen in advance; a gap reported without attribution; no mitigation tested; splits made by admission rather than by patient; subgroup results without confidence intervals or minimum cell sizes; sensitive attributes analysed without a declared purpose; and a policy recommendation the design cannot support.

Frequently asked questions

A fairness audit can form a viable PhD study when it addresses a substantive research gap and goes beyond measurement—for example, by investigating the sources of disparity, evaluating mitigation strategies and producing a defensible methodological contribution. An audit reporting a gap without diagnosing its source is a paper, not a thesis.

The one that matches how the model is used — equalised odds for a fixed-threshold alert, predictive parity for a finite intervention budget, calibration within groups when clinicians read the score as a probability. Fix it before results and report the others anyway.

Yes. Where base rates differ, calibration within groups and equal error rates generally cannot both hold [6], so a model can satisfy one criterion and fail another on the same data.

Then the audit is not feasible as specified, and that is worth knowing before you start. Insurance tier or language may be available where nationality is not; each proxies something different, so the research question changes with the variable.

No. Auditing an existing deployed or published model can be a legitimate and methodologically focused design, provided the model specification can be frozen and an independent evaluation dataset not used for model development or tuning can be obtained.

No. It can report audit findings and discuss what they would imply if confirmed. Decisions about care pathways rest with clinicians and the institution through their own governance.

Request a clinical-AI fairness topic feasibility and methodology review

 PhD Assistance’s research-methodology team can review your proposed topic, audit design and feasibility against this framework.

  1. Initial topic assessment — your question mapped against published readmission and fairness work, and where the novelty claim rests.
  2. Audit design review — subgroup definition, metric selection and justification, and the disparity-decomposition plan.
  3. Feasibility review — data access, subgroup cell sizes and expected precision, governance mapping across all three levels.
  4. Recommendations and action points — a written summary of what is defensible, what needs evidence and what to change, to take into a supervisor meeting.

To start, share your draft topic, your target programme, and what you know about data access and ethics approval. The review supports your own proposal writing and does not replace it, and it does not provide clinical guidance.

Related support: research proposal development · machine learning and biostatistics · systematic review

Reference

  1. Bokhari, S. A., & Javaid, S. F. Development and validation of a clinical prediction model and risk score for 30-day psychiatric readmission in the United Arab Emirates. BMC Psychiatry. 2026;26:514. doi:10.1186/s12888-026-08165-z
  2. van Walraven C, Dhalla IA, Bell C, et al. Derivation and validation of an index to predict early death or unplanned readmission after discharge from hospital to the community. CMAJ. 2010;182(6):551–557. doi:10.1503/cmaj.091117
  3. Donzé J, Aujesky D, Williams D, Schnipper JL. Potentially avoidable 30-day hospital readmissions in medical patients: derivation and validation of a prediction model. JAMA Intern Med. 2013;173(8):632–638. doi:10.1001/jamainternmed.2013.3023
  4. Obermeyer Z, Powers B, Vogeli C, Mullainathan S. Dissecting racial bias in an algorithm used to manage the health of populations. Science. 2019;366(6464):447–453. doi:10.1126/science.aax2342
  5. Hardt M, Price E, Srebro N. Equality of opportunity in supervised learning. Advances in Neural Information Processing Systems 29 (NIPS 2016). 2016:3315–3323.
  6. Chouldechova A. Fair prediction with disparate impact: a study of bias in recidivism prediction instruments. Big Data. 2017;5(2):153–163. doi:10.1089/big.2016.0047
  7. Weerts H, Dudík M, Edgar R, Jalali A, Lutz R, Madaio M. Fairlearn: assessing and improving fairness of AI systems. J Mach Learn Res. 2023;24(257):1–8.
  8. Collins GS, Moons KGM, Dhiman P, et al. TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ. 2024;385:e078378. doi:10.1136/bmj-2023-078378
  9. Moons KGM, Damen JAA, Kaul T, Hooft L, Andaur Navarro C, et al. PROBAST+AI: an updated quality, risk of bias, and applicability assessment tool for prediction models using regression or artificial intelligence methods. BMJ. 2025;388:e082505. doi:10.1136/bmj-2024-082505
  10. Riley RD, Ensor J, Snell KIE, et al. Calculating the sample size required for developing a clinical prediction model. BMJ. 2020;368:m441. doi:10.1136/bmj.m441
  11. United Arab Emirates. Federal Law No. 2 of 2019 Concerning the Use of Information and Communications Technology in Health Fields, Arts 13, 16 and 20. Issued 6 February 2019; Official Gazette No. 647; in force 14 May 2019. https://uaelegislation.gov.ae/en/legislations/1209
  12. United Arab Emirates Cabinet. Cabinet Resolution No. 32 of 2020 Concerning the Executive Regulation of Federal Law No. 2 of 2019. Issued 22 April 2020; Official Gazette No. 677; in force 30 October 2020. https://uaelegislation.gov.ae/en/legislations/1444
  13. Ministry of Health and Prevention, United Arab Emirates. Ministerial Resolution No. 51 of 2021 concerning the cases in which health data may be stored, processed, generated or transferred outside the State. Consult the operative text for the conditions attached to each exception.