Skip to main content

phdassistance

Federated Learning for Privacy-Preserving Sepsis Prediction Across UAE Healthcare Networks

“AI in healthcare” is not a research question, and a survey of clinician perceptions offers a narrower methodological contribution than a study of an unresolved modelling problem. Federated learning sepsis prediction in the UAE sets a harder, examinable question: can a sepsis early-warning model be trained across hospitals in different emirates without patient records leaving any of them, and can a doctoral study prove it accurate, calibrated, fair and governable?

Who this is for: PhD scholars in health informatics, computer science, data science, biomedical engineering and critical-care research; clinician-data scientists; and supervisors scoping AI projects in the UAE and GCC.

Scope and separation of content. This article distinguishes three things throughout: published evidence, cited to source; research hypotheses, labelled as such; and design guidance from PhD Assistance. It is research-methodology guidance only — not clinical advice, and nothing here should inform the treatment of any patient. Data-protection and ethics requirements change; confirm the current position with your ethics committee, your sites and the relevant health authority before submission.

1. Why sepsis early warning is a high-value health-AI PhD target in the UAE

Sepsis suits prediction research for reasons unrelated to fashion. It is time-critical, so firing earlier is an advantage measurable in hours. It has a consensus definition computable from routine data [4]. And in intensive care it is common enough that a doctoral cohort may contain usable numbers of positive cases — but that is a hypothesis about your sites, not a fact you may assume. Event counts differ by unit type, case-mix and admission policy, and the available number of qualifying sepsis events must be established by a site-level feasibility and event-count assessment before the aims are fixed. Section 9 sets out how.

The UAE adds a structural feature most countries do not have. Records are exchanged through emirate-level platforms — Malaffi in Abu Dhabi, operated by Abu Dhabi Health Data Services under the Department of Health [2], and NABIDH in Dubai, operated by the Dubai Health Authority — which integrate into the national Riayati platform. The Ministry of Health and Prevention’s Riayati page reports 4,952 connected facilities and 116,071 connected clinicians, with that page last updated in October 2024; these are platform-reported connection counts, not a 2026 measurement, and not a measure of research data availability [1]. Exchange of records for care delivery, typically through FHIR-based interoperability, is a different thing from a research data pathway.

Infrastructure is not access. It does not follow from any of the above that a doctoral researcher can assemble a pooled, identifiable, multi-emirate ICU time-series dataset for model training. That is a separate question of legal basis, ethics approval, custodianship and site-level agreement, and it is the question this topic is built around.

The research gap and the search that supports it. Our documented search did not identify a published study training and evaluating a sepsis early-warning model federated across UAE sites, with a prespecified label rule, per-site external validation, calibration and alert-burden reporting and a stated privacy-accounting method. A novelty claim of this kind is defensible only as the output of a reproducible search, so the proposal must record it in full:

Search element

What to record

Databases

MEDLINE/PubMed, Embase, Scopus, Web of Science, IEEE Xplore, ACM Digital Library, plus medRxiv/arXiv and a trials register

Search strings

The full Boolean string per database, with field tags, as an appendix

Date run

The exact date of each search, and the date of the pre-submission re-run

Limits

Language and date limits, and the justification for each

Eligibility

Population (hospital or ICU inpatients), index method (federated or distributed learning), outcome (sepsis onset or deterioration), setting (UAE or GCC for the local claim)

Screening

Number screened, number of full texts assessed, exclusions with reasons, and whether screening was duplicated

Limitations

Grey literature, non-indexed theses and non-English sources not retrieved; publication lag; the field is moving quickly

State the claim as the result of that search, not as a statement about all literature.

2. Why sepsis models trained at one hospital fail at another

This is the empirical foundation of the topic, and it is well evidenced. External validation of a widely deployed proprietary sepsis model across 38,455 hospitalisations found an AUROC of 0.63 (95% CI 0.62–0.64) against 0.76–0.83 reported by the developer, missing 1,709 of 2,552 patients with sepsis (67%) while alerting on 18% of hospitalisations [5]. A deep-learning model developed across four international ICU databases reached an internal AUROC of 0.846 but 0.761 at held-out sites [6]. A 2026 analysis of more than 216,000 ICU stays across MIMIC-IV, eICU and HiRID found that 37–60% of examined features differed significantly between datasets [7]. All three are international; none describes UAE hospitals, and none of their numbers should be extrapolated to them.

Four mechanisms drive this, and a review that separates them is already stronger than most: case-mix differs between sites; practice variation changes when cultures are drawn and fluids given, moving the label itself; documentation and sampling differences change what the model sees; and temporal distribution shift degrades unmonitored models over time.

What federated learning does and does not do. Federated learning enables collaborative model training across distributed healthcare datasets without requiring centralisation of raw patient records. However, it does not automatically resolve differences in patient case-mix, clinical practices, outcome definitions or EHR documentation. These sources of heterogeneity require separate data harmonisation, methodological controls and site-specific validation. Temporal distribution shift also requires ongoing evaluation rather than being addressed by federated training alone.

3. Defining the label: Sepsis-3, SOFA change and coding-based outcomes

Label choice most often determines whether results are examinable, and is the decision scholars most often leave vague.

Sepsis-3 defines sepsis as life-threatening organ dysfunction caused by a dysregulated host response to infection, operationalised as suspected infection plus an acute increase of two or more points in the Sequential Organ Failure Assessment score [4]. The definition is not an implementation. Seven decisions must be specified in the protocol, and the same specification must be executable at every site.

Decision

What the protocol must state

Suspected infection

The operational rule — for example a body-fluid culture order paired with an antibiotic administration

Culture timing

The window within which culture collection must fall relative to antibiotic start

Antibiotic timing

Which agents and routes qualify, the minimum duration, and the ordering of the two events

Baseline SOFA

The baseline assumed where no prior measurement exists, and how far back it may be drawn

Missing SOFA components

Whether a missing component scores zero, is imputed, or renders the window non-evaluable

Outcome timing

How the first qualifying event is identified, and the time origin for prediction

Recurrent episodes

Whether later episodes are excluded, censored, or treated as separate events

Harmonisation is not the same as a shared definition. A common rule is necessary but not sufficient: if source variables, units, culture-ordering practice or documentation differ, the same rule measures different things at different sites. Specify the source-variable mapping per site, run the rule on a sample extract at each site before the main analysis, and report the resulting event rate per site as a study characteristic.

Coding-based labels are the alternative, and they are not equivalent. Discharge codes are cheaper to extract but reflect documentation and reimbursement practice as much as physiology, and that practice differs by site — importing site variation directly into the target variable.

The design implication, with the hypothesis labelled as such. Prespecify one primary label rule and report an alternative as a sensitivity analysis. The hypothesis that federated training narrows the between-site performance gap is testable only if the label means the same thing at every site, so harmonising it is part of the method, not preparation for it.

4. Building features from EHR data: time windows, missingness and irregular sampling

Four decisions shape the feature set, and each belongs in the methods chapter with a reason attached.

State the task as a time problem. Specify the prediction horizon, the scoring frequency and the gap between the last observation used and predicted onset. Without that gap, measurements taken in response to deterioration leak in and inflate performance. A defensible protocol compares predefined horizons — for example 6, 12 and 24 hours before the operationally defined onset — and justifies the primary choice on clinical usefulness and data feasibility rather than on which horizon maximises the metric.

Observations are irregular, and the sampling rate is itself informative, because sicker patients are measured more often. Decide whether it is a feature or a confounder, and justify it.

Missingness is not random. State the imputation approach and report results under one alternative.

Features must be computable identically at every site, which in a federated design constrains the set to what all sites hold comparably. A variable held at one site only is a liability, not an advantage.

An illustrative specification, offered as a starting point to adapt rather than a claim that these variables are available across Malaffi-, NABIDH- or Riayati-connected sites:

Feature category

Example variables

Measurement window

Preprocessing and harmonisation

Missingness consideration

Vital signs

Heart rate, respiratory rate, systolic and mean arterial pressure, temperature, SpO₂

Rolling windows to the prediction point, e.g. last value plus 6-hour summary statistics

Define units; cap physiologically implausible values by a prespecified rule

Frequently measured; short gaps forward-filled within a stated limit

Laboratory parameters

Creatinine, platelet count, bilirubin, lactate, white cell count

Most recent value within a stated look-back

Harmonise units and local reference ranges across sites

Irregular; record time since last result as a feature

Neurological indicators

Glasgow Coma Scale

Most recent assessment in the window

Map local scoring conventions to a common scale

Often absent in non-ICU settings; state the rule

Clinical indicators (label-adjacent — see below)

Culture orders, antibiotic administration, suspected-infection flag

Only events completed strictly before the prediction point, outside the blanking window

Map local order codes to a common vocabulary; record the rule per site

Presence or absence is itself a signal, and is also part of the label

Temporal variables

Measurement frequency, time since last observation, time since admission

Computed to the prediction point only

Computed identically at every site

Must not reach past the prediction point — the leakage control

Label leakage is the single most likely way this study produces an inflated result. The Sepsis-3 label in Section 3 is constructed from culture orders and antibiotic administration. Those same events are the clinical indicators a modeller instinctively reaches for as features — so using them naively means predicting the label from the label. A model that appears to detect sepsis hours in advance may simply be detecting that a clinician has already suspected it, which is not an early warning and is not a research contribution.

Four controls, all prespecified:

  • Define the time origin and a blanking window. No feature may use information from the window immediately preceding the labelled onset; the width of that window is stated in the protocol and varied in sensitivity analysis. This separates prediction from documentation of a decision already taken.
  • Treat infection-related variables as label-adjacent, not ordinary features. Report the primary model twice: once excluding culture, antibiotic and suspected-infection variables entirely, and once including only those events completed before the blanking window. The exclusion-based model is the honest estimate of early-warning performance; the difference between the two quantifies how much apparent performance comes from clinician suspicion.
  • Audit for proxy leakage. Variables that are not part of the label can still encode the clinical response to it — a sudden change in measurement frequency, a transfer event, an order set. Examine feature importance for anything behaving as a proxy for the care pathway rather than physiology, and report what was found.
  • State the consequence in the limitations. Any residual leakage inflates performance, and the direction of that bias is known. Say so rather than letting a reviewer discover it.

5. Federated learning explained for clinical researchers

The mechanism is simple. A model is initialised centrally and sent to each site; each site trains on its own data; only the resulting parameter updates return to a central server, which aggregates them into a new global model and sends it back. This repeats over many communication rounds, and the number of rounds is itself a design parameter: each round costs network transfer, coordination and — where differential privacy is used — part of the privacy budget [10].

-sepsis-prediction

Flow diagram of one federated training round. Three site boxes labelled Abu Dhabi, Dubai and Northern Emirates each feed a local training box; arrows cross a dashed site boundary into a secure aggregation box running FedAvg or FedProx; the resulting global model is returned to every site each round and feeds a per-site validation box listing AUROC, AUPRC, calibration, lead time, alert burden and subgroup fairness. A legend states that model parameters and agreed aggregate counts cross the boundary, that patient records, identifiers and raw time series do not, and warns that updates are not inherently private because training content can be reconstructed from gradient updates, so the aggregation step must be protected by secure aggregation, differential privacy with stated clipping, noise, epsilon and delta, encryption in transit and at rest, and access control with audit trails.

A federated training round, and what crosses each boundary. Research-design schematic — not a deployed system, and not a description of any existing data-sharing arrangement.

Report convergence, not just final performance. Record the global and per-site loss by round, the number of rounds to a prespecified stopping criterion, local epochs per round, and the communication volume per round. A model that reaches good aggregate performance only after a round count the sites cannot sustain is not a feasible design.

Choosing between FedAvg and FedProx is a research decision, not a preference.

Research consideration

FedAvg [8]

FedProx [9]

Basic mechanism

Weighted aggregation of local model updates

FedAvg-style optimisation with a proximal regularisation term

Site heterogeneity

Baseline comparison

Intended to help with heterogeneous local optimisation; the benefit is reported in the originating experiments and is not guaranteed in any given dataset

Client drift

May be more sensitive to local drift from the global model

Penalises deviation from the global model; the strength of the penalty is a tuned hyperparameter, not a fixed property

Unequal client capacity

Assumes comparable local work per round

Tolerates variable local computation across sites

Research role

Baseline

Comparator

Evaluation

Global and per-site performance, convergence, computational cost

Global and per-site performance, convergence, computational cost, and the sensitivity of results to the proximal term

Compare them empirically on convergence behaviour, per-site performance and computational cost, and treat the comparison as a finding rather than a formality: FedProx is not uniformly superior to FedAvg, its advantage depends on the degree of heterogeneity and on the proximal term’s tuning, and reporting a case where it does not help is a legitimate result. Do not select an algorithm because it is newer.

6. Handling non-identical data across emirates and hospitals

Federated optimisation assumes nothing about sites holding identically distributed data, and reported experiments find that performance can degrade as distributions diverge, though how much depends on the task, the model and the degree of divergence. Emirate-level sites would be expected to differ in case-mix, demography and documentation — that is a reasonable prior, not a measured finding, and Section 6’s whole purpose is to measure it rather than assume it. In non-IID settings client drift can arise: local models pull toward their own site’s distribution between aggregation rounds, and the averaged global model may then underperform a site’s own local optimum. Three families of response exist, and the choice should be prespecified.

Personalisation keeps a shared global model but fits a site-specific component. Site weighting changes how much each site’s update counts, trading aggregate performance against performance at smaller sites. Proximal regularisation constrains local drift directly [9]. Each carries a cost: personalisation yields several models to validate rather than one; weighting by sample size lets the largest site dominate; strong regularisation can suppress adaptation the data actually support.

Quantify the divergence rather than asserting it. Report per-site distributions for the main features and the outcome rate, treating heterogeneity as a measured study characteristic; the 2026 distribution-shift analysis offers a reporting template [7].

7. Privacy safeguards: secure aggregation, differential privacy and their limits

“Data never leaves the site” is the claim that gets federated proposals approved, and it is the claim an examiner will press hardest.

Records do not move. It does not follow that nothing about them is disclosed: published work shows training data can, under some conditions, be reconstructed from gradient updates [11]. Model parameters are a disclosure surface to be assessed, not a neutral artefact.

Four distinct controls do four distinct jobs, and a protocol should name all four rather than relying on one.

Control

What it protects against

What the protocol must specify

Secure aggregation [16]

The server seeing any individual site’s update

Trust assumptions, collusion tolerance, protocol and implementation

Differential privacy [15]

The influence of a single record (or other privacy unit) on the released model

Privacy unit and adjacency, mechanism, clipping norm, noise multiplier, (ε, δ) and the named accountant

Encryption

Interception in transit or at rest

Encryption in transit and at rest, and key management

Access control

Unauthorised use of updates, logs and research artefacts

Authentication, authorisation and audit trails

Differential privacy needs a reproducible specification, not a label. Stating that noise is added and a budget is fixed is not enough for an examiner or an ethics committee. The protocol should specify:

  • Privacy unit and adjacency: whether the guarantee is defined over a single record, a single patient, or a whole site, and the corresponding adjacency relation — patient-level guarantees are usually what a clinical reviewer means, and are stronger and costlier than record-level ones.
  • Architecture: central differential privacy, where a trusted aggregator adds noise to the aggregate, versus local differential privacy, where each site perturbs its own update before release. They make different trust assumptions and carry very different utility costs.
  • Mechanism and clipping: the noise mechanism (typically Gaussian for an (ε, δ) guarantee), the per-example or per-update gradient clipping norm, and the noise multiplier — clipping is what bounds sensitivity, so without a stated clipping norm the (ε, δ) pair is not meaningful.
  • Accounting: the named accountant used to compose the guarantee across rounds — for example the moments accountant introduced with DP-SGD, or a Rényi differential privacy accountant, both of which compose far more tightly than naive sequential composition — together with the resulting (ε, δ) pair, the number of rounds, the sampling rate and the target δ. The guarantee is an (ε, δ)-differential privacy statement: ε bounds the privacy loss and δ is the probability that the bound fails, so both must be reported, and δ is conventionally set below the reciprocal of the number of records. Because ε accumulates with every round, the round count in Section 5 and the privacy budget are a single joint decision rather than two independent ones [15].
  • Threat model and residual risk: what each control assumes about the server, the sites and any observer, and which risks remain after all four controls.

Treat privacy as a measured dimension rather than a binary property: report performance at more than one (ε, δ) setting, and present the privacy-utility trade-off in the results rather than in a limitations paragraph.

8. Validation examiners expect: discrimination, calibration, lead time and alert burden

Discrimination alone is not enough. AUROC is often described as prevalence-independent, but that holds for the metric’s definition rather than for its practical interpretation: at low event prevalence a respectable AUROC can coexist with an alert stream that is mostly false positives. AUPRC is more informative for a rare outcome, but it too depends on prevalence, so it must always be read against the event rate in the same sample and against a stated baseline.

Metric

What it answers

How to interpret it

AUROC

Discrimination across all thresholds

Report with CIs; do not read it as clinical usefulness at any one threshold

AUPRC

Precision-recall performance

Interpret against the event prevalence and the no-skill baseline in that sample

Calibration slope

Whether predicted risks are systematically too extreme or too conservative

Slope below 1 indicates over-fitted, overly extreme predictions

Calibration-in-the-large

Whether mean predicted risk matches observed event frequency

The first thing to shift when case-mix or event rate differs between sites

Sensitivity

Proportion of qualifying sepsis events identified

Report at the prespecified operating threshold, not the best observed one

Positive predictive value

Proportion of alerts corresponding to qualifying events

Prevalence-dependent; report alongside the event rate

Lead time before onset

Interval from the first qualifying alert to the labelled onset time

Report the distribution, not a single number; define the censoring rule

Alert burden

Alerts per 100 admissions, per 1,000 patient-hours monitored, and per true positive

Decides whether the model is usable at all [5]; depends on the repeat-alert rule

Decision-curve analysis

Net benefit across clinically plausible thresholds

Connects statistical performance to the decision the alert supports [18]

Per-site external validation

Whether performance holds elsewhere

The central question of the thesis

Define these four quantities operationally, because each can be computed several ways.

  • Lead time. The interval between the first alert that crosses the prespecified threshold and the labelled onset time defined in Section 3, computed only among true positives. State the censoring rule: alerts that fire before the eligibility window opens, patients who are discharged or die before onset, and alerts that fire after onset are each handled by a stated rule rather than silently dropped. Report the median and interquartile range, and the proportion of true positives with a lead time below a clinically meaningful floor — a model with a long median lead time and a large minority of near-simultaneous alerts is a different proposition from one with a tight distribution.
  • Alert burden. Report three denominators, because each answers a different question: alerts per 100 admissions (what a unit experiences), alerts per 1,000 patient-hours monitored (comparable across differing lengths of stay), and alerts per true positive (the signal-to-noise ratio the clinician perceives). State the scoring frequency alongside them, since burden scales with how often the model scores.
  • Repeat alerts. A model scoring every hour will re-alert on the same patient repeatedly. Prespecify whether an episode counts once or every time the threshold is crossed, and whether a silencing window suppresses re-alerts for a stated period after a first alert. This choice can change the reported alert burden by an order of magnitude, so it must be stated before results are seen and held constant across sites.
  • Calibration assessment. Report a flexible calibration curve of observed against predicted risk, estimated by a smoother rather than by arbitrary risk deciles; the calibration slope; and calibration-in-the-large. Estimate all three separately at each site, with confidence intervals, and state the minimum number of events below which a site’s calibration curve is not reported.

Calibration should be evaluated per site rather than assumed to behave in any particular way. Calibration behaviour is dataset- and model-dependent: it often shifts across settings with differing case-mix or event rate, but whether it degrades before or after discrimination in any given study is an empirical question to be reported, not a general rule to be asserted.

Plan for drift. Even within the study period, performance may change over time. Report performance by time block as well as by site, and state what ongoing monitoring a deployed version would require — as further work, not as a claim.

Report against TRIPOD+AI [12], mapping the items at protocol stage rather than at write-up: the item set covers the data sources, participant flow, predictor definitions, missing-data handling, model specification, performance measures and fairness assessment this design depends on. For scale, a published international model achieved a median lead time of 3.7 hours at 1.4 false alerts per true positive and 80% sensitivity [6] — figures from another setting, useful for powering the study, not a benchmark to claim.

9. Sample size and statistical power considerations for federated sepsis prediction

This is the step that decides whether the study is feasible, and the one most often missing from health-AI proposals.

This is not a power calculation, and saying so matters. Conventional power analysis sizes a study to detect a specified effect under a null hypothesis at a given type I error rate. A prediction-model study has no such null hypothesis: the quantities of interest are model performance and the precision with which it is estimated. Sample size is therefore driven by the number of outcome events, and is chosen so that the model can be developed with acceptable overfitting — bounded optimism, a shrinkage factor close to one, and a small difference between apparent and adjusted performance — and so that performance can be estimated with a confidence interval narrow enough to be useful [17]. A generic rule such as ten events per variable is not an adequate justification and should not appear in the proposal. Where the thesis does include a genuine comparative test — for example a paired comparison of federated against site-local performance — that comparison needs its own precision justification, stated separately from the development sample size.

Establish the numbers before fixing the aims. The feasibility assessment should produce, per site:

Quantity

How to obtain it

Why it matters

Eligible ICU admissions per year

Retrospective count from the site’s own records

Sets the ceiling on the whole study

Qualifying sepsis events

The label rule run on a sample extract, not a coding count

Coding counts and Sepsis-3 counts differ, often substantially

Event rate and its distribution across sites

Events ÷ eligible admissions, per site

Drives per-site precision and the weighting decision

Candidate predictor count

The feature specification in Section 4

More predictors require more events

Validation requirement

Events available per site for external validation

A site with too few events cannot support a per-site estimate

Two sample-size questions, not one. Development needs enough events to fit a model without excessive optimism. Per-site external validation needs enough events at each site to estimate performance with usable precision — and because the central claim of this thesis is per-site performance, the binding constraint is usually the smallest site, not the total. Report the expected confidence-interval width for the primary metric at each site, and say plainly which sites can support an independent estimate and which can only contribute to training.

What to do when the numbers do not support the design. The honest options are to lengthen the retrospective window, widen the eligible population with the change in case-mix declared, reduce the candidate predictor set, pool the smallest sites for validation with that limitation stated, or move to the simulation fallback in Section 11. Discovering the shortfall after data access is granted is the expensive path.

10. Fairness across sites and patient subgroups

Two distinct fairness questions arise, and conflating them weakens the analysis.

Across sites. A global model can perform well on average while performing poorly at a small site, because aggregate metrics are dominated by large contributors. Report per site, never pooled only.

Across patient subgroups. The UAE population is diverse in nationality, age structure and occupational exposure. Fairness has to be operationalised before analysis, not discussed afterwards:

  • Subgroup definition: the variables and categories, fixed in advance, with the source field named per site.
  • Minimum reporting cell size, defined on events rather than patients: a prespecified minimum number of sepsis events per subgroup below which an estimate is not reported, since a subgroup with many patients but few events cannot support a stable performance estimate, and small cells also carry re-identification risk.
  • Expected precision, calculated in advance: subgroup analyses are almost always underpowered relative to the main analysis. Before data access, compute the expected confidence-interval width for the primary metric in each planned subgroup from the anticipated event counts, and state which subgroups can support an informative estimate and which can only be described. A subgroup difference that is within the overlapping confidence intervals of two imprecise estimates is not evidence of unfairness, and should not be reported as one.
  • Missing subgroup data: whether records with missing subgroup fields are excluded, grouped as unknown, or imputed — stated in advance, with the proportion reported.
  • Metrics per subgroup: AUROC, AUPRC, sensitivity, calibration and alert burden, each with confidence intervals rather than point estimates.
  • Comparisons: which differences are treated as material, and whether the comparison is of discrimination, of calibration, or of alert burden — these can disagree. State whether any multiplicity adjustment is applied across subgroups, or whether the analyses are reported as exploratory without adjustment; either is defensible, but the choice must be made in advance.
  • Reporting granularity: agreed with the ethics committee before analysis rather than after results.

Name the fairness criterion you are using. Statistical parity (equal alert rates across groups), equal opportunity (equal sensitivity across groups) and calibration-based fairness (equally accurate risks across groups) are different objectives, and in general cannot all be satisfied at once. State which the research objectives require, and why. Equalising any of them is a modelling objective, not a guarantee of equitable care, and a thesis claiming the latter from the former will be challenged.

11. Governance, ethics and data access in the UAE — and simulation when live access is not possible

UAE health-data governance operates at three levels, and a proposal that addresses only the first will not survive review.

Federal. Federal Law No. 2 of 2019 on the use of information and communications technology in health fields — issued 6 February 2019, in force 14 May 2019 — provides that health data related to health services provided in the State may not be stored, processed, generated or transferred outside the UAE unless approved by the health authority or the Minister (Art. 13), restricts use of health data for non-health purposes without written consent or a stated exception (Art. 16), and sets a minimum retention period (Art. 20) [3]. Cabinet Resolution No. 32 of 2020, in force 30 October 2020, issues the executive regulation of that law, covering the central health-data system, data standards and the obligations of health-information custodians [13]. The two instruments are separate records on the official legislation portal and should be cited separately; a proposal that cites only the primary law has not engaged with the operative detail. Ministerial Resolution No. 51 of 2021 sets out circumstances in which health data may be transferred, stored or processed outside the State, including for scientific research, each subject to its own conditions [14]. Read the conditions attached to the specific exception you intend to rely on from the instrument itself. Published legal commentary reports that a copy of the data must in any case be retained within the UAE and that the exceptions carry conditions, but this article has not verified the operative text of the Resolution, and the conditions should not be summarised at second hand in a proposal. Obtain the gazette copy, cite the article you rely on, and have the position confirmed in writing.

Emirate. Abu Dhabi’s Department of Health and the Dubai Health Authority set their own research and data-governance requirements for facilities under their jurisdiction, and a study spanning emirates is subject to each.

Institutional. Each participating hospital or university has its own research ethics committee, data custodian and contracting process. In practice a multi-site study needs ethics approval at every site, a data-sharing or collaboration agreement describing exactly what crosses each boundary, and a documented legal basis for the processing.

Do not simplify the legal position. “Federated learning is permitted because records remain local” is the formulation most likely to be challenged. Model updates derived from patient data are themselves a processing activity with recipients, and the applicable legal basis, approvals, custodianship and transfer conditions still have to be evaluated — including where an aggregation server, a cloud region or a collaborator sits. Confirm the current position with your institution and the relevant authority; this is a research-planning summary, not legal advice.

Name the fallback. If multi-site access is delayed, the same architecture can be studied on simulated partitions of a public ICU dataset with heterogeneity induced deliberately. That design is publishable and de-risks the project, provided it states the narrower claim: it evaluates the method, not UAE clinical performance.

12. PhD Assistance federated clinical-AI feasibility matrix

Score the project on six dimensions before writing the aims. The output is the completed matrix, not a single score.

Dimension

What to establish

Evidence to hold before submission

Risk if unresolved

Data access route

Which sites, approvals and custodian

Written expressions of interest; the approval pathway

The commonest cause of abandoned health-AI PhDs

Label definition

One primary rule, one alternative

A computable spec run on sample extracts at each site

Results not comparable across sites or literature

Event-count feasibility

Eligible admissions and qualifying events per site

A retrospective count and an event-based sample-size calculation [17]

A study that cannot estimate per-site performance

Federated strategy

Aggregation, rounds, heterogeneity response

A simulation run showing convergence and cost per round

Convergence failure discovered late

Validation plan

Metrics, per-site reporting, subgroup analysis

A TRIPOD+AI-mapped analysis plan [12]

Reviewers cannot judge clinical usefulness

Governance risk

Legal basis, approvals at three levels, privacy accounting

Approvals and a stated ε, δ and accountant

Analysis complete but unpublishable

How to use it. Any dimension without evidence behind it is either a decision still to be taken or a limitation to declare. Both are acceptable in a proposal; neither is acceptable left unstated.

13. How to turn this blueprint into an approvable PhD proposal

Evidence gap map

Existing evidence

Design

Main contribution

Remaining gap

Wong et al. 2021 [5]

External validation, 38,455 hospitalisations

A deployed model performed far below developer-reported figures

Diagnoses the problem; offers no multi-site solution

Moor et al. 2023 [6]

Four international ICU databases

Internal AUROC 0.846, external 0.761; lead time and alert burden

Centralised pooling, not federated; no UAE data

Tranchellini et al. 2026 [7]

>216,000 ICU stays, MIMIC-IV, eICU and HiRID

Quantifies feature shift; compares adaptation strategies

International datasets; no federated test under governance constraints

Rieke 2020 [10]; McMahan 2017 [8]; Li 2020 [9]

Methodological

Federated learning for health; FedAvg and FedProx

Method papers, not clinical evidence in any Gulf setting

Abadi 2016 [15]; Bonawitz 2017 [16]

Methodological

Privacy accounting and secure aggregation

Not evaluated against a clinical prediction task in this setting

Novelty matrix

Component

What is already known

What remains unknown

What this PhD would add

Transportability

Single-site models degrade elsewhere [5, 6]

Whether federated training closes that gap across UAE sites

Per-site accuracy and calibration under a federated protocol

Heterogeneity

Feature distributions differ markedly across ICUs [7]

The size of between-emirate heterogeneity

A measured heterogeneity profile and a prespecified response

Privacy

Updates can leak training content [11]

The privacy-utility trade-off at useful performance

Performance across stated ε values with a named accountant

Governance

Localisation, consent and transfer rules are published [3, 13, 14]

What a compliant multi-site protocol requires in practice

A documented governance route others can reuse

Objectives, questions and hypotheses. A doctoral blueprint should state what each question would accept as an answer, including an unfavourable one. Several of these questions are estimation problems rather than tests and are labelled as such; where a hypothesis is stated it is directionless, so the design cannot be accused of assuming its own result.

Research question

Objective

Hypothesis or inferential stance

Primary analysis

RQ1 (primary): How does a federated model compare with site-local models?

Estimate per-site AUROC, AUPRC and calibration for both, on the same features and label

H₀: per-site discrimination does not differ between federated and site-local training. The alternative is two-sided — federated training may perform worse at large sites while performing better at small ones

Per-site metrics with CIs; paired per-site comparison, direction not assumed

RQ2: Is the model clinically usable?

Quantify lead time and alert burden at a prespecified threshold

Estimation, not hypothesis testing: the study reports what lead time is achievable at what alert burden, and whether that combination falls inside a prespecified usability range

Lead-time distribution; alerts per 100 admissions and per true positive; decision-curve analysis [18]

RQ3: How heterogeneous are the sites?

Measure between-site distribution and event-rate differences

Descriptive: heterogeneity is characterised, not assumed to take any particular size or to affect calibration more than discrimination

Per-site feature and outcome distributions; FedAvg versus FedProx comparison

RQ4: What does privacy cost?

Measure performance across stated (ε, δ) settings

Estimation: the shape of the privacy-utility curve is reported as observed, not assumed to be monotonic or to contain a usable operating region

Performance at ≥3 (ε, δ) settings with a named accountant [15]

RQ5 (exploratory): Is performance equitable?

Compare subgroup performance after federated training

Exploratory and hypothesis-generating: no direction is assumed, and subgroup estimates are expected to be imprecise

Subgroup metrics with CIs, above the prespecified minimum event count

Expected limitations to state. Retrospective data cannot establish clinical benefit; label harmonisation may itself alter apparent incidence; site count will be small, so between-site inference is descriptive; the smallest site may not support an independent performance estimate; privacy guarantees hold only under their stated threat model; compute and network constraints bound the communication rounds; and no deployment or outcome claim is supportable within a doctoral timeline.

Contribution statement. A completed thesis would contribute (1) per-site accuracy, calibration, lead-time and alert-burden estimates for a federated sepsis early-warning model; (2) a measured between-site heterogeneity profile with a prespecified handling strategy; (3) a privacy-utility characterisation across stated budgets with full accounting; and (4) a documented governance route for multi-site clinical-AI research in the UAE, reusable beyond sepsis.

Design weaknesses that get health-AI proposals sent back: data access assumed rather than evidenced; label rule unspecified or not tested at each site; no event-count feasibility assessment; no gap between last observation and predicted onset; AUROC without calibration or alert burden; performance pooled across sites; privacy claimed rather than accounted; no fallback if access fails; and a deployment claim the design cannot support.

Frequently asked questions

No. It is a method. The examinable question is whether it produces a model that is accurate, calibrated and fair at each site, under stated privacy and governance constraints.

Yes, with a narrower claim. A simulated study on a public ICU dataset tests the architecture and the heterogeneity response; it cannot report UAE clinical performance, and the proposal must say so.

Start with FedAvg as the baseline [8] and justify any alternative by the heterogeneity you measure, not by novelty; FedProx is the usual comparator where sites diverge [9]. Compare convergence, per-site performance and computational cost.

No. It changes what crosses the boundary, not whether a legal basis, approvals at federal, emirate and institutional level, and site agreements are required [3, 13, 14].

That is an event-based calculation, not a rule of thumb. Size the development sample against the candidate predictor count and anticipated performance [17], and size per-site validation against the precision you need at the smallest site.

From clinical usefulness, not from what maximises the metric. Compare predefined horizons such as 6, 12 and 24 hours and report how performance moves across them.

Report it as a finding. Differing event rates affect calibration-in-the-large and predictive values most, which is why both are reported per site rather than pooled.

No. It can report performance and discuss what prospective evaluation would require. Deployment decisions rest with health authorities and the institutions concerned.

A federated-learning sepsis project can look technically feasible on paper but become difficult to defend once the researcher has to define the sepsis label, confirm multi-site data access, prevent feature leakage, justify the sample size, specify the privacy guarantee and obtain approvals across participating institutions.

If you are at this stage, the next step is not simply choosing a federated-learning algorithm. It is testing whether the research question, data pathway, methodology and evaluation plan can work together as a defensible PhD design

How PhD Assistance Can Help

PhD Assistance can review your proposed health-AI study across four practical areas:

  1. Research-question and novelty review — assess whether the proposed question has a defensible contribution against existing sepsis-prediction and federated-learning research.
  2. Data and methodology review — examine the proposed sepsis label, prediction horizon, features, leakage controls, federated-learning strategy and handling of site heterogeneity.
  3. Feasibility, privacy and governance review — identify the data-access, event-count, privacy, ethics and UAE governance requirements that must be resolved before implementation.
  4. Evaluation and proposal-readiness review — assess the validation strategy, per-site reporting, calibration, fairness, sample-size rationale and TRIPOD+AI reporting plan.

What You Receive

The review produces a written set of recommendations identifying what is defensible, what requires supporting evidence, what needs to be changed and what should be addressed first. This gives you a practical methodology roadmap that can be taken into a supervisor or research-ethics discussion.

Start With Your Research Question

If you are developing a PhD proposal on federated learning, privacy-preserving sepsis prediction or multi-site healthcare AI in the UAE, share your draft topic or research questions, target programme, and what you currently know about clinical-data access and ethics approval.

PhD Assistance can then assess the proposed design and identify the methodological and feasibility issues that should be resolved before the study moves into implementation.

This review supports research-proposal development and does not replace institutional ethics, legal or clinical review.

Reference

  1. Ministry of Health and Prevention, United Arab Emirates. About Riayati — National Unified Medical Record. https://mohap.gov.ae/en/riayati/about-numr (page last updated October 2024; accessed 1 October 2026).
  2. Abu Dhabi Health Data Services – SP LLC. Malaffi: frequently asked questions for end users. https://www.malaffi.ae/faq-end-users/general/ (accessed 1 October 2026).
  3. United Arab Emirates. Federal Law No. 2 of 2019 Concerning the Use of Information and Communications Technology in Health Fields, Arts 13, 16 and 20. Issued 6 February 2019; Official Gazette No. 647, 6 February 2019; in force 14 May 2019. https://uaelegislation.gov.ae/en/legislations/1209 (accessed 1 October 2026).
  4. Singer M, Deutschman CS, Seymour CW, et al. The Third International Consensus Definitions for Sepsis and Septic Shock (Sepsis-3). JAMA. 2016;315(8):801–810. doi:10.1001/jama.2016.0287
  5. Wong A, Otles E, Donnelly JP, et al. External validation of a widely implemented proprietary sepsis prediction model in hospitalised patients. JAMA Intern Med. 2021;181(8):1065–1070. doi:10.1001/jamainternmed.2021.2626
  6. Moor M, Bennett N, Plečko D, et al. Predicting sepsis using deep learning across international sites: a retrospective development and validation study. eClinicalMedicine. 2023;62:102124. doi:10.1016/j.eclinm.2023.102124
  7. Tranchellini F, Farag Y, Jutzeler C, Meegahapola L. Evaluating deep learning sepsis prediction models in ICUs under distribution shift: a multi-centre retrospective cohort study. npj Digit Med. 2026;9:306. doi:10.1038/s41746-026-02364-4
  8. McMahan B, Moore E, Ramage D, Hampson S, Agüera y Arcas B. Communication-efficient learning of deep networks from decentralised data. Proc 20th Int Conf Artificial Intelligence and Statistics, PMLR. 2017;54:1273–1282.
  9. Li T, Sahu AK, Zaheer M, Sanjabi M, Talwalkar A, Smith V. Federated optimization in heterogeneous networks. Proc Machine Learning and Systems. 2020;2:429–450.
  10. Rieke N, Hancox J, Li W, et al. The future of digital health with federated learning. npj Digit Med. 2020;3:119. doi:10.1038/s41746-020-00323-1
  11. Zhu L, Liu Z, Han S. Deep leakage from gradients. Adv Neural Inf Process Syst. 2019;32:14774–14784.
  12. Collins GS, Moons KGM, Dhiman P, et al. TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ. 2024;385:e078378. doi:10.1136/bmj-2023-078378
  13. United Arab Emirates Cabinet. Cabinet Resolution No. 32 of 2020 Concerning the Executive Regulation of Federal Law No. 2 of 2019 on the Use of Information and Communications Technology in Health Fields. Issued 22 April 2020; Official Gazette No. 677, 30 April 2020; in force 30 October 2020. https://uaelegislation.gov.ae/en/legislations/1444 (accessed 1 October 2026).
  14. Ministry of Health and Prevention, United Arab Emirates. Ministerial Resolution No. 51 of 2021 concerning the cases in which health data may be stored, processed, generated or transferred outside the State. Consult the operative text for the conditions attached to each exception; at the time of writing the conditions had not been verified against the official gazette copy for this article, and the scholar should obtain it through MoHAP or the UAE legislation portal.
  15. Abadi M, Chu A, Goodfellow I, McMahan HB, Mironov I, Talwar K, Zhang L. Deep learning with differential privacy. Proc 2016 ACM SIGSAC Conf Computer and Communications Security. 2016:308–318. doi:10.1145/2976749.2978318
  16. Bonawitz K, Ivanov V, Kreuter B, et al. Practical secure aggregation for privacy-preserving machine learning. Proc 2017 ACM SIGSAC Conf Computer and Communications Security. 2017:1175–1191. doi:10.1145/3133956.3133982
  17. Riley RD, Ensor J, Snell KIE, et al. Calculating the sample size required for developing a clinical prediction model. BMJ. 2020;368:m441. doi:10.1136/bmj.m441
  18. Vickers AJ, Elkin EB. Decision curve analysis: a novel method for evaluating prediction models. Med Decis Making. 2006;26(6):565–574. doi:10.1177/0272989X06295361