“AI in healthcare” is not a research question, and a survey of clinician perceptions offers a narrower methodological contribution than a study of an unresolved modelling problem. Federated learning sepsis prediction in the UAE sets a harder, examinable question: can a sepsis early-warning model be trained across hospitals in different emirates without patient records leaving any of them, and can a doctoral study prove it accurate, calibrated, fair and governable?
Who this is for: PhD scholars in health informatics, computer science, data science, biomedical engineering and critical-care research; clinician-data scientists; and supervisors scoping AI projects in the UAE and GCC.
Scope and separation of content. This article distinguishes three things throughout: published evidence, cited to source; research hypotheses, labelled as such; and design guidance from PhD Assistance. It is research-methodology guidance only — not clinical advice, and nothing here should inform the treatment of any patient. Data-protection and ethics requirements change; confirm the current position with your ethics committee, your sites and the relevant health authority before submission.
Sepsis suits prediction research for reasons unrelated to fashion. It is time-critical, so firing earlier is an advantage measurable in hours. It has a consensus definition computable from routine data [4]. And in intensive care it is common enough that a doctoral cohort may contain usable numbers of positive cases — but that is a hypothesis about your sites, not a fact you may assume. Event counts differ by unit type, case-mix and admission policy, and the available number of qualifying sepsis events must be established by a site-level feasibility and event-count assessment before the aims are fixed. Section 9 sets out how.
The UAE adds a structural feature most countries do not have. Records are exchanged through emirate-level platforms — Malaffi in Abu Dhabi, operated by Abu Dhabi Health Data Services under the Department of Health [2], and NABIDH in Dubai, operated by the Dubai Health Authority — which integrate into the national Riayati platform. The Ministry of Health and Prevention’s Riayati page reports 4,952 connected facilities and 116,071 connected clinicians, with that page last updated in October 2024; these are platform-reported connection counts, not a 2026 measurement, and not a measure of research data availability [1]. Exchange of records for care delivery, typically through FHIR-based interoperability, is a different thing from a research data pathway.
Infrastructure is not access. It does not follow from any of the above that a doctoral researcher can assemble a pooled, identifiable, multi-emirate ICU time-series dataset for model training. That is a separate question of legal basis, ethics approval, custodianship and site-level agreement, and it is the question this topic is built around.
The research gap and the search that supports it. Our documented search did not identify a published study training and evaluating a sepsis early-warning model federated across UAE sites, with a prespecified label rule, per-site external validation, calibration and alert-burden reporting and a stated privacy-accounting method. A novelty claim of this kind is defensible only as the output of a reproducible search, so the proposal must record it in full:
Search element | What to record |
Databases | MEDLINE/PubMed, Embase, Scopus, Web of Science, IEEE Xplore, ACM Digital Library, plus medRxiv/arXiv and a trials register |
Search strings | The full Boolean string per database, with field tags, as an appendix |
Date run | The exact date of each search, and the date of the pre-submission re-run |
Limits | Language and date limits, and the justification for each |
Eligibility | Population (hospital or ICU inpatients), index method (federated or distributed learning), outcome (sepsis onset or deterioration), setting (UAE or GCC for the local claim) |
Screening | Number screened, number of full texts assessed, exclusions with reasons, and whether screening was duplicated |
Limitations | Grey literature, non-indexed theses and non-English sources not retrieved; publication lag; the field is moving quickly |
State the claim as the result of that search, not as a statement about all literature.
This is the empirical foundation of the topic, and it is well evidenced. External validation of a widely deployed proprietary sepsis model across 38,455 hospitalisations found an AUROC of 0.63 (95% CI 0.62–0.64) against 0.76–0.83 reported by the developer, missing 1,709 of 2,552 patients with sepsis (67%) while alerting on 18% of hospitalisations [5]. A deep-learning model developed across four international ICU databases reached an internal AUROC of 0.846 but 0.761 at held-out sites [6]. A 2026 analysis of more than 216,000 ICU stays across MIMIC-IV, eICU and HiRID found that 37–60% of examined features differed significantly between datasets [7]. All three are international; none describes UAE hospitals, and none of their numbers should be extrapolated to them.
Four mechanisms drive this, and a review that separates them is already stronger than most: case-mix differs between sites; practice variation changes when cultures are drawn and fluids given, moving the label itself; documentation and sampling differences change what the model sees; and temporal distribution shift degrades unmonitored models over time.
What federated learning does and does not do. Federated learning enables collaborative model training across distributed healthcare datasets without requiring centralisation of raw patient records. However, it does not automatically resolve differences in patient case-mix, clinical practices, outcome definitions or EHR documentation. These sources of heterogeneity require separate data harmonisation, methodological controls and site-specific validation. Temporal distribution shift also requires ongoing evaluation rather than being addressed by federated training alone.
Label choice most often determines whether results are examinable, and is the decision scholars most often leave vague.
Sepsis-3 defines sepsis as life-threatening organ dysfunction caused by a dysregulated host response to infection, operationalised as suspected infection plus an acute increase of two or more points in the Sequential Organ Failure Assessment score [4]. The definition is not an implementation. Seven decisions must be specified in the protocol, and the same specification must be executable at every site.
Decision | What the protocol must state |
Suspected infection | The operational rule — for example a body-fluid culture order paired with an antibiotic administration |
Culture timing | The window within which culture collection must fall relative to antibiotic start |
Antibiotic timing | Which agents and routes qualify, the minimum duration, and the ordering of the two events |
Baseline SOFA | The baseline assumed where no prior measurement exists, and how far back it may be drawn |
Missing SOFA components | Whether a missing component scores zero, is imputed, or renders the window non-evaluable |
Outcome timing | How the first qualifying event is identified, and the time origin for prediction |
Recurrent episodes | Whether later episodes are excluded, censored, or treated as separate events |
Harmonisation is not the same as a shared definition. A common rule is necessary but not sufficient: if source variables, units, culture-ordering practice or documentation differ, the same rule measures different things at different sites. Specify the source-variable mapping per site, run the rule on a sample extract at each site before the main analysis, and report the resulting event rate per site as a study characteristic.
Coding-based labels are the alternative, and they are not equivalent. Discharge codes are cheaper to extract but reflect documentation and reimbursement practice as much as physiology, and that practice differs by site — importing site variation directly into the target variable.
The design implication, with the hypothesis labelled as such. Prespecify one primary label rule and report an alternative as a sensitivity analysis. The hypothesis that federated training narrows the between-site performance gap is testable only if the label means the same thing at every site, so harmonising it is part of the method, not preparation for it.
Four decisions shape the feature set, and each belongs in the methods chapter with a reason attached.
State the task as a time problem. Specify the prediction horizon, the scoring frequency and the gap between the last observation used and predicted onset. Without that gap, measurements taken in response to deterioration leak in and inflate performance. A defensible protocol compares predefined horizons — for example 6, 12 and 24 hours before the operationally defined onset — and justifies the primary choice on clinical usefulness and data feasibility rather than on which horizon maximises the metric.
Observations are irregular, and the sampling rate is itself informative, because sicker patients are measured more often. Decide whether it is a feature or a confounder, and justify it.
Missingness is not random. State the imputation approach and report results under one alternative.
Features must be computable identically at every site, which in a federated design constrains the set to what all sites hold comparably. A variable held at one site only is a liability, not an advantage.
An illustrative specification, offered as a starting point to adapt rather than a claim that these variables are available across Malaffi-, NABIDH- or Riayati-connected sites:
Feature category | Example variables | Measurement window | Preprocessing and harmonisation | Missingness consideration |
Vital signs | Heart rate, respiratory rate, systolic and mean arterial pressure, temperature, SpO₂ | Rolling windows to the prediction point, e.g. last value plus 6-hour summary statistics | Define units; cap physiologically implausible values by a prespecified rule | Frequently measured; short gaps forward-filled within a stated limit |
Laboratory parameters | Creatinine, platelet count, bilirubin, lactate, white cell count | Most recent value within a stated look-back | Harmonise units and local reference ranges across sites | Irregular; record time since last result as a feature |
Neurological indicators | Glasgow Coma Scale | Most recent assessment in the window | Map local scoring conventions to a common scale | Often absent in non-ICU settings; state the rule |
Clinical indicators (label-adjacent — see below) | Culture orders, antibiotic administration, suspected-infection flag | Only events completed strictly before the prediction point, outside the blanking window | Map local order codes to a common vocabulary; record the rule per site | Presence or absence is itself a signal, and is also part of the label |
Temporal variables | Measurement frequency, time since last observation, time since admission | Computed to the prediction point only | Computed identically at every site | Must not reach past the prediction point — the leakage control |
Label leakage is the single most likely way this study produces an inflated result. The Sepsis-3 label in Section 3 is constructed from culture orders and antibiotic administration. Those same events are the clinical indicators a modeller instinctively reaches for as features — so using them naively means predicting the label from the label. A model that appears to detect sepsis hours in advance may simply be detecting that a clinician has already suspected it, which is not an early warning and is not a research contribution.
Four controls, all prespecified:
The mechanism is simple. A model is initialised centrally and sent to each site; each site trains on its own data; only the resulting parameter updates return to a central server, which aggregates them into a new global model and sends it back. This repeats over many communication rounds, and the number of rounds is itself a design parameter: each round costs network transfer, coordination and — where differential privacy is used — part of the privacy budget [10].

Flow diagram of one federated training round. Three site boxes labelled Abu Dhabi, Dubai and Northern Emirates each feed a local training box; arrows cross a dashed site boundary into a secure aggregation box running FedAvg or FedProx; the resulting global model is returned to every site each round and feeds a per-site validation box listing AUROC, AUPRC, calibration, lead time, alert burden and subgroup fairness. A legend states that model parameters and agreed aggregate counts cross the boundary, that patient records, identifiers and raw time series do not, and warns that updates are not inherently private because training content can be reconstructed from gradient updates, so the aggregation step must be protected by secure aggregation, differential privacy with stated clipping, noise, epsilon and delta, encryption in transit and at rest, and access control with audit trails.
A federated training round, and what crosses each boundary. Research-design schematic — not a deployed system, and not a description of any existing data-sharing arrangement.
Report convergence, not just final performance. Record the global and per-site loss by round, the number of rounds to a prespecified stopping criterion, local epochs per round, and the communication volume per round. A model that reaches good aggregate performance only after a round count the sites cannot sustain is not a feasible design.
Choosing between FedAvg and FedProx is a research decision, not a preference.
Research consideration | FedAvg [8] | FedProx [9] |
Basic mechanism | Weighted aggregation of local model updates | FedAvg-style optimisation with a proximal regularisation term |
Site heterogeneity | Baseline comparison | Intended to help with heterogeneous local optimisation; the benefit is reported in the originating experiments and is not guaranteed in any given dataset |
Client drift | May be more sensitive to local drift from the global model | Penalises deviation from the global model; the strength of the penalty is a tuned hyperparameter, not a fixed property |
Unequal client capacity | Assumes comparable local work per round | Tolerates variable local computation across sites |
Research role | Baseline | Comparator |
Evaluation | Global and per-site performance, convergence, computational cost | Global and per-site performance, convergence, computational cost, and the sensitivity of results to the proximal term |
Compare them empirically on convergence behaviour, per-site performance and computational cost, and treat the comparison as a finding rather than a formality: FedProx is not uniformly superior to FedAvg, its advantage depends on the degree of heterogeneity and on the proximal term’s tuning, and reporting a case where it does not help is a legitimate result. Do not select an algorithm because it is newer.
Federated optimisation assumes nothing about sites holding identically distributed data, and reported experiments find that performance can degrade as distributions diverge, though how much depends on the task, the model and the degree of divergence. Emirate-level sites would be expected to differ in case-mix, demography and documentation — that is a reasonable prior, not a measured finding, and Section 6’s whole purpose is to measure it rather than assume it. In non-IID settings client drift can arise: local models pull toward their own site’s distribution between aggregation rounds, and the averaged global model may then underperform a site’s own local optimum. Three families of response exist, and the choice should be prespecified.
Personalisation keeps a shared global model but fits a site-specific component. Site weighting changes how much each site’s update counts, trading aggregate performance against performance at smaller sites. Proximal regularisation constrains local drift directly [9]. Each carries a cost: personalisation yields several models to validate rather than one; weighting by sample size lets the largest site dominate; strong regularisation can suppress adaptation the data actually support.
Quantify the divergence rather than asserting it. Report per-site distributions for the main features and the outcome rate, treating heterogeneity as a measured study characteristic; the 2026 distribution-shift analysis offers a reporting template [7].
“Data never leaves the site” is the claim that gets federated proposals approved, and it is the claim an examiner will press hardest.
Records do not move. It does not follow that nothing about them is disclosed: published work shows training data can, under some conditions, be reconstructed from gradient updates [11]. Model parameters are a disclosure surface to be assessed, not a neutral artefact.
Four distinct controls do four distinct jobs, and a protocol should name all four rather than relying on one.
Control | What it protects against | What the protocol must specify |
Secure aggregation [16] | The server seeing any individual site’s update | Trust assumptions, collusion tolerance, protocol and implementation |
Differential privacy [15] | The influence of a single record (or other privacy unit) on the released model | Privacy unit and adjacency, mechanism, clipping norm, noise multiplier, (ε, δ) and the named accountant |
Encryption | Interception in transit or at rest | Encryption in transit and at rest, and key management |
Access control | Unauthorised use of updates, logs and research artefacts | Authentication, authorisation and audit trails |
Differential privacy needs a reproducible specification, not a label. Stating that noise is added and a budget is fixed is not enough for an examiner or an ethics committee. The protocol should specify:
Treat privacy as a measured dimension rather than a binary property: report performance at more than one (ε, δ) setting, and present the privacy-utility trade-off in the results rather than in a limitations paragraph.
Discrimination alone is not enough. AUROC is often described as prevalence-independent, but that holds for the metric’s definition rather than for its practical interpretation: at low event prevalence a respectable AUROC can coexist with an alert stream that is mostly false positives. AUPRC is more informative for a rare outcome, but it too depends on prevalence, so it must always be read against the event rate in the same sample and against a stated baseline.
Metric | What it answers | How to interpret it |
AUROC | Discrimination across all thresholds | Report with CIs; do not read it as clinical usefulness at any one threshold |
AUPRC | Precision-recall performance | Interpret against the event prevalence and the no-skill baseline in that sample |
Calibration slope | Whether predicted risks are systematically too extreme or too conservative | Slope below 1 indicates over-fitted, overly extreme predictions |
Calibration-in-the-large | Whether mean predicted risk matches observed event frequency | The first thing to shift when case-mix or event rate differs between sites |
Sensitivity | Proportion of qualifying sepsis events identified | Report at the prespecified operating threshold, not the best observed one |
Positive predictive value | Proportion of alerts corresponding to qualifying events | Prevalence-dependent; report alongside the event rate |
Lead time before onset | Interval from the first qualifying alert to the labelled onset time | Report the distribution, not a single number; define the censoring rule |
Alert burden | Alerts per 100 admissions, per 1,000 patient-hours monitored, and per true positive | Decides whether the model is usable at all [5]; depends on the repeat-alert rule |
Decision-curve analysis | Net benefit across clinically plausible thresholds | Connects statistical performance to the decision the alert supports [18] |
Per-site external validation | Whether performance holds elsewhere | The central question of the thesis |
Define these four quantities operationally, because each can be computed several ways.
Calibration should be evaluated per site rather than assumed to behave in any particular way. Calibration behaviour is dataset- and model-dependent: it often shifts across settings with differing case-mix or event rate, but whether it degrades before or after discrimination in any given study is an empirical question to be reported, not a general rule to be asserted.
Plan for drift. Even within the study period, performance may change over time. Report performance by time block as well as by site, and state what ongoing monitoring a deployed version would require — as further work, not as a claim.
Report against TRIPOD+AI [12], mapping the items at protocol stage rather than at write-up: the item set covers the data sources, participant flow, predictor definitions, missing-data handling, model specification, performance measures and fairness assessment this design depends on. For scale, a published international model achieved a median lead time of 3.7 hours at 1.4 false alerts per true positive and 80% sensitivity [6] — figures from another setting, useful for powering the study, not a benchmark to claim.
This is the step that decides whether the study is feasible, and the one most often missing from health-AI proposals.
This is not a power calculation, and saying so matters. Conventional power analysis sizes a study to detect a specified effect under a null hypothesis at a given type I error rate. A prediction-model study has no such null hypothesis: the quantities of interest are model performance and the precision with which it is estimated. Sample size is therefore driven by the number of outcome events, and is chosen so that the model can be developed with acceptable overfitting — bounded optimism, a shrinkage factor close to one, and a small difference between apparent and adjusted performance — and so that performance can be estimated with a confidence interval narrow enough to be useful [17]. A generic rule such as ten events per variable is not an adequate justification and should not appear in the proposal. Where the thesis does include a genuine comparative test — for example a paired comparison of federated against site-local performance — that comparison needs its own precision justification, stated separately from the development sample size.
Establish the numbers before fixing the aims. The feasibility assessment should produce, per site:
Quantity | How to obtain it | Why it matters |
Eligible ICU admissions per year | Retrospective count from the site’s own records | Sets the ceiling on the whole study |
Qualifying sepsis events | The label rule run on a sample extract, not a coding count | Coding counts and Sepsis-3 counts differ, often substantially |
Event rate and its distribution across sites | Events ÷ eligible admissions, per site | Drives per-site precision and the weighting decision |
Candidate predictor count | The feature specification in Section 4 | More predictors require more events |
Validation requirement | Events available per site for external validation | A site with too few events cannot support a per-site estimate |
Two sample-size questions, not one. Development needs enough events to fit a model without excessive optimism. Per-site external validation needs enough events at each site to estimate performance with usable precision — and because the central claim of this thesis is per-site performance, the binding constraint is usually the smallest site, not the total. Report the expected confidence-interval width for the primary metric at each site, and say plainly which sites can support an independent estimate and which can only contribute to training.
What to do when the numbers do not support the design. The honest options are to lengthen the retrospective window, widen the eligible population with the change in case-mix declared, reduce the candidate predictor set, pool the smallest sites for validation with that limitation stated, or move to the simulation fallback in Section 11. Discovering the shortfall after data access is granted is the expensive path.
Two distinct fairness questions arise, and conflating them weakens the analysis.
Across sites. A global model can perform well on average while performing poorly at a small site, because aggregate metrics are dominated by large contributors. Report per site, never pooled only.
Across patient subgroups. The UAE population is diverse in nationality, age structure and occupational exposure. Fairness has to be operationalised before analysis, not discussed afterwards:
Name the fairness criterion you are using. Statistical parity (equal alert rates across groups), equal opportunity (equal sensitivity across groups) and calibration-based fairness (equally accurate risks across groups) are different objectives, and in general cannot all be satisfied at once. State which the research objectives require, and why. Equalising any of them is a modelling objective, not a guarantee of equitable care, and a thesis claiming the latter from the former will be challenged.
UAE health-data governance operates at three levels, and a proposal that addresses only the first will not survive review.
Federal. Federal Law No. 2 of 2019 on the use of information and communications technology in health fields — issued 6 February 2019, in force 14 May 2019 — provides that health data related to health services provided in the State may not be stored, processed, generated or transferred outside the UAE unless approved by the health authority or the Minister (Art. 13), restricts use of health data for non-health purposes without written consent or a stated exception (Art. 16), and sets a minimum retention period (Art. 20) [3]. Cabinet Resolution No. 32 of 2020, in force 30 October 2020, issues the executive regulation of that law, covering the central health-data system, data standards and the obligations of health-information custodians [13]. The two instruments are separate records on the official legislation portal and should be cited separately; a proposal that cites only the primary law has not engaged with the operative detail. Ministerial Resolution No. 51 of 2021 sets out circumstances in which health data may be transferred, stored or processed outside the State, including for scientific research, each subject to its own conditions [14]. Read the conditions attached to the specific exception you intend to rely on from the instrument itself. Published legal commentary reports that a copy of the data must in any case be retained within the UAE and that the exceptions carry conditions, but this article has not verified the operative text of the Resolution, and the conditions should not be summarised at second hand in a proposal. Obtain the gazette copy, cite the article you rely on, and have the position confirmed in writing.
Emirate. Abu Dhabi’s Department of Health and the Dubai Health Authority set their own research and data-governance requirements for facilities under their jurisdiction, and a study spanning emirates is subject to each.
Institutional. Each participating hospital or university has its own research ethics committee, data custodian and contracting process. In practice a multi-site study needs ethics approval at every site, a data-sharing or collaboration agreement describing exactly what crosses each boundary, and a documented legal basis for the processing.
Do not simplify the legal position. “Federated learning is permitted because records remain local” is the formulation most likely to be challenged. Model updates derived from patient data are themselves a processing activity with recipients, and the applicable legal basis, approvals, custodianship and transfer conditions still have to be evaluated — including where an aggregation server, a cloud region or a collaborator sits. Confirm the current position with your institution and the relevant authority; this is a research-planning summary, not legal advice.
Name the fallback. If multi-site access is delayed, the same architecture can be studied on simulated partitions of a public ICU dataset with heterogeneity induced deliberately. That design is publishable and de-risks the project, provided it states the narrower claim: it evaluates the method, not UAE clinical performance.
Score the project on six dimensions before writing the aims. The output is the completed matrix, not a single score.
Dimension | What to establish | Evidence to hold before submission | Risk if unresolved |
Data access route | Which sites, approvals and custodian | Written expressions of interest; the approval pathway | The commonest cause of abandoned health-AI PhDs |
Label definition | One primary rule, one alternative | A computable spec run on sample extracts at each site | Results not comparable across sites or literature |
Event-count feasibility | Eligible admissions and qualifying events per site | A retrospective count and an event-based sample-size calculation [17] | A study that cannot estimate per-site performance |
Federated strategy | Aggregation, rounds, heterogeneity response | A simulation run showing convergence and cost per round | Convergence failure discovered late |
Validation plan | Metrics, per-site reporting, subgroup analysis | A TRIPOD+AI-mapped analysis plan [12] | Reviewers cannot judge clinical usefulness |
Governance risk | Legal basis, approvals at three levels, privacy accounting | Approvals and a stated ε, δ and accountant | Analysis complete but unpublishable |
How to use it. Any dimension without evidence behind it is either a decision still to be taken or a limitation to declare. Both are acceptable in a proposal; neither is acceptable left unstated.
Evidence gap map
Existing evidence | Design | Main contribution | Remaining gap |
Wong et al. 2021 [5] | External validation, 38,455 hospitalisations | A deployed model performed far below developer-reported figures | Diagnoses the problem; offers no multi-site solution |
Moor et al. 2023 [6] | Four international ICU databases | Internal AUROC 0.846, external 0.761; lead time and alert burden | Centralised pooling, not federated; no UAE data |
Tranchellini et al. 2026 [7] | >216,000 ICU stays, MIMIC-IV, eICU and HiRID | Quantifies feature shift; compares adaptation strategies | International datasets; no federated test under governance constraints |
Rieke 2020 [10]; McMahan 2017 [8]; Li 2020 [9] | Methodological | Federated learning for health; FedAvg and FedProx | Method papers, not clinical evidence in any Gulf setting |
Abadi 2016 [15]; Bonawitz 2017 [16] | Methodological | Privacy accounting and secure aggregation | Not evaluated against a clinical prediction task in this setting |
Novelty matrix
Component | What is already known | What remains unknown | What this PhD would add |
Transportability | Single-site models degrade elsewhere [5, 6] | Whether federated training closes that gap across UAE sites | Per-site accuracy and calibration under a federated protocol |
Heterogeneity | Feature distributions differ markedly across ICUs [7] | The size of between-emirate heterogeneity | A measured heterogeneity profile and a prespecified response |
Privacy | Updates can leak training content [11] | The privacy-utility trade-off at useful performance | Performance across stated ε values with a named accountant |
Governance | Localisation, consent and transfer rules are published [3, 13, 14] | What a compliant multi-site protocol requires in practice | A documented governance route others can reuse |
Objectives, questions and hypotheses. A doctoral blueprint should state what each question would accept as an answer, including an unfavourable one. Several of these questions are estimation problems rather than tests and are labelled as such; where a hypothesis is stated it is directionless, so the design cannot be accused of assuming its own result.
Research question | Objective | Hypothesis or inferential stance | Primary analysis |
RQ1 (primary): How does a federated model compare with site-local models? | Estimate per-site AUROC, AUPRC and calibration for both, on the same features and label | H₀: per-site discrimination does not differ between federated and site-local training. The alternative is two-sided — federated training may perform worse at large sites while performing better at small ones | Per-site metrics with CIs; paired per-site comparison, direction not assumed |
RQ2: Is the model clinically usable? | Quantify lead time and alert burden at a prespecified threshold | Estimation, not hypothesis testing: the study reports what lead time is achievable at what alert burden, and whether that combination falls inside a prespecified usability range | Lead-time distribution; alerts per 100 admissions and per true positive; decision-curve analysis [18] |
RQ3: How heterogeneous are the sites? | Measure between-site distribution and event-rate differences | Descriptive: heterogeneity is characterised, not assumed to take any particular size or to affect calibration more than discrimination | Per-site feature and outcome distributions; FedAvg versus FedProx comparison |
RQ4: What does privacy cost? | Measure performance across stated (ε, δ) settings | Estimation: the shape of the privacy-utility curve is reported as observed, not assumed to be monotonic or to contain a usable operating region | Performance at ≥3 (ε, δ) settings with a named accountant [15] |
RQ5 (exploratory): Is performance equitable? | Compare subgroup performance after federated training | Exploratory and hypothesis-generating: no direction is assumed, and subgroup estimates are expected to be imprecise | Subgroup metrics with CIs, above the prespecified minimum event count |
Expected limitations to state. Retrospective data cannot establish clinical benefit; label harmonisation may itself alter apparent incidence; site count will be small, so between-site inference is descriptive; the smallest site may not support an independent performance estimate; privacy guarantees hold only under their stated threat model; compute and network constraints bound the communication rounds; and no deployment or outcome claim is supportable within a doctoral timeline.
Contribution statement. A completed thesis would contribute (1) per-site accuracy, calibration, lead-time and alert-burden estimates for a federated sepsis early-warning model; (2) a measured between-site heterogeneity profile with a prespecified handling strategy; (3) a privacy-utility characterisation across stated budgets with full accounting; and (4) a documented governance route for multi-site clinical-AI research in the UAE, reusable beyond sepsis.
Design weaknesses that get health-AI proposals sent back: data access assumed rather than evidenced; label rule unspecified or not tested at each site; no event-count feasibility assessment; no gap between last observation and predicted onset; AUROC without calibration or alert burden; performance pooled across sites; privacy claimed rather than accounted; no fallback if access fails; and a deployment claim the design cannot support.
No. It is a method. The examinable question is whether it produces a model that is accurate, calibrated and fair at each site, under stated privacy and governance constraints.
Yes, with a narrower claim. A simulated study on a public ICU dataset tests the architecture and the heterogeneity response; it cannot report UAE clinical performance, and the proposal must say so.
Start with FedAvg as the baseline [8] and justify any alternative by the heterogeneity you measure, not by novelty; FedProx is the usual comparator where sites diverge [9]. Compare convergence, per-site performance and computational cost.
No. It changes what crosses the boundary, not whether a legal basis, approvals at federal, emirate and institutional level, and site agreements are required [3, 13, 14].
That is an event-based calculation, not a rule of thumb. Size the development sample against the candidate predictor count and anticipated performance [17], and size per-site validation against the precision you need at the smallest site.
From clinical usefulness, not from what maximises the metric. Compare predefined horizons such as 6, 12 and 24 hours and report how performance moves across them.
Report it as a finding. Differing event rates affect calibration-in-the-large and predictive values most, which is why both are reported per site rather than pooled.
No. It can report performance and discuss what prospective evaluation would require. Deployment decisions rest with health authorities and the institutions concerned.
A federated-learning sepsis project can look technically feasible on paper but become difficult to defend once the researcher has to define the sepsis label, confirm multi-site data access, prevent feature leakage, justify the sample size, specify the privacy guarantee and obtain approvals across participating institutions.
If you are at this stage, the next step is not simply choosing a federated-learning algorithm. It is testing whether the research question, data pathway, methodology and evaluation plan can work together as a defensible PhD design
PhD Assistance can review your proposed health-AI study across four practical areas:
What You Receive
The review produces a written set of recommendations identifying what is defensible, what requires supporting evidence, what needs to be changed and what should be addressed first. This gives you a practical methodology roadmap that can be taken into a supervisor or research-ethics discussion.
Start With Your Research Question
If you are developing a PhD proposal on federated learning, privacy-preserving sepsis prediction or multi-site healthcare AI in the UAE, share your draft topic or research questions, target programme, and what you currently know about clinical-data access and ethics approval.
PhD Assistance can then assess the proposed design and identify the methodological and feasibility issues that should be resolved before the study moves into implementation.
This review supports research-proposal development and does not replace institutional ethics, legal or clinical review.