Skip to main content

phdassistance

Unsupervised Domain Adaptation for Diabetic Retinopathy on Handheld Fundus Images: UAE Primary-Care PhD Study Design

“Deep learning for diabetic retinopathy” is a crowded topic that rarely survives a novelty question. This one can: a classifier trained on high-quality tabletop fundus images performs worse on the lower-cost handheld images primary care can afford, and a scholar can measure how much of that loss unsupervised domain adaptation recovers without labelling thousands of images.

Who is this for: PhD scholars in computer vision, biomedical engineering, health informatics and ophthalmology research; clinician-scientists; and supervisors scoping imaging-AI projects in the UAE and GCC.

Scope. This article separates published evidence, cited to source, from research hypotheses, labelled as such, and design guidance. It is research-methodology guidance only. Nothing here is clinical, diagnostic, or screening-policy advice, and no system described here should decide any individual’s care or referral; clinicians, under their institution’s governance, make those decisions.

1. Why diabetic retinopathy screening in UAE primary care is still a research problem

Diabetes is managed in primary care; retinopathy is detected in eye clinics. That gap is the problem. A cross-sectional survey of 513 people with diabetes across eight primary health care centres and two hospitals in Al Ain found retinopathy in 19% (95% CI 15.1–23.5), proliferative disease in 3.8%, and 74% unaware of their eye condition [1]. Its 2003–04 fieldwork shows an unmet detection need was documented in the UAE, not what the figure is today; an updated local estimate is a defensible first study.

Handheld cameras can make screening more feasible in community and primary-care settings, being portable and usable without tabletop-camera infrastructure [18]. The research question is not whether to use them, but what happens to an automated classifier when the images change — and whether that can be fixed without a large labelling budget.

2. DR grading basics: what a screening model must detect

The International Clinical Diabetic Retinopathy scale defines five levels — none, mild, moderate and severe non-proliferative, and proliferative — with a separate scale for diabetic macular oedema [2]. Screening studies rarely predict all five. They collapse the scale into referable DR, usually moderate non-proliferative or worse plus any macular oedema: the threshold at which a screening programme acts.

Two consequences follow. The cutoff must be prespecified, since moving it changes sensitivity and specificity without changing the model. And whether single-field colour photographs support reliable macular-oedema assessment depends on the acquisition protocol, reference standard and target definition, so any DME endpoint should be prespecified and justified against the imaging protocol available.

3. Tabletop vs handheld fundus cameras: what changes in the image

Article - Thumbnail - PA

Schematic comparison titled “Same eye, two cameras”, labelled throughout as a synthetic, procedurally generated schematic that is not a patient image and not for diagnostic interpretation. A left panel shows a wide, evenly illuminated circular retinal field labelled “Tabletop, mydriatic — source domain, model trained here”. A right panel shows a narrower, dimmer field with edge vignetting, a glare patch, a pupil-margin crescent shadow and one-sided blur, labelled “Handheld, non-mydriatic — target domain, model deployed here”. An arrow between them is labelled “domain shift”. Beneath, five labelled differences: field of view, pupil and illumination, resolution and focus, artefacts, and gradability.

Same eye, two cameras — a research-design schematic, not a patient image. Illustrative acquisition differences; actual domain shift should be empirically characterised in the study dataset.

The differences are mechanical. A handheld device captures a narrower field, so peripheral lesions fall outside it. Without dilation, illumination is dimmer and less even, and the pupil margin can intrude as a dark crescent. Lower resolution and motion blur can obscure small, low-contrast lesions, including microaneurysms that contribute to early DR grading, and non-specialist capture produces artefacts a model trained on studio-grade images has never seen.

These are image-level problems, not architecture-level ones. In a controlled comparison, architecture and framework barely moved performance (AUC 0.936–0.944), while imaging factors moved it substantially: AUC fell from 0.936 to 0.891 as file size dropped from 350 KB to 150 KB, and from 0.949 with seven fields to 0.895 with one [3].

4. Domain shift: why a model that scores well on public datasets fails in the clinic

The size of the drop is published rather than hypothetical. On curated tabletop images, performance is high: a widely cited algorithm reported 97.5% sensitivity and 93.4% specificity for referable DR at its high-sensitivity operating point [4]. Moving the same class of system to community screening, and one developed on mydriatic images, applied to non-mydriatic handheld screening, achieved 62.69% sensitivity (95% CI 59.17–66.12), with ungradable images exceeding 30% at some sites [5]. The authors attribute the fall to image quality, acquisition by minimally trained field workers, co-pathology such as cataract, and single-field macula-centred images.

A deployment study in eleven Thai clinics reached the same conclusion from the human side, reporting tension between the model’s data-quality thresholds and the quality obtainable in a resource-constrained environment [6]. Domain shift is the ordinary condition of screening outside research.

State it as a measurable quantity. A proposal should not say a model “may underperform”. It should prespecify what it will measure — referable-DR sensitivity and specificity at a fixed operating point, AUROC, and gradability rate — and report the unadapted model’s target-domain performance as the study’s baseline.

5. Unsupervised domain adaptation methods compared

Unsupervised domain adaptation uses labelled source images and unlabelled target images. That is the economic argument: unlabelled handheld images are a by-product of screening; graded ones cost expert time.

Family

Representative method

What it does

What to watch in a retinal study

Adversarial alignment

DANN, gradient-reversal layer [7]

Learns features a domain classifier cannot separate

May align away lesion signal with camera signal; unstable; report seed stability

Statistical alignment

Deep CORAL [8]

Matches second-order feature statistics

Stable but weaker; may underfit large artefact-driven shifts

Image-to-image translation

CycleGAN [9]

Restyles source images toward target captures

Can hallucinate or erase lesion-sized structures; needs a lesion-preservation check

Self-training

Pseudo-labelling on target images

Uses confident target predictions as labels

Confirmation bias: a biased baseline teaches itself its errors; fix the rule in advance

Test-time adaptation

Tent [10]

Adapts normalisation statistics at inference

Cheap and deployable, but batch-sensitive; adapts to the batch, not the task

Domain adaptation for DR is an active field, and the gap must be stated against it. A source-free active adaptation method transferred colour-fundus knowledge to ultra-wide-field images, raising accuracy by 20.9 points over baseline [11]. Recent work relaxes further constraints: an online, model-agnostic setting assuming neither source-model access nor a static target set, adapting as target data arrives [12]; and a quality-aware approach routing low-quality images to scale-specialised experts rather than discarding them, which bears directly on gradability [13].

That work solves the algorithmic side of adaptation — operating without source data, without a fixed target set, under variable quality. What it does not establish is performance under a specific acquisition shift: these studies adapt between imaging modalities or public benchmarks, and none reports a labelled target test set from a primary-care pathway. This creates a more specific research gap: whether unsupervised adaptation can recover performance when a classifier moves from conventional tabletop fundus imaging to handheld, non-mydriatic images acquired in a UAE primary-care pathway.

Existing work

Domain shift studied

Target data

Labels used

What it does not answer

Source-free active DA [11]

Colour fundus → ultra-wide-field

Unlabelled UWF

Source-free, few active

A handheld primary-care pathway

Model-absent, flowing data [12]

Benchmark DR domains

Streaming public data

Source-free

Handheld, non-mydriatic acquisition

Quality-aware grading [13]

Mixed-quality public datasets

Public DR datasets

Supervised

Whether quality routing survives a device shift

Proposed study

Tabletop → handheld

UAE target images

Zero or few

— the question it answers

Compare families, not one method against published numbers. Fix the backbone, source data, target data and split; vary only the adaptation method. A null result is a finding, and the thesis should remain examinable if that is the outcome.

6. Image quality and gradability: the first model in the pipeline

Gradability should be treated as an explicit prediction and routing task rather than merely as a nuisance variable. When over 30% of community-captured images can be ungradable [5], a system that silently grades them produces confident output on images a human would refuse.

Design it explicitly: a first-stage model classifies each image as gradable or not, and only gradable images reach the DR classifier. Report the gradability rate, the gradability model’s own accuracy, and DR performance under both denominators — gradable images only, and all images presented with ungradable ones handled by a stated rule. Reporting only the first flatters the system. A referral-grade composite, routing ungradable and referable images onward, is also legitimate [5].

7. Building a target test set: graders, adjudication and reference standard

Adaptation is unsupervised; evaluation is not. A labelled target test set is unavoidable, and its quality bounds every claim made.

Reference standard. Grade on the ICDR scale [2] with certified graders masked to the model output and to each other, and a prespecified adjudication route — a third senior grader, or consensus. Report inter-grader agreement before model accuracy. Where graders agree only moderately, the apparent ceiling on performance is partly a property of the reference standard, not of the model.

Sample size. For the primary diagnostic-accuracy endpoint, planning should be driven by the precision required of the sensitivity estimate, while the full design also accounts for specificity, prevalence and the prespecified operating point. Sensitivity is estimated among cases only, so the binding constraint is the expected number of referable cases: work backwards from the interval width the thesis must defend, and state it.

Independence. Split by patient, never by image or eye: two eyes of one patient, or repeat visits, spanning training and test data inflate every number reported.

8. Evaluation examiners expect: referable-DR sensitivity, specificity and gradability

Report at minimum: sensitivity and specificity for referable DR at a prespecified operating point with confidence intervals; AUROC with its interval; gradability rate and its handling; and the same set for the unadapted baseline and each adaptation method on identical target data. Every method is evaluated on the same images, so the comparison is paired and the analysis should say so in advance.

Report subgroup results where cell sizes permit — by operator, site, image-quality band and, where ethically approved, patient factors — with intervals, and treat them as descriptive. Differences inside two overlapping intervals are not evidence.

For diagnostic-accuracy reporting, STARD 2015 is the primary framework [14]. TRIPOD+AI, which supersedes TRIPOD 2015, applies where prediction-model development or evaluation is reported [15]. SPIRIT-AI and CONSORT-AI become relevant only if the work evolves into an AI intervention trial and its protocol [16,17] — a doctoral timeline rarely supports one, which belongs in the proposal rather than at viva.

9. The labelling-cost argument: adaptation vs small-sample fine-tuning

This section is the thesis’s economic spine, and it rests on a published pair of results from the same research programme.

Applied unadapted to handheld community images, the system reached 62.69% sensitivity for referable DR [5]. Trained directly on handheld non-mydriatic images — 32,494 images from 9,778 participants across 20 community sites — a model reached 93.86% sensitivity (95% CI 91.34–96.08) and 96.00% specificity, AUROC 0.99, on two-field images [18].

Read together, those numbers define the research question. Handheld images are not inherently ungradable by machines; the accuracy is recoverable. But recovering it took a very large labelled handheld dataset, which may be impractical for systems without access to large labelled datasets.

So the examinable question is a curve, not a claim. Plot target-domain sensitivity against the number of labelled target images: zero labels (unsupervised adaptation), then small labelled sets for fine-tuning at several sizes, and record how many labelled images unsupervised adaptation is worth. This is a research hypothesis, not a finding: the method may recover most, some or none of the loss, or prove unstable.

10. Ethics, image licensing and approval in UAE research

Three layers; a proposal addressing only the first layer may be insufficient for institutional ethics and data-governance review.

Federal legislation. Federal Law No. 2 of 2019 on the use of information and communications technology in health fields sets out distinct requirements, including Article 13, restricting the storage, processing, generation or transfer outside the State of health data related to services provided in it, subject to the applicable UAE localisation and permitted-transfer framework and the exceptions established under the relevant health authority decisions; and Article 16, requiring consent for non-health purposes [19]. Its executive regulation is Cabinet Resolution No. 32 of 2020 [20], and the permitted cases under Article 13 are set out in Ministerial Resolution No. 51 of 2021, which defines categories — among them approved scientific research — each with its own conditions [21]. Cite the applicable articles and categories individually, read the conditions from the operative text, and confirm the current text and any superseding instrument with the relevant authority before submission.

Institutional. Name the health authority, the ethics committee and the data custodian the study will approach. Retinal images may constitute personal and, depending on the processing and identifiability context, biometric data, so the study should document its de-identification and data-governance approach: stripping embedded capture metadata, stating residual re-identification risk rather than claiming anonymity, and having the approach approved.

Dataset licensing. Public source datasets differ in whether they permit commercial use, redistribution or derivative model release. Read the licence before building on it and record the terms in the methods chapter — a thesis that cannot release its model should know at the start.

11. PhD Assistance imaging-AI domain-shift study design checklist

Dimension

Evidence to hold before submission

Risk if unresolved

Status

Source data

A named dataset licensed for the intended use

Model cannot be released or published

To be confirmed

Target device

A named camera and sites, with a written access route to unlabelled images

The commonest cause of abandoned imaging-AI PhDs

To be confirmed

Adaptation method

Two or more families on a fixed backbone and split [7–11]

A single method reads as a tutorial, not a comparison

To be confirmed

Reference standard

Masked certified graders, adjudication rule, agreement reported first

No credible denominator for accuracy claims

To be confirmed

Evaluation metric

Prespecified operating point; interval widths computed in advance [14,15]

An estimate the sample cannot support

To be confirmed

Feasibility

Compute, grader time and labelled-subset budget costed

Sound design, undeliverable in the timeline

To be confirmed

Any dimension without evidence behind it is a decision still to be taken or a limitation to declare — neither acceptable left unstated.

12. How to turn this blueprint into an approvable PhD proposal

Evidence gap map

Existing evidence

Main contribution

Remaining gap

How the proposed UAE study addresses it

Tabletop DL validation [4]

High accuracy on curated images

Mydriatic, specialist-captured

Measures the same task on handheld images from the intended pathway

Mydriatic-to-community evaluation [5]

Quantifies the drop: 62.69% sensitivity; >30% ungradable

Documents the failure; tests no remedy

Treats that drop as the baseline adaptation must beat

Handheld-trained model [18]

93.86% sensitivity from 32,494 labelled handheld images

Recovery bought with a very large labelled dataset

Asks what recovery is available with zero or few labels

Imaging-factor analysis [3]

Image factors outweigh architecture choice

Identifies causes; adapts no model

Uses them to design the shift diagnostics and ablations

UDA for DR grading [11–13]

Source-free, model-agnostic and quality-aware adaptation

Modality or benchmark shifts, on research datasets

Moves it to tabletop-to-handheld shift in primary care

Deployment field study [6]

Clinic quality thresholds conflict with clinic images

Qualitative; no accuracy target

Converts the constraint into gradability measures

UAE prevalence survey [1]

19% DR, 74% unaware, in UAE primary care

Dated 2003–04; no device comparison

Supplies local justification; motivates an update

Research questions, with the inferential stance stated. RQ1 (primary): how much of the tabletop-to-handheld drop in referable-DR sensitivity does unsupervised adaptation recover against the unadapted baseline on the same target test set? The null is no recovery; the alternative is two-sided, since adaptation can degrade performance. RQ2: How do adaptation families compare on one backbone and split — estimation, not testing. RQ3: How many labelled target images would match the best unsupervised result by fine-tuning? RQ4 (exploratory): does a gradability stage change the comparison?

Contribution. A completed thesis would deliver a measured characterisation of tabletop-to-handheld shift on a named device; a controlled comparison of adaptation families under one protocol; a labelled target test set with a documented reference standard; and a labelling-cost curve. Single-device evidence bounds generalisation and the reference standard bounds measurable accuracy. This design alone would not support a clinical deployment claim; deployment would require further prospective, regulatory and health-system evaluation appropriate to the intended use.

Common weaknesses that may require proposal revision: no named target device or site; novelty resting only on a public dataset; one adaptation method without comparison; a reference standard assumed rather than designed; splits by image rather than patient; gradability ignored; and a deployment claim the design cannot support.

Frequently asked questions

On curated tabletop images, performance is high [4]. On handheld community-screening images, a system developed on mydriatic images reached 62.69% sensitivity for referable DR [5]. The transfer is the unsolved part, and where the contribution remains.

Because the comparison is the research question. One model reached 93.86% sensitivity after training on 32,494 labelled handheld images [18]; most systems have no such dataset. Measuring how far adaptation gets with zero labels, and how many labels would match it, is the contribution.

Yes — for evaluation, not training. Adaptation uses unlabelled target images; the test set must be graded by masked certified graders with an adjudication rule, sized by the precision required for sensitivity among referable cases.

That can still constitute a scientifically informative result, provided the baseline, comparison and analysis were prespecified. Design the thesis so a null outcome still answers RQ1 and RQ3, rather than contributing only on a positive finding.

Existing work is defined by the shift it studies: between imaging modalities, between public benchmarks, or under streaming or variable-quality conditions [11–13]. None report adaptation from tabletop to handheld, non-mydriatic capture in a primary-care pathway, evaluated against a labelled target test set from it with the label-cost curve quantified. The gap is the acquisition shift and the label efficiency measured against it, not domain adaptation as a method.

No. It produces measured evidence about model behaviour under domain shift. Clinical use would require prospective evaluation, regulatory approval and health-system governance that a doctoral project does not deliver, and a thesis should say so explicitly.

Request an imaging-AI topic and methodology review

Start an imaging-AI methodology consultation with PhD Assistance’s research-methodology team: a research-design review of your proposed topic, data plan and evaluation design against this framework, delivered as a written recommendations summary you can take into a supervisor meeting. The consultation advises on methodology; you carry out the research.

  1. Topic assessment — your question mapped against published DR, handheld-camera and domain-adaptation work, and where the novelty claim rests.
  2. Data and reference-standard review — source licensing, target-device access, grader plan, adjudication rule.
  3. Method and evaluation review — adaptation families, split design, operating point, precision of planned estimates.
  4. Recommendations — what is defensible, what needs evidence and what to change.

What must be true before this PhD starts. A named source dataset and licence; a named target camera and site; a written image-access route; an estimated target-image volume; grader availability; an ethics route; a compute budget; the adaptation families to compare; and a costed labelled test-set budget. Each is a feasibility decision, not a formality.

To start, share your draft topic, target programme, and what you know about image access and ethics approval. The review supports your own proposal writing rather than replacing it, and provides no clinical guidance.

Related support: research proposal development · machine learning and statistical analysis · systematic review

Reference

  1. Al-Maskari F, El-Sadig M. Prevalence of diabetic retinopathy in the United Arab Emirates: a cross-sectional survey. BMC Ophthalmol. 2007;7:11. doi:10.1186/1471-2415-7-11
  2. Wilkinson CP, Ferris FL 3rd, Klein RE, et al.; Global Diabetic Retinopathy Project Group. Proposed international clinical diabetic retinopathy and diabetic macular edema disease severity scales. Ophthalmology. 2003;110(9):1677–1682. doi:10.1016/S0161-6420(03)00475-5
  3. Yip MYT, Lim G, Lim ZW, et al. Technical and imaging factors influencing performance of deep learning systems for diabetic retinopathy. npj Digit Med. 2020;3:40. doi:10.1038/s41746-020-0247-1
  4. Gulshan V, Peng L, Coram M, et al. Development and validation of a deep learning algorithm for detection of diabetic retinopathy in retinal fundus photographs. JAMA. 2016;316(22):2402–2410. doi:10.1001/jama.2016.17216
  5. Nunez do Rio JM, Nderitu P, Bergeles C, Sivaprasad S, Tan GSW, Raman R. Evaluating a deep learning diabetic retinopathy grading system developed on mydriatic retinal images when applied to non-mydriatic community screening. J Clin Med. 2022;11(3):614. doi:10.3390/jcm11030614
  6. Beede E, Baylor E, Hersch F, Iurchenko A, Wilcox L, Raumviboonsuk P, Vardoulakis LM. A human-centred evaluation of a deep learning system deployed in clinics for the detection of diabetic retinopathy. Proc 2020 CHI Conference on Human Factors in Computing Systems. 2020:1–12. doi:10.1145/3313831.3376718
  7. Ganin Y, Ustinova E, Ajakan H, et al. Domain-adversarial training of neural networks. J Mach Learn Res. 2016;17(59):1–35.
  8. Sun B, Saenko K. Deep CORAL: correlation alignment for deep domain adaptation. In: Computer Vision – ECCV 2016 Workshops. Lecture Notes in Computer Science, vol 9915. Cham: Springer; 2016:443–450. doi:10.1007/978-3-319-49409-8_35
  9. Zhu J-Y, Park T, Isola P, Efros AA. Unpaired image-to-image translation using cycle-consistent adversarial networks. Proc IEEE International Conference on Computer Vision (ICCV). 2017:2223–2232.
  10. Wang D, Shelhamer E, Liu S, Olshausen B, Darrell T. Tent: fully test-time adaptation by entropy minimization. International Conference on Learning Representations (ICLR). 2021.
  11. Ran J, Zhang G, Xia F, Zhang X, Xie J, Zhang H. Source-free active domain adaptation for diabetic retinopathy grading based on ultra-wide-field fundus images. Comput Biol Med. 2024;174:108418. doi:10.1016/j.compbiomed.2024.108418
  12. Su W, Tang S, Liu X, Yi X, Ye M, Zu C, Li J, Zhu X. Domain adaptive diabetic retinopathy grading with model absence and flowing data. Proc IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2025:28337–28346.
  13. Pan Y, Cai X, Xiong P, Mei F, Wang Z, Hong J. From discarding to leveraging: quality-aware collaborative learning for robust diabetic retinopathy grading. Front Med. 2026;13:1795484. doi:10.3389/fmed.2026.1795484
  14. Bossuyt PM, Reitsma JB, Bruns DE, et al.; STARD Group. STARD 2015: an updated list of essential items for reporting diagnostic accuracy studies. BMJ. 2015;351:h5527. doi:10.1136/bmj.h5527
  15. Collins GS, Moons KGM, Dhiman P, et al. TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ. 2024;385:e078378. https://doi:10.1136/bmj-2023-078378
  16. Cruz Rivera S, Liu X, Chan A-W, Denniston AK, Calvert MJ; SPIRIT-AI and CONSORT-AI Working Group. Guidelines for clinical trial protocols for interventions involving artificial intelligence: the SPIRIT-AI extension. Nat Med. 2020;26(9):1351–1363. doi:10.1038/s41591-020-1037-7
  17. Liu X, Cruz Rivera S, Moher D, Calvert MJ, Denniston AK; SPIRIT-AI and CONSORT-AI Working Group. Reporting guidelines for clinical trial reports for interventions involving artificial intelligence: the CONSORT-AI extension. Nat Med. 2020;26(9):1364–1374. org/10.1038/s41591-020-1034-x
  18. Nunez do Rio JM, Nderitu P, Raman R, et al. Using deep learning to detect diabetic retinopathy on handheld non-mydriatic retinal images acquired by field workers in community settings. Sci Rep. 2023;13:1392. doi:10.1038/s41598-023-28347-z
  19. United Arab Emirates. Federal Law No. 2 of 2019 Concerning the Use of Information and Communications Technology in Health Fields, Arts 13 and 16. Issued 6 February 2019; Official Gazette No. 647; in force 14 May 2019. https://uaelegislation.gov.ae/en/legislations/1209
  20. United Arab Emirates Cabinet. Cabinet Resolution No. 32 of 2020 Concerning the Executive Regulation of Federal Law No. 2 of 2019. Issued 22 April 2020; Official Gazette No. 677; in force 30 October 2020. https://uaelegislation.gov.ae/en/legislations/1444
  21. Ministry of Health and Prevention, United Arab Emirates. Ministerial Resolution No. (51) of 2021 regarding the cases in which health data and information may be stored or transferred outside the country. UAE Public Health Legislation portal. https://uaephl.mohap.gov.ae/en/health-policies-and-legislations-advocacy/health-legislations