Spatial refinement of expert annotations for retinal images using foundation models
Haider, Fasie (2026)
Diplomityö
Haider, Fasie
2026
School of Engineering Science, Laskennallinen tekniikka
Kaikki oikeudet pidätetään.
Julkaisun pysyvä osoite on
https://urn.fi/URN:NBN:fi-fe2026060864782
https://urn.fi/URN:NBN:fi-fe2026060864782
Tiivistelmä
Expert annotations in retinal image datasets are commonly spatially imprecise: lesions can be coarsely marked with geometric primitives like circles, ellipses, and polygons that approximate where a lesion lies but do not follow its boundary. A pipeline is developed and evaluated for refining these coarse annotations into boundary-accurate segmentation masks using promptable foundation models. The coarse markings from the DiaRetDB2 dataset are converted into point, box, and mask prompts and passed to three models zero-shot inference: Segment Anything Model (SAM), medical-domain variant MedSAM, and SAM 2.1, using six prompt strategies. Each refined mask is post-processed and constrained to remain within the expert’s coarse region. The approach is tested in three experiments: refinement against a 2-of-3 expert consensus, refinement of the least-detailed marking toward the more detailed ones, and refinement on a three-channel representation derived from the spectral images by principal component analysis. At aggregate level, no model clearly exceeds the coarse baseline, but a per-class analysis shows that refinement adds real value where the coarse marking overstates the lesion, for small punctate lesions, and degrades diffuse classes whose coarse mask already covers the lesion. The mean pair-wise agreement between experts is well below 0.5 Dice. The results show that foundation models can refine a lesion boundary but cannot correct a misplaced location, that mask- and box-based prompts are the safe default while point-only prompts are damaging, and that the spectral projection does not improve refinement over standard color input.
