HomeOcular & Auditory Diagnostic ImagingOptical Coherence Tomography Retinal Layer Segmentation

🩺 Optical Coherence Tomography Retinal Layer Segmentation

This simulation focuses on optical coherence tomography (OCT) retinal layer segmentation. It allows for the detailed analysis of retinal layers, which is essential in diagnosing and monitoring various eye diseases such as macular degeneration and diabetic retinopathy.

Ocular & Auditory Diagnostic Imaging2DModerate60 FPS
oct-retinal-segmentation ↗ Open standalone

Low-Coherence Interferometry — Measuring Tissue Depth with a Beam of Light

Optical coherence tomography (OCT), introduced by Huang, Swanson, Fujimoto and colleagues (Science, 1991), performs a kind of "optical ultrasound": instead of timing sound echoes, it times reflections of near-infrared light using interference rather than direct electronic timing, because light travels roughly a million times faster than sound — far too fast to time directly with any conventional detector.

  • 1991: OCT introduced (Huang et al., Science; MIT/Fujimoto lab)
  • SLD: Light source (superluminescent diode, 840–1050nm)
  • <750 μW: Typical power at cornea (ANSI Class 1 laser safety limit)
  • ~2 mm: Depth range imaged (full retinal + choroidal thickness)

The Michelson interferometer and coherence gating

OCT exploits low temporal coherence to convert an impossible timing problem into a solvable interference problem:

• A broadband light source (coherence length Lc of only a few micrometers) is split by a beam splitter into a reference arm (fixed or scanned mirror) and a sample arm (directed into the eye) • Light backscattered from each depth in the retina travels back and recombines with the reference beam • Interference fringes form ONLY when the two path lengths match to within the coherence length of the source — this is "coherence gating" • Because Lc is inversely proportional to source bandwidth, a spectrally broad, temporally incoherent source paradoxically gives the sharpest possible depth discrimination • Scanning the reference mirror path length (time-domain OCT, 1991–2002) or Fourier-decoding a fixed-path spectrum (Fourier-domain OCT, 2002–present) reconstructs a full reflectivity profile — the A-scan — along the beam axis

Each A-scan encodes the depth-resolved reflectivity of one point on the retina; scanning the beam laterally across thousands of points builds the two-dimensional B-scan cross-section, and raster-scanning B-scans builds a full 3-D macular cube.

Axial resolution — why bandwidth beats wavelength

The axial (depth) resolution of OCT is governed by the coherence length of the source, not by the numerical aperture of the optics as in conventional microscopy:

Lc ≈ 0.44 · λ0² / Δλ

where λ0 is the center wavelength and Δλ is the spectral bandwidth (FWHM) of the source. A narrowband source (Δλ≈20nm, typical of early time-domain systems) yields Lc≈15μm — barely enough to distinguish major retinal layers. Modern spectral-domain systems (Δλ≈50nm at 840nm) achieve ≈5μm axial resolution, resolving all ten reflective bands. Research-grade ultrahigh-resolution OCT using Ti:sapphire or supercontinuum sources with Δλ>100nm reaches 2–3μm, sufficient to resolve individual photoreceptor outer-segment discs in some conditions.

Critically, axial resolution in OCT is decoupled from lateral (transverse) resolution, which is set instead by the focused spot size of the sample-arm optics (typically 15–20μm in clinical instruments, limited by the pupil aperture) — the opposite trade-off from a conventional camera lens.

A 50nm-bandwidth, 840nm superluminescent diode gives roughly 5.4μm axial resolution — fine enough to separate ten distinct retinal bands within a retina that is itself only 250–300μm thick, a feat equivalent to resolving 50-plus distinguishable layers across the width of a human hair.

Spectral-Domain and Swept-Source Acquisition — Two Ways to Read the Spectrum

Modern clinical OCT is Fourier-domain: instead of mechanically scanning a reference mirror through every depth (slow, ~400 A-scans/s in 1990s time-domain systems), the entire depth profile is recovered in one shot from the spectral interference pattern via a Fourier transform — a change that increased imaging speed roughly 100-fold and made dense 3-D macular cubes routine.

  • 26–85 kHz: SD-OCT A-scan rate (840nm, spectrometer + CCD)
  • 100–600 kHz: SS-OCT A-scan rate (1050nm, swept laser + balanced photodiode)
  • ~400 Hz: Time-domain OCT (1996) (Stratus OCT — first commercial system)
  • lower in SS: Sensitivity roll-off (longer coherence length of swept laser)

Spectral-domain OCT — a diffraction grating replaces the moving mirror

In spectral-domain OCT (SD-OCT), the reference mirror is fixed. The combined sample-and-reference light is instead passed through a diffraction grating that spreads it into its component wavelengths, which fall across a line-scan CCD or CMOS detector — each pixel records one narrow wavelength band of the interference spectrum simultaneously.

An inverse Fourier transform of this captured spectrum yields the complete depth-reflectivity profile (A-scan) in a single camera exposure, eliminating the mechanical mirror scan entirely. Because there is no moving part in the reference arm, SD-OCT dramatically increased both speed (from ~400 A-scans/s to tens of kHz) and phase stability. Commercial 840nm SD-OCT systems (Cirrus, Spectralis, RTVue) are the clinical standard for macular and glaucoma imaging today, typically acquiring 25,000–85,000 A-scans per second.

SD-OCT does suffer "sensitivity roll-off": signal strength falls off with imaging depth because the finite pixel size and optical resolution of the spectrometer blur higher-frequency spectral fringes, which correspond to deeper structures — limiting practical imaging depth to roughly 2mm.

Swept-source OCT — a tunable laser sweeps through wavelength in time

Swept-source OCT (SS-OCT) instead uses a rapidly wavelength-tunable laser that sweeps through its full spectral range (e.g., 1000–1100nm) many thousand times per second. A single fast photodetector (balanced photodiode pair) records the interference signal as a function of time, which corresponds directly to wavelength. This eliminates the spectrometer and its resolution-limiting pixel grid, permitting much higher sweep — and thus A-scan — rates (100–600 kHz in current commercial systems, with research systems exceeding 1 MHz using Fourier-domain mode-locked lasers).

Operating near 1050nm rather than 840nm also reduces both light scattering by the retinal pigment epithelium and absorption, giving SS-OCT deeper, more uniform penetration into the choroid and sclera — clinically valuable for evaluating choroidal thickness in pachychoroid disease, high myopia, and central serous chorioretinopathy. The trade-off is a modest reduction in axial resolution compared to the widest-bandwidth SD-OCT systems, since 1050nm swept lasers currently offer somewhat narrower usable bandwidths than 840nm superluminescent diodes.

Reading the B-Scan — Ten Bands from Vitreous to Choroid

A single high-quality OCT B-scan through the fovea reveals the entire laminar architecture of the neurosensory retina as alternating hyper- and hypo-reflective horizontal bands. The 2014 international IN·OCT consensus nomenclature standardized the naming of ten such bands, ending decades of inconsistent terminology across manufacturers and publications.

  • 10: Consensus bands (2014) (IN·OCT panel, Ophthalmology)
  • ~250 μm: Foveal center thickness (central subfield, normal adult)
  • RPE: Most hyperreflective band (melanin + high scattering)
  • ELM / BM: Thinnest resolvable band (~10–20 μm, near resolution limit)

The laminar anatomy encoded in reflectivity

From the vitreous cavity inward to the choroid, a foveal-region B-scan (outside the foveal pit, where inner layers are present) shows:

1. Retinal Nerve Fiber Layer (RNFL) — moderately hyperreflective, ganglion cell axons converging toward the optic nerve; thickness here is the primary glaucoma biomarker 2. Ganglion Cell Layer (GCL) — hyporeflective, ganglion cell bodies; thins early in glaucoma and some optic neuropathies 3. Inner Plexiform Layer (IPL) — hyperreflective synaptic layer between GCL and INL 4. Inner Nuclear Layer (INL) — hyporeflective, cell bodies of bipolar, horizontal, and amacrine cells; the layer where diabetic macular edema cysts most often accumulate 5. Outer Plexiform Layer (OPL) — hyperreflective synaptic layer (Henle fiber layer obliquely oriented near the fovea) 6. Outer Nuclear Layer (ONL) — hyporeflective, photoreceptor cell bodies 7. External Limiting Membrane (ELM) — a thin hyperreflective line marking junctional complexes between photoreceptor inner segments and Müller cells; ELM integrity predicts visual recovery after macular surgery 8. Ellipsoid Zone (EZ, formerly "IS/OS junction") — a sharply hyperreflective band from mitochondria-dense photoreceptor inner segments; EZ disruption is one of the single best OCT predictors of visual acuity in macular disease 9. Retinal Pigment Epithelium (RPE) — the brightest band, melanin- and lipofuscin-rich; drusen elevate and CNV disrupts this layer in AMD 10. Bruch's Membrane (BM) — a thin band separating RPE from the choriocapillaris, often only resolvable as separate from RPE when the two are pathologically split by sub-RPE fluid or drusen

The ellipsoid zone and RPE together form the two brightest bands in any OCT B-scan because photoreceptor mitochondria and RPE melanosomes are both exceptionally strong backscatterers of near-infrared light — a purely biophysical accident of pigment and organelle density that ophthalmologists now depend on daily to grade disease severity.

Automated Layer Segmentation — From Graph Search to Convolutional Networks

Manually tracing ten boundaries across the roughly 50 million voxels in a modern OCT macular cube is impossible at clinical scale. Two computational generations solved this: graph-theoretic boundary search in the 2000s–2010s, then convolutional deep learning from 2015 onward — the latter now the basis of every major commercial OCT segmentation engine.

  • 2015: U-Net architecture (Ronneberger, Fischer, Brox)
  • 0.93–0.99: ReLayNet layer Dice (Roy et al. 2017, 8 retinal classes)
  • ~50–200 M: Voxels per macular cube (512×128×1024 typical scan)
  • >99%: Manual grading time saved (automated vs. expert boundary tracing)

From graph-cut boundary tracing to encoder–decoder networks

Early automated segmentation (Chiu et al. 2010, Garvin et al. 2009) treated each B-scan as a weighted graph: pixel-to-pixel edges were assigned costs based on intensity gradients, and Dijkstra-style shortest-path or graph-cut algorithms found globally optimal layer boundaries. These methods work well on clean, high-signal scans but degrade sharply in the presence of fluid, hemorrhage, or severe distortion — exactly the pathological cases where clinicians need segmentation most.

U-Net (Ronneberger et al., MICCAI 2015) reframed the problem as dense pixel-wise classification. Its encoder path repeatedly downsamples the image through convolution and max-pooling, building increasingly abstract, large-receptive-field features; a mirrored decoder path upsamples back to full resolution, and skip connections concatenate encoder feature maps directly into the corresponding decoder stage — preserving the fine spatial detail that pure downsampling would destroy. The network outputs a per-pixel probability distribution over layer classes (10 retinal layers + background + optional pathology classes), trained end-to-end against expert-annotated ground truth using a pixel-wise cross-entropy or Dice loss.

ReLayNet (Roy et al., Biomedical Optics Express 2017) was among the first to adapt this architecture specifically to retinal OCT, using a weighted combination of Dice loss and logistic loss with class weighting to counter the huge imbalance between thin layers (ELM, few pixels) and thick ones (ONL, many pixels) — reporting mean Dice coefficients of 0.93–0.99 across nine retinal layers plus fluid class on 100-patient held-out test data.

Training data, loss functions, and failure modes

Clinical-grade segmentation networks are trained on thousands to tens of thousands of expert-annotated B-scans, typically double-graded by two retina specialists with adjudication of disagreements to produce consensus ground truth. Data augmentation (elastic deformation, intensity jitter, simulated speckle noise) compensates for limited labeled data relative to natural-image datasets.

Common loss functions combine per-pixel cross-entropy with a soft Dice or Tversky loss term, which directly optimizes the overlap metric clinicians care about and is more robust to the severe class imbalance between thin and thick layers. 3-D networks that segment full macular cubes (rather than B-scan by B-scan) exploit inter-slice continuity for smoother, more anatomically plausible surfaces.

Failure modes remain clinically important: segmentation networks trained predominantly on one manufacturer's scan characteristics can degrade on another vendor's speckle pattern or signal-to-noise profile (domain shift); dense subretinal hemorrhage, large pigment epithelial detachments, and advanced geographic atrophy can all confuse boundary classification, and every commercial platform still ships a manual correction tool because fully automated output is not yet certified as error-free for every pathology.

From Pixels to Diagnosis — Macular Edema, Glaucoma, and AMD Metrics

Layer segmentation is not an end in itself — it is the substrate for quantitative biomarkers that drive real treatment decisions: how many milliliters of anti-VEGF to inject, whether a glaucoma patient's field of vision is truly deteriorating, and how large a drusen deposit has grown between visits.

  • ~250–275 μm: Normal CST (central subfield) (ETDRS 1mm central circle)
  • >320 μm (M) / >305 μm (F): DME treatment threshold (central subfield thickness, protocol-dependent)
  • ~95–105 μm: Normal avg. RNFL thickness (peripapillary circle scan, age-dependent)
  • ~1–2 μm/yr: Glaucoma RNFL loss rate (vs. ~0.5 μm/yr normal aging)

Diabetic macular edema — central subfield thickness and fluid volume

Diabetic macular edema (DME) affects an estimated 750,000+ people in the United States and remains a leading cause of vision loss in working-age adults worldwide. Segmentation software overlays the ETDRS 9-subfield grid on the macular cube and reports central subfield thickness (CST) — the average retinal thickness (ILM to RPE) within the central 1mm circle — the single most widely used OCT biomarker in retina clinics.

Breakdown of the blood-retinal barrier allows fluid to accumulate intraretinally (IRF, classically pooling in the outer plexiform and inner nuclear layers as cystoid spaces) or beneath the retina (subretinal fluid, SRF, between photoreceptors and RPE). Segmentation algorithms increasingly classify and volumetrically quantify these fluid compartments separately, since IRF and SRF carry different prognostic weight and respond differently to anti-VEGF versus corticosteroid therapy. Landmark trials (RIDE/RISE, VISTA/VIVID, Protocol T) all used OCT-measured CST as a primary or key secondary endpoint, and treat-and-extend or PRN retreatment decisions in routine practice are made largely on month-to-month CST trends.

Glaucoma — RNFL and ganglion cell complex thinning

Glaucoma progressively destroys retinal ganglion cells and their axons, thinning the RNFL and the combined ganglion cell-inner plexiform layer (GCIPL, sometimes called the ganglion cell complex). OCT segmentation generates a circular peripapillary RNFL thickness profile and a macular GCIPL thickness map, each compared automatically against a normative database stratified by age, and color-coded green/yellow/red by percentile.

Normal average RNFL thickness is roughly 95–105μm (varies by device and by quadrant — thickest inferiorly and superiorly, thinnest temporally, following the "ISNT rule": Inferior > Superior > Nasal > Temporal). Healthy aging causes slow physiologic thinning of roughly 0.5μm per year; glaucomatous eyes typically thin 2–4 times faster (roughly 1–2μm/year, more in aggressive or poorly controlled disease). Because a single scan's measurement noise (test-retest variability ~2–3μm) can rival one year of true glaucomatous change, OCT progression analysis relies on trend-based statistics (linear regression of RNFL thickness across many visits, e.g. Guided Progression Analysis) rather than any single measurement — segmentation reproducibility is therefore as clinically important as segmentation accuracy.

Age-related macular degeneration — drusen and neovascular fluid

In age-related macular degeneration (AMD), segmentation of the RPE and Bruch's membrane boundaries allows automated detection and volumetric measurement of drusen — lipid and protein deposits that elevate the RPE band above Bruch's membrane. Drusen volume and area are validated biomarkers for progression risk to advanced AMD (geographic atrophy or neovascular AMD).

In neovascular ("wet") AMD, choroidal neovascular membranes leak fluid that segmentation software again classifies as IRF, SRF, or sub-RPE fluid — the presence of any exudative fluid on OCT is now the standard trigger for anti-VEGF re-treatment in treat-and-extend regimens used worldwide, effectively replacing fluorescein angiography for most day-to-day management decisions. This single application — OCT-guided anti-VEGF dosing — is estimated to account for the majority of all OCT scans performed in retina clinics globally.

De Fauw et al. 2018 — DeepMind, Moorfields, and OCT-Based Referral Triage

The clearest demonstration that OCT segmentation networks can operate at expert clinical level came from a 2018 Nature Medicine collaboration between DeepMind and Moorfields Eye Hospital, London: a two-stage deep learning system that reads raw OCT volumes and recommends a referral pathway with accuracy matching or exceeding senior ophthalmologists and optometrists.

  • 2018: Publication (De Fauw et al., Nature Medicine)
  • 14,884: Training scans (from 7,621 patients, Moorfields)
  • 4: Referral pathways classified (urgent / semi-urgent / routine / observe)
  • 0.992: Urgent-referral AUC (matched top 3 of 8 expert graders)

A two-stage architecture: segment first, then diagnose

Rather than training one network to map raw OCT voxels directly to a clinical decision — an approach vulnerable to shortcut learning and poor generalization across device types — the DeepMind/Moorfields system deliberately split the task in two:

Stage 1 — 3D U-Net tissue segmentation: converts a raw OCT volume (from either Topcon 3D OCT-1000 or Topcon 3D OCT-2000 scanners) into a device-independent tissue-segmentation map, classifying each voxel into one of 15 tissue types (retinal layers plus pathological features such as fluid, hemorrhage, and various forms of neovascular and atrophic change). This intermediate representation strips away device-specific noise characteristics, letting the diagnostic stage generalize across scanner hardware without retraining.

Stage 2 — classification network: takes the segmentation map (not the raw pixels) and outputs both a probability distribution over 53 retinal diagnoses and a referral recommendation across four urgency tiers: urgent, semi-urgent, routine, and observation-only. Framing the intermediate output as clinically interpretable tissue maps also let ophthalmologists visually audit what the network "saw" before trusting its recommendation — an early practical answer to the black-box problem in medical AI.

Performance and validation against expert clinicians

Trained on 14,884 OCT scans from 7,621 patients (retrospectively collected, anonymized, spanning the common sight-threatening conditions seen in a general retina clinic — AMD, diabetic eye disease, retinal vein occlusion, and others), the system was evaluated against a held-out test set graded independently by a panel of retina specialists.

For the highest-stakes decision — identifying patients needing urgent referral — the model achieved an area under the ROC curve of 0.992, with error rates that did not exceed those of the top 3 out of 8 expert retina specialists and optometrists in the reference panel. Performance was maintained even on scans acquired by the second scanner model not represented in most of the training set, supporting the tissue-segmentation intermediate step's role in cross-device generalization.

Because the model never had to be retrained for the second Topcon scanner type, the study offered early evidence that segmentation-first architectures could resist the device-generalization failures that plague many single-stage medical imaging networks — a design pattern subsequently echoed across other clinical deep-learning triage systems.

From research system to deployed clinical tool

The De Fauw et al. system was explicitly designed as a triage aid, not an autonomous diagnostician: its output is a referral urgency recommendation reviewed by a clinician, not a final diagnosis delivered to a patient. This framing reflects the practical regulatory and liability path for most clinical AI in ophthalmology to date — decision support that sits inside, not instead of, the referral pathway.

Commercial OCT segmentation has since become standard equipment: Zeiss, Heidelberg Engineering, Topcon, and Optovue all ship deep-learning-based automated layer segmentation, drusen and fluid quantification, and progression analysis as core software on current-generation devices, and several platforms now include AI-assisted referral flags for suspected AMD, DME, and glaucomatous change directly at the point of imaging — bringing the De Fauw architecture's core idea, segment-then-triage, into everyday high-volume eye clinics screening tens of millions of patients per year.

⚙ Under the hood

This simulation focuses on optical coherence tomography (OCT) retinal layer segmentation. It allows for the detailed analysis of retinal layers, which is essential in diagnosing and monitoring various eye diseases such as macular degeneration and diabetic retinopathy.

CanvasBiomedicine

2D · HTML5 Canvas 2D · 60 FPS target · runs fully client-side, no install

What did you find?

Add reproduction steps (optional)