HomeIMRT/VMAT Treatment PlanningRadiotherapy Plan Quality Assurance Gamma Analysis

🎯 Radiotherapy Plan Quality Assurance Gamma Analysis

This simulation performs gamma analysis for quality assurance in radiotherapy plans, comparing the planned dose distribution with the delivered dose to identify discrepancies and ensure treatment accuracy.

IMRT/VMAT Treatment Planning2DModerate60 FPS
radiotherapy-gamma-analysis-qa ↗ Open standalone

Plan Export & Patient-Specific QA Setup

Before a radiotherapy plan ever touches a patient, it must be verified on a phantom. The treatment planning system (TPS) exports the exact beam parameters — MLC sequences, monitor units, gantry angles — that will later be delivered to the patient, and a 2D detector array is set up in the beam path to measure what the linac actually produces.

  • 5–10 mm: Detector chamber spacing (typical 2D ion-chamber array)
  • IMRT / VMAT: Plans requiring patient-specific QA (all modulated deliveries)
  • Pre-treatment: QA performed (before fraction 1)
  • TPS dose grid: Reference dataset (exported as DICOM RT-Dose)

Why measure what was already calculated?

A treatment planning system computes dose using an approximate physical model of the linac head, patient anatomy, and tissue heterogeneities. The plan is only theoretical until the actual machine reproduces it. Delivery introduces real-world sources of error that no calculation can predict: multi-leaf collimator (MLC) leaves that lag or overshoot their commanded positions, gantry sag, dose-rate fluctuations, and collimator backlash.

Patient-specific QA closes this loop. The plan is delivered to a phantom instead of the patient, dose is measured with a calibrated detector, and the measured distribution is compared against the TPS-predicted one. If they agree within clinically accepted tolerances, the plan is cleared for treatment. If not, the plan or the machine must be investigated before a single fraction is delivered.

Patient-specific QA is mandatory for essentially all modulated deliveries (IMRT, VMAT, stereotactic techniques) precisely because their highly non-uniform fluence patterns are the deliveries most vulnerable to small mechanical and dosimetric errors.

Exporting the reference dose distribution

The TPS calculates a 3D dose grid (typically 1–3 mm voxel spacing) using a validated dose-calculation algorithm — collapsed cone convolution, analytic anisotropic algorithm (AAA), or Monte Carlo. For QA purposes, a 2D or 3D dose plane matching the detector array's measurement geometry is exported, usually as a DICOM RT-Dose object, along with the beam delivery instructions (RT-Plan) that will be sent to the linac control system.

This exported plane becomes the reference — the "planned" dose map — against which every measured point will later be judged. Its accuracy depends entirely on how well the TPS algorithm models beam transport, scatter, and tissue heterogeneity, which is itself a known limitation the QA process is partly designed to catch.

Setting up the detector array

A 2D array of ionization chambers or diodes is placed in a slab or cylindrical phantom, oriented in the plane the plan will be measured in (commonly coronal, for a composite or per-arc measurement). The array must be positioned with sub-millimeter setup accuracy, since a shifted array introduces its own spurious dose-difference and distance-to-agreement errors unrelated to the actual plan.

Modern arrays contain hundreds to over a thousand individual sensors spaced roughly 5–10 mm apart, each capable of registering an absolute dose reading in real time as the beam is delivered — enough spatial density to resolve the dose gradients typical of a modulated field, though still coarser than the sub-millimeter resolution of film.

Detector Array Measurement Delivery

The plan is delivered to the phantom-mounted detector array exactly as it will later be delivered to the patient — same monitor units, same gantry motion, same MLC sequence. As the beam sweeps across the array, each sensor independently records the dose it received, building up a measured dose map cell by cell.

  • 700–1500+: Sensors per array (depending on model)
  • Real-time: Readout (during beam-on)
  • Ion chamber / diode: Typical modalities (array technology)
  • Panel dosimetry: Alternative: EPID (transit / non-transit)

How a 2D detector array measures dose

Each element in the array — typically a small-volume ionization chamber or a silicon diode — independently converts the radiation it absorbs into an electrical signal proportional to dose. The array electronics digitize and calibrate every channel against a cross-calibration factor, correcting for individual detector sensitivity differences, so that all sensors report dose on a common absolute scale.

Because the beam is delivered as a full dynamic sequence (gantry rotating, MLC leaves moving, dose rate modulating for VMAT, or as a series of static IMRT segments), the array integrates signal continuously through the whole delivery, producing a single cumulative 2D measured dose map that captures the net effect of everything the machine actually did — not just what it was commanded to do.

Because ionization chambers and diodes have finite physical volume, they inherently average dose over their sensitive area — a small but real source of spatial resolution loss compared to the calculation grid, and one reason detector spacing (5–10 mm) sets a practical floor on how fine a QA measurement can resolve.

EPID-based dosimetry — an alternative measurement path

Many modern linacs are equipped with an electronic portal imaging device (EPID) — a flat-panel amorphous-silicon detector originally built for patient positioning imaging, now widely repurposed for dosimetry. EPID panels offer far higher spatial resolution (sub-millimeter pixel pitch) than chamber arrays and can be used in two configurations:

• Non-transit dosimetry: the EPID measures the beam directly (no patient/phantom in the path), analogous to a very high-resolution 2D array, used for pure machine-delivery verification. • Transit dosimetry: the EPID measures the beam after it has passed through the patient during actual treatment, enabling true in-vivo, per-fraction verification rather than a pre-treatment phantom measurement.

EPID dosimetry has become central to the move toward more frequent, lower-burden QA, since the imaging panel is already present on every linac and requires no separate phantom setup.

Sources of delivery error captured at this step

This is the step where real machine behavior — as opposed to the idealized commanded plan — is captured. Errors that only manifest during actual delivery include:

• MLC leaf positioning lag or miscalibration during dynamic motion • Dose-rate servo response lag during VMAT gantry-speed/dose-rate modulation • Gantry sag and collimator angle inaccuracies affecting beam geometry • Output (monitor-unit calibration) drift of the linac

None of these are visible in the TPS dose calculation, which assumes perfect, instantaneous execution of the plan. The measured map is therefore the first place these delivery-specific errors become visible as a physical dose distribution, ready to be compared against the calculated one.

Dose Distribution Comparison

With both the planned (TPS) and measured (detector array) dose maps in hand, the two are registered to the same coordinate grid and compared point by point. A simple dose-difference map already reveals where the delivery diverges from the plan — but on its own, dose difference is a poor and overly punishing metric near steep dose gradients.

  • |ΔDose|: Comparison metric (naive) (point-by-point subtraction)
  • Steep gradients: Problem region (penumbra, field edges)
  • 2–3%: Typical dose-diff tolerance (of prescription / max dose)
  • Spatial tolerance: Why gamma is needed (accounts for small shifts)

The limits of pure dose-difference comparison

A naive comparison simply subtracts measured dose from planned dose at every point. In flat, low-gradient regions this works well: a genuine dosimetric error shows up cleanly as a percentage difference. But in high-gradient regions — beam penumbras, field edges, the boundary of a steep IMRT segment — even a sub-millimeter spatial misalignment between the measured and planned grids produces a huge apparent dose difference, even though the delivery was essentially perfect.

This creates a fundamental tension: a tolerance loose enough to avoid false failures at field edges is too loose to catch real errors in flat regions, and a tolerance tight enough for flat regions produces enormous numbers of spurious failures at every gradient. Dose-difference alone cannot distinguish "wrong dose" from "right dose, slightly misplaced."

Building the difference map

Despite its limitations, the raw difference map remains a useful diagnostic first step. Cells exceeding the dose-difference criterion are flagged for closer inspection, giving a quick visual read on where discrepancies cluster — a random scatter of small differences usually points to detector noise or calibration uncertainty, while a spatially clustered patch (as in the highlighted region here) is far more consistent with a systematic delivery error, such as an MLC leaf bank that consistently under- or over-travels in one part of the field.

This clustering pattern is itself diagnostic: physicists reviewing a failed QA report look first at whether failures are randomly scattered (often measurement noise) or geometrically clustered (often a real, correctable machine or calculation error).

Why a combined criterion is required

The solution, formalized by Low and colleagues in 1998, is to evaluate agreement in a two-dimensional space that combines both a dose-difference tolerance and a distance-to-agreement (DTA) tolerance into a single figure of merit. A measured point is allowed to "search" nearby planned points for the best match — if a planned point close by (within the DTA) has a similar dose (within the dose-difference criterion), the point passes, even if the exact same physical location shows a larger raw dose difference.

This combined approach — the gamma index — is what stage 4 computes, transforming the raw difference map into a single pass/fail decision per point that correctly tolerates small spatial shifts in high-gradient regions while still catching genuine dosimetric errors.

Gamma Index Calculation

The gamma index, introduced by Daniel Low, William Harms, Sasa Mutic, and James Purdy in a landmark 1998 Medical Physics paper, combines a dose-difference criterion and a distance-to-agreement criterion into a single unitless number per point. A gamma value of 1 or less means the point passes; a value above 1 means it fails both tolerances simultaneously.

  • Low et al., 1998: Origin (Medical Physics 25(5))
  • 3% / 3 mm: Standard criteria (global) (conventional IMRT / VMAT)
  • 2% / 2 mm: Stricter criteria (SBRT / SRS, per TG-218)
  • γ ≤ 1: Pass condition (fail if γ > 1)

The gamma index formula

For each measured point r_m, the gamma index searches every point r_p in the planned (reference) distribution and finds the minimum combined distance in a normalized dose–distance space:

γ(r_m) = min over r_p of √[ (Δr(r_m,r_p) / DTA_crit)² + (ΔD(r_m,r_p) / ΔD_crit)² ]

where Δr is the physical distance between the two points, ΔD is their dose difference, DTA_crit is the distance-to-agreement criterion (e.g. 3 mm), and ΔD_crit is the dose-difference criterion (e.g. 3% of a normalization dose, often the global maximum or prescription dose).

Geometrically, this defines an ellipsoid around every planned point in dose-distance space. If any planned point falls inside the ellipsoid centered on the measured point, γ ≤ 1 and the point passes. If the nearest acceptable match lies outside every such ellipsoid, γ > 1 and the point fails — meaning no nearby planned point matches its dose closely enough at any tolerable distance.

The genius of the gamma index is that it treats dose error and spatial error as interchangeable within the tolerance ellipsoid: a point can pass either because its dose matches closely at its exact location, or because a very similar dose exists just a short distance away — exactly the behavior needed near steep dose gradients.

Standard clinical criteria and normalization

The 3%/3% dose-difference and 3 mm DTA combination (often written "3%/3mm") has been the long-standing default for conventional fractionated IMRT and VMAT plans, typically applied with global normalization (the dose-difference percentage is relative to a single fixed reference dose, usually the maximum or prescription dose, rather than the local dose at each point) and a low-dose threshold (e.g. 10% of maximum) excluding low-dose background noise from the analysis.

For stereotactic body radiotherapy (SBRT) and stereotactic radiosurgery (SRS) — hypofractionated treatments delivering very high dose per fraction to small targets with steep dose falloff and minimal margin for error — AAPM Task Group 218 (TG-218, Miften et al., 2018) recommends tighter criteria, commonly 2%/2mm, sometimes evaluated with local rather than global normalization, since a small spatial error has a much larger clinical consequence when the target and surrounding critical structures are only millimeters apart.

From per-point gamma to a pass rate

Computing γ for every point in the measured distribution — a computationally intensive nearest-neighbor search performed here as an expanding search radius around each point — produces a full gamma map. Each point is classified pass (γ ≤ 1) or fail (γ > 1), and the overall result is summarized as the gamma pass rate: the percentage of evaluated points with γ ≤ 1.

This single number becomes the basis for the clinical accept/reject decision. It intentionally discards spatial detail about where failures occur when reduced to one percentage, which is why physicists always also inspect the spatial gamma map itself — a low pass rate concentrated in one region tells a very different clinical story than the same pass rate scattered randomly across the field.

Pass Rate Report & Clinical Action

The final gamma pass rate is compared against a pre-defined clinical action threshold. AAPM Task Group 218 formalized these thresholds so that institutions apply a consistent, evidence-based standard rather than an arbitrary in-house number — turning a statistical summary into a binding go/no-go decision for patient treatment.

  • ≥ 95%: TG-218 tolerance limit (3%/3mm, global, 10% threshold)
  • ≥ 90%: TG-218 action limit (below this, treatment halted)
  • Tighter limits: High-risk plans (SRS/SBRT) (2%/2mm, higher pass rate)
  • Growing: Virtual QA adoption (log-file-based verification)

AAPM TG-218 action limits

AAPM Task Group 218 (2018) recommends standardized gamma analysis parameters — 3%/3mm dose-difference/DTA criteria, global normalization to the maximum dose, and a 10% low-dose threshold — together with two pass-rate benchmarks: a 95% "tolerance limit" (results below this should prompt investigation) and a 90% "action limit" (results below this should not proceed to treatment without resolving the discrepancy). For higher-risk, highly conformal techniques such as SRS and SBRT, TG-218 recommends tighter analysis criteria (e.g. 2%/2mm) and correspondingly higher pass-rate expectations, reflecting the reduced margin for error around small, high-dose targets close to critical structures.

These thresholds represent a balance: criteria tight enough to catch clinically meaningful errors, loose enough that ordinary measurement and calculation uncertainty does not trigger false, resource-wasting failures on every plan.

A "PASS" is not proof of a perfect delivery — it certifies that any discrepancies are statistically consistent with expected measurement and calculation uncertainty. A "FAIL" does not always mean the patient plan is unsafe — it means the discrepancy must be investigated before treatment proceeds.

Common causes of QA failure

When a plan fails its gamma criteria, medical physicists investigate several recurring root causes:

• MLC positioning errors: individual leaves that consistently lag, overshoot, or miscalibrate their commanded position, especially during fast dynamic motion in VMAT • Dose calculation algorithm limitations: TPS models that under- or over-estimate scatter, tissue heterogeneity effects, or small-field output factors, especially near interfaces and small apertures • Detector setup error: array tilt, incorrect SSD, or phantom misalignment introducing spurious geometric discrepancy unrelated to the actual delivery • Linac output drift: monitor-unit calibration drift since the last routine machine QA • Gantry sag or collimator angle inaccuracy affecting beam geometry at certain gantry angles

Distinguishing between these requires more than the pass-rate number alone — the spatial pattern of failures (localized vs. scattered, gantry-angle-dependent vs. constant) is usually the key diagnostic clue.

The shift toward log-file-based "virtual" QA

Physical phantom measurement is time-consuming, occupies linac treatment time, and only verifies the plan once, before the first fraction. Modern linacs record detailed trajectory or log files — actual MLC leaf positions, gantry angle, dose rate, and jaw positions at every control point during every single delivered fraction, not just the pre-treatment measurement.

Software systems can reconstruct the actual delivered fluence from these log files and recompute (or directly compare) the resulting dose distribution against the plan, without needing a phantom or detector array at all — enabling "virtual" or log-file-based QA on every fraction of every patient's treatment, rather than a single pre-treatment measurement. This dramatically reduces the measurement burden on QA staff and machine time, though it is complementary rather than a full replacement for physical measurement: because log-file QA typically reuses the same TPS dose-calculation engine (rather than an independent measurement), it is well-suited to catching MLC and delivery errors, but less able to independently catch systematic dose-calculation algorithm errors that a physical detector would reveal.

⚙ Under the hood

This simulation performs gamma analysis for quality assurance in radiotherapy plans, comparing the planned dose distribution with the delivered dose to identify discrepancies and ensure treatment accuracy.

CanvasBiomedicine

2D · HTML5 Canvas 2D · 60 FPS target · runs fully client-side, no install

What did you find?

Add reproduction steps (optional)