From camera pixels to perceived phosphenes — mapping a scene onto an implanted electrode array for blind patients
For patients blinded by outer-retinal degeneration — most commonly retinitis pigmentosa (RP) or advanced age-related macular degeneration (AMD) — the photoreceptor layer (rods and cones) has died, but the inner retina (bipolar cells, ganglion cells, and the optic nerve) often remains functional for years. A retinal prosthesis exploits this by bypassing the dead photoreceptors entirely: a head-mounted camera becomes the new "eye," and its image is electronically routed around the damaged layer straight to surviving neurons.
Retinitis pigmentosa is a group of inherited retinal dystrophies in which rod photoreceptors degenerate first (causing night blindness and progressive peripheral field loss), followed by cone loss and eventual central vision loss — end-stage RP leaves patients with bare light perception or complete blindness. Advanced dry AMD similarly destroys the macular photoreceptor mosaic responsible for central, high-acuity vision.
Critically, both conditions largely spare the inner retinal circuitry — bipolar cells, amacrine cells, and retinal ganglion cells (RGCs) whose axons form the optic nerve — for a substantial period after photoreceptor loss, though some degree of inner-retinal remodeling and cell loss does occur over time and can degrade prosthesis performance in very late-stage disease.
This surviving downstream circuitry is precisely what a retinal prosthesis targets: instead of restoring photoreceptor function, the device electrically stimulates the neurons that photoreceptors would normally have driven, "tricking" the visual system into perceiving light.
Retinal prostheses are only viable when enough inner-retinal neurons survive to be stimulated. Patients are typically screened with electroretinography (ERG) and optical coherence tomography (OCT) to confirm sufficient inner-retinal integrity before implantation is offered.
The visual pipeline begins outside the eye entirely. A miniature CMOS or CCD camera is mounted on a pair of glasses (in first-generation systems like Argus II) or, in newer designs, replaced by a subretinal photodiode array that responds directly to incoming light without an external camera at all (as in PRIMA).
In camera-based systems, the glasses-mounted sensor captures a continuous video stream of the scene in front of the patient — typically at 640×480 pixels or higher, 20–30 frames per second. This raw frame is sent by cable to a small belt-worn or pocket-worn Video Processing Unit (VPU), a wearable computer that performs image enhancement (contrast boosting, edge detection, gain control) before the drastic downsampling required in the next stage.
Because the eventual electrode count is orders of magnitude lower than the camera's native resolution, camera image quality itself is rarely the bottleneck — what matters far more is which few hundred "pixels" of information the processing pipeline chooses to preserve.
Because the camera sits on the glasses rather than inside the eye, image framing is coupled to head movement rather than eye movement — patients must learn to scan a scene by turning their head, since their (now largely non-functional) eye movements no longer redirect the effective "gaze." This is a major rehabilitation and training burden, and is one of the core reasons some newer subretinal designs (PRIMA) instead place a photodiode array directly in the eye so that residual natural eye movements continue to drive scanning, paired with an external camera on smart glasses that projects processed infrared light into the eye.
A modern camera sensor captures hundreds of thousands to millions of pixels per frame. A retinal implant, by contrast, has somewhere between 16 and roughly 1,600 individually addressable electrodes. The video processing unit must compress an enormously rich image down to a tiny brightness matrix — one value per electrode — while trying to preserve the visual information a patient needs most: edges, contrast, and object boundaries.
The core downsampling operation is a spatial average-pooling (or more sophisticated edge-preserving filtering): the camera frame is divided into a coarse grid matching the physical layout of the electrode array — for Argus II, a 6×10 grid corresponding to 60 electrodes. Each grid cell's pixels are averaged (or processed through an edge/contrast-enhancing filter) into a single brightness value, which is then mapped to a stimulation current for the corresponding electrode.
This is analogous to viewing a photograph reduced to a 60-pixel thumbnail — most fine detail, texture, and color information is discarded. What survives is coarse spatial layout: where are the bright and dark regions, where are strong edges (like a doorway, a face outline, or high-contrast text).
Many systems apply real-time image enhancement before downsampling — contrast stretching, edge enhancement, and adaptive gain control — since naively averaging pixels tends to wash out useful boundaries into a uniform gray.
Because so much spatial detail must be discarded, downstream processing choices matter enormously: studies show that patients using edge-enhancement or contour-extraction algorithms outperform patients viewing simple brightness-averaged images at object recognition tasks, even with identical electrode hardware.
Once the coarse brightness grid is computed, each cell value must be converted into stimulation parameters for its electrode: primarily current amplitude, but also pulse width, frequency, and sometimes pulse pattern. Brighter grid cells typically map to higher current amplitude (within safety limits), producing a brighter-appearing phosphene; darker cells receive little or no stimulation.
Modern systems often apply a non-linear brightness-to-current mapping (a "gamma" curve) because perceived phosphene brightness does not scale linearly with injected charge — small increases near an electrode's perceptual threshold produce large brightness changes, while further increases near the safety ceiling produce diminishing returns.
Thresholds vary substantially between electrodes and between patients — some electrodes may require several times more current than neighboring ones to evoke any percept at all, due to variable distance from surviving neurons, scarring, or local tissue impedance. Clinical fitting sessions individually calibrate each electrode's threshold before the device is used functionally.
The full pipeline — camera capture, filtering, downsampling, current mapping, and wireless transmission — must run within roughly 20–50 milliseconds to avoid noticeable lag between head movement and perceived phosphene updates, since a delayed visual stream is disorienting and worsens mobility performance. Video processing units are therefore purpose-built low-power embedded computers rather than general-purpose smartphones, prioritizing deterministic low latency over raw image quality.
No wire can permanently cross the eye wall or scalp without an unacceptable infection risk, so retinal prostheses transmit both the stimulation data and the electrical power needed to drive it wirelessly, using inductive coupling between an external coil (worn on glasses or as a headband/coil unit) and a subdermal or intraocular receiver coil connected to the implanted stimulator chip.
The external unit contains a primary transmitter coil driven by an oscillating current, generating a time-varying magnetic field. A secondary receiver coil, implanted just under the skin near the eye (or, in some designs, within an intraocular device), sits within range of that magnetic field. By Faraday's law of electromagnetic induction, the changing magnetic flux through the receiver coil induces an alternating current in it — power that is then rectified and regulated on-chip to run the implanted electronics.
The same inductive link (or, in many systems, a modulated companion RF channel) also carries the digital stimulation data — the per-electrode current amplitudes computed in Stage 2 — encoded onto the carrier signal, typically via amplitude-shift keying (ASK) or frequency-shift keying (FSK) modulation. The implant demodulates this signal to reconstruct the intended stimulation pattern for the current video frame.
Because this is a fully closed, batteryless, wireless link, the implant contains no internal battery to fail, deplete, or require surgical replacement — power arrives fresh with every video frame while the external unit is worn.
Coil alignment is a real practical constraint: if the external coil drifts more than roughly 1–1.5 cm from the internal coil, or tilts significantly, coupling efficiency drops sharply and stimulation can fail — which is why glasses-mounted or headband-mounted transmitter coils are carefully positioned and secured over the implant site.
On the receiving end sits a hermetically sealed application-specific integrated circuit (ASIC), typically encapsulated in titanium and ceramic to survive decades inside the eye's aqueous, saline environment. This chip performs several functions: rectifying induced AC power into stable DC voltages, demodulating the incoming data stream, decoding it into per-electrode stimulation commands, and driving current-controlled output stages connected to each electrode in the array via a thin, flexible cable (in epiretinal designs) or directly beneath the array (in subretinal designs).
Hermeticity is one of the hardest engineering problems in the field — any moisture ingress into the electronics compartment causes corrosion and device failure, so packaging must maintain a water-tight seal for the patient's entire remaining lifetime, potentially 30+ years.
Transmitted RF power levels are kept well below regulatory specific absorption rate (SAR) limits for tissue heating, and the coil operating frequencies (typically low single-digit to low double-digit MHz) are chosen to balance efficient coupling against tissue absorption and interference with other implanted or external electronics. Because power only flows while the external unit is worn and powered on, the implant is inherently "off" and inert whenever the glasses are removed.
The final electrical step happens directly on retinal tissue: an array of microelectrodes, positioned either on the retina's inner surface (epiretinal), beneath the retina near the degenerated photoreceptor layer (subretinal), or in the suprachoroidal space outside the retina entirely, delivers charge-balanced current pulses that depolarize nearby neurons enough to trigger action potentials — the same electrochemical signal a healthy eye would send toward the brain.
Three anatomical placements dominate current and past clinical devices, each with different tradeoffs:
• Epiretinal (e.g. Argus II, Second Sight): the electrode array is surgically tacked to the inner (vitreous-facing) surface of the retina, directly contacting retinal ganglion cell axons and somas. Stimulation here bypasses essentially all remaining retinal processing and directly drives the retina's output neurons — powerful but less able to exploit any residual retinal signal processing, and axon-of-passage stimulation can cause visual percepts to appear smeared or mislocalized relative to the stimulating electrode.
• Subretinal (e.g. Alpha AMS, PRIMA): the array sits in the space once occupied by degenerated photoreceptors, stimulating bipolar cells and thereby engaging more of the retina's own downstream signal processing (contrast enhancement, temporal filtering) before the signal reaches ganglion cells — potentially yielding more naturalistic percepts, at the cost of a technically harder surgical implantation beneath the retina.
• Suprachoroidal (e.g. some Bionic Vision Australia / Phoenix99 designs): electrodes sit outside the eye wall entirely, in the suprachoroidal space between the choroid and sclera — a substantially safer, shorter surgery, but with electrodes farther from target neurons, requiring higher stimulation currents and yielding lower spatial resolution.
Electrode-to-target distance is one of the biggest determinants of both the current needed and the resolution achievable: epiretinal and subretinal arrays sit within tens to a couple hundred micrometers of their target neurons, while suprachoroidal arrays are millimeters away — explaining why suprachoroidal designs need much higher currents for comparatively coarser vision.
Electrodes deliver current as biphasic pulses — a cathodic (negative) phase followed by an anodic (positive) phase of equal and opposite charge — rather than a simple DC current. This charge balancing is essential: any net charge injection over time drives irreversible electrochemical reactions at the electrode-tissue interface (electrolysis, pH shifts, toxic byproduct generation, and electrode corrosion) that can damage both the tissue and the electrode itself.
Safe stimulation is bounded by the electrode's charge density limit — commonly cited around 0.35 mC/cm² per phase for platinum and platinum-iridium electrodes (the Shannon safety criterion combines charge and charge density limits) — meaning smaller electrodes, which enable higher spatial resolution, can safely deliver less total charge each, fundamentally linking electrode miniaturization to a reduced per-pixel dynamic range.
This is a central engineering tension in the field: shrinking electrodes to pack more of them into an array (improving resolution) reduces the safe current each electrode can deliver, and also increases electrode impedance, both of which push against reliably reaching perceptual thresholds in individual electrodes.
Injected current does not stay perfectly confined to the tissue directly beneath an electrode — it spreads through the conductive retinal and vitreous tissue, activating neurons near, but not exactly at, the target electrode, and can activate axon-of-passage fibers that carry signals from more distant, unrelated retinal locations (particularly problematic in epiretinal designs, where ganglion cell axons run across the retinal surface before converging on the optic nerve). This current spread is a major contributor to the blurred, overlapping, and sometimes spatially inaccurate phosphenes patients report, and is why higher stimulation currents — while more reliably perceived — can paradoxically reduce effective resolution by activating overlapping neural populations between adjacent electrodes.
| Product | Indication | Trial Design | Key Result |
|---|---|---|---|
| Argus II (Second Sight) | Epiretinal, ganglion cells | 60 electrodes, 6×10 array, glasses-mounted camera | First FDA-approved (2013); largest clinical dataset |
| Alpha AMS (Retina Implant AG) | Subretinal, bipolar cells | 1,600 photodiode-electrode pixels, no external camera | Natural eye-movement scanning; higher pixel count |
| PRIMA (Pixium Vision / Science) | Subretinal, photovoltaic array | 378–10,500 pixels, infrared-projected image from smart glasses | Wireless photovoltaic pixels, no percutaneous cable |
| Suprachoroidal arrays (e.g. Phoenix99) | Suprachoroidal space | Electrodes outside sclera/choroid, shorter safer surgery | Lower surgical risk; simpler implantation |
The end product of this entire chain — camera, downsampling, wireless link, and electrode stimulation — is a perceptual experience unlike normal vision: a sparse constellation of small, often flickering points or blobs of light called phosphenes, whose spatial arrangement roughly traces the brightness pattern of the original scene, well enough for many patients to detect motion, locate high-contrast objects like doorways and windows, and follow lines of text in large print, but far short of natural acuity.
Patients across essentially all retinal prosthesis trials describe phosphenes not as clean, uniform pixels but as variable points, streaks, arcs, or diffuse blobs of light — yellow-white, blue-white, or colorless flashes that can flicker, fade over the course of a stimulation train, vary in apparent size and brightness across electrodes even for identical stimulation parameters, and sometimes persist briefly after stimulation ends. This perceptual messiness stems directly from the biological realities covered in Stage 4: variable electrode-to-neuron distance, current spread activating unintended neurons, axon-of-passage activation causing mislocalized percepts, and individual differences in surviving retinal circuitry.
Because of this variability, most patients undergo extended perceptual training and rehabilitation after implantation, learning to interpret their device's particular pattern of phosphenes — head-scanning strategies, associating certain phosphene arrangements with real-world object categories, and integrating flickering, imperfect input into useful behavior, much as the brain's visual cortex must adapt to a fundamentally novel input statistics.
Because phosphene perception is imprecise and does not spatially align in a simple one-to-one grid with the physical electrode layout, researchers now build detailed patient-specific "phosphene models" — mapping each electrode's actual perceived location, size, brightness, and shape from clinical psychophysics testing — and use these models to optimize the image-processing pipeline for each individual patient rather than assuming a uniform pixel grid.
Clinically measured visual acuity with current devices remains far below normal (20/20) vision. Argus II recipients have demonstrated visual acuity around 20/1260 in the best-performing cases using the standard Grating Visual Acuity or Landolt-C tests — well within the legal blindness range (20/200 or worse) but a dramatic functional improvement over the bare light perception or complete blindness patients had before implantation.
PRIMA, with its much higher pixel density (up to 10,500 photovoltaic pixels compared to Argus II's 60 electrodes) and smaller 100 μm pixel pitch, has achieved substantially better letter acuity in trials — around 20/438 to 20/460 in top-performing subjects — approaching the resolution where large-print reading becomes feasible, illustrating the direct link between electrode/pixel density and achievable acuity described in Stage 2.
Critically, acuity metrics only capture part of the functional picture: many patients report the larger benefit is in orientation and mobility tasks — detecting doorways, locating high-contrast objects on a table, following a bright window or a moving person — rather than reading or facial recognition, which typically remain out of reach with current-generation devices.
Ongoing research aims to push electrode counts and effective resolution higher through several converging approaches: smaller, denser photovoltaic pixel arrays (PRIMA-style designs scaling toward tens of thousands of pixels); 3D or penetrating microelectrodes that reach closer to target neurons to reduce current spread; optogenetic approaches that genetically sensitize surviving retinal neurons to light directly, avoiding electrodes altogether; and smarter, patient-personalized image processing algorithms — including deep-learning-based scene simplification — that select which visual information is most worth preserving in a low-bandwidth phosphene display, rather than relying on simple pixel averaging.
Even without further hardware gains, algorithmic advances in Stage 2's downsampling and encoding step have repeatedly been shown, in simulated prosthetic vision studies, to improve object recognition and mobility performance for a fixed electrode count — meaning smarter software, not just more electrodes, is a major lever for improving real-world outcomes.