How sound becomes an electrical stimulation pattern in the cochlea — from microphone to auditory nerve
A cochlear implant (CI) begins its work entirely outside the body. A small behind-the-ear (or off-the-ear) sound processor houses one or more directional microphones that continuously sample the acoustic environment. For the ~466 million people worldwide with disabling hearing loss, this is the first link in a chain that restores access to sound when hearing aids — which merely amplify — are no longer enough.
Cochlear implants are indicated for moderate-to-profound sensorineural hearing loss when hearing aids no longer provide adequate benefit. Candidacy criteria have broadened considerably since the first FDA approval in 1984:
• Audiometric threshold: typically pure-tone average worse than 60–70 dB HL in the implanted ear, with limited open-set sentence recognition (often <50–60%) using best-fit hearing aids • Age: FDA-approved from 9–12 months of age in the US for congenital deafness (early implantation is critical during the sensitive period for auditory pathway development); no upper age limit • Etiology: congenital (genetic, e.g., GJB2/connexin-26 mutations), ototoxic drug exposure, meningitis, presbycusis (age-related), noise trauma, autoimmune inner ear disease • Auditory nerve integrity: confirmed by imaging (CT/MRI) and, where feasible, promontory stimulation testing — the nerve must be present and able to conduct a signal
Unlike hearing aids, which amplify sound for a still-damaged cochlea, a CI bypasses the damaged sensory hair cells entirely and stimulates the auditory nerve directly with electrical current.
Roughly 90–95% of profound hearing loss originates in the cochlea itself — damaged or absent inner hair cells — rather than the auditory nerve. This is precisely why electrical stimulation downstream of the hair cells works: the nerve is usually still there, waiting for a signal.
The externally worn component contains every part responsible for capturing and digitizing sound before any coding takes place:
• Microphones: modern processors use 2–3 omnidirectional or directional MEMS microphones, enabling adaptive beamforming that emphasizes speech from the front while suppressing background noise from the sides and rear • Analog-to-digital converter (ADC): converts the continuous acoustic waveform into a discrete digital stream, typically 16–24 bit resolution at 16–48 kHz sampling rate — well above the ~8 kHz upper limit of useful speech information • Digital signal processor (DSP) chip: a low-power embedded processor running the sound-coding strategy in real time, with a total processing latency usually kept under ~10 ms to preserve lip-sync and natural interaction • Battery: disposable zinc-air or rechargeable lithium cells, sized for 1–2 days of continuous use • Telecoil / Bluetooth streaming: most modern processors also stream audio directly from phones, TVs, and remote microphones, bypassing the acoustic microphone entirely for improved signal-to-noise ratio
Every later stage of the implant — filtering, coding, transmission, and electrical stimulation — can only work with the information present in this first captured signal. Two design choices dominate real-world performance:
• Dynamic range compression: everyday sound spans well over 100 dB SPL (a whisper to a jackhammer), but electrically evoked hearing has a narrow electrical dynamic range of only ~6–20 dB between the threshold (T-level) and uncomfortably loud (C-level) currents. The processor must compress the huge acoustic range into this narrow electrical window using a logarithmic instantaneous-input/output curve, tuned per patient and per electrode • Noise and beamforming: because electric hearing provides coarser spectral detail than normal hearing, background noise disproportionately degrades speech understanding for CI users compared to normal-hearing listeners — making front-facing adaptive beamforming and remote-microphone accessories disproportionately valuable in classrooms and restaurants
Once digitized, the sound signal is decomposed into its frequency content by a bank of bandpass filters — a digital analogue of the mechanical frequency analysis that the biological basilar membrane performs. The number of filters, their spacing, and the coding strategy chosen to convert filter outputs into stimulation pulses are the single biggest determinants of how much spectral detail a CI user can perceive.
The digital audio stream is passed through a bank of bandpass filters (implemented as an FFT or a set of digital IIR/FIR filters), each tuned to a contiguous slice of the audible spectrum roughly spanning 188 Hz to about 8 kHz — the range carrying most speech information. Each filter output is then rectified and low-pass filtered to extract its instantaneous amplitude envelope, discarding the rapid fine-structure oscillation and keeping only the slow envelope that varies with syllables and phonemes.
The number of analysis channels (typically 12–22, matched to the number of physical electrodes) determines spectral resolution — how finely the incoming sound is sliced by frequency before being mapped to specific electrode contacts.
Continuous Interleaved Sampling (CIS, 1991): stimulates every electrode channel in rapid, non-overlapping sequence at a fixed high rate (800–1800 pps per channel), delivering the envelope amplitude from every filter on every cycle. Interleaving (never firing two adjacent electrodes simultaneously) minimizes current-field summation ("channel interaction") between neighboring contacts.
Advanced Combination Encoder (ACE, Cochlear Ltd.): a hybrid "n-of-m" strategy — out of m available spectral channels (e.g., 22), only the n channels with the highest instantaneous energy (e.g., 8–12) are selected for stimulation on each cycle. This spectral-peak-picking approach concentrates stimulation on the perceptually dominant frequencies of speech and allows a higher effective stimulation rate per selected electrode.
SPEAK (Spectral Peak, predecessor to ACE): similar n-of-m peak-picking philosophy but at a lower, fixed overall rate (~250 pps), historically the first widely used peak-picking strategy.
Fine Structure Processing (FSP) and MP3000: newer strategies that also encode timing information from the fine structure of low-frequency channels (below ~1 kHz), aiming to improve pitch and music perception beyond what envelope-only coding provides.
The "n-of-m" principle behind ACE mirrors a form of automatic sparse coding: rather than trying to stimulate all 22 electrodes on every cycle (which would blur together via current spread), the processor stimulates only the handful of electrodes carrying the strongest, most informative energy — closer to how the ear naturally emphasizes spectral peaks like vowel formants.
Once a channel is selected for stimulation, its envelope amplitude is mapped, per-electrode, through a patient-specific compression function bounded by two calibrated current levels:
• T-level (threshold level): minimum current that is just audible • C/M-level (comfort/maximum level): loudest current that remains comfortable
The acoustic dynamic range (often >100 dB) is logarithmically compressed into this narrow electrical window (roughly 6–20 dB wide), then converted into a train of brief symmetric biphasic current pulses (typically 25–100 microseconds per phase) whose amplitude codes loudness and whose electrode location codes pitch.
The coded stimulation pattern must cross an unbroken barrier: intact skin and the skull. Rather than a percutaneous plug (used in early experimental devices and now largely abandoned due to infection risk), modern cochlear implants use inductive radiofrequency (RF) coupling — two coils held in alignment by matching magnets, transmitting both power and digital data transcutaneously with no wires crossing the skin.
A cochlear implant system is split cleanly into two halves that never physically touch:
External (removable, worn like a hearing aid): • Microphone(s) and sound processor (DSP chip + battery) • Transmitter coil, housed in a small disc worn on the scalp, held in place by a magnet that pairs with a matching magnet in the implant beneath the skin
Internal (surgically implanted, permanent): • Receiver-stimulator: a hermetically sealed titanium/ceramic package placed in a shallow bed drilled into the skull bone (mastoidectomy + facial recess approach), containing the receiving antenna, decoding electronics, and an internal magnet • Electrode array: a flexible silicone lead carrying 12–22 platinum electrode contacts, threaded through a cochleostomy or the round window into the scala tympani of the cochlea
Because the internal component has no battery and no moving parts, it is designed to function for decades — most recipients never need internal hardware replaced.
The external and internal coils form a loosely coupled transformer across the skin, typically operating with an RF carrier in the low-MHz range:
• Power transfer: the external coil generates an oscillating magnetic field that induces current in the internal coil, providing all the electrical power needed to run the receiver-stimulator chip and deliver stimulation current — the implant itself carries no battery • Data transfer: the digital stimulation pattern (which electrode, what current amplitude, what pulse timing) is encoded onto the same RF carrier, commonly using amplitude-shift keying (ASK) or similar modulation, and demodulated by the internal chip • Alignment: a pair of matching magnets (one in the external coil housing, one in the internal receiver) keeps the two coils centered over each other through the scalp; magnet strength is chosen to balance secure retention against skin pressure/discomfort, and can often be adjusted or swapped for MRI compatibility
Because the link only needs to cross a few millimeters of skin and subcutaneous tissue, coupling efficiency is high and transmission latency is negligible (well under a millisecond) — imperceptible compared to the processing delay already introduced by sound coding.
Because there is no percutaneous connector, the skin remains completely intact — dramatically reducing the infection risk that plagued early experimental "plug and socket" implants of the 1970s and making the modern transcutaneous design safe for lifelong, 24/7 use, including swimming (with appropriately rated processors).
Because the internal package contains a magnet and electronics, MRI scanning historically required surgical magnet removal. Modern devices are designed with MRI-conditional ratings (commonly 1.5T, and increasingly 3T with a head wrap or magnet designed to rotate in the field rather than dislodge), allowing most recipients to undergo MRI scans without surgery — an important consideration since many CI recipients will need MRI imaging for unrelated medical reasons over a lifetime.
Deep inside the temporal bone, the cochlea is coiled into roughly two and three-quarter turns, spanning about 35 mm from its wide basal end to its narrow apical tip. This shape is not incidental — it is a frequency map. The basilar membrane is stiff and narrow at the base (tuned to high frequencies) and wide and floppy at the apex (tuned to low frequencies), a property called tonotopic organization. The electrode array is threaded into this spiral specifically to exploit that map electrically.
In the healthy cochlea, a traveling wave set up by incoming sound peaks at a location along the basilar membrane determined by frequency: high frequencies (up to ~20 kHz) peak near the stiff basal end close to the round window, while low frequencies (down to ~20 Hz) peak near the floppy apical tip. Inner hair cells at each location translate mechanical displacement into a neural signal that the brain interprets as a specific pitch, purely based on which nerve fibers fired — a spatial code called tonotopy.
A cochlear implant recreates this place code electrically. Each electrode contact on the array sits at a different position along the cochlear spiral, corresponding to a different frequency in the tonotopic map. The frequency-to-electrode assignment in the processor's map is deliberately designed so that a filter analyzing high-frequency content drives an electrode near the base, and a filter analyzing low-frequency content drives an electrode further toward the apex — mimicking, as closely as an array of only a dozen-odd contacts can, the frequency gradient nature built with thousands of inner hair cells.
A normal cochlea has roughly 3,500 inner hair cells providing near-continuous tonotopic resolution. A cochlear implant delivers the same frequency range through only 12–22 physical electrode contacts — meaning electric hearing is inherently a coarser, lower-resolution version of the biological place code, which is the central reason CI users report speech as understandable but music and tonal-language pitch as more difficult.
How far the array is threaded into the cochlea, and how it is shaped, both affect how well electrical tonotopy matches biological tonotopy:
• Insertion depth: most modern arrays reach 17–31 mm into the ~35 mm cochlear duct, covering roughly one and a half to two turns. Deeper insertion accesses more apical, low-frequency territory but carries higher risk of trauma to delicate intracochlear structures • Lateral wall arrays: thin, flexible, hug the outer wall of the scala tympani, generally atraumatic and preserve residual low-frequency acoustic hearing in hybrid/electro-acoustic candidates • Perimodiolar arrays: pre-curved to hug the inner wall (modiolus) where the spiral ganglion cell bodies are concentrated, sitting closer to the neural target and typically requiring lower stimulation current — at the cost of a more complex insertion • Channel interaction: because the cochlea is filled with electrically conductive fluid (perilymph), current injected at one electrode spreads and can also depolarize neurons near neighboring contacts, blurring the intended spatial (place) resolution. This is the primary reason interleaved, non-simultaneous stimulation (as in CIS/ACE) is used — it prevents current fields from two active electrodes summing at once, and why perimodiolar placement (closer to the nerve, less current needed) can sharpen effective spatial selectivity
Modern arrays carry 12–22 electrode contacts, yet perceptual studies consistently show that CI listeners can typically discriminate only about 8–10 truly independent spectral channels, regardless of how many physical electrodes are present — a phenomenon attributed almost entirely to channel interaction and current spread through cochlear fluid. This is why "n-of-m" strategies like ACE, which select only the strongest handful of channels each cycle rather than driving all electrodes simultaneously, can match or outperform strategies that use every available electrode: reducing simultaneous current spread often matters more than raw electrode count.
The final, decisive step happens in biology, not electronics: electrical current pulses delivered by each electrode must depolarize nearby spiral ganglion cell bodies and their peripheral processes, triggering action potentials that travel along the auditory nerve to the cochlear nucleus, brainstem, and ultimately the auditory cortex, where the brain learns — often over months of rehabilitation — to interpret this novel electric code as meaningful sound.
Each biphasic current pulse delivered by an active electrode creates a local electric field in the perilymph and surrounding tissue. Where that field is strong enough to depolarize the membrane of a spiral ganglion cell or its peripheral dendrite past threshold, an action potential is triggered and propagates centrally along the auditory nerve (cranial nerve VIII) to the cochlear nucleus in the brainstem — the first way-station of the central auditory pathway, and from there onward through the superior olivary complex, inferior colliculus, medial geniculate body of the thalamus, and finally the auditory cortex.
Unlike acoustic hearing, where a single tone recruits a graded, overlapping population of hair cells and fibers, electrical stimulation tends to recruit large, relatively synchronous, and somewhat unnatural volleys of nerve fibers near each electrode — one reason electrically evoked sound has a characteristically different, more "buzzy" or mechanical quality, especially in the first months after activation, before the brain adapts.
Two parallel neural codes combine to convey both pitch and loudness:
• Place code: which electrode (and therefore which tonotopic location) is active, conveying spectral/pitch information as described in the previous stage • Temporal/rate code: how rapidly pulses are delivered on a given electrode (stimulation rate, typically several hundred to a few thousand pulses per second) and how their amplitude varies, conveying loudness and some timing/periodicity cues important for voice pitch and speech rhythm
Outcomes vary substantially between recipients, and the single largest biological factor is spiral ganglion cell survival — the population of neurons an electrode array must actually excite. Cell counts decline with duration of deafness, cause of hearing loss (e.g., meningitis causes profound loss and sometimes ossification of the cochlea), and age. Recipients implanted soon after onset of severe hearing loss, and children implanted early during the sensitive period of auditory pathway plasticity (ideally before ~2–3 years of age for congenital deafness), tend to achieve the best outcomes, since the developing brain is still highly capable of learning to interpret the electric code.
The brain's plasticity is doing at least as much work as the electronics. Post-implantation, the auditory cortex must relearn to map an unfamiliar, coarser electric code onto meaningful phonemes and words — a process that typically continues improving for 6–12 months (and longer in prelingually deaf children), which is why structured aural rehabilitation and auditory training are considered essential parts of implant care, not optional extras.
With modern devices and strategies, the average post-lingually deafened adult CI recipient achieves 70–80%+ open-set sentence recognition in quiet within the first year — a dramatic gain from typically <50% pre-operative performance with hearing aids. Performance in noisy environments remains more challenging, since the coarser spectral resolution of electric hearing provides fewer cues to segregate a target voice from competing background sound.
Many recipients have residual hearing in the non-implanted ear, or in the low frequencies of the implanted ear itself (partial/hybrid implantation preserving residual acoustic hearing). Combining electric hearing from the CI with acoustic hearing from a conventional hearing aid in the other ear — bimodal hearing — reliably improves sound localization, speech-in-noise performance, and music/pitch perception compared to a CI alone, because the acoustic ear supplies fine-grained temporal and pitch information the electric ear cannot. For similar reasons, many centers now offer bilateral cochlear implants (one in each ear) to restore some binaural cues directly in the electric domain.