Mapping a ligand's binding footprint via residue-by-residue shifts in a protein's 2D NMR spectrum
Two-dimensional 1H-15N Heteronuclear Single Quantum Coherence (HSQC) spectroscopy correlates the chemical shift of each backbone amide proton with that of its attached nitrogen. Because every non-proline residue contributes exactly one cross-peak, the resulting spectrum is a fingerprint of the folded state — a map with as many dots as there are residues, each one a potential reporter of local structural change.
The HSQC pulse sequence transfers magnetization from amide protons to the directly bonded backbone 15N, then back to protons for detection. The result is a 2D plane in which each peak's coordinates are set by the local electronic environment of one specific N-H pair. Because the amide nitrogen and proton chemical shifts are exquisitely sensitive to hydrogen bonding, ring-current effects from nearby aromatics, and backbone torsion angles, the spectrum is essentially a barcode of the folded structure: two proteins with different folds — or the same protein in two different conformational or binding states — produce visibly different peak patterns.
For a well-behaved, well-folded domain of 100–200 residues, a high quality HSQC can be recorded in 10–20 minutes on a cryoprobe-equipped spectrometer, making it fast enough to repeat at every point of a titration series without prohibitive instrument time.
A single HSQC spectrum reports simultaneously on essentially every residue in the protein in one experiment — no other structural technique gives a residue-resolved readout of the entire chain this quickly.
Natural-abundance 14N is NMR-silent for practical purposes (low gyromagnetic ratio, low natural abundance of the NMR-active 15N isotope is only 0.37%). To perform HSQC experiments, the protein must be recombinantly expressed — almost always in E. coli — in minimal (M9) media where the sole nitrogen source is 15NH4Cl, uniformly incorporating the NMR-active 15N isotope into every amide nitrogen.
For larger proteins or when unambiguous residue-specific assignment is needed, cells are additionally grown on 13C-glucose (13C labeling of the carbon backbone for triple-resonance assignment experiments) and sometimes in D2O-based media for perdeuteration, which removes dipolar relaxation pathways from aliphatic protons and dramatically sharpens linewidths — essential for proteins above ~25–30 kDa, especially when combined with TROSY (Transverse Relaxation-Optimized SpectroscopY) detection.
Before CSP mapping can identify which residue moved, every peak in the apo spectrum must first be assigned to a specific residue number. This is done with a suite of triple-resonance experiments (HNCA, HNCACB, CBCA(CO)NH, etc.) that walk sequentially along the backbone using scalar couplings between 1H, 15N, 13Cα and 13Cβ nuclei. Once this backbone assignment table exists for the apo protein, every peak in every subsequent titration spectrum can be tracked back to its residue of origin, turning a scatter of anonymous dots into a residue-by-residue readout of the protein surface.
CSP mapping is fundamentally a titration experiment. Small, sub-stoichiometric aliquots of ligand (dissolved in DMSO or aqueous buffer) are added stepwise to the 15N-labeled protein sample, and an HSQC is recorded after each addition — typically at molar ratios of 0.1:1, 0.25:1, 0.5:1, 1:1, 2:1 and above, tracing out a full binding isotherm peak by peak.
Because CSP magnitude follows the fraction of protein in the bound state, titration points are chosen to bracket the expected Kd: several sub-stoichiometric points below 1:1 to catch the steepest part of the binding curve, one or two points near 1:1, and a final large excess (often 3–10:1) to approach saturation. Ligand is typically dissolved in DMSO at high stock concentration so that only microliter volumes are added, keeping the total organic co-solvent below ~2–5% to avoid denaturing the protein or perturbing peaks non-specifically.
Not every peak movement reflects specific binding. Careful CSP experiments include a DMSO-only (vehicle) titration control to subtract solvent effects, monitor for pH drift upon ligand addition (protonatable ligands can shift buffer pH and cause global, non-specific shifts), and watch peak intensities for line-broadening from aggregation or non-specific low-affinity "sticking," which produces widespread small perturbations rather than a localized site.
A clean, specific titration shows the large majority of peaks completely stationary while a defined subset move progressively and reproducibly — this pattern, more than any single spectrum, is the signature of genuine site-specific binding.
Fragment-based drug discovery libraries are typically titrated to several millimolar ligand concentration because fragment Kd values of 100 µM–10 mM are common — far weaker than the nanomolar affinities needed for SPR or ITC to give a clean signal.
A major practical advantage of CSP titrations is sample economy: a single 15N-labeled protein sample (300–600 µL, 50–300 µM) can be titrated through an entire series without replenishment, since each HSQC only consumes NMR signal, not material. This makes NMR titration comparatively cheap per data point relative to techniques requiring a fresh sample or chip surface for every measurement, and is one reason NMR remains a first-line screening tool in fragment-based drug discovery campaigns that must triage thousands of candidate fragments.
As ligand is added, a peak's behavior in the spectrum depends on how the exchange rate between free and bound states (kex) compares to the chemical shift difference between those states (Δδ, in angular frequency units). This comparison — not affinity alone — determines whether a peak glides smoothly to a new position or splits into two discrete, intensity-weighted peaks.
When the ligand off-rate (koff) is fast relative to the chemical shift difference between free and bound states, the NMR timescale sees only a population-weighted average of the two environments. As the bound fraction rises across the titration, the observed peak position moves continuously and smoothly along a straight or gently curved trajectory from the apo position to the fully-bound (holo) position. This is the most common and most interpretable regime for weak-to-moderate affinity interactions (roughly Kd in the µM to low mM range), and it is what most fragment-based and small-molecule CSP campaigns are designed to sit within.
Tighter binders with slow koff (often sub-µM to nM affinity) can fall into slow exchange: rather than one peak sliding, two separate peaks are observed simultaneously — one at the apo position, one at the holo position — with their relative intensities tracking the bound and free populations directly. Between these extremes lies intermediate exchange, the least friendly regime for the spectroscopist: peaks can broaden so severely from exchange-induced relaxation that they lose intensity and temporarily vanish into the noise partway through the titration, only to reappear at the new position once exchange has slowed relative to the (by-then larger) shift difference.
The boundary between regimes is not fixed by affinity alone — it depends on the ratio of koff to Δδ (in rad/s), which is why the very same protein-ligand pair can show fast exchange for a residue with a small shift and slow exchange for a residue with a large one.
Only residues whose local chemical environment actually changes upon ligand binding will move — this includes residues in direct contact with the ligand, but also residues undergoing binding-induced conformational or dynamic changes elsewhere in the fold. Residues distant from any structural consequence of binding remain essentially stationary throughout the entire titration, providing an internal negative control within every single spectrum and making the moving subset stand out clearly against a stable background.
To compare shifts across residues fairly, the 1H and 15N shift changes are combined into a single weighted distance metric, then plotted against residue number as a bar chart. A statistical threshold — commonly the mean plus one or two standard deviations of all residue CSPs — separates "significant" from "background" perturbation.
Because 15N chemical shifts span a much wider ppm range than 1H shifts (roughly 25 ppm vs 4 ppm across the amide region), a raw sum of the two would be dominated by nitrogen. The standard combined CSP formula down-weights the nitrogen term:
Δδ_combined = √[(ΔδH)² + (ΔδN / 5)²]
The scaling factor (commonly 5, sometimes 6.5 or a per-residue empirical value derived from the ratio of gyromagnetic ratios) roughly equalizes the two nuclei's contribution to the metric, so that a peak moving mostly in the nitrogen dimension is not automatically over- or under-weighted relative to one moving mostly in the proton dimension.
With a combined CSP value calculated for every assigned residue, the distribution across the whole protein is used to set an objective cutoff — most commonly the mean CSP plus one or two standard deviations, sometimes after removing evident outliers before computing the statistics (to avoid the outliers inflating the very threshold meant to detect them). Residues above this line are called "significantly perturbed"; this operational definition converts a continuous, noisy per-residue measurement into a clean binary map suitable for overlaying on structure.
A well-defined binding site typically produces a tight cluster of 8–15 contiguous or spatially adjacent significant residues — a smoking-gun spatial pattern that is far more convincing than any single residue's CSP value on its own.
Beyond simply flagging significant residues, the CSP value of a single well-resolved reporter residue can be tracked across all titration points and fit to a one-site binding isotherm:
Δδ_obs = Δδ_max × ([P]+[L]+Kd − √(([P]+[L]+Kd)² − 4[P][L])) / (2[P])
Fitting this equation to the observed CSPs versus ligand concentration recovers Kd directly from the NMR titration itself — no separate biophysical assay is required, though results are often cross-validated against ITC or SPR when those techniques are applicable.
The final step overlays every significantly perturbed residue onto the protein's three-dimensional structure (or, when no structure is available, a homology model or even just the topology diagram). A tight, contiguous cluster of surface-exposed perturbed residues typically marks the direct binding pocket; smaller, spatially separated clusters mark allosteric sites where binding at one location propagates a conformational or dynamic signal elsewhere.
When perturbed residues are painted onto the folded structure, genuine direct-contact binding sites reveal themselves as a spatially contiguous patch on the protein surface — residues that are far apart in sequence number but close together in 3D space, consistent with a real pocket rather than a sequence artifact. This spatial clustering test is one of the strongest pieces of evidence that a CSP signal reflects specific, structurally meaningful binding rather than nonspecific surface effects, and it is routinely used as an ambiguous interaction restraint (AIR) to drive protein-ligand or protein-protein docking calculations in programs like HADDOCK.
Because CSP reports on any change in local chemical environment — not just direct steric contact — residues far from the ligand itself can still shift if binding triggers a conformational rearrangement, a change in backbone dynamics, or a shift in a hydrogen-bond network that propagates through the fold. These distal, spatially separated perturbed residues are the fingerprint of allostery: binding at one site altering function or affinity at another. CSP mapping has been used to trace allosteric communication pathways in kinases, PDZ domains, and G-protein coupled receptor fragments that would be essentially invisible to a technique that only reports on direct contacts, such as most crystallographic ligand-soaking experiments.
Because CSP requires no direct contact to report a signal, it is one of very few structural techniques capable of directly detecting allosteric coupling between a ligand-binding event and a distal functional site in a single experiment.
NMR CSP screening is a cornerstone of fragment-based drug discovery (FBDD): fragment libraries (typically 1,000–3,000 compounds, MW <300 Da) are pooled and titrated against a 15N-labeled target, and the resulting CSP patterns both identify which fragments bind and, crucially, where — information essential for structure-guided fragment growing and linking into a larger, higher-affinity lead compound. Because NMR is sensitive to millimolar-affinity binding that techniques such as SPR or crystallographic soaking often miss, it remains a preferred primary or orthogonal screen for exactly the weak, transient interactions that mark a promising fragment starting point.
| Product | Indication | Trial Design | Key Result |
|---|---|---|---|
| NMR Chemical Shift Perturbation | µM–mM Kd; weak & fragment binders | Track 1H/15N HSQC peak movement per residue across a ligand titration | Atomic-resolution site + allostery, in solution, no crystal needed |
| X-ray Crystallography | Any Kd; needs crystallizable complex | Diffraction of ligand-soaked or co-crystallized protein-ligand complex | Highest resolution (<2 Å), definitive atomic binding pose |
| HDX-MS | Any affinity; large or flexible complexes | Measures backbone amide deuterium-uptake rate changes upon binding | No protein size limit, low sample amount, works on membrane proteins |
| Computational Docking | In silico candidate ligands | Scores predicted poses against a rigid or flexible receptor pocket | Virtual screening of millions of compounds before any wet-lab work |