Copying one strand, not both
Transcription is the first step of gene expression: the cell copies a gene's DNA sequence into a single-stranded messenger RNA (mRNA) molecule that can leave the nucleus and be translated into protein. The enzyme RNA polymerase binds at a promoter sequence just upstream of the gene, locally unwinds the DNA double helix into a small 'transcription bubble,' and reads only one of the two strands — the template strand — while leaving the other, the coding/sense strand, displaced and unused for that gene.
RNA polymerase moves along the template strand 3' → 5' (reading it in that direction), synthesising the new RNA strand 5' → 3' by adding one ribonucleotide at a time, each one chosen because it base-pairs with the template: template A pairs with RNA U, template T pairs with RNA A, template G pairs with RNA C, and template C pairs with RNA G. The result is an mRNA strand that is complementary to the template and — critically — the same sequence as the coding strand, except with U substituted everywhere the coding strand has T, since RNA uses uracil in place of thymine.
template (3'→5'): ...T A C G G C A T T... mRNA (5'→3'): ...A U G C C G U A A... (RNA pol reads template, builds complementary RNA) coding (5'→3'): ...A T G C C G T A A... (same sequence as mRNA, T instead of U)
Three phases: initiation, elongation, termination
Initiation begins when RNA polymerase, often assisted by transcription factors that recognise specific promoter sequences (in eukaryotes, commonly a TATA box roughly 25–35 base pairs upstream of the start site), binds the promoter and unwinds the DNA to expose the template strand. Elongation is the repetitive core of the process — polymerase travels along the gene, continuously unwinding DNA ahead of it and rewinding the double helix behind it, extending the growing RNA strand one nucleotide at a time at a rate of roughly tens of nucleotides per second in eukaryotes. Termination ends the process when polymerase reaches a specific terminator sequence, or in eukaryotes a polyadenylation signal, causing polymerase to release the completed RNA transcript and detach from the DNA.
Codons: how the sequence will later specify amino acids
The mRNA sequence is read, during the later process of translation, in non-overlapping three-base groups called codons, each of which specifies one amino acid (or a stop signal) according to the genetic code — a mapping that is essentially universal across all known life, from bacteria to humans, strong evidence for a shared evolutionary origin. Because there are 4³ = 64 possible codons but only 20 standard amino acids, the code is degenerate (redundant): most amino acids are specified by more than one codon, which provides some tolerance against point mutations, particularly at the third ('wobble') position of a codon, where many substitutions still yield the same amino acid.
Eukaryotic processing: the transcript isn't finished yet
In eukaryotes, the raw transcript — pre-mRNA — is not yet ready for translation and undergoes further processing before it leaves the nucleus. A protective 5' cap (a modified guanine nucleotide) is added to the 5' end almost as soon as transcription begins, and a poly-A tail (up to a few hundred adenine nucleotides) is added to the 3' end after termination; both modifications protect the transcript from degradation and assist later steps including nuclear export and translation initiation. Between them, splicing removes non-coding introns and joins the coding exons together — and because a single pre-mRNA can be spliced in more than one way (alternative splicing), a single gene can produce multiple distinct protein products depending on which exons are retained, a major source of protein diversity beyond what the raw gene count alone would suggest. Bacteria, lacking a nucleus and largely lacking introns in most genes, mostly skip this processing step entirely — their mRNA can begin being translated by ribosomes even while transcription is still in progress.
Frequently asked questions
Why does RNA polymerase only copy one of the two DNA strands?
Because only the template strand is read to build a complementary, meaningful mRNA; copying both strands of the same gene simultaneously would produce two different, generally nonsensical RNA sequences. The unused strand, called the coding or sense strand, happens to have the same sequence as the resulting mRNA except with T in place of U.
What is the difference between a codon and a base pair?
A base pair is a single matched pair of nucleotides (like A-T or G-C) between two complementary strands. A codon is a group of three consecutive bases on the mRNA, read together during translation to specify one amino acid or a stop signal — the genetic code is built from codons, not from individual base pairs.
Why does the same gene sometimes produce different proteins?
Because eukaryotic pre-mRNA is processed by splicing, which removes introns and joins exons together, and a single pre-mRNA can be spliced in different ways to include or exclude particular exons. This alternative splicing lets one gene template multiple distinct final mRNAs, and therefore multiple distinct proteins.
Try it live
Everything above runs in your browser — open DNA Transcription — RNA Polymerase Animation and change the parameters while it is running. Nothing is installed, nothing is uploaded, the whole model lives in one tab.
▶ Open DNA Transcription — RNA Polymerase Animation simulation