Transcription and Translation
From Gene to Protein
The central dogma of molecular biology describes the flow of genetic information:
DNA → (transcription) → mRNA → (translation) → Protein
A gene is a sequence of DNA nucleotides that codes for a polypeptide (or a functional RNA molecule). The process of using the information in a gene to synthesise a polypeptide involves two main stages: transcription (DNA → mRNA) and translation (mRNA → polypeptide).
Transcription
Transcription is the synthesis of an mRNA molecule using one strand of DNA as a template. It occurs in the nucleus (in eukaryotes).
Step-by-Step Process
1. Initiation:
- RNA polymerase binds to the promoter region — a specific DNA sequence upstream (before) the gene
- In eukaryotes, transcription factors must first bind to the promoter before RNA polymerase can attach (this is a key point of gene regulation)
- The DNA double helix is unwound and the two strands separate in the region of the gene
2. Elongation:
- RNA polymerase moves along the template strand (also called the antisense strand) in the 3'→5' direction
- It reads the template strand and synthesises a complementary mRNA strand in the 5'→3' direction
- Free RNA nucleotides (containing ribose sugar and uracil instead of thymine) are added by complementary base pairing:
- A on DNA → U on mRNA
- T on DNA → A on mRNA
- G on DNA → C on mRNA
- C on DNA → G on mRNA
- The mRNA sequence is identical to the coding strand (sense strand) of DNA, except with U replacing T
3. Termination:
- RNA polymerase reaches a terminator sequence on the DNA
- The mRNA molecule is released
- The DNA strands re-anneal (reform the double helix)
Post-Transcriptional Modification (Eukaryotes Only)
The initial mRNA produced (called pre-mRNA or primary transcript) undergoes processing before leaving the nucleus:
- 5' capping — a modified guanine nucleotide is added to the 5' end, protecting the mRNA from degradation and aiding ribosome binding
- 3' polyadenylation — a tail of 100-250 adenine nucleotides (poly-A tail) is added to the 3' end, increasing mRNA stability
- Splicing — introns (non-coding sequences) are removed and exons (coding sequences) are joined together by spliceosomes (complexes of snRNPs — small nuclear ribonucleoproteins)
Alternative splicing — different combinations of exons can be joined together, allowing a single gene to code for multiple different polypeptides. This increases the diversity of the proteome beyond the number of genes in the genome. For example, the human genome has ~20,000 genes but produces over 100,000 different proteins.
The mature mRNA then exits the nucleus through nuclear pores and enters the cytoplasm for translation.
The Genetic Code
The sequence of bases on mRNA is read in groups of three called codons. Each codon specifies one amino acid (or a stop signal). The genetic code has the following properties:
| Property | Meaning |
|---|---|
| Triplet | Three bases code for one amino acid |
| Degenerate (redundant) | Most amino acids are coded for by more than one codon (e.g. leucine has 6 codons) |
| Non-overlapping | Each base is part of only one codon |
| Universal | The same codons specify the same amino acids in almost all organisms |
| Start codon | AUG — codes for methionine and signals the start of translation |
| Stop codons | UAA, UAG, UGA — signal the end of translation (no amino acid) |
Translation
Translation is the synthesis of a polypeptide on a ribosome, using the mRNA sequence as a template and transfer RNA (tRNA) molecules to deliver amino acids. It occurs in the cytoplasm (on free ribosomes or rough ER).
Transfer RNA (tRNA)
Each tRNA molecule has:
- A cloverleaf shape (2D) that folds into an L-shape (3D)
- An anticodon — a sequence of three bases complementary to an mRNA codon
- An amino acid attachment site at the 3' end — each tRNA is loaded with its specific amino acid by an aminoacyl-tRNA synthetase enzyme (using ATP)
- There are at least 20 different aminoacyl-tRNA synthetases, one for each amino acid
Ribosomes
Ribosomes are composed of rRNA and protein, with two subunits (large and small). They have three tRNA binding sites:
- A site (aminoacyl) — where the incoming charged tRNA binds
- P site (peptidyl) — where the tRNA carrying the growing polypeptide chain sits
- E site (exit) — where discharged tRNAs leave
Step-by-Step Process
1. Initiation:
- The small ribosomal subunit binds to the 5' end of the mRNA and moves along until it reaches the start codon (AUG)
- The initiator tRNA (carrying methionine, with anticodon UAC) binds to the start codon at the P site
- The large ribosomal subunit joins, completing the ribosome
2. Elongation:
- A second charged tRNA, with an anticodon complementary to the codon at the A site, enters and binds
- A peptide bond is formed between the amino acid at the P site and the amino acid at the A site — this reaction is catalysed by peptidyl transferase (a ribozyme — catalytic rRNA in the large subunit)
- The ribosome moves one codon along the mRNA in the 5'→3' direction (translocation):
- The tRNA at the P site moves to the E site (and is released)
- The tRNA at the A site (now carrying the growing polypeptide) moves to the P site
- A new codon is exposed at the A site for the next tRNA
- This cycle repeats, adding one amino acid per codon
3. Termination:
- A stop codon (UAA, UAG, or UGA) enters the A site
- No tRNA has a complementary anticodon for stop codons
- A release factor protein binds to the stop codon
- The polypeptide is released from the final tRNA
- The ribosome dissociates into its two subunits
Polyribosomes (Polysomes)
Multiple ribosomes can translate the same mRNA molecule simultaneously, each at a different position along the mRNA. This is called a polysome and allows rapid, efficient production of many copies of the same polypeptide.
Post-Translational Modification
The polypeptide chain folds into its secondary, tertiary, and (if applicable) quaternary structure. Additional modifications may include:
- Cleavage — removal of signal sequences or activation of zymogens (e.g. insulin from proinsulin)
- Glycosylation — addition of carbohydrate groups (in the rough ER and Golgi apparatus)
- Phosphorylation — addition of phosphate groups (regulating protein activity)
- Formation of disulfide bonds between cysteine residues
Exam Tips
- AQA expects you to distinguish clearly between transcription (DNA→mRNA, in the nucleus, by RNA polymerase) and translation (mRNA→polypeptide, on ribosomes, in cytoplasm)
- Know which strand of DNA is the template (antisense/3'→5') and which is the coding strand (sense/5'→3') — the mRNA matches the coding strand (with U not T)
- When describing translation, use precise terminology: codon, anticodon, A site, P site, peptide bond, peptidyl transferase
- Splicing and alternative splicing are frequently tested — be able to explain how one gene can produce multiple proteins
- Remember that the code is read in a fixed reading frame starting from AUG — a frameshift mutation changes every codon downstream