Genome Projects and Bioinformatics
The Human Genome Project
The Human Genome Project (HGP) was an international research programme that ran from 1990 to 2003. Its primary goal was to determine the complete nucleotide sequence of human DNA and to identify all human genes.
Key Findings
- The human genome contains approximately 3.2 billion base pairs
- There are approximately 20,000-25,000 protein-coding genes — far fewer than the ~100,000 originally predicted
- Protein-coding sequences make up only about 1.5% of the genome
- Much of the remaining DNA was initially called "junk DNA" but is now known to include regulatory sequences, introns, repetitive elements, transposons, and sequences with roles that are still being discovered
- Humans share approximately 99.9% of their DNA sequence — individual variation comes from the 0.1% that differs
- Significant portions of the human genome are shared with other organisms (e.g. ~96% with chimpanzees, ~85% with mice), reflecting common ancestry
Methods Used
The HGP used whole-genome shotgun sequencing and hierarchical shotgun sequencing:
1. DNA fragmentation — the genome was broken into smaller, overlapping fragments using restriction enzymes or mechanical shearing
2. Cloning — fragments were inserted into bacterial artificial chromosomes (BACs) and replicated
3. Sequencing — each fragment was sequenced using the Sanger (chain-termination) method, which uses dideoxynucleotides (ddNTPs) that terminate the growing DNA strand at specific bases
4. Assembly — computer algorithms aligned overlapping sequences to reconstruct the full genome
Modern sequencing uses next-generation sequencing (NGS) technologies that are massively parallel, faster, and cheaper — the cost of sequencing a human genome has dropped from ~$3 billion (HGP) to under $1,000 today.
Applications of Genome Sequencing
Pharmacogenomics
Pharmacogenomics uses an individual's genomic information to predict their response to drugs:
- Genetic variants affect how patients metabolise drugs (e.g. variations in cytochrome P450 enzymes)
- Some variants predict adverse drug reactions — testing can identify patients at risk before prescribing
- This enables personalised medicine — tailoring drug choice and dosage to the individual's genotype
- Example: testing for *HLA-B5701** before prescribing the HIV drug abacavir prevents a potentially fatal hypersensitivity reaction
Identifying Disease Genes
Genome data enables identification of genes associated with genetic diseases:
- Genome-wide association studies (GWAS) compare DNA sequences of affected and unaffected individuals to find single nucleotide polymorphisms (SNPs) associated with disease risk
- This has identified risk variants for type 2 diabetes, Alzheimer's disease, heart disease, and many cancers
- Understanding the genetic basis of disease can lead to new drug targets and therapeutic strategies
Comparative Genomics and Evolution
Comparing genomes of different species reveals:
- Evolutionary relationships — the degree of sequence similarity reflects how recently species diverged from a common ancestor
- Conserved genes — genes found across many species (e.g. homeobox/Hox genes controlling body plan development) are likely to be functionally essential
- Gene duplication and divergence as a mechanism for generating new gene functions
Genetic Testing and Screening
- Predictive testing — identifying individuals carrying alleles for late-onset genetic disorders (e.g. BRCA1/2 for breast cancer risk, HTT for Huntington's disease)
- Prenatal screening — non-invasive prenatal testing (NIPT) uses cell-free fetal DNA in maternal blood to screen for chromosomal abnormalities
- Carrier screening — identifying heterozygous carriers of recessive conditions (e.g. cystic fibrosis, sickle cell)
- Forensic identification — DNA profiling using short tandem repeats (STRs) for criminal investigations and paternity testing
Bioinformatics
Bioinformatics is the use of computer science, mathematics, and statistics to store, retrieve, analyse, and interpret biological data, particularly large genomic datasets.
Key Tools and Databases
| Resource | Purpose |
|---|---|
| GenBank / EMBL / DDBJ | Public databases storing DNA and protein sequences from all organisms |
| BLAST (Basic Local Alignment Search Tool) | Compares a query sequence against databases to find similar sequences — used to identify genes, find homologues across species, and predict function |
| UniProt | Database of protein sequences and functional information |
| Ensembl | Genome browser for annotated genomes of vertebrates and other eukaryotes |
| Protein Data Bank (PDB) | 3D structures of proteins determined by X-ray crystallography and cryo-EM |
Sequence Alignment and Homology
Sequence alignment compares two or more DNA or protein sequences to identify regions of similarity:
- High sequence similarity (homology) suggests common evolutionary origin
- Conserved regions often indicate functional importance — mutations in these regions are selected against
- Alignment can reveal mutations that cause disease when compared to a reference sequence
Gene Annotation
Once a genome is sequenced, it must be annotated — identifying where genes are, what they code for, and their regulatory elements:
- Open reading frames (ORFs) — sequences beginning with a start codon (ATG) and ending with a stop codon, of sufficient length to encode a protein
- Expressed sequence tags (ESTs) — short sequences derived from mRNA, used to confirm that an ORF is actually transcribed
- Functional annotation — assigning a function to a gene based on similarity to genes of known function in other organisms
Proteomics
Proteomics is the large-scale study of the complete set of proteins (proteome) expressed by a cell, tissue, or organism at a given time. While the genome is essentially fixed, the proteome varies between:
- Different cell types (a liver cell and a neurone express different proteins)
- Different developmental stages
- Different environmental conditions or disease states
Techniques include mass spectrometry (identifying proteins by their peptide fragments) and 2D gel electrophoresis (separating proteins by charge and mass).
Metabolomics and Systems Biology
- Metabolomics studies the complete set of metabolites in a biological sample
- Systems biology integrates genomic, proteomic, and metabolomic data to model and understand biological systems as a whole — how genes, proteins, and metabolites interact in networks
Ethical, Social, and Legal Issues
| Issue | Details |
|---|---|
| Genetic privacy | Who has access to an individual's genetic data? Risk of discrimination by employers or insurers |
| Informed consent | Participants in genome studies must understand how their data will be used |
| Insurance and employment | The UK Genetic Insurance Moratorium (2019, extended) prevents insurers from requiring predictive genetic test results (with limited exceptions) |
| Genetic determinism | Risk that society overemphasises genetic causes and underestimates environmental/lifestyle factors |
| Ownership | Can genes or sequences be patented? The US Supreme Court (2013) ruled that naturally occurring DNA sequences cannot be patented, but cDNA (artificially synthesised) can |
| Equity | Benefits of genomic medicine should be accessible to all, not just wealthy nations or individuals |
Exam Tips
- AQA expects you to discuss the applications AND ethical implications of genome projects — always include both in extended answers
- When describing sequencing, mention fragmentation, sequencing of fragments, and computational assembly
- Be able to explain why there are fewer genes than proteins (alternative splicing, post-translational modification)
- Bioinformatics questions often ask about comparing sequences — state that high similarity suggests common ancestry and/or shared function
- Know the distinction between genome (DNA), transcriptome (mRNA), and proteome (proteins) — and why the proteome varies between cells even though the genome does not