Genome Projects and Bioinformatics

A-Level Biology · Gene Expression and Regulation

The Human Genome Project

The Human Genome Project (HGP) was an international research programme that ran from 1990 to 2003. Its primary goal was to determine the complete nucleotide sequence of human DNA and to identify all human genes.

Key Findings

  • The human genome contains approximately 3.2 billion base pairs
  • There are approximately 20,000-25,000 protein-coding genes — far fewer than the ~100,000 originally predicted
  • Protein-coding sequences make up only about 1.5% of the genome
  • Much of the remaining DNA was initially called "junk DNA" but is now known to include regulatory sequences, introns, repetitive elements, transposons, and sequences with roles that are still being discovered
  • Humans share approximately 99.9% of their DNA sequence — individual variation comes from the 0.1% that differs
  • Significant portions of the human genome are shared with other organisms (e.g. ~96% with chimpanzees, ~85% with mice), reflecting common ancestry

Methods Used

The HGP used whole-genome shotgun sequencing and hierarchical shotgun sequencing:

1. DNA fragmentation — the genome was broken into smaller, overlapping fragments using restriction enzymes or mechanical shearing

2. Cloning — fragments were inserted into bacterial artificial chromosomes (BACs) and replicated

3. Sequencing — each fragment was sequenced using the Sanger (chain-termination) method, which uses dideoxynucleotides (ddNTPs) that terminate the growing DNA strand at specific bases

4. Assembly — computer algorithms aligned overlapping sequences to reconstruct the full genome

Modern sequencing uses next-generation sequencing (NGS) technologies that are massively parallel, faster, and cheaper — the cost of sequencing a human genome has dropped from ~$3 billion (HGP) to under $1,000 today.

Applications of Genome Sequencing

Pharmacogenomics

Pharmacogenomics uses an individual's genomic information to predict their response to drugs:

  • Genetic variants affect how patients metabolise drugs (e.g. variations in cytochrome P450 enzymes)
  • Some variants predict adverse drug reactions — testing can identify patients at risk before prescribing
  • This enables personalised medicine — tailoring drug choice and dosage to the individual's genotype
  • Example: testing for *HLA-B5701** before prescribing the HIV drug abacavir prevents a potentially fatal hypersensitivity reaction

Identifying Disease Genes

Genome data enables identification of genes associated with genetic diseases:

  • Genome-wide association studies (GWAS) compare DNA sequences of affected and unaffected individuals to find single nucleotide polymorphisms (SNPs) associated with disease risk
  • This has identified risk variants for type 2 diabetes, Alzheimer's disease, heart disease, and many cancers
  • Understanding the genetic basis of disease can lead to new drug targets and therapeutic strategies

Comparative Genomics and Evolution

Comparing genomes of different species reveals:

  • Evolutionary relationships — the degree of sequence similarity reflects how recently species diverged from a common ancestor
  • Conserved genes — genes found across many species (e.g. homeobox/Hox genes controlling body plan development) are likely to be functionally essential
  • Gene duplication and divergence as a mechanism for generating new gene functions

Genetic Testing and Screening

  • Predictive testing — identifying individuals carrying alleles for late-onset genetic disorders (e.g. BRCA1/2 for breast cancer risk, HTT for Huntington's disease)
  • Prenatal screening — non-invasive prenatal testing (NIPT) uses cell-free fetal DNA in maternal blood to screen for chromosomal abnormalities
  • Carrier screening — identifying heterozygous carriers of recessive conditions (e.g. cystic fibrosis, sickle cell)
  • Forensic identification — DNA profiling using short tandem repeats (STRs) for criminal investigations and paternity testing

Bioinformatics

Bioinformatics is the use of computer science, mathematics, and statistics to store, retrieve, analyse, and interpret biological data, particularly large genomic datasets.

Key Tools and Databases

ResourcePurpose
GenBank / EMBL / DDBJPublic databases storing DNA and protein sequences from all organisms
BLAST (Basic Local Alignment Search Tool)Compares a query sequence against databases to find similar sequences — used to identify genes, find homologues across species, and predict function
UniProtDatabase of protein sequences and functional information
EnsemblGenome browser for annotated genomes of vertebrates and other eukaryotes
Protein Data Bank (PDB)3D structures of proteins determined by X-ray crystallography and cryo-EM

Sequence Alignment and Homology

Sequence alignment compares two or more DNA or protein sequences to identify regions of similarity:

  • High sequence similarity (homology) suggests common evolutionary origin
  • Conserved regions often indicate functional importance — mutations in these regions are selected against
  • Alignment can reveal mutations that cause disease when compared to a reference sequence

Gene Annotation

Once a genome is sequenced, it must be annotated — identifying where genes are, what they code for, and their regulatory elements:

  • Open reading frames (ORFs) — sequences beginning with a start codon (ATG) and ending with a stop codon, of sufficient length to encode a protein
  • Expressed sequence tags (ESTs) — short sequences derived from mRNA, used to confirm that an ORF is actually transcribed
  • Functional annotation — assigning a function to a gene based on similarity to genes of known function in other organisms

Proteomics

Proteomics is the large-scale study of the complete set of proteins (proteome) expressed by a cell, tissue, or organism at a given time. While the genome is essentially fixed, the proteome varies between:

  • Different cell types (a liver cell and a neurone express different proteins)
  • Different developmental stages
  • Different environmental conditions or disease states

Techniques include mass spectrometry (identifying proteins by their peptide fragments) and 2D gel electrophoresis (separating proteins by charge and mass).

Metabolomics and Systems Biology

  • Metabolomics studies the complete set of metabolites in a biological sample
  • Systems biology integrates genomic, proteomic, and metabolomic data to model and understand biological systems as a whole — how genes, proteins, and metabolites interact in networks

Ethical, Social, and Legal Issues

IssueDetails
Genetic privacyWho has access to an individual's genetic data? Risk of discrimination by employers or insurers
Informed consentParticipants in genome studies must understand how their data will be used
Insurance and employmentThe UK Genetic Insurance Moratorium (2019, extended) prevents insurers from requiring predictive genetic test results (with limited exceptions)
Genetic determinismRisk that society overemphasises genetic causes and underestimates environmental/lifestyle factors
OwnershipCan genes or sequences be patented? The US Supreme Court (2013) ruled that naturally occurring DNA sequences cannot be patented, but cDNA (artificially synthesised) can
EquityBenefits of genomic medicine should be accessible to all, not just wealthy nations or individuals

Exam Tips

  • AQA expects you to discuss the applications AND ethical implications of genome projects — always include both in extended answers
  • When describing sequencing, mention fragmentation, sequencing of fragments, and computational assembly
  • Be able to explain why there are fewer genes than proteins (alternative splicing, post-translational modification)
  • Bioinformatics questions often ask about comparing sequences — state that high similarity suggests common ancestry and/or shared function
  • Know the distinction between genome (DNA), transcriptome (mRNA), and proteome (proteins) — and why the proteome varies between cells even though the genome does not
Don't understand a part?

Sign in and ask our AI tutor to explain any passage in plain English.

Try AI explanations →

More on Gene Expression and Regulation

Protein Synthesis: Transcription and Translation Epigenetics and Gene Regulation

← All A-Level Biology notes