🧬 Biology · Graduate · BIO 480

Molecular Biology

A rigorous graduate treatment of how genetic information is stored, copied, repaired, read and turned into protein, built on primary literature and free full-text sources. The course opens with chromatin as a physical object, from the histone octamer to loop domains and compartments, then works through replication, the repair pathways and the diseases that exposed each of them, transcription in…

Start the interactive course (quizzes, progress, videos) →

Free forever. No sign-up, no ads. 18 lessons. The full lesson text is below so you can read it right here.

Module 1: The Genome as a Physical Object

Chromatin from the histone octamer to loop domains, and the marks and remodellers that decide which stretches of it a polymerase can reach.

From the Nucleosome to the Folded Genome

  • Explain what the micrococcal nuclease ladder and the 2.8 angstrom crystal structure each established about nucleosome organisation.
  • Describe the histone octamer, linker histone and histone variants, and how DNA is held against the protein core.
  • Distinguish what is well established about higher-order chromatin folding from what remains contested.

A ladder where a smear should have been

In 1973 Dean Hewish and Leigh Burgoyne incubated rat liver nuclei with a trace of the endogenous nuclease already present in them, stopped the reaction after a few minutes, and ran the extracted DNA on a gel. Naked DNA cut by a nuclease gives a smear: cuts fall anywhere, so fragment lengths are continuous. What they got instead was a ladder. Bands sat at roughly 200 base pairs, 400, 600, 800, marching up the gel in even steps.

A ladder means the nuclease could only cut at regularly spaced positions, which means something was sitting on the DNA at regular intervals and shielding everything in between. Within a year Roger Kornberg had assembled the chemistry, the X-ray scattering and the histone cross-linking data into a proposal that now sounds obvious and then did not: chromatin is a repeating particle of about 200 base pairs of DNA wrapped around an octamer containing two copies each of histones H3, H4, H2A and H2B, with histone H1 sitting outside.

Key idea: The 200 base pair periodicity was not inferred from a picture of chromatin. It was inferred from what a nuclease could not reach, which is the shape of most good structural evidence before crystallography arrives.

What the 2.8 angstrom map settled

Twenty-three years later, Karolin Luger and colleagues in Timothy Richmond's laboratory published a crystal structure of the nucleosome core particle at 2.8 angstrom resolution, built from recombinant Xenopus histones and a defined 146 base pair palindromic DNA. At that resolution you are no longer arguing about models. You can see which amino acid touches which phosphate.

The core particle holds 147 base pairs of DNA in 1.65 left-handed superhelical turns around the octamer. The DNA does not bind uniformly. It contacts the protein at fourteen points, spaced roughly every ten base pairs, wherever the minor groove faces inward, and at each of those points an arginine side chain inserts into the minor groove. Nothing grips the bases. The contacts are to the phosphodiester backbone and to the deoxyribose, which is exactly why a nucleosome can form on almost any sequence and why it will not read the message it is packaging.

The proteins themselves are built from one repeated motif, the histone fold: a long central helix flanked by two shorter ones, which pairs head-to-tail with a partner fold. H3 pairs with H4, H2A pairs with H2B. Two H3-H4 dimers form a tetramer that organises the central 60 base pairs, and an H2A-H2B dimer binds on each side to take the remaining DNA in. This is why nucleosome assembly in a cell is ordered: the H3-H4 tetramer is deposited first, then the two H2A-H2B dimers, and disassembly runs the other way. When a polymerase pushes through a nucleosome, it is the H2A-H2B dimers that come off first.

The N-terminal tails of all four histones project out through and past the DNA gyres. They are disordered in the crystal, which is a polite way of saying they were not resolved, and they are the substrate for essentially every covalent modification discussed in the next lesson.

The point: The nucleosome grips the DNA backbone at fourteen minor-groove contacts and never touches the bases, so packaging is sequence-tolerant by construction.

Positioning is a bias, not an address

If the histones do not read the sequence, why is nucleosome position reproducible across cells at many loci? Because bending is not sequence-neutral even when binding is. Wrapping 147 base pairs around a protein forces the double helix into a tight curve, and some dinucleotide steps bend more easily than others. AA, TT and TA steps compress the minor groove; GC steps favour the opposite orientation. A sequence that repeats those preferences with the helical period of about ten base pairs bends into the correct superhelix more cheaply than a random sequence.

Jonathan Widom's laboratory exploited this to select the 601 sequence, an artificial 147 base pair fragment that positions a nucleosome so strongly that it is now the standard substrate for in vitro chromatin work. That is the top of the range. Across a genome the effect is real but modest: sequence preference sets a probability, and the actual position is then set by competition with sequence-specific binding proteins, by ATP-dependent remodellers, and by statistical packing against a fixed barrier. A promoter with a nucleosome-depleted region does not have one because the DNA there refuses to bend. It has one because a pioneer factor and a remodeller cleared it.

Variants: the same particle with different parts

The four core histones have variants that swap into specific locations and change what the particle does.

VariantReplacesWhere it goesWhat changes
H2A.ZH2AFlanking nucleosome-depleted regions at promotersAlters nucleosome stability and helps poise genes for activation
H3.3H3Actively transcribed bodies, replication-independentDeposited outside S phase, marks turnover
CENP-AH3Centromeres onlyDefines the kinetochore assembly site epigenetically, not by sequence
macroH2AH2AEnriched on the inactive X chromosomeAssociated with stable repression

CENP-A is the one to sit with. Human centromeres contain alpha satellite repeats, but the repeat sequence is neither necessary nor sufficient: neocentromeres form on sequences with no satellite DNA at all, and satellite arrays exist that carry no centromere. What marks a centromere is the presence of CENP-A nucleosomes, propagated through division by a template-copying mechanism rather than by the underlying sequence. If you want a clean example of information carried in chromatin rather than in DNA, this is a better one than most of the ones usually offered.

What happened to the 30 nanometre fibre

Textbook diagrams for forty years showed the nucleosome string coiling into a 30 nanometre fibre, either as a one-start solenoid or a two-start zigzag, then into progressively thicker structures. The 30 nanometre fibre is real. It forms reliably in vitro from purified, regularly spaced nucleosome arrays at moderate ionic strength, and cryo-electron microscopy has resolved the zigzag arrangement in reconstituted arrays.

The problem is finding it inside a cell. Small-angle X-ray scattering of intact nuclei and cryo-electron tomography of vitrified sections have repeatedly failed to detect a dominant 30 nanometre population in interphase nuclei of most cell types. The ChromEMT imaging work of Ou and colleagues in 2017 described chromatin in situ as disordered chains of nucleosomes between roughly 5 and 24 nanometres wide, packed at varying density, rather than as a uniform fibre. The honest summary is that regular higher-order fibres form under defined conditions in vitro, are seen in a few unusually uniform chromatin types such as avian erythrocyte nuclei, and are not the general organising principle of interphase chromatin in a mammalian nucleus. Textbook figures have been slow to catch up.

Worth holding on to: When a structure is easy to make in a tube and hard to find in a cell, the cautious position is that the tube conditions selected it.

The organisation that did survive: loops and compartments

What replaced the fibre hierarchy came from a different kind of experiment. In chromosome conformation capture, you fix cells with formaldehyde so that DNA segments in physical contact become cross-linked, digest with a restriction enzyme, ligate under dilute conditions so that only cross-linked ends join, and sequence the junctions. Every read tells you that two genomic positions were near each other. Erez Lieberman-Aiden and colleagues scaled this to the whole genome in 2009 and found two things.

First, contacts partition the genome into two large classes, called the A and B compartments, which correspond closely to open, gene-rich, transcriptionally active chromatin and to closed, gene-poor, late-replicating chromatin. Sequences within a compartment contact each other far more than they contact the other compartment, even across tens of megabases.

Second, contact frequency falls off with genomic distance far more slowly than a random coil predicts, which is the signature of a polymer that is compacted without being knotted.

By 2014, Suhas Rao and colleagues had pushed the resolution to about a kilobase and could see the individual loops. Roughly ten thousand loops in a human cell, most of them under a megabase, with CTCF binding sites at both anchors. The striking detail is orientation: at the great majority of loop anchors the two CTCF motifs point toward each other. Sequence-specific binding alone cannot explain a rule about orientation. Loop extrusion can: the cohesin ring is loaded onto chromatin, reels DNA through itself to form a growing loop, and is halted by a CTCF molecule bound in the correct orientation, like a knot that only catches one way round. Cut the CTCF site out, or invert it, and the loop moves or disappears.

Loops assemble into contact domains, usually called topologically associating domains, within which regulatory elements find their targets and across whose boundaries they largely do not. This is why a deletion that removes a boundary can cause disease without touching a coding sequence: an enhancer that was penned into one domain is released into the next and switches on a gene it was never meant to reach.

Numbers worth carrying

  • A diploid human cell contains about 2 metres of DNA in a nucleus roughly 6 micrometres across. That is a linear compaction of order ten thousand fold.
  • At about 200 base pairs of DNA per nucleosome, roughly 30 million nucleosomes package one diploid genome.
  • The core particle holds 147 base pairs; the linker between particles runs from about 20 to 90 base pairs depending on cell type, which is where the 200 base pair average comes from.
  • A loop domain in the Rao data averages under a megabase; compartments run to tens of megabases.

Common misconceptions

  • "Chromatin exists to compact DNA." Compaction is necessary but it is not the interesting part. A protein-free 2 metre polymer could be condensed many ways. Chromatin is a regulated access system, and every level of it is built to be locally undone.
  • "The 30 nanometre fibre is the second level of packing in cells." It is the second level of packing in a test tube. In situ imaging finds irregular, variably packed nucleosome chains instead.
  • "Nucleosomes sit wherever the sequence tells them to." Sequence sets a bending preference of a few kilocalories at most. Remodellers, transcription factors and packing against barriers set the actual positions.
  • "Histones are inert spools." Their tails are the most heavily modified protein surfaces in the nucleus, and swapping one variant for another changes centromere identity.
  • "A TAD is a physical box." It is a statistical statement about contact frequency in a population of cells. Single-cell contact maps are far messier than the ensemble average suggests.

Pulling it together

  • A nuclease ladder in 1973 revealed a repeating particle; Kornberg named its composition in 1974; Luger and colleagues resolved it at 2.8 angstroms in 1997.
  • 147 base pairs wrap 1.65 left-handed turns around an octamer of H3, H4, H2A and H2B, held at fourteen minor-groove arginine contacts to the backbone, never to the bases.
  • Assembly is ordered: H3-H4 tetramer first, then two H2A-H2B dimers, which are also the first to leave.
  • Sequence biases nucleosome position through bendability; remodellers and binding proteins determine it.
  • Variants specialise the particle, and CENP-A defines centromere identity independently of sequence.
  • The regular 30 nanometre fibre is an in vitro structure; in situ chromatin is irregular.
  • Hi-C revealed A and B compartments and, at kilobase resolution, roughly ten thousand CTCF-anchored loops with convergent motif orientation, consistent with cohesin-driven loop extrusion.

Sources

  1. Luger, K., Mader, A. W., Richmond, R. K., Sargent, D. F., & Richmond, T. J. (1997). Crystal structure of the nucleosome core particle at 2.8 A resolution. Nature, 389(6648), 251-260. pubmed.ncbi.nlm.nih.gov
  2. Kornberg, R. D. (1974). Chromatin structure: A repeating unit of histones and DNA. Science, 184(4139), 868-871. pubmed.ncbi.nlm.nih.gov
  3. Alberts, B., Johnson, A., Lewis, J., Raff, M., Roberts, K., & Walter, P. (2002). Chromosomal DNA and its packaging in the chromatin fiber. In Molecular biology of the cell (4th ed.). Garland Science. ncbi.nlm.nih.gov
  4. Lieberman-Aiden, E., van Berkum, N. L., Williams, L., Imakaev, M., Ragoczy, T., Telling, A., et al. (2009). Comprehensive mapping of long-range interactions reveals folding principles of the human genome. Science, 326(5950), 289-293. pubmed.ncbi.nlm.nih.gov
  5. Rao, S. S. P., Huntley, M. H., Durand, N. C., Stamenova, E. K., Bochkov, I. D., Robinson, J. T., et al. (2014). A 3D map of the human genome at kilobase resolution reveals principles of chromatin looping. Cell, 159(7), 1665-1680. pubmed.ncbi.nlm.nih.gov
  6. National Human Genome Research Institute. (n.d.). Nucleosome. Talking Glossary of Genomic and Genetic Terms. genome.gov
Key terms
Nucleosome core particle
147 base pairs of DNA wrapped in 1.65 left-handed superhelical turns around an octamer of histones H3, H4, H2A and H2B.
Histone fold
The paired helix-strand-helix motif shared by all four core histones, which mediates their heterodimerisation.
Micrococcal nuclease ladder
The regularly spaced fragment pattern produced by limited nuclease digestion of chromatin, the original evidence for a repeating particle.
CENP-A
A histone H3 variant that marks centromeres and propagates centromere identity independently of the underlying DNA sequence.
Hi-C
A genome-wide chromosome conformation capture method in which cross-linked, ligated DNA junctions are sequenced to measure contact frequency between loci.
Loop extrusion
The model in which cohesin reels chromatin through its ring to form a loop that is halted by CTCF bound in a convergent orientation.
Topologically associating domain
A genomic interval within which contacts are enriched and across whose boundary they are depleted, defined statistically from population contact maps.
601 sequence
An artificially selected 147 base pair DNA fragment with unusually strong nucleosome positioning, used as a standard in vitro substrate.

Histone Marks, Remodellers, and What Epigenetic Inheritance Claims

  • Describe the writer, eraser and reader logic of histone modification and name the marks associated with active and repressed chromatin.
  • Explain how ATP-dependent remodellers and DNA methylation change access to a promoter.
  • Distinguish mitotic propagation of a chromatin state from transgenerational epigenetic inheritance, and assess the evidence for each.

A gel band that changed the argument

Histone acetylation had been correlated with active transcription since Vincent Allfrey reported it in 1964, and for thirty years the correlation went nowhere, because nobody could say whether acetylation caused transcription, followed it, or merely accompanied it. In 1996 James Brownell, working with David Allis, purified the enzyme responsible from Tetrahymena macronuclei using an in-gel assay, sequenced it, and found it was a homolog of yeast Gcn5p.

Gcn5p was not a mystery protein. It was already known to geneticists as a transcriptional coactivator, isolated years earlier in screens for factors required to switch genes on. The enzyme that acetylates histones and the protein that activates transcription turned out to be the same molecule. That single identification converted a thirty-year correlation into a mechanism and started the field that this lesson is about.

Why this matters: The decisive experiment was not a better correlation. It was finding that two separately named things were one thing.

Writers, erasers, readers

Every covalent histone modification involves three kinds of protein, and keeping them straight prevents most of the confusion in this area.

  • Writers add the mark. Histone acetyltransferases such as Gcn5, p300 and CBP transfer an acetyl group from acetyl-CoA to a lysine. Histone methyltransferases such as SET7 and EZH2 transfer methyl groups from S-adenosylmethionine.
  • Erasers remove it. Histone deacetylases (HDACs, including the NAD-dependent sirtuins) strip acetyl groups; demethylases such as LSD1 and the Jumonji-domain enzymes remove methyl groups, which for two decades was thought impossible.
  • Readers bind the mark and do something. A bromodomain binds acetylated lysine; a chromodomain binds methylated lysine; PHD fingers, tudor domains and others recognise particular states.

Acetylation has a direct physical effect as well as a signalling one: it neutralises the positive charge on a lysine, weakening the electrostatic grip of the tail on the DNA backbone and loosening the array. Methylation does not change charge at all. A methylated lysine is a docking site and nothing more, which is why the same modification can mean opposite things depending on which reader is present and where the residue sits.

MarkTypical locationUsual association
H3K4me3Promoters of active genesActive initiation
H3K27acActive enhancers and promotersActive, used to call enhancers
H3K36me3Bodies of transcribed genesElongation, suppresses cryptic starts
H3K27me3Polycomb domainsFacultative repression, reversible
H3K9me3Pericentric heterochromatin, repeatsConstitutive repression

The last two are worth separating carefully. H3K27me3 is written by PRC2 through its catalytic subunit EZH2 and read by PRC1; it marks genes that are off now but could be on later, which is why it dominates developmental regulator loci. H3K9me3 is read by HP1 through its chromodomain, and because HP1 also recruits the methyltransferase that writes the mark, the state can spread along the fibre and copy itself. That self-propagating loop is the best-characterised mechanism by which a chromatin state persists.

Remodellers: the machines that move nucleosomes

Marks change the surface. Remodellers change the position. ATP-dependent chromatin remodelling complexes all contain a related ATPase motor domain and use the energy of ATP hydrolysis to translocate DNA across the histone octamer surface, but they differ in what they do with that capability.

  • SWI/SNF family complexes slide and evict nucleosomes, opening promoters and enhancers. Subunits of the human complex (ARID1A, SMARCB1, SMARCA4) are among the most frequently mutated genes in cancer, which tells you how much regulation runs through them.
  • ISWI family complexes space nucleosomes evenly, tidying arrays rather than opening them.
  • CHD family complexes both slide and, in the NuRD form, couple remodelling to deacetylation.
  • INO80 and SWR1 exchange histone dimers, with SWR1 swapping H2A-H2B for H2A.Z-H2B at promoter-flanking positions.

In short: Writers and erasers change what a nucleosome looks like to a reader; remodellers change where it is. Transcriptional activation usually needs both.

DNA methylation, and the one thing it does well

In vertebrates, methylation of cytosine at the 5 position occurs almost entirely in the CpG dinucleotide. DNMT3A and DNMT3B establish new patterns; DNMT1 copies existing ones, because it prefers hemimethylated CpG and therefore restores the mark on the daughter strand after replication. That preference is the whole mechanism of heritability for this particular mark, and it is genuinely templated: the parental strand carries the information, the enzyme copies it.

Removal is less tidy. The TET enzymes oxidise 5-methylcytosine to 5-hydroxymethylcytosine and further, feeding into base excision repair for eventual replacement with unmodified cytosine. There is no simple demethylase that lifts the methyl group off.

CpG dinucleotides are depleted across the vertebrate genome, because methylated cytosine deaminates to thymine and the resulting T:G mismatch is repaired inefficiently, so CpG has been eroding for hundreds of millions of years. The exceptions are CpG islands, stretches of a few hundred to a few thousand base pairs with near-expected CpG density, sitting at roughly 60 percent of human promoters. Those islands are usually unmethylated. When one is methylated, the gene is usually silenced stably, which is exactly what happens to tumour suppressors such as MLH1 and CDKN2A in many cancers, and it is why hypomethylating agents such as azacitidine exist as drugs.

Imprinting and X inactivation: the two clean cases

Two phenomena show unambiguously that identical DNA sequences can carry stably different expression states.

Genomic imprinting: about 100 human genes are expressed from only one parental allele. At the IGF2/H19 locus on chromosome 11, an imprinting control region between the two genes is methylated on the paternal allele and unmethylated on the maternal one. Unmethylated, it binds CTCF, which blocks the shared downstream enhancers from reaching IGF2, so the maternal allele expresses H19. Methylated, CTCF cannot bind, the enhancers reach IGF2, and the paternal allele expresses IGF2. Loss of imprinting at this locus is one route into Beckwith-Wiedemann syndrome. Notice that the mechanism is entirely conventional: methylation blocks a protein, the protein blocks an enhancer.

X inactivation: in female mammals one X is silenced early in development, coated by the XIST long non-coding RNA transcribed from the inactive X itself, which recruits repressive complexes and eventually macroH2A and dense DNA methylation. Once established, the choice is clonally inherited through every subsequent mitosis, which is why a calico cat's coat is patchy rather than blended.

The dispute: what does "epigenetic inheritance" mean

Here the field genuinely divides, and the division is partly about evidence and partly about definitions.

Mitotic inheritance is not in dispute. A chromatin state can survive DNA replication and be re-established on both daughter chromatids. The mechanisms are known and at least partly templated: DNMT1 copying hemimethylated CpG, recycling of parental H3-H4 tetramers onto both daughter strands so that a reader-writer pair such as HP1 with its methyltransferase can re-mark the new nucleosomes, and self-reinforcing PRC2 recruitment to existing H3K27me3. A differentiated cell keeps being that cell type through hundreds of divisions largely by these means.

Transgenerational inheritance in mammals, meaning a chromatin state transmitted through the germ line to grandchildren and beyond without a DNA sequence change, is a much weaker claim. The obstacle is reprogramming: mammalian germ cells and early embryos erase and rewrite most DNA methylation twice, once in the primordial germ cells and once after fertilisation. Marks that survive both waves are rare, and the strongest survivors are precisely the imprinted loci and certain repeat elements.

The two examples most often cited deserve to be described accurately. In the agouti viable yellow mouse, an intracisternal A particle retrotransposon upstream of the Agouti gene drives ectopic expression when it is hypomethylated, giving yellow, obese offspring; when it is methylated the mice are brown and lean. Robert Waterland and Randy Jirtle showed in 2003 that supplementing the mothers' diet with methyl donors shifted the distribution of coat colours in the offspring. That is a real, replicated, dose-dependent effect on a metastable epiallele. It is also a mouse strain built around an unusual retrotransposon, and generalising from it to human traits is a long jump.

In humans, Bastiaan Heijmans and colleagues found in 2008 that individuals conceived during the Dutch Hunger Winter of 1944 to 1945 had, six decades later, slightly lower methylation at the imprinted IGF2 differentially methylated region than their unexposed same-sex siblings. The comparison to siblings is what makes the study strong. What it shows is a persistent mark associated with a prenatal exposure in the exposed individual. It does not show transmission to the next generation, and the effect size is small.

Outside mammals the picture is different and stronger. In Caenorhabditis elegans, RNA-mediated silencing can persist for many generations. In plants, which have no sequestered germ line and much weaker reprogramming, methylation epialleles can be inherited stably for many generations; the peloric variant of Linaria vulgaris that Linnaeus described in 1749 was found in 1999 to be caused by methylation of the Lcyc gene, not by a sequence change.

Bottom line: Say "a chromatin state that is propagated through mitosis" when that is what you mean, and reserve "transgenerational epigenetic inheritance" for germ line transmission with the reprogramming problem addressed. Most published use of the phrase does not clear that bar in mammals.

Common misconceptions

  • "Epigenetics means inheritance without DNA change." The word is used for at least three things: chromatin marks generally, mitotic propagation of a cell state, and germ line transmission. Papers routinely slide between them.
  • "Histone methylation represses and acetylation activates." Acetylation is reliably activating. Methylation depends entirely on residue and degree: H3K4me3 is active, H3K9me3 is repressive, and the same lysine can mean different things at mono, di and tri levels.
  • "DNA methylation silences genes." It silences promoters and CpG islands. Methylation in gene bodies is a normal feature of highly transcribed genes.
  • "The histone code is a code." The word implies a lookup table. What exists is a set of overlapping, context-dependent binding preferences, and attempts to read combinations as discrete instructions have not held up well.
  • "Lifestyle epigenetics can be measured from a cheek swab." Methylation differs between cell types more than it differs between most exposures, so any measurement in a mixed tissue is dominated by cell composition unless that is explicitly corrected for.

Where this leaves us

  • The Gcn5 identification in 1996 tied histone acetylation directly to transcriptional activation and launched the writer, eraser and reader framework.
  • Acetylation both loosens the array electrostatically and creates bromodomain docking sites; methylation only creates docking sites.
  • H3K4me3 and H3K27ac mark activity; H3K27me3 marks reversible Polycomb repression; H3K9me3 with HP1 forms a self-propagating repressive loop.
  • Remodellers of the SWI/SNF, ISWI, CHD and INO80 families slide, evict, space and exchange nucleosomes; SWI/SNF subunits are heavily mutated in cancer.
  • DNMT1's preference for hemimethylated CpG makes DNA methylation genuinely templated across replication; TET-mediated removal is indirect.
  • Imprinting at IGF2/H19 and X inactivation are the two unambiguous cases of stable, heritable expression differences on identical sequence.
  • Mitotic propagation is well supported; mammalian transgenerational inheritance is limited by two waves of germ line reprogramming and rests on a small number of special cases.

Sources

  1. Brownell, J. E., Zhou, J., Ranalli, T., Kobayashi, R., Edmondson, D. G., Roth, S. Y., & Allis, C. D. (1996). Tetrahymena histone acetyltransferase A: A homolog to yeast Gcn5p linking histone acetylation to gene activation. Cell, 84(6), 843-851. pubmed.ncbi.nlm.nih.gov
  2. Waterland, R. A., & Jirtle, R. L. (2003). Transposable elements: Targets for early nutritional effects on epigenetic gene regulation. Molecular and Cellular Biology, 23(15), 5293-5300. pubmed.ncbi.nlm.nih.gov
  3. Heijmans, B. T., Tobi, E. W., Stein, A. D., Putter, H., Blauw, G. J., Susser, E. S., Slagboom, P. E., & Lumey, L. H. (2008). Persistent epigenetic differences associated with prenatal exposure to famine in humans. PNAS, 105(44), 17046-17049. pubmed.ncbi.nlm.nih.gov
  4. National Human Genome Research Institute. (n.d.). Epigenomics fact sheet. National Institutes of Health. genome.gov
  5. Clark, M. A., Douglas, M., & Choi, J. (2018). Eukaryotic epigenetic gene regulation. In Biology 2e (Section 16.3). OpenStax. openstax.org
  6. Brown, T. A. (2002). Accessing the genome. In Genomes (2nd ed., Chapter 8). Wiley-Liss. ncbi.nlm.nih.gov
Key terms
Writer, eraser, reader
The three functional classes of protein that add, remove and bind a histone modification respectively.
Bromodomain
A protein module that binds acetylated lysine, coupling histone acetylation to recruitment of other factors.
PRC2
Polycomb repressive complex 2, whose EZH2 subunit writes H3K27me3 at developmentally regulated genes.
CpG island
A promoter-associated stretch of several hundred base pairs with near-expected CpG density, normally unmethylated.
DNMT1
The maintenance DNA methyltransferase, which prefers hemimethylated CpG and so copies the methylation pattern onto the daughter strand.
Metastable epiallele
An allele whose expression state is set stochastically early in development and then maintained, as at the agouti viable yellow locus.
Genomic imprinting
Parent-of-origin-specific expression of a gene, established by differential methylation of an imprinting control region in the germ line.
Germ line reprogramming
The two waves of genome-wide methylation erasure and re-establishment, in primordial germ cells and the early embryo, that limit transgenerational transmission of marks.

Module 2: Replicating the Genome

What the replisome actually does at a moving fork, how a cell licenses origins so that every base is copied exactly once, and why linear chromosomes need a separate solution at their ends.

Mechanics at the Fork, and How a Cell Decides Where to Start

  • Explain what the Meselson and Stahl density-shift experiment excluded and what it could not distinguish.
  • Name the components of the replisome and state the problem each one solves at a moving fork.
  • Describe origin licensing and firing, and explain how a cell prevents any origin from firing twice in one cycle.

Two bands, then three, then two again

In 1957 Matthew Meselson and Franklin Stahl grew Escherichia coli for fourteen generations in medium whose only nitrogen source was ammonium chloride made with the heavy isotope nitrogen-15. Every base in every DNA molecule was therefore heavy. They then washed the cells into ordinary nitrogen-14 medium and sampled the culture as it divided, spinning each sample to equilibrium in a caesium chloride density gradient, where DNA settles at the point where its buoyant density matches the salt gradient.

Before the shift, one band, at heavy density. After one generation, one band, at exactly the midpoint between heavy and light. After two generations, two bands, one at the midpoint and one fully light, in equal amounts. After three, the light band was three times the hybrid band.

Read those numbers. Conservative replication, in which the parent duplex stays intact and an entirely new duplex is made, would have given heavy and light bands after one generation and no hybrid band ever. It was eliminated immediately. Dispersive replication, in which parental and new DNA are interspersed along both strands, would have given a single band moving steadily toward light and never resolving into two. It was eliminated at generation two. What remains is semiconservative replication: each daughter duplex contains one intact parental strand and one new one.

Remember: The experiment is celebrated for its elegance, but the reason it worked is that it produced qualitatively different predictions, not merely quantitatively different ones. Design your experiments that way when you can.

The problem list at a moving fork

Semiconservative copying sounds simple until you write down what a fork has to solve simultaneously. Each item on this list is solved by a different protein.

  1. The duplex must be opened. Helicase does this, using ATP: DnaB in bacteria, translocating on the lagging-strand template, and the CMG complex (Cdc45, MCM2-7, GINS) in eukaryotes, translocating on the leading-strand template.
  2. Opening a helix ahead of the fork over-winds the DNA in front of it. Topoisomerases relieve that torsional strain; DNA gyrase in bacteria actively introduces negative supercoils, and topoisomerase II handles the intertwined daughters at the end.
  3. Separated single strands would re-anneal or form hairpins. Single-strand binding protein (SSB in bacteria, RPA in eukaryotes) coats them.
  4. No DNA polymerase can start a chain from nothing. Primase lays down a short RNA primer, about 10 to 12 nucleotides, that provides a 3 prime hydroxyl.
  5. Polymerase falls off DNA after a few dozen nucleotides on its own. A sliding clamp (the beta clamp in bacteria, PCNA in eukaryotes) is a ring loaded around the duplex by an ATP-driven clamp loader, and it tethers the polymerase so it can copy thousands of bases without dissociating.
  6. DNA polymerase can only extend a chain 5 prime to 3 prime, so one template can be read continuously and the other cannot. The lagging strand is made in fragments.
  7. The RNA primers must be removed and the nicks sealed. RNase H and, in bacteria, DNA polymerase I with its 5 prime to 3 prime exonuclease do the removal; in eukaryotes FEN1 cleaves the displaced flap. DNA ligase seals.

The core of it: Every component of the replisome exists because of a specific physical or chemical constraint. If you can state the constraint, you can reconstruct the component list from memory.

Okazaki fragments and the trombone

Reiji and Tsuneko Okazaki tested the discontinuity prediction directly in the 1960s by pulse-labelling phage-infected E. coli with tritiated thymidine for a few seconds and denaturing the DNA before sizing it. Most of the label appeared in short pieces, around 1000 to 2000 nucleotides. Extend the pulse and the label moved into high molecular weight DNA, as the short pieces were joined. Eukaryotic Okazaki fragments are much shorter, roughly 100 to 200 nucleotides, which is close to the nucleosome repeat length and not a coincidence: the nascent lagging strand is being handed to chromatin assembly as it goes.

The two strands are copied by polymerases that are physically coupled, which raises a geometry question: how can two polymerases moving in opposite directions along antiparallel templates travel together? The trombone model answers it. The lagging-strand template loops out through the polymerase so that both enzymes move in the same physical direction along the fork; the loop grows as a fragment is synthesised and collapses when the polymerase releases and re-engages at the next primer, like the slide of a trombone.

In eukaryotes the labour is divided further. Pol alpha, which carries the primase activity, lays down an RNA primer and extends it with about 20 nucleotides of DNA, then hands off. Pol epsilon takes the leading strand, Pol delta the lagging strand. Pol alpha has no proofreading exonuclease, which means the first 20 or so nucleotides of every Okazaki fragment are made by an error-prone enzyme; most of that DNA is subsequently displaced and removed during maturation, which is a neat piece of quality control by deletion.

How accurate is accurate

Fidelity comes from three layers and it is worth holding the numbers.

LayerMechanismError rate after this layer
Base selectionGeometry of the polymerase active site favours correct Watson-Crick pairsabout 1 in 105
Proofreading3 prime to 5 prime exonuclease excises a misinserted base before the next one is addedabout 1 in 107
Mismatch repairPost-replicative correction, covered in Module 3about 1 in 109 to 1010

At 1 error in 109 to 1010 and 6 billion base pairs per diploid replication, a human cell division introduces on the order of a few new mutations. That is the number to hold when you read claims about mutation rates: it is small per division and enormous per organism per lifetime.

Bacterial forks move at roughly 1000 nucleotides per second. Eukaryotic forks move at roughly 50 nucleotides per second, twenty times slower. E. coli copies 4.6 million base pairs from a single origin; a human cell copies 6 billion in a comparable S phase. The arithmetic forces the conclusion: eukaryotes must use many origins.

Where replication starts, and the once-only problem

E. coli has one origin, oriC, about 245 base pairs containing DnaA boxes and an AT-rich region that melts first. DnaA loaded with ATP oligomerises on the boxes, bends and opens the duplex, and loads the DnaB helicase. Initiation is coupled to growth rate, and fast-growing cells begin a new round before the previous one finishes, so a rapidly dividing cell can contain more than two copies of the origin-proximal genome.

Eukaryotic origins were harder to find, because in most eukaryotes they are not defined by a strong consensus sequence. The breakthrough was biochemical. In 1992 Stephen Bell and Bruce Stillman used a footprinting assay on yeast replicator sequences to purify a six-subunit complex that bound them in an ATP-dependent manner, and named it the origin recognition complex, ORC. ORC is the landing pad, present at origins throughout the cycle.

The licensing scheme that follows is the elegant part, and it is built entirely around one rule: an origin can only be loaded when CDK activity is low, and can only fire when CDK activity is high. Those two conditions cannot be satisfied at the same time, so no origin can be loaded twice in one cycle.

  • Licensing, in G1 (low CDK). ORC, with CDC6 and CDT1, loads inactive MCM2-7 double hexamers around origin DNA. This is the licence.
  • Firing, at the G1 to S transition (high CDK plus DDK). S-phase CDK and the Dbf4-dependent kinase Cdc7 recruit Cdc45 and GINS, converting MCM2-7 into the active CMG helicase. The double hexamer splits and two forks set off in opposite directions.
  • Blocking re-licensing, through S, G2 and M (high CDK). CDK phosphorylates CDC6 for export or degradation, phosphorylates ORC, and in metazoans the protein geminin binds and sequesters CDT1. Licensing cannot restart until CDK activity collapses at the end of mitosis.

The upshot: Once-per-cycle replication is not enforced by counting. It is enforced by making loading and firing require mutually exclusive conditions, which is a far more robust design.

Dormant origins and the shape of S phase

A human cell licenses far more origins than it uses, on the order of tens of thousands of MCM loading sites, of which perhaps one in ten fires in a given cycle. The excess is not waste. When a fork stalls, a nearby dormant origin can fire and rescue the unreplicated region, and cells with reduced MCM loading are hypersensitive to replication stress. Origins also fire on a reproducible schedule: gene-rich, open, A-compartment chromatin replicates early, and gene-poor, B-compartment and lamina-associated chromatin replicates late. Replication timing is one of the most reproducible genome-wide measurements there is, and it changes with cell type.

When a fork does stall, on a lesion, a tightly bound protein, or a depleted nucleotide pool, the uncoupling of helicase from polymerase generates long stretches of RPA-coated single-stranded DNA. That is the signal for the ATR kinase, which stabilises the fork, suppresses distant origin firing, and delays mitosis. Loss of that response, rather than the stalling itself, is what turns replication stress into chromosome breakage.

Common misconceptions

  • "The lagging strand is copied backwards." Every polymerase on the planet extends 5 prime to 3 prime. The lagging strand is copied in the same chemical direction as the leading strand, just in short pieces initiated repeatedly as the fork exposes new template.
  • "Primers are DNA." They are RNA, laid down by primase, and must be removed and replaced. This is not an accident: RNA primers are chemically marked as low-fidelity and are guaranteed to be excised.
  • "Meselson and Stahl proved semiconservative replication in one step." The first generation eliminated conservative replication. It took the second generation, with its two bands, to eliminate dispersive replication.
  • "Eukaryotic origins are defined by a sequence like oriC." Budding yeast has sequence-defined replicators, but in metazoans origin location is determined largely by chromatin context, transcription and ORC binding preference rather than by a consensus motif.
  • "More origins fired means faster S phase." Origin firing is limited by rate-limiting firing factors; firing everything at once would exhaust them and leave no dormant reserve for stalled forks.

Putting it together

  • Density-shift data gave one hybrid band at generation one and hybrid plus light at generation two, which eliminates conservative and then dispersive replication.
  • The replisome is a list of solutions: helicase opens, topoisomerase relieves torsion, SSB or RPA protects, primase starts, the clamp tethers, ligase seals.
  • Lagging-strand synthesis is discontinuous: about 1000 to 2000 nucleotide fragments in bacteria, 100 to 200 in eukaryotes, with the template looped out in the trombone arrangement.
  • Eukaryotic division of labour: Pol alpha primes, Pol epsilon takes the leading strand, Pol delta the lagging strand, and Pol alpha's error-prone output is largely removed during maturation.
  • Fidelity is layered: 10-5 from selection, 10-7 with proofreading, 10-9 to 10-10 after mismatch repair.
  • ORC marks origins; CDC6 and CDT1 load MCM2-7 in G1 under low CDK; CDK and DDK convert them to active CMG helicases; high CDK then blocks re-licensing until mitosis ends.
  • Most licensed origins never fire, and that dormant reserve is what rescues stalled forks under replication stress.

Sources

  1. Meselson, M., & Stahl, F. W. (1958). The replication of DNA in Escherichia coli. PNAS, 44(7), 671-682. pubmed.ncbi.nlm.nih.gov
  2. Okazaki, T., & Okazaki, R. (1969). Mechanism of DNA chain growth, IV: Direction of synthesis of T4 short DNA chains as revealed by exonucleolytic degradation. PNAS, 64(4), 1242-1248. pubmed.ncbi.nlm.nih.gov
  3. Bell, S. P., & Stillman, B. (1992). ATP-dependent recognition of eukaryotic origins of DNA replication by a multiprotein complex. Nature, 357(6374), 128-134. pubmed.ncbi.nlm.nih.gov
  4. Alberts, B., Johnson, A., Lewis, J., Raff, M., Roberts, K., & Walter, P. (2002). DNA replication mechanisms. In Molecular biology of the cell (4th ed.). Garland Science. ncbi.nlm.nih.gov
  5. Alberts, B., Johnson, A., Lewis, J., Raff, M., Roberts, K., & Walter, P. (2002). The initiation and completion of DNA replication in chromosomes. In Molecular biology of the cell (4th ed.). Garland Science. ncbi.nlm.nih.gov
  6. Cooper, G. M. (2000). DNA replication. In The cell: A molecular approach (2nd ed.). Sinauer Associates. ncbi.nlm.nih.gov
Key terms
Semiconservative replication
Each daughter duplex contains one intact parental strand and one newly synthesised strand.
Okazaki fragment
A short DNA piece synthesised discontinuously on the lagging strand, about 1000 to 2000 nucleotides in bacteria and 100 to 200 in eukaryotes.
Sliding clamp
A ring-shaped processivity factor, the beta clamp or PCNA, loaded around DNA to tether the polymerase to its template.
CMG helicase
The active eukaryotic replicative helicase formed from Cdc45, MCM2-7 and GINS at the moment an origin fires.
Origin recognition complex
The six-subunit complex, identified by Bell and Stillman in 1992, that binds origins and serves as the platform for helicase loading.
Licensing
The loading of inactive MCM2-7 double hexamers at origins during G1, permitted only while cyclin-dependent kinase activity is low.
Dormant origin
A licensed origin that does not fire in a normal cycle but can be activated to rescue a region served by a stalled fork.
Replication timing
The reproducible schedule by which open, gene-rich chromatin replicates early in S phase and closed, gene-poor chromatin replicates late.

Telomeres, Telomerase and the End-Replication Problem

  • State the end-replication problem precisely and explain why it applies to linear but not circular chromosomes.
  • Describe telomere structure, shelterin, and how a chromosome end avoids being read as a double-strand break.
  • Relate telomerase deficiency to the telomere biology disorders and evaluate claims made about telomere length as a biomarker.

Three signs in a twelve-year-old

A boy is brought to clinic with fingernails and toenails that have thinned and ridged and are beginning to disappear, white patches on the tongue that will not scrape off, and a lacy, net-like grey-brown pigmentation across the neck and upper chest. His full blood count is falling in all three lineages. The triad of nail dystrophy, oral leukoplakia and reticulate pigmentation with bone marrow failure is dyskeratosis congenita, and it is caused by mutations in genes whose only shared function is maintaining the ends of chromosomes.

The tissues that fail are the ones that divide fastest: bone marrow, skin, mucosa, and in adults, lung and liver. That pattern is the clinical fingerprint of a replicative rather than a metabolic defect. To understand why, you have to start with an arithmetic problem that has no solution within the replisome.

The problem, stated exactly

Take a linear duplex. Each parental strand is copied by a polymerase that extends only 5 prime to 3 prime and cannot start without a primer. On one daughter, the leading strand, a single primer at the far end is enough and synthesis runs to the very end of the template. On the other daughter, the lagging strand, synthesis proceeds in fragments, and the last fragment requires a primer laid down somewhere internal to the 3 prime end of the template. When that RNA primer is removed there is nothing upstream to fill the gap, because filling it would require extending a chain in the 3 prime to 5 prime direction.

So one daughter duplex is shorter than its parent at that end. James Watson pointed this out in 1972 and Alexey Olovnikov, in a 1973 paper he called a theory of marginotomy, worked out the consequences and proposed that this progressive shortening limits the number of divisions a cell can undergo. A circular chromosome has no ends and no problem, which is why bacteria are untroubled by any of this.

The shortening is not only the primer gap. Nucleolytic resection to generate a 3 prime single-stranded overhang adds to it. In human somatic cells the net loss is roughly 50 to 100 base pairs per division.

Key idea: The end-replication problem follows from two facts you already know, unidirectional synthesis and the requirement for a primer. Nothing extra is needed to derive it.

What is at the end

Elizabeth Blackburn and Joseph Gall sequenced the ends of the extrachromosomal ribosomal DNA of Tetrahymena in 1978 and found a simple repeat rather than a unique sequence. In 1988 Robert Moyzis and colleagues showed that human chromosomes carry the repeat TTAGGG, tandemly reiterated, at every end. Human telomeres run to about 10 to 15 kilobases at birth and shorten with age and division. The 3 prime G-rich strand extends beyond the C-rich strand as a single-stranded overhang of roughly 50 to 300 nucleotides.

That overhang creates a problem more urgent than shortening. A single-stranded 3 prime end at a chromosome terminus is molecularly indistinguishable from a resected double-strand break, and a cell that treats it as one will either arrest permanently or ligate two chromosomes together end to end. The solution is a six-protein complex called shelterin.

  • TRF1 and TRF2 bind double-stranded TTAGGG repeats.
  • POT1 binds the single-stranded G overhang, hiding it from RPA and therefore from ATR.
  • TIN2 and TPP1 bridge the two halves of the complex, and TPP1 also recruits telomerase.
  • RAP1 associates with TRF2 and contributes to suppressing end joining.

TRF2 additionally promotes a structural solution: the single-stranded overhang invades the upstream duplex repeats to form a large lasso called the t-loop, with the invading strand displacing one strand of the duplex to form a small displacement loop. The end is now tucked inside the telomere rather than exposed. Remove TRF2 experimentally and chromosome ends fuse within a cell cycle, producing dicentric chromosomes that break at the next anaphase.

What matters here: Shelterin does not repair the end. It conceals it, by keeping ATM, ATR, non-homologous end joining and homologous recombination all switched off at a structure that would otherwise trip every one of them.

Telomerase

In 1985 Carol Greider and Elizabeth Blackburn incubated a Tetrahymena extract with a synthetic telomeric primer and found that it added telomeric repeats to the primer's 3 prime end. The activity was sensitive to RNase, which is the observation that gave the game away: the enzyme carries its own RNA.

Telomerase is a reverse transcriptase. Its catalytic protein subunit, TERT, uses a short template region within an integral RNA subunit, TERC in humans, to add TTAGGG repeats onto the 3 prime end, then translocates and repeats. The C-rich strand is then filled in by the ordinary lagging-strand machinery. It is the only enzyme in the human cell that adds DNA using an RNA template that it brings with it.

Expression is restricted. Germ line cells, stem cell compartments and activated lymphocytes have telomerase; most somatic cells have very little. This is why cultured human fibroblasts stop dividing after a characteristic number of population doublings, the limit Leonard Hayflick described in 1965 and which is now understood as replicative senescence triggered by critically short telomeres signalling a persistent DNA damage response. In 1998 Andrea Bodnar and colleagues introduced hTERT into normal human cells and they kept dividing well past the limit, which closed the argument about causation.

The same mechanism sits under most cancers. Nam Kim and colleagues reported in 1994 that telomerase activity was detectable in about 90 percent of tumour samples and immortal cell lines and absent from most normal somatic tissues. Reactivation is frequently caused by point mutations in the TERT promoter that create a new ETS transcription factor binding site, which are among the most common non-coding mutations in human cancer. The remaining tumours use alternative lengthening of telomeres, a recombination-based mechanism associated with loss of the chromatin remodeller ATRX and with long, heterogeneous telomeres.

When maintenance fails: the telomere biology disorders

Dyskeratosis congenita and its relatives are caused by mutations in telomerase components and in the machinery that assembles and delivers them: DKC1 (X-linked, encoding dyskerin, which stabilises TERC), TERC, TERT, TINF2 (encoding TIN2), RTEL1, and others. Presentations range from the classic childhood triad to adults presenting with idiopathic pulmonary fibrosis, aplastic anaemia or cryptogenic liver cirrhosis and no skin findings at all. The unifying laboratory finding is very short telomeres for age, usually below the first centile by flow-FISH.

Autosomal dominant families show genetic anticipation: each generation presents earlier and more severely. The mechanism is unusually clean. A child inheriting a defective TERC allele inherits not only the mutation but also the already-shortened telomeres that the affected parent's germ line passed on. The starting point, not just the rate, is worse in each generation.

These patients also illustrate why the sequencing result and the functional result are both needed. A missense variant in TERT means little on its own; short telomeres for age make it interpretable.

Telomere length as a biomarker, honestly

Telomere length is measured three ways, and they do not agree well.

MethodWhat it reportsMain limitation
Terminal restriction fragment Southern blotMean length in kilobases, including subtelomeric DNANeeds micrograms of DNA; overestimates because subtelomere is included
Quantitative PCR (T/S ratio)Telomere signal relative to a single-copy geneRelative, not absolute; coefficient of variation often 5 to 10 percent, which is comparable to a decade of age-related change
Flow-FISHLength per cell type, in kilobases, against an age-matched referenceRequires fresh viable cells; available in few laboratories

For diagnosing a telomere biology disorder, flow-FISH on lymphocyte subsets against age-matched normal ranges is the clinically validated test. For consumer or epidemiological claims that a lifestyle intervention lengthened telomeres by some percentage, the usual qPCR measurement error is of the same size as the reported effect, the measurement is confounded by leukocyte composition, and a single tissue at one time point is being used to speak for an organism. Treat the association literature with the same caution you would apply to any noisy biomarker measured once.

So what?: Telomere shortening is a real, mechanistically understood limit on division. It is not a general-purpose clock for ageing, and the gap between those two statements is where most overreach happens.

Common misconceptions

  • "Telomeres shorten because DNA polymerase falls off the end." The leading strand is copied to the end. The loss comes from the final lagging-strand primer, plus deliberate resection to make the overhang.
  • "Telomerase is an anti-ageing enzyme." It is a maintenance enzyme whose reactivation is required by most cancers. Systemic activation is not obviously desirable.
  • "Bacteria have telomeres too." Most bacterial chromosomes are circular and have no end-replication problem. The organisms that do have linear chromosomes, such as Borrelia and Streptomyces, solve it with covalently closed hairpins or terminal proteins instead.
  • "Short telomeres cause senescence by running out of DNA." No coding sequence is lost. A critically short telomere loses the ability to form a protected end structure, and the exposed end signals a persistent DNA damage response through ATM and p53.
  • "Telomere length tells you your biological age." Between-individual variation at any age is far wider than the average change per decade, and measurement error in the commonest assay is of the same magnitude as the signal.

The short version

  • The end-replication problem follows from unidirectional synthesis plus the primer requirement, and applies only to linear chromosomes; net loss is roughly 50 to 100 base pairs per human somatic division.
  • Human telomeres are TTAGGG repeats, about 10 to 15 kilobases at birth, ending in a 3 prime overhang of 50 to 300 nucleotides.
  • Shelterin, with TRF1, TRF2, POT1, TIN2, TPP1 and RAP1, hides the end and enables t-loop formation; losing TRF2 causes end-to-end fusion within one cycle.
  • Telomerase is a reverse transcriptase carrying its own RNA template; Greider and Blackburn identified it in 1985 by its RNase sensitivity.
  • Replicative senescence, the Hayflick limit, is caused by critically short telomeres; hTERT expression abolishes it.
  • About 90 percent of cancers reactivate telomerase, often through TERT promoter mutations; the rest use recombination-based alternative lengthening.
  • Telomere biology disorders present as marrow failure, pulmonary fibrosis or liver disease with very short telomeres for age, and show anticipation because shortened telomeres are inherited along with the mutation.

Sources

  1. Greider, C. W., & Blackburn, E. H. (1985). Identification of a specific telomere terminal transferase activity in Tetrahymena extracts. Cell, 43(2 Pt 1), 405-413. pubmed.ncbi.nlm.nih.gov
  2. Olovnikov, A. M. (1973). A theory of marginotomy: The incomplete copying of template margin in enzymic synthesis of polynucleotides and biological significance of the phenomenon. Journal of Theoretical Biology, 41(1), 181-190. pubmed.ncbi.nlm.nih.gov
  3. Moyzis, R. K., Buckingham, J. M., Cram, L. S., Dani, M., Deaven, L. L., Jones, M. D., et al. (1988). A highly conserved repetitive DNA sequence, (TTAGGG)n, present at the telomeres of human chromosomes. PNAS, 85(18), 6622-6626. pubmed.ncbi.nlm.nih.gov
  4. Bodnar, A. G., Ouellette, M., Frolkis, M., Holt, S. E., Chiu, C. P., Morin, G. B., et al. (1998). Extension of life-span by introduction of telomerase into normal human cells. Science, 279(5349), 349-352. pubmed.ncbi.nlm.nih.gov
  5. Kim, N. W., Piatyszek, M. A., Prowse, K. R., Harley, C. B., West, M. D., Ho, P. L., et al. (1994). Specific association of human telomerase activity with immortal cells and cancer. Science, 266(5193), 2011-2015. pubmed.ncbi.nlm.nih.gov
  6. Savage, S. A., & Niewisch, M. R. (2023). Dyskeratosis congenita and related telomere biology disorders. In M. P. Adam et al. (Eds.), GeneReviews. University of Washington, Seattle. ncbi.nlm.nih.gov
  7. National Human Genome Research Institute. (n.d.). Telomere. Talking Glossary of Genomic and Genetic Terms. genome.gov
Key terms
End-replication problem
The unavoidable loss of terminal sequence on the lagging strand of a linear chromosome, because the final RNA primer leaves a gap that cannot be filled.
Telomere
The tandem TTAGGG repeat array capping each vertebrate chromosome end, ending in a single-stranded 3 prime G-rich overhang.
Shelterin
The six-protein complex (TRF1, TRF2, POT1, TIN2, TPP1, RAP1) that binds telomeric DNA and suppresses the DNA damage response at chromosome ends.
T-loop
The lasso structure formed when the 3 prime overhang invades upstream duplex telomeric repeats, sequestering the end.
Telomerase
A reverse transcriptase, TERT plus its integral RNA subunit TERC, that adds telomeric repeats using its own internal template.
Replicative senescence
The permanent proliferative arrest triggered when telomeres become too short to be protected, signalling through ATM and p53.
Alternative lengthening of telomeres
A recombination-based, telomerase-independent maintenance mechanism used by a minority of tumours, associated with ATRX loss.
Flow-FISH
The clinically validated assay that measures telomere length per leukocyte subset against age-matched reference ranges.

Module 3: Damage, Repair and Recombination

The chemistry that attacks DNA every day, the pathways that reverse it, and the human disorders that revealed each pathway by removing it.

Direct Reversal, Base Excision and Nucleotide Excision Repair

  • Quantify the daily spontaneous damage load on a mammalian genome and name the chemistry behind each major lesion.
  • Contrast direct reversal, base excision repair and nucleotide excision repair by the lesions each handles and the steps each takes.
  • Explain why xeroderma pigmentosum causes cancer and Cockayne syndrome does not, given that both are nucleotide excision repair disorders.

A child who cannot go outside

A two-year-old is brought in with severe blistering sunburn after twenty minutes of ordinary sun exposure, followed over the next year by dense freckling on every sun-exposed surface and, before she is ten, her first basal cell carcinoma. Patients with xeroderma pigmentosum carry roughly a thousandfold increased risk of skin cancer and a median age at first skin cancer of about nine years, against the mid-sixties in the general population.

In 1968 James Cleaver took fibroblasts from such patients, irradiated them with ultraviolet light, and measured repair replication, the incorporation of labelled thymidine outside S phase, which reports on gap filling after damage removal. Normal cells did it. Xeroderma pigmentosum cells did not. One assay, one clean negative, and the connection between a repair pathway and a cancer syndrome was made. What follows in this lesson is the machinery that Cleaver's patients were missing, and the two other pathways that keep DNA readable in the meantime.

The daily damage census

DNA is not a stable molecule under physiological conditions. Tomas Lindahl worked out the spontaneous decay rates in the 1970s and summarised the case in a 1993 review that is still the reference point. Per human cell per day, the approximate loads are:

  • Depurination: roughly 10,000 purines lost by spontaneous hydrolysis of the glycosidic bond, leaving abasic sites.
  • Cytosine deamination: on the order of 100 to 500 conversions of cytosine to uracil. Deamination of 5-methylcytosine gives thymine instead, which is far worse because thymine is a legitimate base.
  • Oxidative damage: thousands of lesions from reactive oxygen species, of which 8-oxoguanine is the most studied because it pairs with adenine as readily as with cytosine and therefore fixes G to T transversions.
  • Alkylation: from endogenous methyl donors as well as environmental agents, with O6-methylguanine the mutagenic one because it pairs with thymine.

To that add exogenous insults: ultraviolet B produces cyclobutane pyrimidine dimers and 6-4 photoproducts between adjacent pyrimidines, ionising radiation produces breaks and clustered oxidative damage, and chemical carcinogens such as benzo[a]pyrene diol epoxide form bulky covalent adducts.

The point: The genome is under continuous chemical attack from its own solvent. Repair is not an emergency response, it is routine maintenance, and the mutation rate you actually observe is the residue left after that maintenance.

Direct reversal: the cheapest option

Some lesions can simply be undone, with no excision and no resynthesis.

Photolyase absorbs a photon of blue light and uses the energy to split a cyclobutane pyrimidine dimer back into two pyrimidines. It is present in bacteria, plants and many animals, and absent from placental mammals. If humans had photolyase, xeroderma pigmentosum would be a much milder disease.

MGMT, O6-methylguanine-DNA methyltransferase, transfers the offending methyl group from guanine onto one of its own cysteine residues. Having done so it is irreversibly inactivated and degraded, so it is a suicide protein rather than an enzyme: one molecule, one lesion. This has a direct clinical consequence. Glioblastomas whose MGMT promoter is methylated, and which therefore make little MGMT, respond substantially better to the alkylating drug temozolomide, and MGMT promoter methylation status is used to guide treatment. The tumour's own repair deficiency is the therapeutic opening.

Base excision repair: small lesions, one base at a time

Base excision repair handles damaged, deaminated or oxidised bases that do not greatly distort the helix. It runs as a fixed relay.

  1. A DNA glycosylase flips the target base out of the helix and hydrolyses the glycosidic bond, leaving an abasic site. Glycosylases are lesion-specific: uracil-DNA glycosylase (UNG) removes uracil, OGG1 removes 8-oxoguanine paired with cytosine, MUTYH removes an adenine that has been misinserted opposite 8-oxoguanine.
  2. APE1, an AP endonuclease, cuts the backbone 5 prime to the abasic site.
  3. DNA polymerase beta removes the remaining sugar-phosphate and inserts one correct nucleotide (short-patch repair). If the end is not easily processed, Pol delta or epsilon with PCNA and FEN1 replace 2 to 10 nucleotides instead (long-patch repair).
  4. DNA ligase III with its partner XRCC1, or ligase I in the long-patch route, seals the nick.

The MUTYH example repays a moment. 8-oxoguanine mispairs with adenine during replication. If OGG1 fails to remove the lesion before replication, the daughter strand now has A opposite 8-oxoG. MUTYH removes that adenine, giving replication a second chance to insert cytosine. People with biallelic MUTYH loss develop MUTYH-associated polyposis with hundreds of colorectal adenomas, and the tumours carry a characteristic excess of G to T transversions, which is precisely the mutation signature the missing enzyme was there to prevent. Signature and mechanism agree, which is the kind of internal consistency that makes a mechanistic claim believable.

Worth holding on to: Base excision repair is glycosylase-driven, so its specificity comes from a family of small enzymes each of which knows one lesion. Nucleotide excision repair works the opposite way.

Nucleotide excision repair: one machine, many lesions

Nucleotide excision repair removes bulky, helix-distorting lesions: pyrimidine dimers, 6-4 photoproducts, chemical adducts, cisplatin intrastrand crosslinks. It does not recognise the chemistry of the lesion at all. It recognises the distortion, which is why one pathway handles a chemically diverse set of substrates.

There are two ways in.

  • Global genome repair surveys the whole genome. The XPC-RAD23B complex detects the destabilised base pairing that a bulky lesion creates.
  • Transcription-coupled repair is triggered when RNA polymerase II stalls at a lesion in a transcribed strand. CSA and CSB, plus UVSSA, remodel or displace the stalled polymerase and recruit the same downstream machinery. This route is faster, and it prioritises the strand that is actually being read.

From there the pathways converge. TFIIH, the same ten-subunit complex that serves transcription initiation, is recruited; its XPB and XPD helicase subunits open about 25 base pairs around the lesion and XPD verifies that a genuine lesion is present by scanning the strand. XPA confirms the geometry, RPA coats the undamaged strand, and then two structure-specific nucleases cut: XPF-ERCC1 incises 5 prime to the lesion and XPG incises 3 prime. The excised oligonucleotide is roughly 24 to 32 nucleotides long. Pol delta, epsilon or kappa fills the gap using the intact strand, and ligase seals it.

The XP proteins are named for the complementation groups defined by fusing cells from different patients: fuse XP-A cells with XP-C cells and repair is restored, because each supplies what the other lacks; fuse two XP-A lines and it is not. That genetic logic, done before any of the genes were cloned, is how the pathway was decomposed into steps.

Two diseases, one pathway, opposite outcomes

Here is the puzzle worth sitting with. Xeroderma pigmentosum and Cockayne syndrome are both nucleotide excision repair disorders. Xeroderma pigmentosum patients develop skin cancers in childhood. Cockayne syndrome patients develop severe growth failure, progressive neurological degeneration, cataracts and a characteristic aged appearance, and they do not have a markedly raised cancer risk at all.

The difference tracks which branch is broken.

Xeroderma pigmentosumCockayne syndrome
GenesXPA to XPG; XPV encodes Pol etaERCC6 (CSB), ERCC8 (CSA)
Branch lostGlobal genome repair, usually both branchesTranscription-coupled repair only
Lesions persist inThe whole genome, including replicating DNATranscribed strands
Cellular outcomeMutations fixed at replicationPersistent transcriptional arrest, then apoptosis
Clinical outcomeMassive skin cancer excessDevelopmental and neurodegenerative disease, little cancer excess

Unrepaired lesions cause cancer when a replication fork copies past them and installs a mutation. Cells with only transcription-coupled repair lost do repair their genome globally, so replication is comparatively safe; what they cannot do is clear a lesion out of the way of a stalled polymerase, and a permanently stalled polymerase is a potent apoptotic signal. A cell that dies does not become a tumour. Cancer risk and degeneration risk are therefore not two severities of the same thing; they are what happens when you break repair in front of a polymerase versus in front of a replisome.

The eighth complementation group, XP variant, is different again. These patients have completely normal nucleotide excision repair. What they lack is DNA polymerase eta, the translesion polymerase that copies accurately across a cyclobutane pyrimidine dimer. Without it, a different and error-prone polymerase does the job, and the mutation load rises. Same clinical syndrome, entirely different mechanism, which is a useful warning against reasoning backwards from phenotype to pathway.

Common misconceptions

  • "DNA damage is mainly caused by the environment." Spontaneous hydrolysis and oxidation dominate the daily load. Roughly 10,000 depurinations per cell per day happen with no external agent at all.
  • "Damage and mutation are the same thing." A lesion is chemistry on the DNA and is reversible. A mutation is a change in sequence that has been copied and is not. Most damage never becomes mutation.
  • "Nucleotide excision repair recognises specific chemical adducts." It recognises helix distortion, which is why one pathway handles ultraviolet dimers, cisplatin crosslinks and polycyclic aromatic adducts alike.
  • "Any repair defect raises cancer risk." Cockayne syndrome shows otherwise. If the consequence of the defect is cell death rather than mutation, the cancer risk does not rise.
  • "MGMT is an enzyme that turns over." It transfers the methyl group to itself and is consumed. Stoichiometry, not catalysis.

What to carry forward

  • The spontaneous load is large and continuous: about 10,000 depurinations, hundreds of deaminations and thousands of oxidative lesions per cell per day.
  • Direct reversal is cheapest: photolyase for pyrimidine dimers, absent in placental mammals, and MGMT as a single-use suicide protein whose promoter methylation predicts temozolomide response.
  • Base excision repair is glycosylase-initiated and lesion-specific: UNG for uracil, OGG1 for 8-oxoguanine, MUTYH for the adenine opposite it, then APE1, Pol beta and ligase III with XRCC1.
  • MUTYH-associated polyposis produces exactly the G to T transversion signature the missing enzyme prevents.
  • Nucleotide excision repair detects distortion, not chemistry, through global genome and transcription-coupled routes converging on TFIIH, XPA, RPA and the XPF-ERCC1 and XPG incisions, removing 24 to 32 nucleotides.
  • Losing global repair gives cancer; losing only transcription-coupled repair gives degeneration, because stalled transcription kills the cell before it can mutate.
  • XP variant has intact excision repair and lacks polymerase eta, which is a different mechanism reaching a similar phenotype.

Sources

  1. Cleaver, J. E. (1968). Defective repair replication of DNA in xeroderma pigmentosum. Nature, 218(5142), 652-656. pubmed.ncbi.nlm.nih.gov
  2. Lindahl, T. (1993). Instability and decay of the primary structure of DNA. Nature, 362(6422), 709-715. pubmed.ncbi.nlm.nih.gov
  3. Alberts, B., Johnson, A., Lewis, J., Raff, M., Roberts, K., & Walter, P. (2002). DNA repair. In Molecular biology of the cell (4th ed.). Garland Science. ncbi.nlm.nih.gov
  4. Kraemer, K. H., DiGiovanna, J. J., & Tamura, D. (2022). Xeroderma pigmentosum. In M. P. Adam et al. (Eds.), GeneReviews. University of Washington, Seattle. ncbi.nlm.nih.gov
  5. Laugel, V. (2024). Cockayne syndrome. In M. P. Adam et al. (Eds.), GeneReviews. University of Washington, Seattle. ncbi.nlm.nih.gov
  6. The Nobel Foundation. (2015). The Nobel Prize in Chemistry 2015: Tomas Lindahl, Paul Modrich and Aziz Sancar. nobelprize.org
Key terms
Depurination
Spontaneous hydrolysis of the glycosidic bond releasing a purine base and leaving an abasic site; about 10,000 events per human cell per day.
8-oxoguanine
The commonest mutagenic oxidative lesion, which pairs with adenine as well as cytosine and therefore drives G to T transversions.
DNA glycosylase
An enzyme that flips a damaged base out of the helix and cleaves its glycosidic bond, initiating base excision repair.
MGMT
O6-methylguanine-DNA methyltransferase, a suicide protein that transfers a methyl group to its own cysteine and is then destroyed.
Global genome repair
The nucleotide excision repair branch initiated by XPC-RAD23B detecting helix distortion anywhere in the genome.
Transcription-coupled repair
The nucleotide excision repair branch triggered by RNA polymerase II stalling at a lesion, requiring CSA and CSB.
Complementation group
A class defined by cell fusion: cells from two patients that restore function when fused are in different groups and therefore have different defective genes.
Polymerase eta
The translesion polymerase that copies accurately across cyclobutane pyrimidine dimers; its loss causes xeroderma pigmentosum variant.

Double-Strand Breaks: End Joining, Homologous Recombination and the PARP Trap

  • Compare non-homologous end joining and homologous recombination by their requirements, accuracy and cell cycle availability.
  • Trace the resection and strand invasion steps of homologous recombination and the factors that decide which pathway is used.
  • Explain the synthetic lethality between PARP inhibition and BRCA deficiency, and the routes by which tumours escape it.

A result that looked like a mistake

On 14 April 2005 two independent groups published back to back in Nature. Both had taken cells lacking BRCA1 or BRCA2 and treated them with a small molecule inhibiting poly(ADP-ribose) polymerase. Both found the same thing: the BRCA-deficient cells were killed at concentrations up to a thousandfold lower than the concentrations that troubled their otherwise identical wild-type counterparts.

On the face of it this makes no sense. PARP1 works on single-strand breaks. BRCA1 and BRCA2 work on double-strand breaks. Inhibiting a repair enzyme in one pathway should not selectively annihilate cells that have lost a different pathway. That the result held up, and became the first approved cancer therapy designed on the principle of synthetic lethality, means the two pathways are connected in a way that the textbook diagram does not show. Working out how is the spine of this lesson.

Why a double-strand break is a different category of problem

Every repair pathway in the previous lesson depends on the same thing: an intact complementary strand to copy from. A double-strand break removes that. There is no template on the broken molecule, and the two ends can physically drift apart, which is why a single unrepaired double-strand break is enough to trigger permanent arrest or apoptosis in a mammalian cell.

Breaks arise from ionising radiation and radiomimetic drugs, from topoisomerase II reactions that fail to religate (which is exactly how etoposide kills), and above all from replication: a fork that runs into a nick generates a one-ended break, and a collapsed fork must be rebuilt. Some breaks are made deliberately, by RAG1 and RAG2 during V(D)J recombination, by SPO11 in meiosis, and by AID-initiated processing during antibody class switching. The immunology course in this catalogue treats the first and third of those; here they matter as evidence that the cell will risk a double-strand break when it needs one.

Detection is fast. The MRN complex, MRE11 with RAD50 and NBS1, binds the ends and activates the ATM kinase. ATM phosphorylates the histone variant H2AX on serine 139 across up to a megabase of flanking chromatin, producing the gamma-H2AX foci that you can count under a microscope one hour after irradiation. MDC1 binds gamma-H2AX and amplifies the signal; 53BP1 and BRCA1 arrive and begin the argument about which pathway will be used.

Non-homologous end joining: fast, always available, occasionally lossy

The Ku70-Ku80 heterodimer is a ring that threads onto a broken DNA end within seconds. It is present at extremely high abundance, it binds ends without regard to sequence, and it recruits the catalytic subunit DNA-PKcs, forming the DNA-PK holoenzyme that tethers the two ends together.

If the ends are chemically clean and compatible, XRCC4 with DNA ligase IV and XLF simply ligates them, and the repair is accurate. If they are not, and after ionising radiation they usually are not, end-processing enzymes act first: Artemis trims overhangs and hairpins, polynucleotide kinase phosphatase fixes the chemistry at the termini, and polymerases mu and lambda fill short gaps without a proper template. Each of those steps can lose or add a few nucleotides. So end joining is not intrinsically error-prone; it becomes error-prone in proportion to how damaged the ends are.

End joining works in every phase of the cell cycle, needs no partner molecule, and is the dominant pathway in G1 and in non-dividing cells, which is most of the cells in an adult human.

Key idea: End joining is not a lower-quality version of recombination. It is the only option available when there is no sister chromatid, which is most of the time in most tissues.

Homologous recombination: accurate, but only in S and G2

Homologous recombination copies the missing information from an identical sequence, in practice the sister chromatid, which exists only after replication. That single constraint explains its cell cycle restriction.

  1. Resection. MRN with CtIP nicks and trims the 5 prime strands, then EXO1 or the BLM-DNA2 pair extends resection for hundreds to thousands of nucleotides. The product is long 3 prime single-stranded tails.
  2. RPA coating. RPA binds the single-stranded DNA, removing secondary structure. It also activates ATR, which is why resection converts an ATM signal into an ATR signal.
  3. RAD51 loading. RPA has to be replaced by RAD51 to make a strand-invasion filament, and it does not leave willingly. BRCA2 is the mediator that does this: it binds RAD51 through its eight BRC repeats and delivers it onto RPA-coated single-stranded DNA. That is BRCA2's principal molecular job, and it is why BRCA2 loss abolishes homologous recombination.
  4. Strand invasion. The RAD51 nucleoprotein filament searches the genome for homology, invades the sister duplex and displaces one strand, forming a displacement loop.
  5. Synthesis and resolution. A polymerase extends the invading 3 prime end using the sister as template. In synthesis-dependent strand annealing the extended strand is then displaced and anneals back to the other broken end, giving a non-crossover product; this is the usual outcome in somatic cells. Alternatively both ends engage, producing a double Holliday junction, which can be dissolved by BLM helicase with topoisomerase III alpha to give non-crossovers, or resolved by nucleases to give crossovers.

Somatic cells work hard to avoid crossovers, because a crossover between homologous chromosomes rather than sisters produces loss of heterozygosity. Bloom syndrome makes the point: loss of BLM helicase removes the dissolution route, and the diagnostic laboratory finding is a tenfold elevation in sister chromatid exchanges, with a corresponding broad cancer predisposition.

Who decides which pathway

The choice is made at the step of resection, because once long single-stranded tails exist, Ku cannot load and end joining is off the table.

  • 53BP1, with its downstream shieldin complex, blocks resection and therefore channels repair to end joining.
  • BRCA1, with CtIP, promotes resection and therefore channels repair to homologous recombination. BRCA1 antagonises 53BP1 directly.
  • Cyclin-dependent kinase activity licenses the whole thing: CDK phosphorylation of CtIP and EXO1 in S and G2 is what permits resection at all, which is how the pathway choice is tied to whether a sister chromatid exists.

There is a third route worth naming because it explains drug resistance. Microhomology-mediated end joining, also called alternative end joining or theta-mediated end joining, uses PARP1 for end synapsis and DNA polymerase theta to anneal short microhomologies revealed by limited resection. It always leaves a deletion, and it is heavily used by homologous-recombination-deficient tumours, which makes polymerase theta an active drug target.

Back to the 2005 result

Now the synthetic lethality can be assembled. PARP1 detects single-strand breaks and base excision repair intermediates and signals for their repair. Inhibit it and unrepaired single-strand breaks persist into S phase. When a replication fork meets a nick, the fork collapses and generates a one-ended double-strand break. In a normal cell homologous recombination restarts that fork. In a BRCA1 or BRCA2 mutant cell it cannot, so the break is either left or repaired by microhomology-mediated end joining with a deletion, and after enough of these the cell dies.

That is the classical account and it is broadly right, but the more important mechanism turned out to be different. Catalytically inhibited PARP1 does not simply stop working; it becomes trapped on the DNA, because the inhibitor blocks the auto-PARylation that normally releases it. A trapped PARP1-DNA complex is itself a replication block, and clinical PARP inhibitors differ far more in trapping potency than in catalytic potency. Talazoparib traps strongly and is used at very low doses; veliparib traps weakly and disappointed in trials that assumed catalytic inhibition was what mattered.

Why this matters: The drugs were designed on one mechanism and work mainly by another. The clinical dosing only makes sense once you know that.

Tumours escape, and the escape routes are diagnostic of the mechanism. Secondary reversion mutations restore the BRCA reading frame. Loss of 53BP1 or shieldin partially restores resection and therefore recombination even without BRCA1. Upregulation of drug efflux pumps removes the inhibitor. Each of these is now looked for in resistant disease.

The disorders that mapped the pathway

DisorderGeneStep lostHallmark
Ataxia-telangiectasiaATMBreak signallingCerebellar degeneration, ocular telangiectasia, radiosensitivity, lymphoid malignancy, raised alpha-fetoprotein
Nijmegen breakage syndromeNBNMRN sensing and resectionMicrocephaly, immunodeficiency, lymphoma
LIG4 syndromeLIG4End joining ligationMicrocephaly, marrow failure, combined immunodeficiency
Fanconi anaemiaFANC genes, including BRCA2 as FANCD1Interstrand crosslink repair feeding into recombinationMarrow failure, radial ray defects, chromosome breakage on diepoxybutane testing
Bloom syndromeBLMHolliday junction dissolutionGrowth restriction, sun-sensitive facial erythema, tenfold excess sister chromatid exchanges

The Fanconi entry deserves a note. Biallelic BRCA2 mutation does not give a mild BRCA phenotype; it gives Fanconi anaemia complementation group D1, with childhood cancers. Heterozygous carriers have the familiar adult breast, ovarian, prostate and pancreatic risk. One gene, two dose-dependent syndromes.

Common misconceptions

  • "Non-homologous end joining is inherently sloppy." Ligation of two clean, compatible ends is accurate. Nucleotide loss comes from the processing that dirty ends require, and from the fact that damaged ends are the common case.
  • "Homologous recombination uses the homologous chromosome." In somatic cells it overwhelmingly uses the sister chromatid, and the machinery actively suppresses crossovers to avoid loss of heterozygosity.
  • "PARP inhibitors work by blocking base excision repair." The dominant mechanism is trapping PARP1 on DNA, which is why clinical potency tracks trapping rather than catalytic inhibition.
  • "BRCA1 and BRCA2 do the same thing." BRCA1 acts early, promoting resection against 53BP1. BRCA2 acts later, loading RAD51 onto RPA-coated DNA. Their mutation spectra and tumour profiles differ accordingly.
  • "A gamma-H2AX focus is a double-strand break." It is a chromatin domain of up to a megabase marking one break. Foci count breaks, but focus size says nothing about lesion size.

Looking back

  • A double-strand break removes the template that every other repair pathway relies on, so it must be either rejoined or copied from a sister.
  • MRN senses the break and activates ATM, which spreads gamma-H2AX over up to a megabase of flanking chromatin.
  • End joining uses Ku70-Ku80, DNA-PKcs, Artemis and XRCC4-LIG4-XLF; it works in every cell cycle phase and loses nucleotides only in proportion to how damaged the ends are.
  • Homologous recombination requires resection, RPA coating, BRCA2-mediated RAD51 loading, strand invasion, and resolution that usually avoids crossover.
  • Pathway choice is decided at resection: 53BP1 and shieldin block it, BRCA1 and CtIP promote it, and CDK activity permits it only in S and G2.
  • PARP inhibition kills BRCA-deficient cells because trapped PARP1 blocks replication and the resulting one-ended breaks cannot be repaired by recombination.
  • Resistance arises through BRCA reversion, 53BP1 or shieldin loss, and efflux, each of which confirms the mechanism.

Sources

  1. Farmer, H., McCabe, N., Lord, C. J., Tutt, A. N. J., Johnson, D. A., Richardson, T. B., et al. (2005). Targeting the DNA repair defect in BRCA mutant cells as a therapeutic strategy. Nature, 434(7035), 917-921. pubmed.ncbi.nlm.nih.gov
  2. Bryant, H. E., Schultz, N., Thomas, H. D., Parker, K. M., Flower, D., Lopez, E., et al. (2005). Specific killing of BRCA2-deficient tumours with inhibitors of poly(ADP-ribose) polymerase. Nature, 434(7035), 913-917. pubmed.ncbi.nlm.nih.gov
  3. Alberts, B., Johnson, A., Lewis, J., Raff, M., Roberts, K., & Walter, P. (2002). General recombination. In Molecular biology of the cell (4th ed.). Garland Science. ncbi.nlm.nih.gov
  4. Veenhuis, S., van Os, N., Weemaes, C., Kamsteeg, E.-J., & Willemsen, M. (2025). Ataxia-telangiectasia. In M. P. Adam et al. (Eds.), GeneReviews. University of Washington, Seattle. ncbi.nlm.nih.gov
  5. Mehta, P. A., & Ebens, C. (2026). Fanconi anemia. In M. P. Adam et al. (Eds.), GeneReviews. University of Washington, Seattle. ncbi.nlm.nih.gov
  6. Langer, K., Cunniff, C. M., & Kucine, N. (2023). Bloom syndrome. In M. P. Adam et al. (Eds.), GeneReviews. University of Washington, Seattle. ncbi.nlm.nih.gov
Key terms
MRN complex
MRE11, RAD50 and NBS1, the sensor that binds double-strand break ends, activates ATM and initiates resection.
Gamma-H2AX
Histone H2AX phosphorylated on serine 139 by ATM, spreading over up to a megabase around a break and visible as a nuclear focus.
Ku70-Ku80
The abundant ring-shaped heterodimer that binds broken DNA ends within seconds and commits repair to non-homologous end joining.
Resection
Nucleolytic removal of 5 prime strands at a break to expose 3 prime single-stranded tails, the committed step toward homologous recombination.
BRCA2
The mediator that loads RAD51 onto RPA-coated single-stranded DNA through its BRC repeats, enabling strand invasion.
53BP1
A break-associated factor that, with shieldin, blocks resection and thereby directs repair toward end joining.
Synthetic lethality
The situation in which loss of either of two genes is tolerated but loss of both is lethal, as with BRCA deficiency plus PARP inhibition.
PARP trapping
The retention of inhibited PARP1 on DNA, which creates a replication block and accounts for most of the cytotoxicity of clinical PARP inhibitors.

Mismatch Repair, and Recombination as a General-Purpose Tool

  • Explain how mismatch repair identifies which of two strands carries the error, in bacteria and in eukaryotes.
  • Connect mismatch repair loss to microsatellite instability, Lynch syndrome, and the response of such tumours to checkpoint blockade.
  • Describe the uses to which cells and experimenters put recombination beyond repair: meiosis, site-specific systems and transposition.

Bands in the wrong place

In May 1993 Stephen Thibodeau and colleagues published a simple experiment. They amplified short dinucleotide repeat tracts from matched tumour and normal tissue in 90 colorectal cancers and ran the products side by side. In most pairs the tumour bands sat exactly where the normal bands did. In 28 of the 90 they did not: the tumour had gained or lost repeat units, producing extra bands at shifted positions.

Repeat tracts change length only when the polymerase slips and the slippage is not corrected. A tumour whose repeats have wandered has lost the correction step. Within months the responsible genes were identified as human homologues of the bacterial mutator genes mutS and mutL, and a hereditary cancer syndrome that had been described clinically since Aldred Warthin's family studies in 1913 acquired a molecular cause.

The strand discrimination problem

Mismatch repair has to solve something the other pathways do not face. When base excision repair meets a uracil, it knows uracil does not belong in DNA. When mismatch repair meets a G opposite a T, both bases are legitimate. One of them is a replication error and the other is the original. Removing the wrong one converts an error into a fixed mutation, which is worse than doing nothing.

So the pathway must identify the newly synthesised strand, and the two solutions are different.

In Escherichia coli, the Dam methyltransferase methylates adenine in GATC sequences, but it lags a minute or two behind the fork. Immediately after replication the parental strand is methylated and the daughter is not, so the duplex is transiently hemimethylated. MutS binds the mismatch, MutL couples it to MutH, and MutH nicks the unmethylated strand at the nearest GATC. Exonucleases then degrade from the nick past the mismatch and the gap is refilled. The methyl mark is the strand label, and the mechanism is called methyl-directed mismatch repair for that reason.

In eukaryotes, there is no Dam and no MutH. The label is the discontinuity of the new strand itself: the nicks between Okazaki fragments on the lagging strand, and the 3 prime terminus on the leading strand. The MutL alpha complex, MLH1 with PMS2, carries a latent endonuclease that is activated by PCNA, and because PCNA is loaded at a nick with a defined orientation, it directs incision onto the nascent strand. Strand discrimination is a consequence of replication geometry rather than of a chemical mark.

The recognition step is split by lesion size:

  • MutS alpha (MSH2 with MSH6) recognises single base mismatches and one to two base insertion or deletion loops.
  • MutS beta (MSH2 with MSH3) recognises larger loops, up to about a dozen bases.

This division explains a clinical detail. MSH6 mutation gives a milder, later-onset phenotype than MSH2 mutation, because MSH2 is a partner in both complexes and MSH6 only in one; loss of MSH6 leaves MutS beta intact.

In short: Mismatch repair is the only repair pathway that has to decide which correct-looking base is wrong, and everything about its architecture follows from that.

Lynch syndrome, and what microsatellite instability tells you

Lynch syndrome is autosomal dominant germ line loss of one allele of MLH1, MSH2, MSH6 or PMS2, or a deletion of the 3 prime end of EPCAM that silences the adjacent MSH2 by read-through methylation. The tumour arises when the second allele is lost. Lifetime colorectal cancer risk is high but gene-dependent, roughly 40 to 60 percent for MLH1 and MSH2 carriers and considerably lower for PMS2; endometrial cancer risk in women is comparable to the colorectal risk, and the syndrome also raises risk of ovarian, gastric, small bowel, urothelial and sebaceous tumours.

Two laboratory tests screen for it, and they answer slightly different questions. Immunohistochemistry for the four proteins asks which one is missing, and because MLH1 and PMS2 are obligate partners, loss of MLH1 takes PMS2 down with it, giving a paired staining pattern that points at the gene. Microsatellite instability testing, by PCR or now from sequencing data, asks whether the functional consequence is present.

A positive result is not a diagnosis. About 12 to 15 percent of sporadic colorectal cancers are microsatellite unstable, almost always because of somatic MLH1 promoter hypermethylation rather than a germ line mutation, and those tumours frequently carry the BRAF V600E mutation. So the reflex algorithm is: absent MLH1 on immunohistochemistry leads to BRAF and MLH1 methylation testing, and only if both are negative does germ line testing follow.

There is a striking therapeutic consequence. A mismatch repair deficient tumour accumulates thousands of mutations, many of them frameshifts in coding repeat tracts, and frameshifts produce entirely novel peptide sequences. In 2015 Dung Le and colleagues reported that pembrolizumab, an antibody blocking PD-1, produced objective responses in 40 percent of mismatch repair deficient colorectal cancers and in none of the mismatch repair proficient ones. That result led in 2017 to the first tissue-agnostic drug approval in oncology: a marker, not an organ, defines eligibility. The immunology course in this catalogue takes up checkpoint blockade in detail; the point here is that a repair defect changed a cancer's visibility to T cells.

Biallelic germ line mismatch repair loss is a separate and much more severe entity, constitutional mismatch repair deficiency, presenting in childhood with cafe-au-lait macules that mimic neurofibromatosis, brain tumours, haematological malignancy and early colorectal cancer.

Recombination is not only for repair

The homologous recombination machinery of the previous lesson gets used for several jobs that have nothing to do with accidental damage. Recognising that changes how you read the pathway.

Meiosis. Recombination in meiosis is not a response to damage; the cell manufactures the damage on purpose. SPO11, a topoisomerase-like protein, introduces on the order of 200 to 300 double-strand breaks per human meiosis. Most are repaired as non-crossovers or gene conversions; a minority mature into crossovers. In most mammals the positions are not random: the zinc finger protein PRDM9 binds a sequence motif, trimethylates nearby H3K4 and H3K36, and directs SPO11 to those hotspots. PRDM9 is one of the fastest-evolving genes in the genome, and its zinc finger array differs between human populations, which is why hotspot locations are not shared between humans and chimpanzees.

Every bivalent must receive at least one crossover, the obligate crossover, because the resulting chiasma is what holds homologues together until anaphase I. Bivalents that fail to get one segregate at random. This is the mechanistic route from a recombination failure to aneuploidy, and it is a substantial contributor to trisomy 21 of maternal origin.

Crossover interference then ensures the crossovers that do occur are spaced apart rather than clustered. The net result in humans is roughly one to three crossovers per bivalent and around 25 to 30 per gamete on the maternal side, with fewer on the paternal side.

Gene conversion. When an invading strand is extended and then returns, a short tract of the donor sequence can be copied into the recipient, non-reciprocally. This is why the products of meiosis at a heterozygous site are occasionally three to one rather than two to two, and it is the mechanism by which pseudogenes corrupt their functional relatives, as happens in the CYP21A2 gene in congenital adrenal hyperplasia.

Site-specific recombination. Bacteriophage lambda integrase joins a specific attP site in the phage to a specific attB site in the E. coli chromosome, with no homology requirement beyond a short core. The same logic gave molecular biology two of its most useful tools: Cre recombinase acting on loxP sites, and FLP acting on FRT sites. A conditional knockout mouse is nothing more than a gene flanked by loxP sites plus Cre expressed from a tissue-specific promoter, and the entire strategy is borrowed from a bacteriophage.

Transposition. Roughly half the human genome derives from transposable elements. DNA transposons move by cut and paste and are largely inactive in humans; retrotransposons move through an RNA intermediate. LINE-1 elements encode their own reverse transcriptase and endonuclease and are the only autonomously active family left in humans, with perhaps a hundred retrotransposition-competent copies per genome; Alu elements are non-autonomous and borrow the LINE-1 machinery. In 1988 Haig Kazazian and colleagues sequenced the factor VIII gene in two unrelated boys with sporadic haemophilia A and found a new LINE-1 insertion disrupting exon 14, absent from both parents. That was the first demonstration that transposition causes human disease in real time.

The upshot: The same strand-exchange chemistry that fixes a break also shuffles alleles in meiosis, integrates a phage, builds a conditional mouse and lets a retroelement insert itself into a clotting factor gene.

Common misconceptions

  • "Mismatch repair fixes damaged bases." Both bases in a mismatch are chemically normal. The pathway corrects replication errors, which is why strand discrimination is its central problem.
  • "Microsatellite instability means Lynch syndrome." Most microsatellite unstable colorectal cancers are sporadic, caused by somatic MLH1 promoter methylation and often carrying BRAF V600E.
  • "Eukaryotes use methylation to mark the new strand." That is the bacterial solution. Eukaryotes use strand discontinuities and PCNA orientation.
  • "Meiotic recombination happens because DNA gets damaged during meiosis." SPO11 makes the breaks deliberately, in the hundreds, and their repair as crossovers is required for correct segregation.
  • "Transposons are junk that no longer moves." LINE-1 remains active in humans and causes disease by insertion; roughly one in twenty newborns carries a new retrotransposition event somewhere in the genome.

What you now know

  • Mismatch repair corrects replication errors and must identify the new strand: bacteria use transient hemimethylation of GATC and MutH; eukaryotes use strand nicks and PCNA-directed MutL alpha endonuclease.
  • MutS alpha (MSH2-MSH6) handles small mismatches, MutS beta (MSH2-MSH3) handles larger loops, which explains why MSH6 disease is milder than MSH2 disease.
  • Lynch syndrome arises from germ line loss in MLH1, MSH2, MSH6, PMS2 or EPCAM, with high colorectal and endometrial risk that varies by gene.
  • Immunohistochemistry names the missing protein and microsatellite instability reports the functional consequence; sporadic cases are excluded by BRAF V600E and MLH1 promoter methylation.
  • Mismatch repair deficient tumours carry heavy frameshift neoantigen loads and respond to PD-1 blockade, which produced oncology's first tissue-agnostic approval.
  • Meiosis uses SPO11 to make 200 to 300 deliberate breaks, positioned by PRDM9, with an obligate crossover per bivalent whose failure causes nondisjunction.
  • Site-specific recombination gave us Cre-lox and FLP-FRT; LINE-1 retrotransposition still causes human disease, first shown in haemophilia A in 1988.

Sources

  1. Thibodeau, S. N., Bren, G., & Schaid, D. (1993). Microsatellite instability in cancer of the proximal colon. Science, 260(5109), 816-819. pubmed.ncbi.nlm.nih.gov
  2. Fishel, R., Lescoe, M. K., Rao, M. R., Copeland, N. G., Jenkins, N. A., Garber, J., Kane, M., & Kolodner, R. (1993). The human mutator gene homolog MSH2 and its association with hereditary nonpolyposis colon cancer. Cell, 75(5), 1027-1038. pubmed.ncbi.nlm.nih.gov
  3. Le, D. T., Uram, J. N., Wang, H., Bartlett, B. R., Kemberling, H., Eyring, A. D., et al. (2015). PD-1 blockade in tumors with mismatch-repair deficiency. New England Journal of Medicine, 372(26), 2509-2520. pubmed.ncbi.nlm.nih.gov
  4. Kazazian, H. H., Jr., Wong, C., Youssoufian, H., Scott, A. F., Phillips, D. G., & Antonarakis, S. E. (1988). Haemophilia A resulting from de novo insertion of L1 sequences represents a novel mechanism for mutation in man. Nature, 332(6160), 164-166. pubmed.ncbi.nlm.nih.gov
  5. Alberts, B., Johnson, A., Lewis, J., Raff, M., Roberts, K., & Walter, P. (2002). Site-specific recombination. In Molecular biology of the cell (4th ed.). Garland Science. ncbi.nlm.nih.gov
  6. Brown, T. A. (2002). Mutation, repair and recombination. In Genomes (2nd ed., Chapter 14). Wiley-Liss. ncbi.nlm.nih.gov
Key terms
Strand discrimination
The step by which mismatch repair identifies which strand carries the replication error, since both bases in a mismatch are chemically normal.
Methyl-directed mismatch repair
The bacterial system in which transient hemimethylation of GATC sites lets MutH nick the unmethylated nascent strand.
MutS alpha
The MSH2-MSH6 heterodimer that recognises single base mismatches and very small insertion or deletion loops.
Microsatellite instability
Length variation of short repeat tracts between tumour and normal DNA, the functional signature of mismatch repair loss.
Lynch syndrome
Dominantly inherited cancer predisposition from germ line loss of MLH1, MSH2, MSH6, PMS2 or EPCAM, with high colorectal and endometrial risk.
SPO11
The topoisomerase-like protein that deliberately introduces the several hundred double-strand breaks that initiate meiotic recombination.
Obligate crossover
The requirement that every bivalent receive at least one crossover, whose chiasma holds homologues together until anaphase I.
LINE-1
The only autonomously active human retrotransposon family, encoding its own reverse transcriptase and endonuclease.

Module 4: Transcription and its Control

What RNA polymerases actually do, step by step, in bacteria and in eukaryotes, and the regulatory grammar of promoters, enhancers and the proteins that read them.

Bacterial RNA Polymerase and the Logic of the Operon

  • Describe the bacterial transcription cycle from promoter recognition through termination, naming the subunit responsible at each step.
  • Explain negative and positive control of the lac operon and why the two together produce a logical AND gate.
  • Contrast repression, attenuation and riboswitching as mechanisms for regulating a bacterial operon.

A growth curve with a step in it

Jacques Monod grew Escherichia coli in a medium containing both glucose and lactose and plotted optical density against time. The culture grew exponentially, then stopped for about an hour, then grew exponentially again. He named the pattern diauxie. Nothing had been added and nothing had run out except the glucose. During the pause, the cells were building an enzyme they had not previously needed.

The obvious explanation, that lactose selects for pre-existing mutants, is wrong: the shift happens in the whole population within an hour, far too fast and far too uniform for selection. What the pause represents is a genetic programme being switched on. Twenty years of work on that pause produced the operon model, the first mechanistic account of how a gene is turned on and off, and the 1965 Nobel Prize for Jacob, Lwoff and Monod.

The enzyme, subunit by subunit

Bacterial RNA polymerase core enzyme is five polypeptides: two copies of alpha, one beta, one beta prime, and one omega. Its shape is often described as a crab claw, with the two large subunits forming the pincers and the catalytic magnesium ion held at their base. The core can elongate RNA but it cannot find a promoter.

For that it needs a sigma factor. Core plus sigma is the holoenzyme, and sigma does three things: it recognises promoter sequences, it lowers non-specific DNA binding so the enzyme can slide, and it helps melt the duplex. The housekeeping sigma in E. coli is sigma-70, which reads two hexamers upstream of the start site: TTGACA centred about 35 base pairs upstream, and TATAAT centred about 10 base pairs upstream, separated by a spacer of 17 base pairs give or take one.

The consensus is worth taking seriously as a quantitative object. Almost no real promoter matches it exactly. The closer a promoter comes to consensus, the more strongly the holoenzyme binds and the more often it initiates; promoter strength is a continuous variable encoded in how far the sequence deviates. That is a very different design from a switch.

Key idea: Sigma is not part of the catalytic machine. It is an interchangeable targeting module, and swapping it redirects the same polymerase to an entirely different set of genes.

Bacteria exploit that. Sigma-32 is induced by heat and directs the polymerase to chaperone and protease genes. Sigma-54 requires an activator that hydrolyses ATP and controls nitrogen metabolism. Bacillus subtilis sporulation runs on an ordered cascade of sigma factors, each one switching on the genes for the next stage. One polymerase, many programmes.

The transcription cycle

  1. Closed complex. Holoenzyme binds double-stranded promoter DNA.
  2. Open complex. About 13 base pairs melt around the start site, forming the transcription bubble. This step is often rate limiting and is where most regulation acts.
  3. Abortive initiation. The enzyme synthesises and releases short RNAs of 2 to 10 nucleotides, repeatedly, while remaining anchored to the promoter by sigma. This looks like failure and is in fact routine; the enzyme scrunches downstream DNA into itself to reach the first phosphodiester bonds without letting go.
  4. Promoter escape. Once the transcript exceeds roughly 10 nucleotides, the sigma contacts break, sigma is released, and the enzyme commits to elongation.
  5. Elongation. At about 50 to 100 nucleotides per second, with a bubble of about 13 base pairs and an RNA-DNA hybrid of 8 or 9 base pairs travelling with it.
  6. Termination. Two mechanisms. In intrinsic termination, the RNA folds into a GC-rich hairpin followed by a run of uridines; the hairpin destabilises the complex and the rU-dA hybrid is the weakest possible, so the transcript falls off. In Rho-dependent termination, the Rho helicase loads onto a C-rich unstructured stretch of the RNA, translocates along it using ATP, catches the paused polymerase and pulls the transcript out.

One structural fact separates bacteria from eukaryotes more than any other: there is no nuclear envelope, so ribosomes load onto the 5 prime end of an mRNA while its 3 prime end is still being made. Transcription and translation are physically coupled, which is why bacterial mRNAs can be polycistronic and why a mechanism like attenuation can work at all.

The lac operon as a logic gate

The lac operon is three genes in one transcription unit: lacZ (beta-galactosidase), lacY (lactose permease) and lacA (a transacetylase). One promoter, one mRNA, three proteins, stoichiometrically produced. Upstream sits lacI, encoding a repressor that is made constitutively.

Negative control. The LacI tetramer binds the operator, which overlaps the promoter, and blocks transcription. Allolactose, an isomer of lactose produced as a side reaction by beta-galactosidase, binds LacI and changes its conformation so that it releases the operator. Note the bootstrapping problem this creates: you need beta-galactosidase to make the inducer that switches on beta-galactosidase. It works because repression is leaky, so a few molecules of enzyme are always present.

Positive control. Removing the repressor is not sufficient. The lac promoter is intrinsically weak. It needs the catabolite activator protein, CAP, also called CRP, which binds cAMP and then binds a site just upstream of the promoter, where it bends the DNA by about 90 degrees and recruits the polymerase through direct contact with an alpha subunit. Glucose transport lowers cAMP; no glucose means high cAMP, means CAP is loaded, means the promoter is strong.

GlucoseLactoseLacI on operatorCAP-cAMP boundTranscription
PresentAbsentYesNoOff
PresentPresentNoNoVery low
AbsentAbsentYesYesOff
AbsentPresentNoYesHigh

Read the table as a logical AND: lactose present AND glucose absent. Neither input alone produces meaningful expression. This is the diauxic pause, at the level of DNA: while glucose remains, cAMP is low and CAP cannot help, so even in the presence of lactose the operon stays nearly silent.

The point: Bacterial promoters are routinely governed by two or more independent inputs whose combination, not whose sum, sets the output. That principle scales all the way to metazoan enhancers.

Attenuation: regulation by ribosome position

The trp operon encodes tryptophan biosynthesis and is repressed by a tryptophan-activated repressor, but that is only part of the control. Between the promoter and the first structural gene lies a leader region encoding a short peptide with two adjacent tryptophan codons, and four segments of RNA that can pair in alternative ways: segments 3 and 4 form a terminator hairpin, while segments 2 and 3 form an antiterminator that prevents it.

Because translation is coupled to transcription, the position of the leading ribosome decides which structure forms.

  • Tryptophan scarce. The ribosome stalls at the two tryptophan codons in segment 1, leaving segment 2 free to pair with segment 3. The terminator cannot form. Transcription proceeds into the structural genes.
  • Tryptophan plentiful. The ribosome reads straight through the leader peptide and physically covers segment 2. Segment 3 pairs with segment 4 instead, the terminator forms, and transcription stops before the structural genes.

The cell is measuring charged tryptophanyl-tRNA concentration by using a ribosome as the sensor. No dedicated regulatory protein is involved.

Riboswitches take the same idea further and remove the ribosome as well. A riboswitch is a structured element in the 5 untranslated region that binds a small molecule directly, such as thiamine pyrophosphate, flavin mononucleotide, S-adenosylmethionine or guanine, and changes fold on binding to expose or occlude a terminator or a ribosome binding site. RNA is doing sensing, computation and control with no protein at all, which is one of the better arguments for RNA-based regulation being ancient.

Remember: Bacteria regulate at initiation with proteins, during elongation with attenuation, and at both with RNA structure. When you meet a new operon, ask which of the three you are looking at before assuming it is a repressor.

Common misconceptions

  • "Lactose is the inducer of the lac operon." Allolactose is, produced from lactose by the enzyme the operon encodes. The laboratory workhorse IPTG is a non-hydrolysable analogue of allolactose, which is precisely why it does not get consumed.
  • "Removing the repressor turns the operon on." It removes the brake. Without cAMP-CAP the promoter is too weak to give substantial expression.
  • "Bacteria have no positive regulation." CAP, sigma-54 activators and dozens of activator families are positive regulators; the lac operon is simply the most famous negative one.
  • "Sigma factor is part of the catalytic core." It is a dissociable specificity subunit released shortly after promoter escape, which is what allows one core enzyme to serve many regulons.
  • "Attenuation is a kind of repression." Repression blocks initiation. Attenuation acts after initiation, on a polymerase that has already started, and it works only because translation is coupled to transcription.

The takeaway

  • Core polymerase (two alpha, beta, beta prime, omega) elongates; sigma confers promoter specificity and is released after escape.
  • Sigma-70 reads the TTGACA at minus 35 and TATAAT at minus 10; deviation from consensus tunes promoter strength continuously.
  • The cycle runs closed complex, open complex, abortive initiation, promoter escape, elongation at 50 to 100 nucleotides per second, then intrinsic or Rho-dependent termination.
  • Transcription and translation are physically coupled in bacteria, which enables polycistronic messages and attenuation.
  • The lac operon is an AND gate: LacI must be released by allolactose and CAP-cAMP must be loaded, and the diauxic pause is what that gate looks like in a growth curve.
  • Attenuation at the trp operon uses ribosome stalling to choose between antiterminator and terminator hairpins.
  • Riboswitches perform the same sensing and switching with RNA alone.

Sources

  1. Jacob, F., & Monod, J. (1961). Genetic regulatory mechanisms in the synthesis of proteins. Journal of Molecular Biology, 3, 318-356. pubmed.ncbi.nlm.nih.gov
  2. The Nobel Foundation. (1965). The Nobel Prize in Physiology or Medicine 1965: Francois Jacob, Andre Lwoff and Jacques Monod. nobelprize.org
  3. Alberts, B., Johnson, A., Lewis, J., Raff, M., Roberts, K., & Walter, P. (2002). From DNA to RNA. In Molecular biology of the cell (4th ed.). Garland Science. ncbi.nlm.nih.gov
  4. Clark, M. A., Douglas, M., & Choi, J. (2018). Prokaryotic gene regulation. In Biology 2e (Section 16.2). OpenStax. openstax.org
  5. Clark, M. A., Douglas, M., & Choi, J. (2018). Prokaryotic transcription. In Biology 2e (Section 15.2). OpenStax. openstax.org
Key terms
Sigma factor
A dissociable subunit that confers promoter specificity on bacterial RNA polymerase and is released shortly after promoter escape.
Open complex
The state in which about 13 base pairs of promoter DNA have melted to form the transcription bubble, often the rate-limiting step.
Abortive initiation
Repeated synthesis and release of very short RNAs while the polymerase remains anchored at the promoter.
Intrinsic termination
Termination driven by a GC-rich RNA hairpin followed by a uridine run, without any accessory protein.
Rho-dependent termination
Termination in which the Rho helicase loads on unstructured RNA, translocates and removes the transcript from a paused polymerase.
Operon
A set of genes transcribed from one promoter into a single polycistronic mRNA under shared regulation.
Catabolite activator protein
CAP or CRP, which binds cAMP, bends promoter DNA and recruits polymerase, providing positive control that is high only when glucose is low.
Attenuation
Control after initiation, in which ribosome position on a leader peptide determines whether a terminator hairpin forms in the nascent RNA.

The Eukaryotic Transcription Cycle

  • Distinguish RNA polymerases I, II and III by product, location and alpha-amanitin sensitivity.
  • Assemble the polymerase II preinitiation complex in order and state what each general factor contributes.
  • Explain promoter-proximal pausing and the role of the carboxy-terminal domain in coupling transcription to RNA processing.

A mushroom that kills two days later

Someone eats Amanita phalloides. For six to twelve hours nothing happens. Then violent gastrointestinal illness, which resolves over the next day into an apparent recovery, and the patient feels well. On day three to five, liver failure. The delay and the false recovery are the classic and lethal features of death cap poisoning, and they are entirely explained by the toxin's target.

Alpha-amanitin is a bicyclic octapeptide that binds in the cleft of RNA polymerase II near the bridge helix and slows the enzyme from thousands of nucleotides per minute to a few. Nothing is destroyed. Existing mRNAs continue to be translated and existing proteins continue to work, so the cell functions normally until its messages decay and cannot be replaced. Hepatocytes, with their high secretory protein output, run out first. The clinical latency is the half-life of the transcriptome.

That same differential sensitivity was the tool Robert Roeder and William Rutter used in 1969, when they fractionated nuclear extracts by ion exchange chromatography and resolved three distinct peaks of DNA-dependent RNA polymerase activity. One was unaffected by alpha-amanitin, one was abolished at very low concentrations, and one required much higher concentrations. Three peaks, three enzymes, three jobs.

Three polymerases, three products

EnzymeProductsLocationAlpha-amanitin
Pol IThe 45S pre-rRNA, processed to 28S, 18S and 5.8SNucleolusResistant
Pol IIAll mRNA, most snRNA, miRNA precursors, most lncRNANucleoplasmInhibited at about 1 microgram per litre
Pol IIItRNA, 5S rRNA, U6 snRNA, 7SL RNA, Alu transcriptsNucleoplasmInhibited only at high concentration

Do not read this as three enzymes of equal importance. By mass, Pol I dominates: in a growing cell, rRNA synthesis is roughly 60 percent of all transcription, running from a few hundred tandem rDNA repeats clustered on the acrocentric chromosomes. Pol II makes a small fraction of the RNA by mass and essentially all of the regulatory interest.

Promoter architecture differs correspondingly. Pol III genes carry many of their promoter elements inside the transcribed region: tRNA genes use internal A and B boxes bound by TFIIIC. That is an odd arrangement until you realise it means a tRNA gene carries its own promoter with it wherever it goes, which is exactly what Alu elements exploit to be transcribed after inserting somewhere new.

What Pol II looks like, and what the parts do

Roger Kornberg's laboratory published the structure of yeast Pol II at 2.8 angstrom resolution in 2001, work recognised by the 2006 Nobel Prize in Chemistry. Twelve subunits, arranged around a positively charged cleft that holds the DNA. At the base of the cleft sits the active site with two magnesium ions. A pore beneath it admits nucleoside triphosphates from below rather than through the cleft, and doubles as the exit route for backtracked RNA during proofreading. The bridge helix spans the cleft next to the active site and cycles between straight and bent conformations with each nucleotide addition; the trigger loop closes over a correctly paired substrate and is the main selectivity element. A clamp closes over the DNA to give processivity, and a wall forces the RNA-DNA hybrid to turn and the RNA to exit through a separate channel.

The feature with no counterpart in bacteria is the carboxy-terminal domain of the largest subunit: a tandem repeat of the heptapeptide tyrosine-serine-proline-threonine-serine-proline-serine, 26 repeats in yeast and 52 in humans, extending from the RNA exit channel as a long flexible tail. Its residues are phosphorylated and dephosphorylated in a defined pattern across the transcription cycle, and the pattern is read by the enzymes of RNA processing.

Why this matters: The carboxy-terminal domain is the reason capping, splicing and polyadenylation happen co-transcriptionally and in the right order. It is a moving scaffold whose modification state announces which stage the polymerase has reached.

Building the preinitiation complex

Pol II cannot recognise a promoter unaided, and instead of one sigma factor it needs six general transcription factors, assembled in a canonical order that was worked out in vitro.

  1. TFIID binds first. Its TBP subunit inserts phenylalanine side chains into the minor groove of a TATA box and bends the DNA by about 80 degrees. Its dozen or so TAF subunits contact other core promoter elements and, importantly, allow TFIID to bind promoters that have no TATA box at all.
  2. TFIIA stabilises TBP on DNA and counteracts negative cofactors.
  3. TFIIB bridges TBP and polymerase, and its reader and finger domains help set the exact start site.
  4. Pol II with TFIIF is recruited as a unit; TFIIF suppresses non-specific binding, much as sigma does in bacteria.
  5. TFIIE arrives and recruits the last factor.
  6. TFIIH completes the complex. It is a ten-subunit machine with two activities that matter here: the XPB subunit is a translocase that pumps DNA to open the duplex, and the CDK7 kinase phosphorylates serine 5 of the carboxy-terminal domain repeats.

Notice XPB. This is the same TFIIH, with the same XPB and XPD subunits, that performs nucleotide excision repair in Module 3. A single complex serves initiation and repair, which is why some TFIIH mutations give a repair disease and others give a developmental disorder such as trichothiodystrophy in which the transcriptional function is compromised.

The TATA box itself deserves deflating. It appears in perhaps 10 to 20 percent of human promoters. Most human promoters are CpG island promoters with dispersed start sites, no TATA box, and reliance on TAF contacts and on general factors recruited through sequence-specific activators. If your mental model of a promoter is a TATA box 25 base pairs upstream of a single start site, it fits a minority of the genome.

Mediator, and how an activator reaches the polymerase

A transcription factor bound at an enhancer tens of kilobases away has to communicate with the preinitiation complex. Mediator is the coactivator that does it: about 26 subunits in humans, organised into head, middle and tail modules, with a dissociable kinase module. Activators bind the tail; the head and middle contact Pol II and the general factors, including the carboxy-terminal domain. Mediator does not bind DNA sequence-specifically. It is an adaptor, converting a wide variety of activator surfaces into a common effect on the polymerase.

Pausing: the step that regulation actually uses

For years the assumption was that recruiting the polymerase was the regulated step, as it broadly is in bacteria. Work on the Drosophila heat shock genes upset this. On uninduced hsp70, polymerase was already there, engaged, with about 25 nucleotides of RNA made, and stopped. Heat shock did not recruit it; heat shock released it.

Genome-wide measurement showed this is the rule rather than a curiosity. A large fraction of human genes, including most developmentally regulated ones, carry a polymerase paused 20 to 60 nucleotides downstream of the start site. Two factors hold it there: DSIF (SPT4 and SPT5) and NELF. Release requires P-TEFb, whose CDK9 subunit phosphorylates NELF-E, the SPT5 subunit of DSIF, and serine 2 of the carboxy-terminal domain. NELF leaves, DSIF converts into a positive elongation factor, and the polymerase enters productive elongation.

Two consequences are worth carrying. First, pausing keeps a promoter nucleosome-free and the gene poised, which is how a cell responds to a signal in seconds rather than minutes. Second, it makes P-TEFb a control point that viruses and cancers exploit: HIV Tat recruits P-TEFb to the viral promoter, and BET inhibitors targeting BRD4, which delivers P-TEFb, were developed as anticancer agents on exactly this logic.

Worth holding on to: In metazoans, the regulated step is frequently pause release, not recruitment. A ChIP experiment that finds polymerase at a silent gene has not found a contradiction.

Elongation and termination

Elongating polymerase must traverse nucleosomes. FACT destabilises the nucleosome ahead by removing an H2A-H2B dimer and helps reassemble it behind, while SPT6 acts as an H3-H4 chaperone. The net effect is that chromatin is transiently disassembled and put back, which is why H3K36 methylation deposited in gene bodies during elongation matters: it recruits deacetylases that suppress spurious internal initiation on the reassembled chromatin.

Termination on Pol II genes is coupled to 3 prime end processing. The cleavage and polyadenylation machinery recognises the AAUAAA signal in the emerging transcript and cuts. The 5 prime end of the downstream RNA, still attached to the polymerase, is then degraded by the XRN2 exonuclease, which catches up with the polymerase and dislodges it. That is the torpedo model, and it fits the observation that XRN2 depletion causes read-through for tens of kilobases past the normal end.

In short: Pol II does not stop at a sequence. It stops because a nuclease chases it down after the transcript has been cut away.

Common misconceptions

  • "Most eukaryotic promoters have a TATA box." Only 10 to 20 percent of human promoters do. CpG island promoters with dispersed start sites are the commonest type.
  • "Transcription is regulated at recruitment, as in bacteria." In metazoans, pause release by P-TEFb is frequently the regulated step, and polymerase is often already loaded at a gene that appears silent.
  • "The general transcription factors are analogous to sigma." Functionally they overlap, but sigma is one dissociable subunit while the eukaryotic system uses six multi-subunit complexes plus Mediator, and it accommodates regulation from far away.
  • "Alpha-amanitin kills cells immediately." It stops new mRNA synthesis. Cells die when existing messages and proteins run out, which is why death cap poisoning has a three to five day latency.
  • "Pol II makes most of the RNA in a cell." By mass, ribosomal RNA made by Pol I dominates. Pol II makes most of the regulatory diversity.

Summing up

  • Three nuclear polymerases were resolved by chromatography and distinguished by alpha-amanitin sensitivity: Pol I resistant, Pol II highly sensitive, Pol III intermediate.
  • Pol II is a twelve-subunit enzyme with a two-metal active site, a bridge helix and trigger loop for catalysis and selectivity, and a separate RNA exit channel.
  • The carboxy-terminal domain heptad repeat is phosphorylated on serine 5 at initiation by CDK7 and on serine 2 during elongation by CDK9, coupling capping, splicing and polyadenylation to the right stage.
  • The preinitiation complex assembles as TFIID, TFIIA, TFIIB, Pol II with TFIIF, TFIIE and TFIIH; TFIIH opens the duplex through XPB and is the same complex used in nucleotide excision repair.
  • Mediator adapts sequence-specific activators to the general machinery without binding DNA itself.
  • Promoter-proximal pausing by DSIF and NELF, released by P-TEFb, is a principal regulated step in metazoans and a target for HIV Tat and for BET inhibitors.
  • Termination follows cleavage at the polyadenylation signal, with XRN2 degrading the downstream RNA and dislodging the polymerase.

Sources

  1. Roeder, R. G., & Rutter, W. J. (1969). Multiple forms of DNA-dependent RNA polymerase in eukaryotic organisms. Nature, 224(5216), 234-237. pubmed.ncbi.nlm.nih.gov
  2. Cramer, P., Bushnell, D. A., & Kornberg, R. D. (2001). Structural basis of transcription: RNA polymerase II at 2.8 angstrom resolution. Science, 292(5523), 1863-1876. pubmed.ncbi.nlm.nih.gov
  3. The Nobel Foundation. (2006). The Nobel Prize in Chemistry 2006: Roger D. Kornberg. nobelprize.org
  4. Clark, M. A., Douglas, M., & Choi, J. (2018). Eukaryotic transcription. In Biology 2e (Section 15.3). OpenStax. openstax.org
  5. Cooper, G. M. (2000). Regulation of transcription in eukaryotes. In The cell: A molecular approach (2nd ed.). Sinauer Associates. ncbi.nlm.nih.gov
Key terms
Alpha-amanitin
The death cap toxin that binds near the bridge helix of RNA polymerase II and slows it drastically; its differential potency defined the three nuclear polymerases.
Carboxy-terminal domain
The heptad repeat tail of the largest Pol II subunit, 52 repeats in humans, whose phosphorylation pattern recruits RNA processing machinery.
TFIID
The general factor containing TBP and the TAFs, which recognises core promoter elements and bends TATA DNA by about 80 degrees.
TFIIH
The ten-subunit factor whose XPB translocase opens the promoter and whose CDK7 kinase phosphorylates serine 5 of the carboxy-terminal domain; also the core of nucleotide excision repair.
Mediator
A large coactivator complex that adapts sequence-specific activators to the polymerase and general factors without binding DNA itself.
Promoter-proximal pausing
The arrest of engaged polymerase 20 to 60 nucleotides downstream of the start site by DSIF and NELF, released by P-TEFb.
P-TEFb
The CDK9-containing kinase that phosphorylates NELF, SPT5 and serine 2 of the carboxy-terminal domain to license productive elongation.
Torpedo model
The account of Pol II termination in which XRN2 degrades the downstream transcript after cleavage and dislodges the polymerase.

Promoters, Enhancers and Transcription Factors

  • State the defining experimental properties of an enhancer and how they were established.
  • Describe the main classes of DNA-binding domain and explain how transcription factors achieve specificity despite short recognition motifs.
  • Evaluate what a reporter assay, a ChIP-seq peak and a genomic deletion each do and do not establish about a candidate regulatory element.

A 72 base pair repeat in the wrong place

In 1981 Julian Banerji, Sandro Rusconi and Walter Schaffner were testing what parts of simian virus 40 DNA were needed to express a linked gene. They put a rabbit beta-globin gene into a plasmid with various pieces of SV40 attached and measured transcript levels after transfection. One fragment, a tandem 72 base pair repeat, raised expression by roughly two hundredfold.

That alone would have been unremarkable. What made it a discovery was the control experiments. The fragment worked when placed upstream of the gene and when placed downstream of it. It worked in either orientation. It worked from more than a kilobase away. No promoter element known at the time behaved like that, because promoters are position-dependent and orientation-dependent by definition. Schaffner's group had found a new category of sequence, and they called it an enhancer.

Key idea: An enhancer is defined operationally, by orientation and position independence, not by a sequence signature. Every later method for identifying enhancers is a proxy for that behaviour.

How a protein reads a base pair

Transcription factors distinguish sequences by contacting the edges of base pairs exposed in the major groove, where the pattern of hydrogen bond donors, acceptors and methyl groups differs for each of the four possible base pairs. The minor groove carries much less discriminating information, which is why most sequence-specific recognition happens in the major groove and why nucleosomes, which bind through the minor groove and the backbone, are sequence-tolerant.

A handful of structural solutions dominate.

DomainHow it bindsExamples
Helix-turn-helix and homeodomainA recognition helix lies in the major groove; the homeodomain adds an N-terminal arm in the minor grooveLac repressor, HOX proteins, PIT1
C2H2 zinc fingerEach finger contacts about three base pairs; fingers are strung in arrays for longer sitesSP1, EGR1, CTCF, PRDM9
Basic zipperA coiled-coil dimerises two subunits; the basic regions grip opposite halves of a palindromeCREB, FOS-JUN, ATF4
Basic helix-loop-helixSimilar dimerisation logic, binding E-box sequencesMYC-MAX, MYOD
Nuclear receptor zinc moduleDimers on direct or inverted repeats, with ligand-dependent conformational changeGlucocorticoid receptor, oestrogen receptor

The zinc finger array is worth a second look because it is modular: more fingers means a longer and therefore rarer recognition site. CTCF uses eleven. That modularity is exactly what made zinc fingers the first programmable gene editing platform, before TALENs and then CRISPR displaced them.

The specificity problem, stated with numbers

Suppose a factor recognises a 6 base pair motif. In 3 billion base pairs of haploid sequence, that motif occurs by chance about 3,000,000,000 divided by 4 to the sixth power, which is roughly 700,000 times on each strand. Even a 10 base pair motif occurs about 3,000 times by chance. Yet a typical factor occupies a few thousand sites and regulates far fewer genes than that. Something other than the motif is doing most of the work.

Three things, in fact.

  • Cooperativity. Factors bind in combination. Two proteins each binding weakly but contacting each other bind together far more tightly than either alone, and the combination is much rarer in the genome than either motif. This is why enhancers cluster binding sites for several factors, and why the same factor can activate different genes in different cell types depending on its partners.
  • Accessibility. Most of the genome is wrapped in nucleosomes that occlude the major groove. A factor can only sample the few percent of sites that are in open chromatin, which is why ATAC-seq maps of accessibility predict factor binding better than motif scans do.
  • Pioneer activity. Some factors, notably the FOXA family, can engage their motif on nucleosomal DNA and open the region for others. FOXA1's winged helix fold resembles linker histone H1 closely enough to compete with it. Pioneer factors are how a closed locus gets opened in the first place, and they are the reason a small set of factors can reprogram cell identity.

The upshot: Motif presence is a weak predictor of binding, and binding is a weak predictor of regulation. Treat each arrow in that chain as a hypothesis.

What the other end of the protein does

A transcription factor is modular: a DNA-binding domain plus one or more effector domains, and the classic demonstration is the domain swap. Fuse the DNA-binding domain of one factor to the activation domain of another and you get a protein that activates the first factor's target genes. This is the entire basis of the yeast two-hybrid assay and of GAL4-UAS systems in flies.

Activation domains work by recruitment. They contact Mediator, they recruit histone acetyltransferases such as p300 and CBP, and they recruit chromatin remodellers. They are often intrinsically disordered and unusually tolerant of sequence change, which fits recruitment through low-affinity multivalent contacts rather than a lock and key.

Repressors act by several distinct routes worth keeping separate: competing for the same site as an activator, binding an adjacent site and quenching the activator's contact surface, or recruiting corepressor complexes such as NCoR and SMRT with HDAC3, or Groucho and TLE. A nuclear receptor without ligand recruits corepressors and with ligand recruits coactivators, which is one protein switching between the two modes.

Reaching a promoter a megabase away

The SHH locus provides the cleanest example in human genetics. The enhancer that drives SHH expression in the developing limb bud, the ZRS, sits about one megabase away from SHH, inside intron 5 of an unrelated gene, LMBR1. Leah Lettice and colleagues showed in 2003 that point mutations in the ZRS cause preaxial polydactyly, extra digits on the thumb side, in humans and in mice.

Read what that means. A single base change, in an intron of a gene that has nothing to do with the phenotype, a million bases from the gene it actually controls, produces a limb malformation. Any variant interpretation pipeline that only looks at coding sequence would miss it entirely, and any pipeline that assigns non-coding variants to the nearest gene would assign it to LMBR1 and be wrong.

Contact is achieved by looping within a topologically associating domain, as in Module 1. CTCF and cohesin set the boundaries, and the boundary is what keeps the ZRS talking to SHH rather than to something else. Deleting a boundary lets enhancers reach genes they should not: rearrangements at the EPHA4 locus that disrupt a boundary allow limb enhancers to activate neighbouring genes and cause distinct limb malformations depending on which gene is captured.

What each method actually shows

This is where most misreading of the regulatory literature happens, so it is worth tabulating.

MethodWhat it demonstratesWhat it does not
Reporter assay (plasmid or integrated)The fragment is sufficient to drive expression in that contextThat it is used, or necessary, at its native locus in its native chromatin
ChIP-seq peakThe factor was crosslinked near that sequence in that cell stateThat binding is functional; many peaks have no effect when the site is mutated
ATAC-seqThe chromatin is accessibleWhich factor is bound, or what it does
Genomic deletion, no phenotypeThe element is dispensable under those conditionsThat it does nothing; shadow enhancers commonly provide redundancy, and phenotypes often appear only under stress
CRISPR interference at the elementThe element contributes to expression of a nearby geneThe distance and specificity, since the repressive domain spreads over kilobases

Two artefacts are worth naming because they have generated real literature. Highly expressed genes and certain repeats appear as peaks in almost any ChIP experiment, including negative controls, which is why an input or IgG control is not optional. And a large fraction of enhancers identified by chromatin marks give no phenotype when deleted, because developmental loci are buffered by multiple partially redundant enhancers with overlapping activity. That redundancy is a real biological finding, not a failure of the deletion experiment.

Bottom line: Sufficiency and necessity are different claims, and the standard enhancer toolkit tests sufficiency far more often than necessity.

Common misconceptions

  • "Enhancers are upstream regulatory sequences." They work upstream, downstream, inside introns, in either orientation, and across a megabase. That is the definition.
  • "A transcription factor binds where its motif is." Motifs occur hundreds of thousands of times; occupancy is set by cooperativity and chromatin accessibility, and most motif occurrences are never bound.
  • "The nearest gene is the target." The ZRS controls SHH from a megabase away while sitting inside LMBR1. Assigning non-coding variants to the nearest gene is a known error mode.
  • "If deleting an enhancer gives no phenotype it was not an enhancer." Redundant and shadow enhancers routinely buffer each other; the phenotype may appear only when a partner is also removed or under environmental stress.
  • "Activation domains have conserved recognition sequences." Many are intrinsically disordered and tolerate extensive mutation, consistent with multivalent low-affinity recruitment rather than a defined interface.

Where this leaves us

  • The SV40 72 base pair repeat defined enhancers in 1981 by working in either orientation, upstream or downstream, from a distance.
  • Sequence-specific recognition happens mainly in the major groove, using helix-turn-helix, zinc finger, basic zipper, basic helix-loop-helix and nuclear receptor folds.
  • Short motifs occur hundreds of thousands of times, so specificity comes from combinatorial binding, chromatin accessibility and pioneer factor activity.
  • Factors are modular: swapping a DNA-binding domain onto a different activation domain redirects the same regulatory output.
  • Repression works by competition, quenching or corepressor recruitment, and a nuclear receptor switches between corepressor and coactivator on ligand binding.
  • The ZRS controls SHH from a megabase inside another gene, and point mutations there cause preaxial polydactyly.
  • Reporter assays show sufficiency, ChIP-seq shows crosslinking, deletion shows necessity, and only the last of these tests whether an element matters at its own locus.

Sources

  1. Banerji, J., Rusconi, S., & Schaffner, W. (1981). Expression of a beta-globin gene is enhanced by remote SV40 DNA sequences. Cell, 27(2 Pt 1), 299-308. pubmed.ncbi.nlm.nih.gov
  2. Lettice, L. A., Heaney, S. J. H., Purdie, L. A., Li, L., de Beer, P., Oostra, B. A., et al. (2003). A long-range Shh enhancer regulates expression in the developing limb and fin and is associated with preaxial polydactyly. Human Molecular Genetics, 12(14), 1725-1735. pubmed.ncbi.nlm.nih.gov
  3. Alberts, B., Johnson, A., Lewis, J., Raff, M., Roberts, K., & Walter, P. (2002). DNA-binding motifs in gene regulatory proteins. In Molecular biology of the cell (4th ed.). Garland Science. ncbi.nlm.nih.gov
  4. Alberts, B., Johnson, A., Lewis, J., Raff, M., Roberts, K., & Walter, P. (2002). How genetic switches work. In Molecular biology of the cell (4th ed.). Garland Science. ncbi.nlm.nih.gov
  5. Clark, M. A., Douglas, M., & Choi, J. (2018). Eukaryotic transcription gene regulation. In Biology 2e (Section 16.4). OpenStax. openstax.org
Key terms
Enhancer
A regulatory element that increases transcription of a linked gene independently of its orientation and of its position relative to the promoter.
Major groove readout
Sequence recognition by contacting the distinctive hydrogen bonding and methyl pattern presented by each base pair in the major groove.
Pioneer factor
A transcription factor able to engage its motif on nucleosomal DNA and open the locus for other factors, as FOXA1 does.
Cooperativity
Mutual stabilisation of factors binding adjacent sites, which sharpens specificity because the combination is far rarer than either motif alone.
Corepressor
A complex such as NCoR or SMRT with HDAC3, recruited by a repressor to reduce transcription without itself binding DNA sequence-specifically.
ZRS
The limb enhancer of SHH, located about one megabase away inside an intron of LMBR1, whose point mutations cause preaxial polydactyly.
Shadow enhancer
A second, partially redundant enhancer with overlapping activity, which buffers a locus so that deleting one element gives little or no phenotype.
Sufficiency versus necessity
The distinction between an element being able to drive expression in a reporter and being required for expression at its native locus.

Module 5: The Life of an RNA

From the cap that goes on at nucleotide thirty to the decay pathway that finally removes the message, including the splicing decisions that make one gene into many proteins.

Capping, Splicing and Alternative Splicing

  • Describe the chemistry of the two transesterification steps of splicing and the role of each snRNP in the cycle.
  • Explain how splice sites are recognised, and why weak consensus sequences require additional regulatory input.
  • Analyse a splicing-based disease mechanism and the therapeutic strategies that exploit it.

Four matches and three loops

In 1977 Louise Chow, Richard Gelinas, Thomas Broker and Richard Roberts hybridised adenovirus 2 late messenger RNA to single-stranded viral DNA and looked at the products by electron microscopy. If a gene were colinear with its message, the hybrid would be a single continuous duplex with single-stranded tails. What they saw instead was a duplex interrupted three times, with large loops of unpaired DNA bulging out at each interruption. The message matched the genome in four separated blocks. Susan Berget, Claire Moore and Phillip Sharp reached the same conclusion independently by a different route the same year.

Neither group set out to discover introns. Both were doing careful mapping of viral transcripts and found that the map made no sense under the assumption everyone held. Colinearity was not a hypothesis anyone was testing; it was the background assumption of molecular biology, imported wholesale from bacteria. The 1993 Nobel Prize in Physiology or Medicine went to Roberts and Sharp for the finding.

The cap goes on first

Long before splicing is finished, the 5 prime end has been modified. When the nascent transcript reaches about 25 to 30 nucleotides and emerges from the exit channel, the capping enzymes act. A triphosphatase removes the terminal phosphate, a guanylyltransferase adds GMP through an unusual 5 prime to 5 prime triphosphate bridge, and a methyltransferase methylates that guanine at N7. The result is 7-methylguanosine joined backwards to the first transcribed nucleotide.

The enzymes are recruited by serine 5 phosphorylation of the polymerase carboxy-terminal domain, which is why capping happens at initiation and not later. The cap does four things: it blocks 5 prime exonucleases, it binds the initiation factor eIF4E for translation, it promotes splicing of the first intron, and it licenses nuclear export. A transcript that fails to be capped is degraded quickly, which is a quality control step disguised as a modification.

Key idea: The backwards 5 prime to 5 prime linkage is the point. No ordinary exonuclease can start on a structure with no free 5 prime end, so the cap is a chemical dead end for degradation and a specific handle for everything else.

Splicing is two transesterifications and no ATP in the chemistry

The reaction itself is elegantly economical.

  1. The 2 prime hydroxyl of a specific adenosine within the intron, the branch point, attacks the phosphate at the 5 prime splice site. The 5 prime exon is released and the intron 5 prime end becomes joined to the branch adenosine through a 2 prime to 5 prime bond, forming a lariat.
  2. The freed 3 prime hydroxyl of the 5 prime exon attacks the phosphate at the 3 prime splice site. The exons are ligated and the lariat intron is released, later debranched and degraded.

Both steps are transesterifications: the number of phosphodiester bonds is unchanged, so the chemistry is energetically neutral. ATP is consumed in abundance, but by the RNA helicases that drive the conformational rearrangements of the spliceosome, not by bond formation.

The signals are short and depressingly degenerate. In the major spliceosome the intron begins GU and ends AG, with a longer consensus around each and a branch point adenosine typically 18 to 40 nucleotides upstream of the 3 splice site, preceded by a polypyrimidine tract. Almost none of that information is specific enough on its own. A GU dinucleotide occurs every sixteen nucleotides by chance.

The spliceosome and its cycle

Five small nuclear ribonucleoproteins, each an snRNA plus proteins, assemble and rearrange on every intron.

  • U1 base pairs with the 5 prime splice site. U2AF binds the polypyrimidine tract and 3 prime site, and U2 base pairs at the branch point, bulging out the branch adenosine so its 2 prime hydroxyl is exposed and positioned. This is complex A.
  • The U4-U6-U5 tri-snRNP joins, giving complex B.
  • Massive rearrangement: U1 and U4 leave, U6 replaces U1 at the 5 prime splice site and pairs with U2, forming the catalytic core. U5 holds the two exon ends together.
  • Catalysis proceeds through complex C, exons are ligated, the lariat is released and the snRNPs are recycled.

The catalytic centre is made of RNA. U6 and U2 snRNA coordinate two magnesium ions in a geometry that closely resembles the active site of self-splicing group II introns, which perform the identical two-step chemistry with no protein at all. The spliceosome is best understood as a group II intron that has outsourced its structural work to proteins while keeping the chemistry in RNA.

A minor spliceosome, using U11, U12, U4atac and U6atac, handles a small class of introns, historically called AT-AC introns. Mutations in its RNU4ATAC gene cause microcephalic osteodysplastic primordial dwarfism type 1, which is a striking demonstration that fewer than one percent of introns can carry a severe phenotype.

How real splice sites are chosen

Since the core signals are weak, selection depends on context.

Exon definition. In organisms with long introns and short exons, which describes humans, the spliceosome initially recognises exons rather than introns: U1 at the downstream 5 site and U2AF at the upstream 3 site pair across the exon. Human introns average many kilobases while exons average about 140 nucleotides, so this is the tractable direction.

Splicing regulatory elements. Short sequences within exons and introns act as enhancers or silencers. SR proteins, with RS domains rich in serine and arginine, bind exonic splicing enhancers and promote inclusion; hnRNP proteins such as PTBP1 and hnRNP A1 bind silencers and promote skipping. The balance of SR and hnRNP concentrations differs between tissues, and that ratio is a principal determinant of tissue-specific splicing.

Co-transcriptional context. Splicing happens while the transcript is still being made, so polymerase speed matters. A slow polymerase gives a weak upstream exon more time to be recognised before a competing downstream site appears, which is why elongation rate mutants change splice isoform ratios. Nucleosome occupancy is higher over exons than introns, which slows the polymerase at exactly those positions.

The point: Splice site choice is a kinetic competition decided by weak signals, protein concentration ratios and transcription speed, which is why it is so easily perturbed by mutation and so amenable to drugs.

Alternative splicing, and one gene that makes 38,016 proteins

Roughly 95 percent of human multi-exon genes produce more than one splice isoform. The patterns are conventionally sorted into cassette exon inclusion or skipping, mutually exclusive exons, alternative 5 prime or 3 prime splice sites, and intron retention, with alternative polyadenylation and alternative promoters adding further diversity at the ends.

The extreme case is Drosophila Dscam1, an axon guidance receptor. It contains four clusters of mutually exclusive alternative exons, with 12, 48, 33 and 2 variants respectively. Multiplying gives 38,016 possible mRNA isoforms from one gene, which exceeds the number of genes in the fly genome. Neurons express different combinations, giving each neuron a distinctive surface identity used for self-avoidance.

Be careful how you generalise from this. Detecting an isoform in RNA-seq data is not the same as showing it makes a stable protein or that it does anything. Many low-abundance isoforms are splicing noise, and proteomic support for the great majority of annotated human isoforms is thin. The honest statement is that alternative splicing is pervasive, that a well-characterised subset is functionally important, and that isoform-level annotation is far ahead of isoform-level validation.

When splicing goes wrong, and how to fix it

Estimates vary, but a substantial fraction of disease-causing mutations act through splicing, including many that look synonymous or intronic and are therefore ignored by naive variant filters.

Beta-thalassaemia supplies textbook examples. Some alleles mutate the normal 5 prime splice site of intron 1. Others do something subtler: the IVS1-110 G to A mutation creates a new 3 prime splice site 19 nucleotides upstream of the normal one, and the cell uses the new site most of the time, inserting intronic sequence and shifting the reading frame. The normal site is intact. The mutation added a competitor.

Spinal muscular atrophy is the most instructive case, because it produced a drug. Patients lack functional SMN1. They retain SMN2, a nearly identical duplicate that differs by a translationally silent C to T change in exon 7. That change disrupts an exonic splicing enhancer, so exon 7 is skipped in most SMN2 transcripts and the truncated protein is unstable. Copy number of SMN2 modifies severity, which was the clue.

Nusinersen is an antisense oligonucleotide that binds an intronic splicing silencer downstream of exon 7, called ISS-N1, blocking hnRNP binding and shifting SMN2 splicing toward exon 7 inclusion. In the 2017 trial reported by Richard Finkel and colleagues, infants with infantile-onset disease receiving intrathecal nusinersen showed motor milestone responses in 41 percent against none in the control group, with improved event-free survival. The drug does not repair a gene or replace a protein. It changes which exon a spliceosome chooses.

So what?: A silent mutation that changes nothing about the protein sequence can still be the cause of a fatal disease, and a drug that changes nothing about the sequence can still treat it.

The 3 prime end

Polyadenylation is also co-transcriptional and also carboxy-terminal-domain coupled. CPSF recognises the AAUAAA hexamer, CstF binds a downstream GU-rich element, and cleavage occurs between them, typically 10 to 30 nucleotides after the hexamer. Poly(A) polymerase then adds roughly 200 adenosines without a template, and nuclear poly(A) binding protein coats the tail. The tail supports export, stability and translation initiation through the closed-loop interaction between poly(A) binding protein and eIF4G.

Many genes have more than one polyadenylation site. Shifting to a proximal site shortens the 3 untranslated region and removes microRNA binding sites, which typically increases stability and output; proliferating cells systematically favour proximal sites. This is regulation without changing a single codon.

Common misconceptions

  • "Splicing requires ATP for the chemistry." Both steps are transesterifications with no net change in bond number. ATP powers the helicases that rearrange the spliceosome, not the ligation.
  • "The GU-AG rule is enough to find introns." A GU occurs every sixteen nucleotides by chance. Real site selection needs exon definition, SR and hnRNP proteins, and transcription kinetics.
  • "The spliceosome is a protein enzyme." The catalytic centre is RNA, built from U2 and U6, and closely resembles the group II intron active site.
  • "Synonymous mutations are silent." They frequently disrupt or create splicing regulatory elements, as at SMN2 exon 7.
  • "Every annotated isoform is a functional protein." Detection in RNA-seq is not evidence of stable protein or of function, and proteomic support for most annotated isoforms is limited.

What to remember

  • Electron microscopy of adenovirus RNA-DNA hybrids in 1977 showed the message matching the genome in separated blocks, which is how introns were discovered.
  • The 7-methylguanosine cap is added at about nucleotide 25 to 30 through a 5 prime to 5 prime linkage, recruited by serine 5 phosphorylation, and it protects, exports and initiates.
  • Splicing runs as two transesterifications through a lariat intermediate, with the branch adenosine 2 prime hydroxyl as the first nucleophile.
  • U1, U2, U4, U5 and U6 assemble and rearrange so that U6 and U2 form an RNA catalytic core resembling a group II intron.
  • Weak consensus signals are supplemented by exon definition, SR and hnRNP proteins, and polymerase elongation rate.
  • About 95 percent of human multi-exon genes are alternatively spliced, and Dscam1 can generate 38,016 isoforms, but isoform detection is not isoform validation.
  • Splicing disease is common and often invisible to naive filters; nusinersen treats spinal muscular atrophy by blocking a silencer and restoring exon 7 inclusion in SMN2.

Sources

  1. Chow, L. T., Gelinas, R. E., Broker, T. R., & Roberts, R. J. (1977). An amazing sequence arrangement at the 5 prime ends of adenovirus 2 messenger RNA. Cell, 12(1), 1-8. pubmed.ncbi.nlm.nih.gov
  2. Berget, S. M., Moore, C., & Sharp, P. A. (1977). Spliced segments at the 5 prime terminus of adenovirus 2 late mRNA. PNAS, 74(8), 3171-3175. pubmed.ncbi.nlm.nih.gov
  3. Finkel, R. S., Mercuri, E., Darras, B. T., Connolly, A. M., Kuntz, N. L., Kirschner, J., et al. (2017). Nusinersen versus sham control in infantile-onset spinal muscular atrophy. New England Journal of Medicine, 377(18), 1723-1732. pubmed.ncbi.nlm.nih.gov
  4. The Nobel Foundation. (1993). The Nobel Prize in Physiology or Medicine 1993: Richard J. Roberts and Phillip A. Sharp. nobelprize.org
  5. Clark, M. A., Douglas, M., & Choi, J. (2018). RNA processing in eukaryotes. In Biology 2e (Section 15.4). OpenStax. openstax.org
  6. Cooper, G. M. (2000). RNA processing and turnover. In The cell: A molecular approach (2nd ed.). Sinauer Associates. ncbi.nlm.nih.gov
Key terms
5 prime cap
7-methylguanosine joined to the first transcribed nucleotide by a 5 prime to 5 prime triphosphate bridge, added at about nucleotide 25 to 30.
Branch point adenosine
The intronic adenosine whose 2 prime hydroxyl attacks the 5 prime splice site, creating the lariat intermediate.
Lariat
The looped intron intermediate formed when the intron 5 prime end is joined to the branch adenosine through a 2 prime to 5 prime bond.
snRNP
A small nuclear RNA plus its associated proteins; U1, U2, U4, U5 and U6 assemble and rearrange on each intron.
Exon definition
Initial recognition of an exon rather than an intron by factors pairing across it, the dominant mode where introns are long and exons short.
SR protein
A serine and arginine rich splicing factor that binds exonic enhancers and promotes exon inclusion, opposed by hnRNP proteins at silencers.
ISS-N1
The intronic splicing silencer downstream of SMN2 exon 7 that nusinersen blocks in order to restore exon inclusion.
Alternative polyadenylation
Use of a different cleavage and polyadenylation site, changing 3 untranslated region length and therefore microRNA regulation and stability.

RNA Stability and Turnover

  • Order the steps of deadenylation-dependent mRNA decay and identify the rate-limiting one.
  • Explain how AU-rich elements, microRNAs and nonsense-mediated decay each target a transcript for destruction.
  • Choose an appropriate method for measuring mRNA half-life and state its artefacts.

A mouse missing 69 nucleotides of untranslated region

In 1999 Dimitris Kontoyiannis and colleagues made a mouse in which the AU-rich element was deleted from the 3 untranslated region of the tumour necrosis factor gene. The coding sequence was untouched. The promoter was untouched. All they removed was a stretch of regulatory RNA sequence that is never translated.

The mice developed chronic inflammatory polyarthritis and an inflammatory bowel disease resembling Crohn ileitis. Tumour necrosis factor was being made in the right cells in response to the right stimuli, but the message could no longer be turned off promptly, and the sustained low-level excess was enough to produce two human-like autoimmune diseases.

Hold that alongside the way transcription is usually taught. The amount of a protein in a cell is set by two rates, not one, and this experiment changed only the second.

Why this matters: Steady-state abundance equals synthesis rate divided by decay rate. Any account of gene regulation that discusses only transcription is describing half the equation.

How long does an mRNA last

Mammalian mRNA half-lives span more than three orders of magnitude. Immediate early gene transcripts such as FOS and MYC turn over in 10 to 30 minutes. Cytokine messages are similarly short-lived. Housekeeping transcripts sit around for many hours, and the most stable, such as globin messages in erythroid precursors, persist for a day or more, which is what allows an enucleated reticulocyte to keep making haemoglobin.

Measuring this is harder than it looks, and the two main approaches fail in different ways.

MethodHow it worksFailure mode
Actinomycin D chaseBlock transcription, sample over time, fit decay curvesBlocking all transcription is a severe perturbation; it depletes short-lived decay factors and artefactually stabilises many transcripts, and cells begin dying
Metabolic labelling with 4-thiouridineLabel new RNA, then chase; chemical conversion makes labelled positions read as a different base in sequencingThiouridine itself is mildly toxic and can perturb ribosome biogenesis; conversion chemistry is incomplete, so rates need careful modelling
Transcriptional shut-off of a single inducible geneTurn off one promoter, follow that transcriptClean for that gene, not generalisable, and requires an engineered system

Published median half-life estimates differ severalfold between methods, which is itself the useful lesson: when two methods for the same quantity disagree by a factor of three, the number is method-dependent and should be reported as such.

The default pathway: deadenylate, then destroy

Almost all mRNA decay in eukaryotes starts by shortening the poly(A) tail.

  1. Deadenylation. PAN2-PAN3 trims the tail from about 200 adenosines down to roughly 80, then the CCR4-NOT complex takes it the rest of the way. This step is slow and is the rate-limiting one for most messages, which means it is where regulation acts.
  2. Loss of the closed loop. As the tail shortens, poly(A) binding protein leaves and the circular arrangement in which the tail communicates with the cap through eIF4G collapses. Translation initiation falls, which is why deadenylation reduces output before the message is destroyed.
  3. Decapping. The LSM1-7 ring binds the short oligo(A) end and recruits DCP2 with DCP1 and enhancers such as EDC4 and DDX6. DCP2 hydrolyses the cap, leaving a free 5 prime monophosphate.
  4. Exonucleolytic destruction. XRN1 degrades the body 5 prime to 3 prime. In parallel or alternatively, the cytoplasmic exosome degrades 3 prime to 5 prime and the remaining cap is cleared by DCPS.

Note the logic. The cap and the tail are both protective, and both must be removed before the message can be attacked from either end. The message is not destroyed by breaking it in the middle; it is uncapped and unprotected and then eaten from the ends.

Three ways to be targeted

AU-rich elements. Sequences matching AUUUA repeated within a uridine-rich context occur in the 3 untranslated regions of many cytokine, proto-oncogene and cell cycle messages. They are read by competing proteins. Tristetraprolin (ZFP36) and KSRP bind and recruit deadenylases, shortening half-life dramatically; HuR (ELAVL1) binds overlapping sites and stabilises. The outcome for a given transcript in a given cell depends on the ratio of these proteins and on their phosphorylation state, which is how a signalling pathway can change mRNA stability within minutes. Tristetraprolin knockout mice develop an inflammatory syndrome driven by excess tumour necrosis factor, which is the same phenotype from the other side of the same interaction as the element deletion described above.

MicroRNAs. A microRNA loaded into Argonaute base pairs with a partially complementary site, usually in the 3 untranslated region, through a seed of nucleotides 2 to 8. Argonaute recruits TNRC6 proteins, which recruit the PAN2-PAN3 and CCR4-NOT deadenylases. So in animals the dominant outcome of microRNA targeting is accelerated deadenylation and decay rather than translational silencing with the message intact, which was the original expectation. The next lesson takes up microRNA biology properly.

Nonsense-mediated decay. This is a surveillance pathway, and its logic is worth working through because it explains a large class of genotype-phenotype relationships.

During splicing, the exon junction complex is deposited on the message about 20 to 24 nucleotides upstream of each exon-exon junction. The first, pioneer round of translation displaces these complexes as the ribosome passes. If a stop codon is reached while an exon junction complex still lies more than about 50 to 55 nucleotides downstream, the ribosome terminates in an abnormal context; UPF1 at the terminating ribosome contacts UPF2 and UPF3 on the complex, SMG1 phosphorylates UPF1, and the message is destroyed.

The 50 to 55 nucleotide rule has direct clinical consequences. A nonsense mutation in an internal exon triggers decay, so no truncated protein is made and the allele behaves as a null: the phenotype is haploinsufficiency, and in a recessive disease a heterozygote is healthy. A nonsense mutation in the last exon, or within the last 50 nucleotides of the penultimate exon, escapes decay; a truncated protein is made and may act as a dominant negative or gain a toxic function. Two mutations of the same class in the same gene can therefore give recessive and dominant disease depending only on position. Beta-thalassaemia shows exactly this: early nonsense alleles are recessive, while some alleles in exon 3 escape decay, produce an unstable truncated beta chain and cause dominantly inherited thalassaemia intermedia.

Nonsense-mediated decay is not only a surveillance system. It also regulates a substantial set of normal transcripts, including many produced by alternative splicing events that deliberately introduce a premature stop, a mechanism sometimes called regulated unproductive splicing and translation. Several splicing factors autoregulate this way.

The upshot: Before you predict the effect of a nonsense variant, find out which exon it is in and how far it is from the last junction.

Granules: where decay happens, and what that does not mean

Decapping and 5 prime to 3 prime decay factors concentrate in cytoplasmic processing bodies, and translationally stalled messages with initiation factors accumulate in stress granules under stress. Both are membraneless condensates formed largely through multivalent interactions among intrinsically disordered proteins and RNA.

It is tempting to conclude that decay happens in processing bodies and that dissolving them would stabilise mRNA. Experiments say otherwise: preventing processing body formation does not generally block decay, and decay proceeds on polysomes. The current reading is that these bodies are consequences of the accumulation of translationally repressed messenger ribonucleoproteins rather than obligatory decay compartments, and possibly storage depots from which messages can return. Be careful with the causal direction here, because a lot of published language is looser than the evidence.

Two other surveillance pathways

  • Non-stop decay handles messages lacking a stop codon, where the ribosome translates into the poly(A) tail. Ski7 and the exosome degrade the message and the ribosome is rescued.
  • No-go decay handles ribosomes stalled by strong structure or damaged RNA: the message is cleaved endonucleolytically and the stalled ribosome is dismantled by the ribosome quality control complex, which also tags the incomplete nascent chain for degradation.

All three surveillance systems share a design: they detect an abnormal ribosome state, not an abnormal sequence. The ribosome is the sensor.

Common misconceptions

  • "Gene expression is set by transcription." Abundance is synthesis over decay, and deleting a stability element alone was sufficient to cause arthritis and ileitis in mice.
  • "mRNAs are cut in the middle to destroy them." The dominant route is deadenylation, then decapping, then exonucleolytic digestion from an end. Endonucleolytic cleavage is the exception.
  • "Actinomycin D gives a clean half-life measurement." Blocking all transcription depletes short-lived decay factors and artefactually stabilises many messages; metabolic labelling is preferable where possible.
  • "MicroRNAs act mainly by blocking translation." In animal cells the dominant measured effect is recruitment of deadenylases and mRNA destabilisation.
  • "All nonsense mutations are loss-of-function nulls." Those in the last exon escape nonsense-mediated decay and can produce dominant-negative truncated proteins.

The short version

  • Deleting only the AU-rich element from the tumour necrosis factor 3 untranslated region caused arthritis and ileitis in mice, with no change to coding sequence or promoter.
  • Half-lives span from about 10 minutes for immediate early transcripts to over a day for globin messages, and published medians are strongly method-dependent.
  • The default decay route is PAN2-PAN3 then CCR4-NOT deadenylation, which is rate-limiting, followed by LSM-dependent decapping and XRN1 or exosome digestion.
  • AU-rich elements are read competitively by destabilising tristetraprolin and KSRP against stabilising HuR, with the balance set by signalling.
  • MicroRNA targeting works largely by recruiting the same deadenylases through TNRC6.
  • Nonsense-mediated decay uses exon junction complexes and the 50 to 55 nucleotide rule, which is why nonsense position determines whether an allele is a null or a dominant negative.
  • Processing bodies and stress granules concentrate decay and stalled translation machinery but are not required for decay to occur.

Sources

  1. Kontoyiannis, D., Pasparakis, M., Pizarro, T. T., Cominelli, F., & Kollias, G. (1999). Impaired on/off regulation of TNF biosynthesis in mice lacking TNF AU-rich elements: Implications for joint and gut-associated immunopathologies. Immunity, 10(3), 387-398. pubmed.ncbi.nlm.nih.gov
  2. Nagy, E., & Maquat, L. E. (1998). A rule for termination-codon position within intron-containing genes: When nonsense affects RNA abundance. Trends in Biochemical Sciences, 23(6), 198-199. pubmed.ncbi.nlm.nih.gov
  3. Alberts, B., Johnson, A., Lewis, J., Raff, M., Roberts, K., & Walter, P. (2002). Posttranscriptional controls. In Molecular biology of the cell (4th ed.). Garland Science. ncbi.nlm.nih.gov
  4. Cooper, G. M. (2000). RNA processing and turnover. In The cell: A molecular approach (2nd ed.). Sinauer Associates. ncbi.nlm.nih.gov
  5. Clark, M. A., Douglas, M., & Choi, J. (2018). Eukaryotic post-transcriptional gene regulation. In Biology 2e (Section 16.5). OpenStax. openstax.org
Key terms
Deadenylation
Progressive shortening of the poly(A) tail by PAN2-PAN3 and then CCR4-NOT, the rate-limiting first step of most mRNA decay.
Decapping
Hydrolysis of the 5 prime cap by DCP2 with DCP1, following LSM1-7 binding to the shortened tail, which exposes the message to XRN1.
AU-rich element
An AUUUA-containing motif in a 3 untranslated region read competitively by destabilising factors such as tristetraprolin and stabilising factors such as HuR.
Exon junction complex
A protein assembly deposited about 20 to 24 nucleotides upstream of each exon-exon junction during splicing and displaced by the first translating ribosome.
Nonsense-mediated decay
Degradation of transcripts whose stop codon is followed by a remaining exon junction complex more than about 50 to 55 nucleotides downstream.
Closed loop
The circular arrangement in which poly(A) binding protein contacts eIF4G at the cap, coupling tail length to translation initiation efficiency.
Processing body
A cytoplasmic condensate enriched in decapping and 5 prime to 3 prime decay factors, associated with but not required for decay.
No-go decay
The surveillance route that cleaves messages on which ribosomes have stalled and dismantles the stalled ribosome and its nascent chain.

Non-Coding RNA, and the Limits of the Evidence

  • Trace microRNA biogenesis from primary transcript to loaded Argonaute and describe how target repression is achieved.
  • Distinguish microRNA, small interfering RNA and PIWI-interacting RNA pathways by origin, machinery and biological role.
  • Apply an evidential standard to claims about long non-coding RNA function, and explain what the ENCODE dispute was actually about.

A gene with no protein in it

A Caenorhabditis elegans larva carrying a mutation in lin-4 repeats the first larval stage instead of moving on to the second, and goes on repeating it. In 1993 Rosalind Lee, Rhonda Feinbaum and Victor Ambros cloned that locus and set about finding the protein it encoded. There was no open reading frame worth the name. What the locus produced were two small RNAs, about 22 and 61 nucleotides long.

The pair then noticed something that made the finding mechanistic rather than merely odd: the 22 nucleotide RNA was partially complementary to seven repeated sequences in the 3 untranslated region of lin-14, the gene lin-4 was known genetically to repress. Genetic epistasis and base pairing agreed. Regulation was being carried out by an RNA using complementarity, and no protein was involved on the regulator's side.

For seven years this looked like a worm curiosity. Then let-7 was found to be conserved from worms to humans, homologous small RNAs turned up everywhere, and the field became one of the largest in biology.

MicroRNA biogenesis, step by step

  1. RNA polymerase II transcribes a primary microRNA, capped and polyadenylated like an mRNA, containing one or more hairpins. Many sit in introns of protein-coding genes and are processed out of the pre-mRNA.
  2. In the nucleus the microprocessor, Drosha with its partner DGCR8, cuts the base of the hairpin, releasing a precursor of about 70 nucleotides with a characteristic two nucleotide 3 prime overhang.
  3. Exportin-5 with RanGTP carries the precursor to the cytoplasm.
  4. Dicer, an RNase III enzyme, removes the loop, leaving a duplex of about 22 base pairs.
  5. The duplex is loaded into an Argonaute protein. One strand, the guide, is retained on the basis of thermodynamic asymmetry at the duplex ends; the passenger strand is discarded. The result is the RNA-induced silencing complex.

Targeting depends mainly on the seed, nucleotides 2 to 8 of the guide, pairing with a site usually in the 3 untranslated region. Six to eight base pairs is not much specificity, and it shows: a single microRNA has hundreds of predicted targets, and prediction algorithms have high false positive rates. Argonaute recruits TNRC6, which recruits the deadenylases from the previous lesson, so the outcome is mainly destabilisation.

Effect sizes matter here and are routinely overstated. In careful measurements, most individual microRNA-target relationships change protein output by less than twofold. The biological argument for their importance is not that any one interaction is strong; it is that a microRNA nudges a whole module of functionally related transcripts at once, and that such tuning buffers noise. DGCR8 lies within the region deleted in 22q11.2 deletion syndrome, which is one line of evidence that global microRNA dosage has phenotypic consequences.

Bottom line: A microRNA is a rheostat acting on many targets, not a switch acting on one. Papers claiming a single microRNA-target pair explains a phenotype should be read with that in mind.

Small interfering RNA, and how the field got here

Andrew Fire and Craig Mello injected double-stranded RNA into C. elegans in 1998 and found it silenced the matching gene far more potently than either single strand alone. The effect was catalytic, spread between cells, and was inherited for a generation or two. That is the paper that named RNA interference and won the 2006 Nobel Prize in Physiology or Medicine.

The mechanistic overlap with microRNAs is nearly complete: Dicer processes long double-stranded RNA into 21 to 23 nucleotide duplexes, which load into Argonaute. The difference is complementarity. A small interfering RNA pairs perfectly with its target, which positions the target for cleavage by the endonuclease activity of Ago2, the only human Argonaute that retains slicer function. A microRNA pairs partially and directs deadenylation instead.

Because the pathway is a general antiviral and anti-transposon defence in plants and invertebrates, and is co-opted for endogenous regulation in mammals, it has been a straightforward target for engineering. Patisiran, licensed in 2018 for hereditary transthyretin amyloidosis, was the first approved small interfering RNA drug; delivery, not mechanism, was the hard part, and lipid nanoparticles and GalNAc conjugation for liver targeting are what made it feasible.

A third small RNA class deserves naming. PIWI-interacting RNAs are 24 to 31 nucleotides, are produced without Dicer, and load into PIWI-clade Argonautes in the germ line. They derive from piRNA cluster loci that act as a genetically encoded archive of transposon sequences, and they amplify through a ping-pong cycle in which sense and antisense species reciprocally direct each other's production. Their job is transposon silencing in the germ line, and loss of the pathway causes sterility. Here the evidence for function is unusually clean, because the phenotype is dramatic and the sequences map to the elements they silence.

Long non-coding RNA: the strong case and the weak cases

A long non-coding RNA is operationally any Pol II transcript over 200 nucleotides with no substantial open reading frame. Tens of thousands are annotated in humans. The number that have been shown to do something is far smaller, and the gap is where the interesting methodological problems live.

The strong case: XIST. Carolyn Brown and colleagues reported in 1991 a gene at the X inactivation centre expressed exclusively from the inactive X. XIST RNA is transcribed from the chromosome it silences, spreads in cis over that chromosome, and recruits repressive machinery. Deleting it prevents inactivation; expressing it from an autosome in the right context silences that autosome. Necessity, sufficiency, cis-action and a mechanism, all demonstrated. This is the standard other claims should be measured against.

A contested case: HOTAIR. This transcript from the HOXC cluster was reported to act in trans on the HOXD cluster by recruiting PRC2, and became one of the most cited lncRNAs. Then mouse knockouts gave inconsistent results: different laboratories deleting overlapping regions reported different, and in some cases minimal, effects on HOXD expression, and it became clear that deleting a transcript's DNA also deletes whatever regulatory elements are embedded in it. The disagreement has not been fully resolved, and it is a good illustration of why the field now insists on separating three distinct hypotheses.

  • The RNA product does the work. Test by rescuing the knockout with the RNA expressed from elsewhere, in trans.
  • The act of transcription does the work, for instance by displacing nucleosomes or interfering with a neighbouring promoter. Test by inserting a premature polyadenylation signal that truncates the transcript while leaving initiation intact.
  • A DNA element inside the locus does the work, and the transcript is incidental. Test with small deletions of the element that do not remove the promoter, and with CRISPR interference targeted to the promoter versus the body.

MALAT1 is another instructive case: an abundant, highly conserved nuclear lncRNA implicated in splicing regulation and metastasis, whose knockout mice are viable and fertile with mild phenotypes. Abundance and conservation raised expectations that loss-of-function did not meet.

The core of it: For a long non-coding RNA, ask which of the three hypotheses the experiment discriminates. Most published experiments discriminate none of them.

What the ENCODE argument was about

In September 2012 the ENCODE consortium published a summary paper reporting that its assays assigned a biochemical function to roughly 80 percent of the human genome. That number was widely reported as the end of junk DNA.

Dan Graur and colleagues replied in 2013 in a paper with an unusually pointed title, arguing that the 80 percent figure rested on a definition of function so permissive that it counted any reproducible biochemical event, including transcription at very low levels and transcription factor binding with no consequence. Their central point was evolutionary: sequence under purifying selection can be measured, and only around 5 to 10 percent of the human genome shows it. If most of the transcribed genome were doing something that mattered, mutations in it would be removed by selection, and they are not.

Manolis Kellis and ENCODE authors published a more measured analysis in 2014 that essentially conceded the terminological point: they distinguished biochemical signatures from evolutionary and genetic evidence of function, and noted that a biochemical signature is a hypothesis rather than a conclusion.

Two things are true at once and it is worth holding both. Pervasive transcription is real: most of the genome is transcribed somewhere at some level. And most of those transcripts show no evidence of being under selection, are present at less than one copy per cell, and have no demonstrated function. Neither the claim that all of it is functional nor the claim that all of it is noise survives contact with the data.

Remember: Biochemical activity, evolutionary conservation and experimental necessity are three different standards of evidence. The ENCODE dispute was a disagreement about which one the word function should name.

Common misconceptions

  • "A microRNA switches its target off." Typical effects are under twofold on any single target; the biological role is collective tuning of many transcripts.
  • "MicroRNA target prediction identifies real targets." A six to eight nucleotide seed match occurs frequently by chance, so predictions require experimental support such as crosslinking data plus a measured effect on protein output.
  • "Small interfering RNA and microRNA are different machineries." They share Dicer and Argonaute. The difference is complementarity, which decides between Ago2 slicing and deadenylation.
  • "Deleting a long non-coding RNA gene tests the RNA." It also deletes the promoter and any enhancer inside the locus, which is why trans rescue and premature termination experiments are needed.
  • "ENCODE showed there is no junk DNA." ENCODE showed pervasive biochemical activity. Estimates of the fraction of the genome under purifying selection remain in the range of 5 to 10 percent.

Recap

  • lin-4 in 1993 was the first regulatory gene shown to work as an RNA, base pairing with sites in the lin-14 3 untranslated region.
  • MicroRNA biogenesis runs Pol II transcript, Drosha and DGCR8, exportin-5, Dicer, Argonaute loading with guide strand selection by thermodynamic asymmetry.
  • Seed pairing at nucleotides 2 to 8 gives limited specificity; repression is mainly by TNRC6-recruited deadenylation and is usually less than twofold per target.
  • Small interfering RNAs differ from microRNAs mainly in being fully complementary, which licenses Ago2 slicing; patisiran in 2018 was the first approved drug of this class.
  • PIWI-interacting RNAs silence transposons in the germ line, amplified by a ping-pong cycle, with sterility as the loss-of-function phenotype.
  • XIST meets the full evidential standard for a functional long non-coding RNA; HOTAIR and MALAT1 illustrate how hard that standard is to meet.
  • The ENCODE 80 percent figure counted biochemical activity, while roughly 5 to 10 percent of the genome shows evidence of purifying selection; the dispute was about the definition of function.

Sources

  1. Lee, R. C., Feinbaum, R. L., & Ambros, V. (1993). The C. elegans heterochronic gene lin-4 encodes small RNAs with antisense complementarity to lin-14. Cell, 75(5), 843-854. pubmed.ncbi.nlm.nih.gov
  2. Fire, A., Xu, S., Montgomery, M. K., Kostas, S. A., Driver, S. E., & Mello, C. C. (1998). Potent and specific genetic interference by double-stranded RNA in Caenorhabditis elegans. Nature, 391(6669), 806-811. pubmed.ncbi.nlm.nih.gov
  3. Brown, C. J., Ballabio, A., Rupert, J. L., Lafreniere, R. G., Grompe, M., Tonlorenzi, R., & Willard, H. F. (1991). A gene from the region of the human X inactivation centre is expressed exclusively from the inactive X chromosome. Nature, 349(6304), 38-44. pubmed.ncbi.nlm.nih.gov
  4. ENCODE Project Consortium. (2012). An integrated encyclopedia of DNA elements in the human genome. Nature, 489(7414), 57-74. pubmed.ncbi.nlm.nih.gov
  5. Graur, D., Zheng, Y., Price, N., Azevedo, R. B. R., Zufall, R. A., & Elhaik, E. (2013). On the immortality of television sets: Function in the human genome according to the evolution-free gospel of ENCODE. Genome Biology and Evolution, 5(3), 578-590. pubmed.ncbi.nlm.nih.gov
  6. Kellis, M., Wold, B., Snyder, M. P., Bernstein, B. E., Kundaje, A., Marinov, G. K., et al. (2014). Defining functional DNA elements in the human genome. PNAS, 111(17), 6131-6138. pubmed.ncbi.nlm.nih.gov
Key terms
Microprocessor
The nuclear Drosha and DGCR8 complex that cuts the base of a primary microRNA hairpin to release the precursor.
Seed sequence
Nucleotides 2 to 8 of a microRNA guide strand, which provide most of the target specificity through base pairing.
RNA-induced silencing complex
Argonaute loaded with a guide strand, which base pairs with targets and recruits effectors.
Slicer activity
The endonuclease function retained by human Ago2, which cleaves targets that are fully complementary to the guide.
piRNA ping-pong cycle
The amplification loop in which sense and antisense PIWI-interacting RNAs reciprocally direct each other's production during germ line transposon silencing.
XIST
The long non-coding RNA transcribed from and coating the inactive X chromosome, the best-evidenced example of lncRNA function.
Pervasive transcription
The observation that most of the genome is transcribed at some level in some cell type, which is not by itself evidence of function.
Purifying selection
Removal of deleterious mutations over evolutionary time; the fraction of the human genome showing it is estimated at roughly 5 to 10 percent.

Module 6: Protein Output

How a message becomes a polypeptide, how that step is throttled by stress and nutrient signalling, and what happens to the chain once it leaves the ribosome.

Translation and its Regulation

  • Describe eukaryotic initiation, elongation and termination and identify the step at which most regulation acts.
  • Explain how eIF2 alpha phosphorylation and mTORC1 signalling change translational output in opposite ways.
  • Interpret a ribosome profiling experiment, including its principal artefacts.

No protein within 18 angstroms

In August 2000 Poul Nissen, Jeffrey Hansen, Nenad Ban, Peter Moore and Thomas Steitz published the structure of the large ribosomal subunit of Haloarcula marismortui with transition state analogues bound, at a resolution good enough to identify individual atoms near the catalytic site. The question everyone wanted answered was which residue performs the chemistry of peptide bond formation.

The answer was that no amino acid is close enough to be a candidate. The nearest polypeptide lay about 18 angstroms from the site where the peptide bond forms. The active site is built entirely of ribosomal RNA. The ribosome is a ribozyme, and the largest and most important enzyme in the cell turns out to be RNA using protein as scaffolding. The 2009 Nobel Prize in Chemistry went to Venkatraman Ramakrishnan, Steitz and Ada Yonath for the ribosome structures.

Key idea: The catalytic core of translation is RNA, as is the catalytic core of splicing. Both are hard to explain except as survivals from a period when RNA did the chemistry.

The machine and its sites

A eukaryotic 80S ribosome is a 60S subunit (28S, 5.8S and 5S rRNAs plus about 47 proteins) and a 40S subunit (18S rRNA plus about 33 proteins). Bacteria run a 70S ribosome of 50S and 30S subunits. Three tRNA binding sites span the interface: the A site receives the incoming aminoacyl-tRNA, the P site holds the tRNA carrying the growing chain, and the E site is the exit. The small subunit holds the codon-anticodon duplex and does the decoding; the large subunit does the chemistry and provides the exit tunnel through which the nascent chain leaves.

Initiation, which is where regulation lives

Bacterial initiation is straightforward: the Shine-Dalgarno sequence in the message base pairs with the 3 prime end of the 16S rRNA, positioning the start codon in the P site, and initiation factors deliver formylmethionyl-tRNA. Because there is no scanning, an internal Shine-Dalgarno sequence works as well as a 5 prime one, which is exactly what makes polycistronic messages possible.

Eukaryotic initiation is longer and offers far more control points.

  1. eIF2 with GTP binds initiator methionyl-tRNA to form the ternary complex, which joins the 40S subunit with eIF1, eIF1A, eIF3 and eIF5 to make the 43S preinitiation complex.
  2. eIF4F assembles at the cap: eIF4E binds the 7-methylguanosine, eIF4G is the scaffold, eIF4A is the helicase that unwinds 5 prime structure. eIF4G also binds poly(A) binding protein, closing the loop.
  3. The 43S complex is recruited and scans 5 prime to 3 prime until it meets an AUG in acceptable context. Marilyn Kozak defined that context in 1986 by systematic point mutation: a purine at position minus 3 and a guanine at plus 4 matter most.
  4. eIF5 triggers GTP hydrolysis on eIF2, factors leave, and eIF5B mediates joining of the 60S subunit to make an elongation-competent 80S.

Two features of this sequence are where nearly all translational regulation acts, and they are worth separating.

eIF4E availability is controlled by mTORC1. The 4E-binding proteins bind eIF4E and prevent it recruiting eIF4G. When mTORC1 is active, it phosphorylates 4E-BP1, which releases eIF4E and permits cap-dependent initiation. Rapamycin and its analogues inhibit mTORC1 and therefore reduce translation, which is a large part of how they act as immunosuppressants and antiproliferative agents. Messages with long structured 5 untranslated regions, including many encoding growth machinery, are the most sensitive.

Ternary complex availability is controlled by phosphorylation of eIF2 alpha on serine 51. This deserves care, because the mechanism is counterintuitive. eIF2 must be recycled from the GDP to the GTP form by the exchange factor eIF2B, which is present at lower abundance than eIF2. Phosphorylated eIF2 alpha binds eIF2B tightly and does not let go, so a modest fraction of phosphorylated eIF2 sequesters most of the exchange factor and shuts down recycling for the whole pool. A small input produces a large, switch-like output.

Four kinases phosphorylate that serine, each reporting a different problem, and their convergence is called the integrated stress response.

KinaseActivated by
GCN2Uncharged tRNA, that is, amino acid starvation
PERKUnfolded protein load in the endoplasmic reticulum
PKRDouble-stranded RNA, typically viral
HRIHaem deficiency, chiefly in erythroid cells

Here is the elegant part. Global initiation falls, but a small set of messages goes up, and ATF4 is the canonical example. Its 5 untranslated region carries two upstream open reading frames. Under normal conditions a scanning ribosome translates uORF1, resumes scanning, and reinitiates at uORF2, which overlaps the ATF4 start codon out of frame, so ATF4 is not made. When ternary complex is scarce, reinitiation takes longer, the scanning ribosome travels further before it can acquire a new initiator tRNA, and it therefore skips uORF2 and lands on the ATF4 start instead. Scarcity of the initiation machinery is converted into selective translation of a transcription factor that then turns on the stress programme.

The point: Translational control is not a volume knob. The same signal that lowers global output raises specific outputs, and the mechanism is a consequence of scanning kinetics.

Elongation, and why accuracy is lower than you expect

eEF1A with GTP delivers aminoacyl-tRNA to the A site. Correct codon-anticodon pairing closes the decoding centre of the 16S or 18S rRNA around the minor groove of the duplex, and only then is GTP hydrolysis triggered efficiently. A second proofreading window follows before accommodation. Peptide bond formation is catalysed by the rRNA, and eEF2 with GTP then translocates the ribosome by one codon.

Translation makes roughly one error in 103 to 104 codons, which is thousands of times worse than DNA replication. That is not a design failure. A mistranslated protein is one defective molecule among many copies and is degraded; a DNA error is copied into every descendant cell. The cell spends its accuracy budget where errors are heritable.

Elongation is also a drug target and a toxin target, which is useful evidence about mechanism. Diphtheria toxin ADP-ribosylates a modified histidine residue in eEF2 called diphthamide, freezing translocation, and a single molecule reaching the cytosol can kill a cell. Antibiotics hit the bacterial ribosome at defined places: aminoglycosides bind the 30S decoding site and cause misreading, tetracyclines block the A site, chloramphenicol blocks the peptidyl transferase centre, and macrolides plug the exit tunnel. Their selectivity comes from the structural differences between bacterial and eukaryotic ribosomes, which is also why they have mitochondrial toxicity, since mitochondrial ribosomes are bacterial in origin.

Termination and recycling

Eukaryotes use a single release factor, eRF1, which recognises all three stop codons and, with eRF3, triggers hydrolysis of the peptidyl-tRNA ester bond. ABCE1 then splits the subunits for reuse. Bacteria use two codon-specific factors, RF1 and RF2, plus RF3.

Stop codon readthrough is not always an error. Selenocysteine is inserted at specific UGA codons directed by a SECIS element in the 3 untranslated region with a dedicated elongation factor, which is why selenoproteins such as glutathione peroxidase exist at all. Aminoglycosides and designed readthrough compounds have been trialled for nonsense-mutation diseases on exactly this principle.

Iron, and one signal read two ways

Iron regulatory proteins bind iron-responsive elements, stem-loops in untranslated regions, when cellular iron is low. Position determines the effect.

  • Ferritin, which stores iron, has an element in its 5 untranslated region. Bound protein blocks the scanning ribosome, so ferritin translation stops when iron is scarce.
  • Transferrin receptor, which imports iron, has elements in its 3 untranslated region. Bound protein masks endonuclease sites and stabilises the message, so receptor levels rise when iron is scarce.

Same protein, same signal, opposite outcomes, decided entirely by which end of the message the element sits in. It is the cleanest example in the cell of position determining mechanism.

Ribosome profiling, and how to misread it

Nicholas Ingolia and Jonathan Weissman published ribosome profiling in 2009. Treat lysate with nuclease so that only the roughly 28 nucleotide fragments protected by ribosomes survive, purify the monosomes, and sequence the protected fragments. You get a genome-wide map of ribosome positions at close to codon resolution, from which you can infer which open reading frames are being translated, including upstream ones, and estimate relative synthesis rates.

Its artefacts are worth knowing because they generated a literature of their own. Pretreating cells with cycloheximide to freeze ribosomes blocks elongation faster than it blocks initiation, so ribosomes pile up near start codons and the resulting 5 prime peak is partly an artefact; flash freezing without drug is now preferred. Nuclease digestion conditions change apparent pause sites. And a high ribosome density means slow movement as readily as it means high output, so density alone does not measure synthesis rate.

Worth holding on to: Ribosome density is occupancy, not flux. Distinguishing a busy codon from a stalled one requires more than a coverage track.

Common misconceptions

  • "Peptide bond formation is catalysed by a ribosomal protein." The nearest protein is about 18 angstroms away; the catalytic centre is ribosomal RNA.
  • "Translation is as accurate as replication." It is about a million times less accurate, which is rational because translational errors are not inherited.
  • "Phosphorylating eIF2 alpha shuts down all protein synthesis." It lowers global initiation but selectively increases translation of messages such as ATF4 through upstream open reading frames.
  • "Eukaryotic ribosomes can initiate internally like bacterial ones." Scanning from the cap is the rule; internal initiation requires a specialised structure, and many published internal ribosome entry sites turned out to be cryptic promoters or splice sites in the reporter construct.
  • "A tall ribosome profiling peak means high protein output there." It means high occupancy, which is equally consistent with a stall.

Summing up

  • The peptidyl transferase centre is made of rRNA, with no protein within about 18 angstroms, so the ribosome is a ribozyme.
  • Bacterial initiation uses Shine-Dalgarno pairing and permits polycistronic messages; eukaryotic initiation uses cap recognition, scanning and Kozak context.
  • mTORC1 controls eIF4E availability through 4E-BP phosphorylation; rapamycin lowers cap-dependent initiation.
  • Phosphorylation of eIF2 alpha sequesters the limiting exchange factor eIF2B, giving a switch-like global shutdown, with four kinases reporting four different stresses.
  • Upstream open reading frames convert that scarcity into selective ATF4 translation.
  • Translation errs about once per 103 to 104 codons, and antibiotics and diphtheria toxin act at defined elongation steps.
  • Iron regulatory proteins illustrate position-determined mechanism: a 5 prime element blocks translation, a 3 prime element stabilises the message.
  • Ribosome profiling maps occupancy at codon resolution but is distorted by cycloheximide pretreatment and cannot by itself distinguish output from stalling.

Sources

  1. Nissen, P., Hansen, J., Ban, N., Moore, P. B., & Steitz, T. A. (2000). The structural basis of ribosome activity in peptide bond synthesis. Science, 289(5481), 920-930. pubmed.ncbi.nlm.nih.gov
  2. Kozak, M. (1986). Point mutations define a sequence flanking the AUG initiator codon that modulates translation by eukaryotic ribosomes. Cell, 44(2), 283-292. pubmed.ncbi.nlm.nih.gov
  3. Harding, H. P., Novoa, I., Zhang, Y., Zeng, H., Wek, R., Schapira, M., & Ron, D. (2000). Regulated translation initiation controls stress-induced gene expression in mammalian cells. Molecular Cell, 6(5), 1099-1108. pubmed.ncbi.nlm.nih.gov
  4. Ingolia, N. T., Ghaemmaghami, S., Newman, J. R. S., & Weissman, J. S. (2009). Genome-wide analysis in vivo of translation with nucleotide resolution using ribosome profiling. Science, 324(5924), 218-223. pubmed.ncbi.nlm.nih.gov
  5. The Nobel Foundation. (2009). The Nobel Prize in Chemistry 2009: Venkatraman Ramakrishnan, Thomas A. Steitz and Ada E. Yonath. nobelprize.org
  6. Alberts, B., Johnson, A., Lewis, J., Raff, M., Roberts, K., & Walter, P. (2002). From RNA to protein. In Molecular biology of the cell (4th ed.). Garland Science. ncbi.nlm.nih.gov
Key terms
Peptidyl transferase centre
The catalytic site of the large ribosomal subunit, built entirely from ribosomal RNA with no protein within about 18 angstroms.
Ternary complex
eIF2 with GTP bound to initiator methionyl-tRNA, the species whose availability limits eukaryotic initiation.
Kozak context
The sequence around a start codon, with a purine at minus 3 and guanine at plus 4 contributing most, that determines how efficiently a scanning ribosome initiates.
4E-BP
A binding protein that sequesters eIF4E until mTORC1 phosphorylates it, thereby coupling nutrient signalling to cap-dependent initiation.
Integrated stress response
The convergence of GCN2, PERK, PKR and HRI on serine 51 of eIF2 alpha, lowering global initiation while raising translation of specific messages.
Upstream open reading frame
A short reading frame in a 5 untranslated region whose translation modulates initiation at the main start codon, as in ATF4.
Iron-responsive element
A stem-loop bound by iron regulatory proteins when iron is low; in a 5 untranslated region it blocks translation, in a 3 untranslated region it stabilises the message.
Ribosome profiling
Sequencing of nuclease-protected ribosome footprints to map ribosome occupancy across the transcriptome at near-codon resolution.

Folding, Chaperones and Degradation

  • State Anfinsen's conclusion and explain why it does not remove the need for chaperones in a cell.
  • Distinguish the Hsp70, chaperonin and Hsp90 systems by substrate, mechanism and client type.
  • Describe ubiquitin-dependent proteolysis and how two classes of drug exploit it.

Unfolding a protein and getting it back

Christian Anfinsen worked with bovine pancreatic ribonuclease A: 124 residues, four disulfide bonds, a small robust enzyme with an easy assay. He denatured it completely in 8 molar urea with beta-mercaptoethanol, which unfolds the chain and reduces all four disulfides to free thiols. The eight cysteines can pair in 105 different ways, only one of which is correct, and the denatured material had no detectable activity.

Then he removed the urea and the reducing agent and let it sit under oxidising conditions. Activity returned, close to fully. The chain had found its native fold and the correct one of 105 disulfide arrangements, with nothing present but buffer.

The conclusion, for which Anfinsen took a share of the 1972 Nobel Prize in Chemistry and which he set out in a 1973 review, is the thermodynamic hypothesis: for a small protein under physiological conditions, the native structure is the thermodynamic minimum and all the information needed to reach it is in the amino acid sequence.

Why this matters: Anfinsen showed that folding needs no external information. He did not show that it needs no help, and the rest of this lesson is about the difference.

Two problems Anfinsen's tube did not have

Cyrus Levinthal pointed out the first. Consider a 100 residue chain with, conservatively, three accessible conformations per residue. That is 3 to the hundredth power, about 5 times 10 to the 47th, possible conformations. Sampling each for a picosecond would take longer than the age of the universe. Yet proteins fold in milliseconds to seconds. Folding therefore cannot be a random search; the energy landscape must be funnelled, so that partially correct structures are on average lower in energy and the chain is steered rather than searching.

The second problem is the cell itself. Anfinsen's protein refolded at low concentration, in dilute buffer, as a complete chain. A cell interior holds 300 to 400 grams of macromolecule per litre. A nascent chain emerges from the ribosome exit tunnel amino terminus first, at 5 to 10 residues per second, so its carboxy-terminal half does not exist while its amino-terminal half is already exposed. Hydrophobic surfaces that would be buried in the finished protein are on display, in a crowded environment full of other partly folded chains with the same problem. Aggregation is second order in concentration; folding is first order. Crowding therefore favours aggregation.

Chaperones exist to lose that race on purpose: they bind exposed hydrophobic segments and release them in cycles, keeping the chain out of aggregates while it explores its funnel.

Three chaperone systems, three jobs

SystemExamplesMechanismTypical client
Ribosome-associatedTrigger factor in bacteria, NAC in eukaryotesSits at the exit tunnel and shields the emerging chainEvery nascent chain
Hsp70DnaK, BiP, cytosolic Hsp70ATP-driven cycles of binding and release of short hydrophobic segments, with J-domain proteins delivering substrate and stimulating hydrolysis and nucleotide exchange factors resetting the cycleNascent and stress-denatured proteins, broadly
ChaperoninGroEL with GroES in bacteria, TRiC or CCT in eukaryotesEncloses a single substrate in a cage whose interior becomes hydrophilic on ATP binding, giving the chain a protected compartment to fold inA defined subset, including actin and tubulin for TRiC
Hsp90Hsp90 alpha and beta, with co-chaperones such as Cdc37 and p23Late-acting; stabilises near-native conformations, often holding a client in an activatable stateKinases, steroid hormone receptors, telomerase
Small heat shock proteinsHSPB1, alphaB-crystallinATP-independent holdases that trap unfolding intermediates for later refoldingStress-destabilised proteins

The Hsp90 row explains a class of drugs. Because Hsp90 clients are enriched for signalling proteins that many tumours depend on, including mutant kinases, Hsp90 inhibitors such as geldanamycin derivatives were developed as anticancer agents. They work in the sense of degrading many clients at once, and that lack of selectivity is also why they have been difficult to use clinically.

When the folding load exceeds capacity, HSF1 trimerises, enters the nucleus and switches on the heat shock genes. That regulatory loop is what makes chaperone genes heat shock proteins by name.

The secretory pathway has its own rules

Proteins entering the endoplasmic reticulum fold in an oxidising compartment, so disulfide bonds form there and not in the cytosol. Protein disulfide isomerase both catalyses formation and, crucially, shuffles incorrect pairings, which is the enzymatic version of what Anfinsen's tube did slowly.

Glycosylation is used as a folding tag. N-linked glycans are added to asparagine in the sequon Asn-X-Ser or Thr, and trimming of the glycan is read by calnexin and calreticulin, which bind monoglucosylated glycans and retain the protein for another folding attempt. UGGT re-adds a glucose to incompletely folded proteins, so the cycle repeats until the protein passes or is given up on. This is a timer implemented in sugar chemistry.

Proteins that fail are retrotranslocated to the cytosol, ubiquitylated and degraded, a process called ER-associated degradation. If unfolded protein accumulates faster than it can be cleared, the unfolded protein response fires through three sensors: IRE1, which splices XBP1 mRNA unconventionally in the cytosol; PERK, which phosphorylates eIF2 alpha as in the previous lesson; and ATF6, which is cleaved in the Golgi to release a transcription factor. Together they raise chaperone capacity, lower the incoming load and, if the stress persists, trigger apoptosis.

Cystic fibrosis makes the pathway clinical. The common F508del allele encodes a CFTR protein that has essentially normal channel function if it can reach the membrane, but folds inefficiently and is captured by ER quality control and degraded. It is a trafficking disease produced by a folding defect. Correctors such as lumacaftor and tezacaftor bind the protein and improve folding so that more reaches the surface, and a potentiator, ivacaftor, increases channel open probability once it is there. Two drugs, two different steps, one mutation.

Alpha-1 antitrypsin deficiency shows the other failure mode. The Z variant folds into a conformation that polymerises within hepatocytes; the retained polymers cause liver disease, and the resulting shortage in plasma causes emphysema because neutrophil elastase in the lung is no longer inhibited. A single misfolding event produces two organ diseases by two different mechanisms, one of gain and one of loss.

Ubiquitin and the proteasome

Degradation is as regulated as synthesis, and the 2004 Nobel Prize in Chemistry to Aaron Ciechanover, Avram Hershko and Irwin Rose recognised the discovery that it is ATP-dependent and tag-directed.

Ubiquitin, 76 residues, is attached through its carboxy-terminal glycine to a lysine side chain of the substrate by a three-enzyme relay: E1 activates ubiquitin using ATP, E2 carries it, and E3 provides substrate specificity. Humans have two E1s, about forty E2s and more than six hundred E3s, and that architecture is the point. Specificity is concentrated in the largest and most diverse layer.

Chain topology carries meaning. Lysine 48 linked chains of four or more target the substrate to the proteasome. Lysine 63 linked chains signal in DNA repair and trafficking without causing degradation. Monoubiquitylation regulates histones and endocytosis. Same modifier, different messages, decided by which lysine of ubiquitin is used in the chain.

The 26S proteasome is a 20S barrel of four stacked heptameric rings, whose inner beta rings carry three distinct peptidase activities in a chamber that only unfolded chains can enter, capped at one or both ends by 19S regulatory particles that recognise the ubiquitin chain, remove it for recycling, and use ATP-driven unfoldase activity to thread the substrate in. Compartmentalising the proteolytic sites inside a barrel is how the cell keeps a powerful protease from destroying everything.

In short: The proteasome is a shredder with a narrow slot. Nothing gets destroyed that has not been tagged and actively fed in.

Two drug classes exploit this. Bortezomib inhibits the chymotrypsin-like activity of the 20S core and is effective in multiple myeloma, which is plausibly because plasma cells produce enormous quantities of immunoglobulin and therefore live close to the limit of their degradative capacity. Second, and more interesting mechanistically, thalidomide and its analogues bind the E3 substrate receptor cereblon and change its surface so that it recruits proteins it would never normally touch, notably the transcription factors IKZF1 and IKZF3, and destroys them. That is not inhibition; it is redirection. The molecular glue and PROTAC field is built on the observation that you can degrade a protein rather than block it, which brings targets without a druggable pocket into range.

Aggregates too large for the proteasome, and whole organelles, are handled by autophagy instead, delivered inside double-membraned autophagosomes to the lysosome.

What structure prediction did and did not solve

AlphaFold2, reported by John Jumper and colleagues in 2021, predicts three-dimensional structure from sequence at accuracy comparable to experiment for a large fraction of single domains, and the associated database now covers most known protein sequences. That is a genuine transformation of structural biology.

It is worth being exact about what it does. It predicts a static structure, not a folding pathway; it says nothing about how the chain gets there or how long it takes. It is trained on the Protein Data Bank plus multiple sequence alignments, so it performs best where evolutionary information is rich and worst on orphan sequences and designed proteins. It handles intrinsically disordered regions by reporting low confidence, which is informative but not a structure. And it is not a reliable predictor of the effect of a point mutation on stability or function, because it was not trained to model energetics. Anfinsen's question, how the chain finds the minimum, is not the question AlphaFold answers.

The upshot: Sequence to structure is largely solved for well-represented folds. Sequence to stability, to dynamics, and to the effect of a variant is not.

Common misconceptions

  • "Anfinsen showed proteins fold without help." He showed a small, disulfide-bonded protein refolds in dilute buffer. In a cell at 300 grams per litre, with chains emerging half-made from ribosomes, chaperones are required.
  • "Chaperones contain folding information." They do not specify structure. They prevent unproductive interactions and give the chain repeated chances to fold according to its own sequence.
  • "Heat shock proteins only matter under heat shock." Hsp70 and TRiC are needed constitutively for normal nascent chain folding; the heat shock name reflects how they were discovered.
  • "Ubiquitylation means degradation." Only lysine 48 linked polyubiquitin reliably signals proteasomal destruction. Lysine 63 chains and monoubiquitin signal repair, trafficking and transcription.
  • "AlphaFold solved protein folding." It predicts static structures well. Folding pathways, stability changes on mutation, and conformational dynamics remain open.

What to carry forward

  • Ribonuclease A refolds to full activity from complete denaturation, so the native state is the thermodynamic minimum encoded by sequence.
  • Levinthal's counting argument rules out random search, requiring a funnelled landscape; cellular crowding and co-translational exposure require chaperones.
  • Hsp70 cycles on short hydrophobic segments with J-protein and exchange factor help; chaperonins enclose single substrates; Hsp90 acts late on signalling clients; small heat shock proteins hold without ATP.
  • The endoplasmic reticulum adds disulfide isomerisation and a glycan-based calnexin timer, with failure routed to ER-associated degradation and the unfolded protein response.
  • F508del CFTR is a folding and trafficking disease treated by a corrector plus a potentiator; alpha-1 antitrypsin Z causes liver disease by polymer gain and lung disease by plasma loss.
  • E1, E2 and more than six hundred E3s attach ubiquitin, and chain linkage decides the message; lysine 48 chains route to the 26S proteasome.
  • Bortezomib inhibits the proteasome; thalidomide analogues redirect the cereblon E3 to destroy IKZF1 and IKZF3, which founded targeted protein degradation.
  • AlphaFold predicts structures, not folding pathways, stabilities or variant effects.

Sources

  1. Anfinsen, C. B. (1973). Principles that govern the folding of protein chains. Science, 181(4096), 223-230. pubmed.ncbi.nlm.nih.gov
  2. Hartl, F. U., Bracher, A., & Hayer-Hartl, M. (2011). Molecular chaperones in protein folding and proteostasis. Nature, 475(7356), 324-332. pubmed.ncbi.nlm.nih.gov
  3. Jumper, J., Evans, R., Pritzel, A., Green, T., Figurnov, M., Ronneberger, O., et al. (2021). Highly accurate protein structure prediction with AlphaFold. Nature, 596(7873), 583-589. pubmed.ncbi.nlm.nih.gov
  4. The Nobel Foundation. (2004). The Nobel Prize in Chemistry 2004: Aaron Ciechanover, Avram Hershko and Irwin Rose. nobelprize.org
  5. Cooper, G. M. (2000). Protein folding and processing. In The cell: A molecular approach (2nd ed.). Sinauer Associates. ncbi.nlm.nih.gov
  6. Alberts, B., Johnson, A., Lewis, J., Raff, M., Roberts, K., & Walter, P. (2002). The shape and structure of proteins. In Molecular biology of the cell (4th ed.). Garland Science. ncbi.nlm.nih.gov
Key terms
Thermodynamic hypothesis
Anfinsen's conclusion that the native structure of a small protein is its thermodynamic minimum, specified entirely by its sequence.
Levinthal's paradox
The observation that random conformational search would take astronomically long, implying a funnelled energy landscape.
Hsp70 cycle
ATP-driven binding and release of short hydrophobic segments, assisted by J-domain proteins and nucleotide exchange factors.
Chaperonin
A double-ring complex such as GroEL-GroES or TRiC that encloses a single substrate in a protected cage for folding.
Calnexin cycle
The endoplasmic reticulum quality control loop in which glycan trimming and re-glucosylation by UGGT retain incompletely folded glycoproteins for further attempts.
ER-associated degradation
Retrotranslocation of terminally misfolded secretory proteins to the cytosol for ubiquitylation and proteasomal destruction.
K48 polyubiquitin chain
A chain of four or more ubiquitins linked through lysine 48, the canonical signal for delivery to the 26S proteasome.
Molecular glue
A compound such as thalidomide that alters an E3 ligase surface so that it recruits and degrades a protein it would not normally touch.

Module 7: The Methods Bench

The instruments of molecular biology taught with their failure modes attached: what each one measures, what it cannot see, and the control that separates a result from an artefact.

Cutting and Copying: Cloning, PCR, qPCR and Blots

  • Explain restriction digestion, ligation and modern assembly methods, and the failure modes of each.
  • Design a PCR and a quantitative PCR experiment with the controls that make the result interpretable.
  • Carry out a delta delta Cq calculation and state the assumptions it rests on.

A gel full of bands and no way to ask a question

By the mid 1970s you could digest genomic DNA with a restriction enzyme and separate the fragments by size on an agarose gel. What you saw was a smear: hundreds of thousands of fragments, indistinguishable. There was no way to ask whether a particular sequence was present, or in which band.

Edwin Southern's 1975 solution was mechanical and complete. Denature the DNA in the gel to single strands, lay a nitrocellulose membrane on top, and let buffer wick through the gel into a stack of dry paper towels. The DNA moves with the flow and sticks to the membrane, preserving the spatial arrangement of the gel. Then hybridise a labelled single-stranded probe, wash away what has not paired, and expose to film. Every band that lights up contains sequence complementary to your probe.

The idea generalised immediately: the same procedure applied to RNA became the northern blot, and applied to protein with an antibody in place of a probe became the western blot. Those names are a joke about Southern's surname that has outlived every attempt to replace it.

Remember: A blot answers "is this sequence present, and at what size, and roughly how much". It does not answer "what is the sequence", and no amount of quantification of a blot turns it into a sequencing result.

Cutting DNA on purpose

Restriction endonucleases are bacterial defences against phage. Type II enzymes, the useful ones, recognise short palindromic sites and cut within them: EcoRI cuts GAATTC between G and A on both strands, leaving four-base 5 prime overhangs, while SmaI cuts CCCGGG in the middle, leaving blunt ends. The host protects its own genome by methylating the same site, which is also why a plasmid grown in a Dam-positive strain may be uncuttable by an enzyme blocked by that methylation, a failure that has cost many people an afternoon.

Two more failure modes are worth knowing.

  • Star activity. At high glycerol, high enzyme concentration, wrong buffer or extended incubation, several enzymes relax their specificity and cut related sites. A digest that produces more bands than the map predicts should raise this before it raises a cloning error.
  • Partial digestion. Incomplete cutting produces a ladder of intermediate products that can be mistaken for extra sites.

Classical cloning ligates a cut insert into a cut vector with T4 DNA ligase, transforms bacteria, and selects on antibiotic. Modern assembly has largely moved on: Gibson assembly joins fragments with 20 to 40 base pair overlapping ends using an exonuclease, polymerase and ligase in one isothermal reaction, and Golden Gate uses type IIS enzymes such as BsaI, which cut outside their recognition sequence and therefore leave user-chosen four-base overhangs, allowing many fragments to be assembled in a defined order in one tube. Both remove the requirement that convenient restriction sites exist where you want them.

PCR, and why the exponential is an idealisation

The 1988 paper by Randall Saiki, David Gelfand and colleagues is the one that made PCR practical, because it substituted the thermostable polymerase from Thermus aquaticus for the Klenow fragment that had to be re-added after every denaturation step. Kary Mullis shared the 1993 Nobel Prize in Chemistry for the concept.

The cycle is three steps. Denature at 94 to 98 degrees; anneal at a temperature set by the primers, typically 50 to 65 degrees; extend at 72 degrees, at roughly 1 kilobase per minute for Taq. Two primers pointing toward each other define the product, and after the first few cycles the dominant species is the discrete fragment bounded by them.

Yield is often written as 2 to the n. It is not, and the discrepancy is where diagnostic quantification lives.

  • Efficiency per cycle is below 1 in practice; a well-optimised reaction runs at 90 to 100 percent, so amplification is (1 plus E) to the n.
  • Reactions plateau. Polymerase is inactivated, nucleotides and primers are consumed, and abundant product reanneals to itself faster than primers can bind. Endpoint yield is therefore nearly independent of starting quantity, which is exactly why endpoint gel band brightness is not a measure of input.

Failure modes worth designing against. Primer dimers, where the two primers pair with each other and amplify a short product that outcompetes the target. Hairpins in a primer or template, particularly in GC-rich regions, which need additives such as DMSO or betaine. Non-specific priming at an annealing temperature that is too low. Allele dropout, where a variant under a primer site prevents amplification of one allele and makes a heterozygote look homozygous, which matters clinically. And polymerase error: Taq lacks a proofreading exonuclease and makes roughly one error per 104 to 105 bases incorporated, so cloning a PCR product without a high fidelity polymerase means sequencing several clones.

The most consequential failure is contamination. A PCR product is present at 1012 or more copies per tube, so a single aerosol droplet carries enough template to produce a convincing false positive. Laboratories that do diagnostic PCR use physically separate pre-amplification and post-amplification areas with one-way workflow, and often substitute dUTP for dTTP so that carried-over product can be destroyed by uracil-DNA glycosylase before the next reaction. A no-template control is not a formality; it is the only thing standing between you and this failure.

Quantitative PCR, worked properly

Quantitative PCR measures fluorescence at every cycle and reads out the cycle at which signal crosses a threshold, the quantification cycle Cq, formerly called Ct. Because the reaction is exponential in its early phase, Cq is linear in the logarithm of starting quantity. Halving the input moves Cq by one cycle at perfect efficiency.

Two chemistries are common. SYBR Green binds any double-stranded DNA, which is cheap and requires a melt curve at the end to confirm a single product, since primer dimers fluoresce just as well as your amplicon. Hydrolysis probes such as TaqMan carry a fluorophore and quencher on an internal oligonucleotide that is degraded by the polymerase only when the correct sequence is amplified, which adds specificity and allows multiplexing.

Now the arithmetic. The comparative method, delta delta Cq, computes:

  • Delta Cq for each sample: Cq of the target minus Cq of a reference gene, which normalises for input amount.
  • Delta delta Cq: delta Cq of the treated sample minus delta Cq of the control sample.
  • Fold change: 2 to the power of minus delta delta Cq.

A worked example. Control sample: target Cq 24.0, reference Cq 20.0, so delta Cq is 4.0. Treated sample: target Cq 21.5, reference Cq 20.2, so delta Cq is 1.3. Delta delta Cq is 1.3 minus 4.0, which is minus 2.7. Fold change is 2 to the power of 2.7, about 6.5-fold up.

That calculation contains three assumptions, and each is a place experiments go wrong.

  1. Both amplicons amplify at close to 100 percent efficiency. Using 2 as the base is only valid then. Measure efficiency from the slope of a standard curve across a dilution series: efficiency equals 10 to the power of minus 1 over slope, minus 1, and a slope of minus 3.32 corresponds to 100 percent. If efficiencies differ between target and reference, the comparative method is invalid and an efficiency-corrected model is required.
  2. The reference gene is genuinely stable across the conditions. This is violated more often than it is checked. Actin and GAPDH both change with proliferation, hypoxia and glucose, which are exactly the conditions people study. The right approach, set out by Jo Vandesompele and colleagues in 2002, is to test several candidates and normalise to the geometric mean of the most stable ones.
  3. The signal comes from RNA, not DNA. For reverse transcription quantitative PCR, a no-reverse-transcriptase control is mandatory, because a primer pair that sits within one exon will amplify contaminating genomic DNA perfectly well. Designing primers to span an exon-exon junction largely removes the problem.

The MIQE guidelines, published by Stephen Bustin and colleagues in 2009, exist because so much published quantitative PCR omitted efficiency, reference gene validation and control descriptions that results could not be evaluated or reproduced.

So what?: A fold change from quantitative PCR is a ratio of two exponentials estimated from two threshold crossings. Every assumption in that sentence needs a control, and the method is only as good as the weakest one.

Digital PCR sidesteps the problem for absolute quantification: partition the reaction into twenty thousand droplets so that most contain zero or one template molecule, amplify to endpoint, count positive droplets and apply a Poisson correction. No standard curve and no reference gene, at the cost of throughput and dynamic range.

Reading a western blot honestly

Western blots deserve a paragraph because they are the most frequently over-interpreted figure in molecular biology. The antibody is the assay, and antibody validation is a recognised problem: many commercial antibodies detect additional bands, and a band at the expected molecular weight is not proof of identity. The persuasive controls are a knockout or knockdown lane showing the band disappears, and a peptide competition. Chemiluminescent detection saturates, so densitometry on an overexposed film is not quantitative, and a loading control run on the same blot must itself be in the linear range. When you see a single tight cropped band with no molecular weight markers, the figure is telling you less than it appears to.

Common misconceptions

  • "PCR yield doubles every cycle." Efficiency is below 1 and reactions plateau as reagents deplete and product reanneals, which is why endpoint band brightness does not measure input.
  • "A no-template control is a formality." Amplicon carryover is the commonest source of false positives in any laboratory that runs the same assay repeatedly.
  • "GAPDH is a stable reference gene." It changes with glucose, hypoxia and proliferation. Reference genes must be validated for the specific comparison being made.
  • "Delta delta Cq works whatever the efficiencies." Using 2 as the base assumes both amplicons run near 100 percent efficiency; unequal efficiencies require an efficiency-corrected calculation.
  • "A band of the right size on a western blot identifies the protein." The antibody is the assay; without a knockdown or knockout lane, the band is a hypothesis.

Pulling it together

  • Southern's capillary transfer preserved gel geometry on a membrane and made sequence-specific detection possible; northern and western blots are the same idea for RNA and protein.
  • Type II restriction enzymes cut palindromes, are blocked by host methylation, and can show star activity under wrong conditions; Gibson and Golden Gate assembly removed the dependence on convenient sites.
  • Thermostable Taq made thermal cycling practical, but amplification is (1 plus E) to the n and plateaus, so endpoint yield does not report input.
  • Primer dimers, GC structure, allele dropout, polymerase error and above all amplicon contamination are the routine failure modes.
  • Quantitative PCR reads a threshold crossing, with SYBR requiring a melt curve and hydrolysis probes adding specificity.
  • Delta delta Cq assumes equal near-perfect efficiencies, a validated stable reference, and no genomic DNA carryover; digital PCR avoids the first two for absolute counts.
  • Western blot interpretation depends entirely on antibody validation and on staying within the linear range of detection.

Sources

  1. Southern, E. M. (1975). Detection of specific sequences among DNA fragments separated by gel electrophoresis. Journal of Molecular Biology, 98(3), 503-517. pubmed.ncbi.nlm.nih.gov
  2. Saiki, R. K., Gelfand, D. H., Stoffel, S., Scharf, S. J., Higuchi, R., Horn, G. T., Mullis, K. B., & Erlich, H. A. (1988). Primer-directed enzymatic amplification of DNA with a thermostable DNA polymerase. Science, 239(4839), 487-491. pubmed.ncbi.nlm.nih.gov
  3. Bustin, S. A., Benes, V., Garson, J. A., Hellemans, J., Huggett, J., Kubista, M., et al. (2009). The MIQE guidelines: Minimum information for publication of quantitative real-time PCR experiments. Clinical Chemistry, 55(4), 611-622. pubmed.ncbi.nlm.nih.gov
  4. Vandesompele, J., De Preter, K., Pattyn, F., Poppe, B., Van Roy, N., De Paepe, A., & Speleman, F. (2002). Accurate normalization of real-time quantitative RT-PCR data by geometric averaging of multiple internal control genes. Genome Biology, 3(7), RESEARCH0034. pubmed.ncbi.nlm.nih.gov
  5. Alberts, B., Johnson, A., Lewis, J., Raff, M., Roberts, K., & Walter, P. (2002). Isolating, cloning, and sequencing DNA. In Molecular biology of the cell (4th ed.). Garland Science. ncbi.nlm.nih.gov
  6. Cooper, G. M. (2000). Detection of nucleic acids and proteins. In The cell: A molecular approach (2nd ed.). Sinauer Associates. ncbi.nlm.nih.gov
Key terms
Southern blot
Capillary transfer of size-separated DNA from a gel to a membrane, followed by hybridisation with a labelled probe.
Star activity
Relaxed specificity of a restriction enzyme under non-optimal conditions, producing cuts at sites related to but not identical with its recognition sequence.
Golden Gate assembly
Ordered multi-fragment assembly using type IIS enzymes that cut outside their recognition site, leaving user-defined overhangs.
Amplification efficiency
The per-cycle multiplication factor of a PCR, measured from the slope of a standard curve; a slope of minus 3.32 corresponds to 100 percent.
Quantification cycle
The cycle at which fluorescence crosses a set threshold, linear in the logarithm of starting template quantity.
Delta delta Cq
The comparative quantification method, valid only when target and reference amplify with similar near-perfect efficiencies.
No-reverse-transcriptase control
A reaction omitting reverse transcriptase, which reveals amplification from contaminating genomic DNA.
Digital PCR
Partitioning of a reaction into thousands of droplets so that endpoint positives can be counted and converted to absolute copy number by Poisson statistics.

Sequencing, from Sanger to Long Reads

  • Explain chain-termination sequencing and the situations in which it remains the appropriate method.
  • Describe short-read sequencing by synthesis, its error profile, and the variant classes it systematically misses.
  • Compare long-read platforms and justify a platform choice from the biological question.

One missing oxygen

Frederick Sanger's 1977 method rests on a single chemical difference. A dideoxynucleoside triphosphate has hydrogen where a deoxynucleotide has a 3 prime hydroxyl. A polymerase will incorporate it happily, and then cannot form the next phosphodiester bond, because the 3 prime hydroxyl that would attack the incoming nucleotide is not there. Chain extension stops at that base.

Run four reactions, each with all four normal nucleotides plus a trace of one dideoxy terminator, and each reaction produces a nested set of fragments ending at every occurrence of that base. Separate them by length at single-nucleotide resolution and read the sequence off the ladder. Within the same year Sanger's group used it to determine the complete 5,386 nucleotide genome of bacteriophage phi X174, the first complete genome of any organism, and the 1980 Nobel Prize in Chemistry followed.

Modern Sanger sequencing puts a different fluorophore on each of the four terminators, so one reaction suffices, and separates by capillary electrophoresis. Read length is 500 to 1000 bases with per-base accuracy above 99.99 percent in the good middle region.

Key idea: Sanger sequencing reads a population of molecules at once and gives you their average. That is why it is superb for a clonal template and poor for a mixture.

The consequences follow directly. Sanger confirmation of a variant found by another method is still standard clinical practice because per-base accuracy is high and the trace is interpretable by eye. But a variant present in less than roughly 15 to 20 percent of molecules disappears into the baseline, which is why Sanger is unsuitable for detecting low-frequency somatic mutations or minority pathogen populations. And a heterozygous insertion or deletion causes the two alleles to run out of register from that point on, producing a superimposed double trace that is unreadable without additional work.

Sequencing by synthesis: the short read era

The 2008 paper by David Bentley and colleagues described whole human genome sequencing by reversible terminator chemistry, and it is a good anchor for how the dominant platform works.

  1. Fragment the DNA and ligate adapters, creating a library. This step usually includes PCR, which is where several biases enter.
  2. Amplify single molecules in situ on a flow cell into clusters of identical copies, so that the signal from one original molecule is bright enough to image.
  3. Each cycle, add all four nucleotides, each carrying a distinct fluorophore and a reversible 3 prime block. One is incorporated per cluster. Image the flow cell. Cleave the fluorophore and the block. Repeat.
  4. Read the second end of the fragment the same way, giving paired-end reads whose separation is known approximately from the library insert size.

Read length is 100 to 300 bases, raw error is around 0.1 to 1 percent and is dominated by substitutions rather than insertions and deletions, and throughput is enormous. The error profile is worth understanding rather than memorising. Within a cluster the copies drift out of step: some molecules fail to incorporate in a cycle and lag behind, others lose the block and run ahead. This phasing and prephasing degrades signal purity as cycles accumulate, which is why quality scores fall along the read and why longer reads on this chemistry are harder than they look.

Two more artefacts matter in practice. Library PCR amplifies GC-balanced fragments better than extreme ones, so coverage dips in GC-rich promoters and AT-rich regions, and PCR duplicates must be marked and removed before variant calling or the same original molecule is counted many times as independent evidence. And on patterned flow cells with exclusion amplification, free adapters can cause index hopping, assigning reads to the wrong sample in a multiplexed run, which unique dual indexing was introduced to control.

How much coverage, and why

Coverage is the average number of reads spanning a base. Reads land approximately at random, so depth at any given base is roughly Poisson distributed around the mean. At 30-fold mean coverage, the chance that a given position gets fewer than about 10 reads is small, and 10 reads is roughly what is needed to call a heterozygote confidently, since each allele is sampled binomially and you must distinguish a genuine 50 percent allele fraction from sampling noise and from sequencing error. That is where the conventional 30-fold standard for germline whole genome sequencing comes from; it is a statement about the tail of a distribution, not a round number.

Exome sequencing typically targets 100-fold because capture efficiency is uneven across targets. Tumour sequencing needs far more, often 500-fold to several thousand-fold for targeted panels, because a subclonal mutation in 5 percent of cells at a diploid locus is present in roughly 2.5 percent of reads and must be distinguished from a sequencing error rate of similar magnitude.

What matters here: Choose depth from the allele fraction you need to detect and the error rate you must beat, not from convention.

What short reads cannot see

This list is the reason long reads exist.

  • Repeats longer than the fragment. A 150 base read inside a 6 kilobase repeat cannot be placed uniquely. Segmental duplications, centromeres and the acrocentric short arms were simply absent from the reference for two decades.
  • Structural variation. Inversions, translocations and large insertions are inferred indirectly from anomalous read pair orientation and split reads, with high false discovery rates.
  • Tandem repeat expansions. If the expanded allele is longer than the read, its length cannot be measured, which matters for a whole class of neurological disease.
  • Phasing. Determining which variants are on the same chromosome requires reads spanning both, so haplotypes beyond a few hundred bases must be inferred statistically.
  • Base modifications. Amplification erases methylation, so it must be interrogated separately by bisulfite conversion or enzymatic methods.

Long reads: two chemistries, different trade-offs

PacBio single molecule real timeOxford Nanopore
PrincipleA polymerase immobilised in a zero-mode waveguide is watched incorporating fluorescent nucleotides in real timeA single strand is ratcheted through a protein pore by a motor enzyme; the ionic current change reports the bases in the pore
Typical read length15 to 25 kilobases for circular consensus reads10 kilobases to over a megabase with ultra-long protocols
AccuracyCircular consensus, resequencing the same circularised molecule many times, gives above 99.9 percentHigher raw error, historically in the percent range, improved substantially by newer pores and basecallers; homopolymers remain the weak point
ModificationsDetectable from polymerase kineticsDetectable directly from the current signal, with no separate protocol
Practical constraintInstrument cost, lower throughput per dollar than short readsRequires high molecular weight, undamaged input DNA; read length is set by how gently the DNA was extracted

The completion of the human genome makes the case concretely. The 2001 draft and every version until recently left roughly 8 percent of the sequence unresolved: centromeric satellite arrays, the short arms of the acrocentric chromosomes, segmental duplications, ribosomal DNA arrays. Sergey Nurk and colleagues reported the first truly complete assembly in 2022, using long reads from both platforms on a cell line homozygous across the genome, which removed the additional problem of separating two haplotypes. That work added nearly 200 million base pairs and almost 2,000 predicted genes to the reference. It could not have been done with short reads at any depth, because the missing information was not a coverage problem.

In short: Short reads fail on repeats for a reason that more of them does not fix. When the question is about structure, length or phase, read length is the variable that matters.

Choosing a platform

Match the method to the question, not to the budget line.

  • Confirming a single known variant in one sample: Sanger.
  • Calling single nucleotide variants and small indels across many samples: short reads, 30-fold for germline.
  • Detecting a 2 percent subclonal mutation: deep targeted short-read sequencing with unique molecular identifiers to distinguish real variants from PCR and sequencing error.
  • Resolving a structural rearrangement, a repeat expansion, or haplotype phase: long reads.
  • Methylation together with sequence, without a separate library: nanopore.
  • A reference-quality assembly of a new genome: long reads, with short reads or Hi-C for scaffolding and polishing.

One more caution applies to all of them. Aligning to a single linear reference introduces reference bias: reads carrying alleles absent from the reference align less well or not at all, which systematically under-detects variation in populations less represented in the reference. Pangenome graph references are the current response to that problem.

Common misconceptions

  • "Sanger sequencing is obsolete." It remains the standard for confirming a single variant, because its per-base accuracy is high and its output is directly interpretable.
  • "Deeper sequencing fixes any problem." A repeat longer than the read is unresolvable at any depth, because the limitation is read length and not sampling.
  • "A 30-fold genome means every base has 30 reads." Coverage is roughly Poisson around the mean, and the standard exists to keep the low tail above the threshold for confident heterozygote calling.
  • "PCR duplicates are harmless." They inflate apparent evidence for a variant by counting one original molecule many times, which is why they are marked and excluded.
  • "Nanopore is inaccurate." Raw error has fallen substantially and consensus accuracy is high; the residual weakness is homopolymer length, and input DNA integrity limits read length more than chemistry does.

Looking back

  • Chain termination works because a dideoxynucleotide lacks the 3 prime hydroxyl needed for the next bond; it gave the first complete genome, phi X174, in 1977.
  • Sanger reads a population average, so it confirms clonal variants well and misses minority alleles below roughly 15 to 20 percent.
  • Sequencing by synthesis clusters and images reversible terminators; phasing and prephasing degrade quality along the read, and library PCR adds GC bias and duplicates.
  • Coverage requirements follow from the allele fraction to be detected: 30-fold germline, about 100-fold exome, far deeper for subclonal tumour variants.
  • Short reads systematically miss long repeats, structural variants, repeat expansions, phase and base modifications.
  • PacBio circular consensus gives long high-accuracy reads; nanopore gives the longest reads and direct modification calling, limited by input DNA quality and homopolymers.
  • The complete human genome in 2022 required long reads and added nearly 200 megabases that no depth of short-read sequencing could have resolved.

Sources

  1. Sanger, F., Nicklen, S., & Coulson, A. R. (1977). DNA sequencing with chain-terminating inhibitors. PNAS, 74(12), 5463-5467. pubmed.ncbi.nlm.nih.gov
  2. Sanger, F., Air, G. M., Barrell, B. G., Brown, N. L., Coulson, A. R., Fiddes, C. A., et al. (1977). Nucleotide sequence of bacteriophage phi X174 DNA. Nature, 265(5596), 687-695. pubmed.ncbi.nlm.nih.gov
  3. Bentley, D. R., Balasubramanian, S., Swerdlow, H. P., Smith, G. P., Milton, J., Brown, C. G., et al. (2008). Accurate whole human genome sequencing using reversible terminator chemistry. Nature, 456(7218), 53-59. pubmed.ncbi.nlm.nih.gov
  4. Jain, M., Koren, S., Miga, K. H., Quick, J., Rand, A. C., Sasani, T. A., et al. (2018). Nanopore sequencing and assembly of a human genome with ultra-long reads. Nature Biotechnology, 36(4), 338-345. pubmed.ncbi.nlm.nih.gov
  5. Nurk, S., Koren, S., Rhie, A., Rautiainen, M., Bzikadze, A. V., Mikheenko, A., et al. (2022). The complete sequence of a human genome. Science, 376(6588), 44-53. pubmed.ncbi.nlm.nih.gov
  6. National Human Genome Research Institute. (n.d.). DNA sequencing fact sheet. National Institutes of Health. genome.gov
Key terms
Dideoxynucleotide
A nucleotide analogue lacking the 3 prime hydroxyl, so that incorporation terminates chain extension.
Reversible terminator
A nucleotide carrying a removable fluorophore and 3 prime block, allowing one base to be added and imaged per cycle.
Phasing and prephasing
Loss of synchrony within a cluster as molecules lag or run ahead, degrading base quality as cycles accumulate.
PCR duplicate
Multiple reads deriving from the same original library molecule, which must be marked so they are not counted as independent evidence.
Index hopping
Misassignment of reads to the wrong sample in a multiplexed run, controlled by unique dual indexing.
Circular consensus sequencing
Repeated passes of a polymerase around a circularised template, so that a long read can also be highly accurate.
Reference bias
Systematic under-detection of alleles absent from a single linear reference, addressed by pangenome graph references.
Unique molecular identifier
A random barcode added before amplification, allowing PCR and sequencing errors to be distinguished from genuine low-frequency variants.

CRISPR, its Off-Target Problem, and the Single Cell

  • Explain Cas9 target recognition, the role of the PAM, and how repair outcome determines the type of edit obtained.
  • Assess off-target and on-target risks of genome editing and the experimental methods used to measure them.
  • Interpret single-cell RNA sequencing data with its dropout, doublet, ambient RNA and dissociation artefacts in view.

Two RNAs, one nuclease, one tube

In August 2012 Martin Jinek, Krzysztof Chylinski, Emmanuelle Charpentier, Jennifer Doudna and colleagues reported a purely in vitro result. They took Cas9 from Streptococcus pyogenes, showed it required two RNAs to cut DNA, fused those two into a single chimeric guide, and then demonstrated that changing 20 nucleotides in that guide changed where the enzyme cut. No cells, no genetics, just protein, RNA and plasmid DNA in a tube.

What made it consequential was the programmability. Zinc finger nucleases and TALENs could already cut chosen sequences, but retargeting either meant re-engineering a protein, which took weeks of work by specialists. Retargeting Cas9 means ordering a different oligonucleotide. The 2020 Nobel Prize in Chemistry went to Charpentier and Doudna for this.

How Cas9 finds its target

The search problem is severe: one enzyme, three billion base pairs, and a 20 nucleotide target. Cas9 does not scan for the guide match directly. It scans for the protospacer adjacent motif, which for Streptococcus pyogenes Cas9 is NGG, immediately 3 prime of the target. NGG occurs roughly every eight base pairs, so there are hundreds of millions of candidate sites, but checking a three base motif is fast. Only after PAM engagement does the enzyme locally unwind the duplex and test guide complementarity, starting from the PAM-proximal end and propagating outward. If pairing holds, an R-loop forms with the guide RNA paired to the target strand and the non-target strand displaced.

Cutting then uses two nuclease domains: HNH cuts the strand paired to the guide, RuvC cuts the displaced strand. The result is a blunt double-strand break three base pairs upstream of the PAM.

Two consequences follow directly from that mechanism. First, the PAM is obligatory, so not every sequence is targetable by SpCas9, which is why other Cas proteins with different PAM requirements and engineered PAM-flexible variants were developed. Second, because pairing is checked from the PAM end outward, mismatches close to the PAM are poorly tolerated while mismatches at the distal end of the guide are often tolerated. That asymmetry is the origin of the off-target problem.

The core of it: Cas9 makes a break. What kind of edit you get is decided afterwards, by which repair pathway the cell uses, and you have limited control over that.

Editing outcomes are repair outcomes

  • Non-homologous end joining dominates. It rejoins the ends, often exactly, but repeated cutting and rejoining eventually produces small insertions or deletions that persist because they destroy the target site. Frameshifting indels give a knockout. This is efficient and works in any cell cycle phase.
  • Homology-directed repair with a supplied donor template installs a defined sequence, but it requires S or G2 phase and is typically much less efficient than end joining. Getting precise edits in non-dividing cells such as neurons or hepatocytes has been the central practical obstacle.
  • Microhomology-mediated end joining produces predictable deletions between short repeats, and because it is predictable, the indel spectrum from a given guide is substantially reproducible and can be forecast from the sequence.

Off-target editing, measured rather than predicted

Early computational off-target prediction, scoring genomic sites by mismatch count and position, turned out to be a poor guide to what actually happens. Empirical genome-wide methods replaced it.

MethodPrincipleLimitation
GUIDE-seqA short double-stranded oligonucleotide is captured into breaks in living cells and its integration sites are sequencedRequires efficient oligonucleotide uptake; sensitivity depends on editing efficiency in that cell type
CIRCLE-seq and Digenome-seqPurified genomic DNA is cut in vitro by the ribonucleoprotein and cut sites are sequencedVery sensitive, but naked DNA lacks chromatin, so it over-reports sites that are inaccessible in cells
DISCOVER-seqChromatin immunoprecipitation of the repair factor MRE11 at Cas9-induced breaks in cells or tissueDetects only sites cut often enough to accumulate repair factors

Tsai and colleagues, introducing GUIDE-seq in 2015, found off-target sites that no prediction algorithm had ranked highly, and also found guides with very few detectable off-targets. The practical lesson is that off-target behaviour is guide-specific and must be measured for each guide in the relevant cell type, not inferred.

Mitigation strategies are worth knowing because they follow from the mechanism. High-fidelity variants such as eSpCas9 and SpCas9-HF1 weaken non-specific contacts to the DNA backbone, raising the energetic threshold for cleavage so that imperfectly paired sites fail. Delivering preassembled ribonucleoprotein rather than a plasmid limits the time the nuclease is present, and exposure time multiplies off-target events. Paired nickases require two adjacent guides to produce a double-strand break, so a single off-target nick is repaired without an indel.

The on-target problems, which are underappreciated

Attention to off-target sites has sometimes obscured that the intended cut site also causes trouble.

  • Large deletions and complex rearrangements at the cut site, extending kilobases, which short amplicon sequencing across the target is blind to because the primer sites themselves are deleted. Measuring only a 200 base pair amplicon around the guide will report a clean edit while missing a 10 kilobase loss.
  • Chromosomal loss and chromothripsis when a break triggers micronucleus formation, particularly with multiple simultaneous cuts.
  • Translocations when two loci are cut at once, which is relevant to multiplexed editing of T cells.
  • p53 activation. A double-strand break activates p53 and arrests the cell, so cells that edit and survive are enriched for cells with impaired p53 signalling. This is a selection effect that operates silently.

Why this matters: The right assay for editing outcome is long-range: long-read sequencing or karyotyping, not a short amplicon around the cut site that can only see what it was designed to see.

Editing without a double-strand break

Two developments address the fact that most of the trouble comes from the break itself.

Base editing, reported by Alexis Komor and colleagues in 2016, fuses a catalytically impaired Cas9 to a deaminase. A cytosine base editor uses an APOBEC-family cytidine deaminase to convert C to U within a small window of the R-loop, plus a uracil glycosylase inhibitor to stop the cell reverting it, and a nickase to bias repair toward the edited strand; replication then fixes the change as C to T. Adenine base editors, built on an evolved tRNA adenosine deaminase, give A to G. No break, no donor, and efficiency in non-dividing cells that homology-directed repair cannot approach. The characteristic problems are bystander editing of other identical bases in the window, and guide-independent deamination of DNA and RNA by the deaminase acting on its own, which no guide design can prevent.

Prime editing, reported by Andrew Anzalone and colleagues in 2019, fuses a Cas9 nickase to a reverse transcriptase and uses an extended guide, the pegRNA, that carries both the targeting sequence and a template for the desired edit. The nicked strand primes reverse transcription off that template, writing the new sequence directly into the genome. It can install all twelve base-to-base changes and small insertions and deletions, without a donor and without a double-strand break. Efficiency is generally lower and it requires more optimisation per site, but the byproduct spectrum is far cleaner.

Clinically, the first approved CRISPR therapy, exagamglogene autotemcel for sickle cell disease and transfusion-dependent beta-thalassaemia, was authorised at the end of 2023. Its design is instructive: rather than correcting the disease mutation, it disrupts an erythroid-specific enhancer of BCL11A in the patient's own haematopoietic stem cells, de-repressing fetal haemoglobin. A knockout is much easier to achieve than a precise correction, so the therapy was designed around what editing does well.

Single-cell measurement, and what it hides

Fuchou Tang and colleagues sequenced the transcriptome of a single cell in 2009. Scale arrived with droplet methods; Evan Macosko and colleagues published Drop-seq in 2015, and the commercial descendants of that design now routinely profile tens of thousands of cells per run.

The scheme is worth understanding because every artefact follows from it. Cells are partitioned into droplets with barcoded beads. Each bead carries oligo-dT primers bearing a cell barcode shared by all primers on that bead and a unique molecular identifier that differs on every primer. Reverse transcription inside the droplet tags every captured transcript with both. Droplets are then broken and everything is amplified and sequenced together. The cell barcode says which cell; the unique molecular identifier lets you count original molecules rather than amplified copies.

Now the failure modes.

  • Dropout. Only a fraction of transcripts in a cell, commonly 10 to 30 percent, is captured. A gene that is genuinely expressed at moderate level will read as zero in many cells. A zero in the matrix is therefore not evidence of absence, and any analysis that treats it as such will find spurious differences.
  • Doublets. Two cells in one droplet share a barcode and appear as one cell with a hybrid profile. The rate scales with loading concentration and is around 0.8 percent per thousand cells recovered on common platforms, so a run of 10,000 cells carries roughly 8 percent doublets. Doublets form apparent intermediate cell states, and more than one reported transitional population has turned out to be one.
  • Ambient RNA. RNA released by cells damaged during dissociation is distributed into every droplet, so highly expressed genes from abundant cell types appear at low level everywhere. Haemoglobin transcripts in non-erythroid cells are the classic tell.
  • Dissociation artefacts. Enzymatic dissociation at 37 degrees induces immediate early and heat shock genes, FOS, JUN and the HSPs, within minutes. Solid tissue and neurons suffer most, and cell types differ in how well they survive dissociation, so the observed cell type proportions may reflect fragility rather than biology. Nuclei rather than cells, from frozen tissue, avoid much of this at the cost of losing cytoplasmic RNA.
  • Batch effects. Samples processed on different days or chips differ systematically, and integration algorithms that remove those differences can also remove real biological differences if the design confounds batch with condition.

One statistical trap deserves separate mention because it is nearly universal in the literature. Clusters are defined from the expression data, and then differential expression is tested between those clusters using the same data. The clustering step already maximised the separation the test is asked to evaluate, so p values from that comparison are anticonservative by a large margin. This is sometimes called double dipping, and the remedies are data splitting, dedicated selective inference procedures, or validating cluster markers in an independent dataset or by another method entirely.

Bottom line: A single-cell experiment measures a sparse, noisy, dissociation-perturbed sample of polyadenylated RNA. It is superb for discovering heterogeneity and generating hypotheses, and weak as a standalone quantitative claim about any one gene in any one cell.

Common misconceptions

  • "Cas9 finds its target by scanning for the guide match." It scans for the PAM first and only then tests guide complementarity, from the PAM-proximal end outward.
  • "CRISPR edits the genome." CRISPR makes a break. The cell's repair pathways make the edit, which is why knockouts are easy and precise corrections are hard.
  • "Off-target sites can be predicted computationally." Empirical methods routinely find sites that algorithms rank low, so off-targets must be measured per guide and per cell type.
  • "A clean amplicon sequencing result means a clean edit." Large deletions and rearrangements remove the primer sites and are therefore invisible to the assay that reports success.
  • "A zero in a single-cell matrix means the gene is off." Capture efficiency is 10 to 30 percent, so most zeros are dropouts rather than absence.

The takeaway

  • Cas9 was shown to be programmable by a single chimeric guide in vitro in 2012, and retargeting requires only a new oligonucleotide.
  • Targeting requires an NGG PAM, and pairing is verified from the PAM-proximal end outward, which is why distal mismatches are tolerated and off-targets exist.
  • End joining gives indels and knockouts efficiently; homology-directed repair gives precise edits but only in S and G2 and at low efficiency.
  • GUIDE-seq, CIRCLE-seq and DISCOVER-seq measure off-targets empirically because prediction is unreliable; high-fidelity variants, ribonucleoprotein delivery and paired nickases reduce them.
  • On-target hazards include kilobase deletions invisible to short amplicon assays, chromothripsis, translocations and selection for impaired p53.
  • Base editors convert C to T or A to G without a break, at the cost of bystander and guide-independent deamination; prime editing writes arbitrary small edits from a pegRNA template with a cleaner byproduct profile.
  • Droplet single-cell RNA sequencing uses a cell barcode and a unique molecular identifier per transcript, and its dropouts, doublets, ambient RNA and dissociation-induced genes must be handled explicitly.
  • Testing differential expression between clusters derived from the same data is anticonservative and requires data splitting or independent validation.

Sources

  1. Jinek, M., Chylinski, K., Fonfara, I., Hauer, M., Doudna, J. A., & Charpentier, E. (2012). A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity. Science, 337(6096), 816-821. pubmed.ncbi.nlm.nih.gov
  2. Tsai, S. Q., Zheng, Z., Nguyen, N. T., Liebers, M., Topkar, V. V., Thapar, V., et al. (2015). GUIDE-seq enables genome-wide profiling of off-target cleavage by CRISPR-Cas nucleases. Nature Biotechnology, 33(2), 187-197. pubmed.ncbi.nlm.nih.gov
  3. Komor, A. C., Kim, Y. B., Packer, M. S., Zuris, J. A., & Liu, D. R. (2016). Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage. Nature, 533(7603), 420-424. pubmed.ncbi.nlm.nih.gov
  4. Anzalone, A. V., Randolph, P. B., Davis, J. R., Sousa, A. A., Koblan, L. W., Levy, J. M., et al. (2019). Search-and-replace genome editing without double-strand breaks or donor DNA. Nature, 576(7785), 149-157. pubmed.ncbi.nlm.nih.gov
  5. Macosko, E. Z., Basu, A., Satija, R., Nemesh, J., Shekhar, K., Goldman, M., et al. (2015). Highly parallel genome-wide expression profiling of individual cells using nanoliter droplets. Cell, 161(5), 1202-1214. pubmed.ncbi.nlm.nih.gov
  6. Tang, F., Barbacioru, C., Wang, Y., Nordman, E., Lee, C., Xu, N., et al. (2009). mRNA-Seq whole-transcriptome analysis of a single cell. Nature Methods, 6(5), 377-382. pubmed.ncbi.nlm.nih.gov
  7. The Nobel Foundation. (2020). The Nobel Prize in Chemistry 2020: Emmanuelle Charpentier and Jennifer A. Doudna. nobelprize.org
Key terms
Protospacer adjacent motif
The short sequence, NGG for SpCas9, that must sit next to the target and is checked before guide complementarity is tested.
Single guide RNA
The engineered fusion of crRNA and tracrRNA whose 20 nucleotide spacer specifies the Cas9 target.
GUIDE-seq
A cell-based method that captures a tag oligonucleotide into double-strand breaks and sequences its integration sites to map off-target cleavage.
High-fidelity Cas9 variant
An engineered nuclease such as eSpCas9 or SpCas9-HF1 with weakened non-specific DNA contacts, raising the threshold for cleaving mismatched sites.
Base editor
A catalytically impaired Cas9 fused to a deaminase, converting C to T or A to G within a small window without a double-strand break.
Prime editing
A Cas9 nickase fused to reverse transcriptase, using a pegRNA that both targets the site and templates the desired edit.
Dropout
A zero count for a gene that is actually expressed, caused by the 10 to 30 percent capture efficiency of single-cell chemistry.
Doublet
Two cells captured under one barcode, producing a hybrid profile that can masquerade as an intermediate cell state.

Open the interactive version with quizzes and progress →