🧬 Biology · Undergraduate · BIO 310

Genetics

A complete undergraduate course in genetics, from Gregor Mendel's garden peas to CRISPR and whole-genome sequencing. You will learn how traits are inherited and predicted, how chromosomes move during meiosis, how genes are mapped, how DNA is copied and read into proteins, how gene expression is controlled, and how allele frequencies shift across whole populations. Every idea is taught in full on…

Start the interactive course (quizzes, progress, videos) →

Free forever. No sign-up, no ads. 18 lessons. The full lesson text is below so you can read it right here.

Module 1: Mendelian Inheritance and Its Extensions

How discrete factors pass from parents to offspring, how to predict crosses with Punnett squares, and the ways real inheritance departs from simple dominance.

Mendel's Laws and the Monohybrid Cross

  • State Mendel's law of segregation and define allele, genotype, and phenotype.
  • Build a Punnett square for a monohybrid cross and read off the ratios.
  • Distinguish homozygous from heterozygous individuals.

The big picture

Gregor Mendel crossed pea plants with purple flowers and plants with white flowers. Their first-generation offspring had purple flowers. When those offspring reproduced, white flowers appeared again. The white-flower form had not disappeared, even though he could not see it in the previous generation.

That observation gives us two separate questions: what does a plant inherit, and what does that inherited information make the plant look like? Keeping those questions separate is the first step in solving a genetics problem.

Mendel began with true-breeding lines: plants that consistently produced the same form of the trait when they self-fertilized. A true-breeding purple line continued to produce purple flowers. That known starting point helped him interpret what changed after crossing two lines.

Start with the gene, then name its versions

Genetics is the study of inheritance and variation. A gene is a stretch of DNA whose sequence contributes to a functional product, such as a protein or a working RNA molecule. Proteins carry out many jobs in cells. RNA can help make proteins or perform other cellular tasks.

A gene can have different versions, called alleles. The letters in a genetics problem label those versions. For our pea-flower example, P labels an allele associated with purple flowers and p labels an allele associated with white flowers. P and p are versions of the same gene, not two separate genes.

This example follows one gene whose variants have a clear effect on flower color. It is a useful starting model, not a claim that every trait has one controlling gene. Many traits depend on several genes and the environment.

Peas are diploid: their cells normally contain two sets of chromosomes, one set from each parent. A chromosome is DNA packaged with proteins. At the gene we are following, a pea plant therefore has two allele copies. The pair might be PP, Pp, or pp.

The two-copy rule here applies to the diploid gene we are studying. It is not a rule for every cell or organism. Reproductive cells have one chromosome set, and other organisms or chromosome regions can require different models.

Three genotypes, two flower colors

The genotype is the allele combination, such as Pp. The phenotype is an observable characteristic, such as purple flowers. Observable does not only mean visible: an enzyme's measured activity can also be a phenotype.

Homozygous means that the two alleles at the gene are the same. PP and pp are both homozygous. Heterozygous means they differ, as in Pp. You can identify either state from the letters before you know anything about the plant's appearance.

GenotypeAllele relationshipFlower-color phenotype in this model
PPHomozygousPurple
PpHeterozygousPurple
ppHomozygousWhite

P is dominant to p for this phenotype: the heterozygote Pp has the same flower-color category as PP. The white-flower phenotype is recessive in this relationship and appears in pp plants. This pattern is called complete dominance.

Dominance describes the heterozygote's phenotype. It does not mean that P destroys p, that purple is more common, or that purple plants are better adapted. A Pp plant still carries p and can pass it on.

Remember: Read the genotype first. Then use the stated relationship between the alleles to predict the phenotype. PP and Pp differ genetically even when their flowers look alike.

How two allele copies become one

A gamete is a reproductive cell, such as an egg or a sperm cell, that contributes a chromosome set at fertilization. Mendel's law of segregation says that the two alleles at a gene separate during the process that produces gametes. Each gamete receives one allele of the pair.

For a Pp parent, about half the gametes carry P and half carry p under the usual segregation model. A gamete does not receive half of P and half of p. It receives one allele copy. A PP parent produces gametes carrying P at this gene; a pp parent produces gametes carrying p.

The chromosome explanation involves meiosis, the cell divisions that reduce the number of chromosome sets. The two corresponding chromosomes separate, taking their allele copies with them. At fertilization, an egg and a sperm unite, restoring two copies in the offspring.

Follow the information through one possible fertilization: a Pp parent contributes P, and the other Pp parent contributes p. The offspring is Pp. Neither parent's whole two-letter genotype enters a single gamete.

Build the Pp x Pp square

A monohybrid cross follows inheritance at one gene. A Punnett square lists possible combinations of parental gametes. For Pp x Pp, work through these steps before counting flower colors:

  1. Write P and p across the top for one parent's gamete types.
  2. Write P and p down the left for the other parent's gamete types.
  3. Fill each box with the allele from its column and the allele from its row.
  4. Group matching genotypes. Pp and pP describe the same allele pair in this example, so write both as Pp.
Gamete from parent 2P from parent 1p from parent 1
PPPPp
pPppp

Each parent contributes either allele with probability 1/2. With random fertilization, each box has probability 1/2 x 1/2 = 1/4. The two Pp boxes are different routes to the same genotype, so both count.

The expected genotype ratio is 1 PP : 2 Pp : 1 pp. To find the phenotype ratio, combine PP and Pp because both are purple. That gives 3 purple : 1 white. The change from 1:2:1 to 3:1 comes from grouping appearances, not from changing the alleles.

The point: The square predicts genotype probabilities. A 3:1 phenotype ratio follows only after we apply complete dominance to those genotypes.

A ratio is a prediction, not a quota

A 1/4 chance of white flowers applies to each offspring of this Pp x Pp cross. Four offspring do not have to include exactly one white plant. All four could be purple, all four could be white, or the colors could be mixed.

If successive fertilizations are independent, the probability that the first two offspring are both white is 1/4 x 1/4 = 1/16. The result of the first fertilization does not use up the chance of white flowers in the next one.

For a large group, expected counts help you compare the model with the observations. Out of 240 offspring, the model predicts 240 x 3/4 = 180 purple and 240 x 1/4 = 60 white on average. Actual counts will usually differ somewhat because of chance.

These predictions assume ordinary segregation, random fertilization, and no genotype-related difference in survival before counting. If a genotype reduces survival, for example, the plants you count may not match the ratio at fertilization. State the assumptions when using the square to interpret real data.

Use a test cross to investigate a purple parent

A purple plant might be PP or Pp. Crossing it with a white pp plant is a test cross. The pp parent contributes p every time, which makes the two possibilities easier to distinguish.

Possible crossOffspring genotypes predictedFlower colors predicted
PP x ppAll PpAll purple
Pp x pp1/2 Pp and 1/2 pp1/2 purple and 1/2 white

A white offspring must have received p from both parents. Under this model, finding a white offspring tells you that the purple parent carried p and therefore was Pp.

The reverse conclusion needs care. Four purple offspring do not prove that the parent was PP. A Pp x pp cross has a 1/2 chance of purple at each fertilization, so four purple offspring in a row have probability (1/2)4 = 1/16. A small sample can miss a possible phenotype.

More all-purple offspring provide stronger evidence for PP, but the result remains a statistical inference. This is why an expected ratio and an observed small family should not be treated as the same thing.

Work backward from the offspring

Suppose two purple plants produce 178 purple offspring and 62 white offspring. Begin with the white plants. A white plant is pp in our model, so both parents must have contributed p.

Each parent is purple, so each also needs P. The parents must therefore be Pp and Pp. Now check your conclusion: their predicted counts among 240 offspring are 180 purple and 60 white, close to the observed counts.

Notice the order of the reasoning. The white offspring establish that each parent carries p; the parents' purple phenotype supplies the other allele. The approximate 3:1 ratio supports the result. You do not need to guess a ratio first and force the parents to match it.

Try the same reasoning with tall and short peas. If T is completely dominant, a Tt plant produces T or t gametes, while a short tt plant produces only t. Combining them gives Tt or tt with equal probability. This is the 1:1 result you will use in the activity.

Common misconceptions

  • Two letters mean two genes. Pp is two alleles at one gene. A problem following two genes needs a second pair of letters.
  • Recessive means absent. A Pp plant carries p even though its flower color is purple. Inheritance and visible expression are different questions.
  • Dominant means common or stronger. Dominance describes a phenotype in a heterozygote. It does not describe population frequency or value.
  • Every four offspring must match the four boxes. The boxes represent possible outcomes and their probabilities, not four assigned places in a family.
  • An all-purple test cross proves PP. A limited sample from Pp x pp can also be all purple. Interpret the number of offspring as well as their colors.

What to carry forward

For a new single-gene cross, write the parental genotypes, list each parent's gametes, combine them, and only then group the phenotypes. Explain what each fraction means for one offspring and which assumptions make the calculation valid.

Key idea: Segregation keeps the inherited allele copies separate. Dominance determines how their combination appears. Keeping those two processes distinct lets you explain both the 1:2:1 genotype ratio and the 3:1 flower-color ratio.

Sources

  1. Abbott, S., & Fairbanks, D. J. (2016). Experiments on plant hybrids by Gregor Mendel. Genetics, 204(2), 407-422. Translation of Mendel's primary report; single-character crosses and subsequent generations. Genetics
  2. National Human Genome Research Institute. (n.d.). Gene. Talking Glossary of Genomic and Genetic Terms. Definition and narration on protein-coding and RNA genes. NHGRI
  3. Clark, M. A., Douglas, M., & Choi, J. (2018). Biology 2e, section 12.2, Characteristics and traits. Genotypes, Punnett squares and test crosses. OpenStax. OpenStax
Key terms
Allele
One of the alternative versions of a gene.
Genotype
The specific combination of alleles an organism carries for a gene.
Phenotype
The observable trait produced by the genotype and environment.
Homozygous
Having two identical alleles at the diploid gene being studied, such as PP or pp.
Heterozygous
Having two different alleles at the diploid gene being studied, such as Pp.
Law of segregation
In the diploid inheritance model, paired alleles separate so each gamete receives one.

Dihybrid Crosses and Independent Assortment

  • State Mendel's law of independent assortment.
  • Predict the 9:3:3:1 ratio from a dihybrid cross.
  • Use the multiplication rule to combine probabilities of separate genes.

The big picture

Mendel counted 556 seeds in a cross following seed shape and color together. He recorded 315 round yellow seeds, 108 round green, 101 wrinkled yellow, and 32 wrinkled green. Why did these four combinations occur in such different numbers?

We can build the prediction from two familiar single-gene crosses. The extra step is deciding whether the two genes are inherited independently. Once that assumption is clear, the arithmetic stays manageable.

Read RrYy one gene at a time

Use R for round seeds and r for wrinkled seeds. Use Y for yellow seeds and y for green seeds. In this model, R is completely dominant to r, and Y is completely dominant to y.

A plant with genotype RrYy is heterozygous at both genes. Read it as Rr for shape, then Yy for color. It has two allele copies at each of two genes, not four copies at one gene.

A dihybrid is an individual heterozygous at two genes. A dihybrid cross follows two genes; the standard example here is RrYy x RrYy. We will also change one parent later, because following two genes does not make every cross give the same ratio.

Find the four gamete types

A gamete receives one allele from the R/r pair and one from the Y/y pair. It cannot receive Rr as its shape contribution: those are the two alternatives that separate during gamete formation.

Independent assortment means that knowing which allele a gamete receives at one gene does not change its probability of receiving either allele at the other gene. In this example, a gamete that receives R still has a 1/2 chance of receiving Y and a 1/2 chance of receiving y.

The physical explanation starts with chromosomes. During meiosis, each pair of corresponding chromosomes can orient independently of the other pairs before separating. Genes on different chromosome pairs therefore usually follow independent assortment. We will distinguish that situation from linked genes below.

  1. Pair R with Y to make RY.
  2. Pair R with y to make Ry.
  3. Pair r with Y to make rY.
  4. Pair r with y to make ry.

Each of these four gamete types has probability 1/2 x 1/2 = 1/4 under the model. The fractions add to one, so the list accounts for all the possibilities.

Remember: Segregation separates the two alleles at one gene. Independent assortment describes how the choices at two different genes combine.

Build and read the sixteen-box square

Both RrYy parents can contribute RY, Ry, rY, or ry. Put those gametes across the top and down the side of a Punnett square. Each box combines one row gamete with one column gamete.

GameteRYRyrYry
RYRRYYRRYyRrYYRrYy
RyRRYyRRyyRrYyRryy
rYRrYYRrYyrrYYrrYy
ryRrYyRryyrrYyrryy

Check one box slowly. Ry from one parent and rY from the other give Rr at the shape gene and Yy at the color gene. The offspring is RrYy. Group the alleles by gene when writing the answer.

Each box has probability 1/4 x 1/4 = 1/16. Now group the genotypes by appearance. Any genotype with at least one R is round, and any genotype with at least one Y is yellow.

We can write R_ to mean either RR or Rr: the blank is an allele whose identity does not change this phenotype. Likewise, Y_ means YY or Yy. The blank is not a missing gene.

PhenotypeGenotype categoryExpected fraction
Round and yellowR_ Y_9/16
Round and greenR_ yy3/16
Wrinkled and yellowrr Y_3/16
Wrinkled and greenrr yy1/16

This is the 9:3:3:1 phenotype ratio. It depends on double-heterozygous parents, independent assortment, complete dominance at each gene, and the stated relationship between genotype and phenotype. Random fertilization and comparable survival are also part of the prediction.

Reach the same answer with multiplication

The multiplication rule, or product rule, says that the probability of two independent events both occurring is the product of their probabilities. Here, split the cross into Rr x Rr and Yy x Yy.

Rr x Rr gives a 3/4 probability of round and a 1/4 probability of wrinkled. Yy x Yy gives a 3/4 probability of yellow and a 1/4 probability of green. Combine the fractions for the outcome you want.

Outcome requestedCalculationAnswer
Round AND yellow3/4 x 3/49/16
Round AND green3/4 x 1/43/16
Wrinkled AND yellow1/4 x 3/43/16
Wrinkled AND green1/4 x 1/41/16

The activity asks for wrinkled and yellow offspring. Choose the rr outcome at the shape gene, with probability 1/4, and the Y_ outcome at the color gene, with probability 3/4. Their combined probability is 3/16. You have used the same reasoning as the square without writing all sixteen boxes.

Be precise about whether a problem asks for a genotype or a phenotype. For RrYy specifically, the calculation is 1/2 x 1/2 = 1/4. For round and yellow, it is 3/4 x 3/4 = 9/16. Several genotypes share that appearance.

The point: Choose each single-gene probability from the actual question, then multiply if the events are independent.

When to add instead

Two outcomes are mutually exclusive if the same observation cannot belong to both. A seed cannot be both round green and wrinkled yellow in our four-category model. To find the chance of either category, add: 3/16 + 3/16 = 6/16 = 3/8.

Do not add probabilities for overlapping categories without correcting for their overlap. For example, a round yellow seed is both round and yellow. Simply adding the chance of round to the chance of yellow would count that seed's category twice.

The product rule also extends to three independent genes. In AaBbCc x AaBbCc, with complete dominance at each gene, the chance of all three dominant phenotypes is 3/4 x 3/4 x 3/4 = 27/64. The chance of the exact genotype AaBbCc is different: 1/2 x 1/2 x 1/2 = 1/8.

Change one parent: the dihybrid test cross

Now cross RrYy with rryy. The first parent still makes four equally frequent gamete types if the genes assort independently. The second parent can contribute only ry.

  1. RY with ry gives RrYy, round yellow.
  2. Ry with ry gives Rryy, round green.
  3. rY with ry gives rrYy, wrinkled yellow.
  4. ry with ry gives rryy, wrinkled green.

Each outcome has probability 1/4, giving a 1:1:1:1 ratio. The recessive parent contributes alleles at both genes. Because those contributions are known, each offspring category reveals the gamete received from the heterozygous parent.

That is why this test cross is useful for investigating linkage. A departure from four equal categories may tell you that some gamete combinations occur more often than others. It is evidence to investigate, not an automatic diagnosis of one mechanism.

When the independent shortcut fails

Linked genes are genes whose chromosome positions make their alleles tend to be inherited together. Genes close together on the same chromosome often show this pattern. In that case, you cannot assume that RY, Ry, rY, and ry each have probability 1/4.

Crossing over is an exchange of DNA between corresponding chromosomes during meiosis. It can separate alleles that started on the same chromosome. Genes far apart on one chromosome can consequently behave approximately as though they assort independently.

The independent product rule does not apply when the relevant outcomes are dependent. You then need information about their joint probabilities, such as observed gamete frequencies, or conditional probabilities that account for the dependence. The arithmetic is not forbidden; the independence assumption is what must change.

New combinations such as round green do not by themselves require new mutations. Existing alleles can be inherited in combinations different from those of the original parents.

Compare a prediction with Mendel's counts

Return to the 556 seeds at the start. A goodness-of-fit test checks how far observed category counts depart from a specified probability model. The chi-square statistic adds the squared difference between observed and expected counts, divided by the expected count, for each category.

Our null hypothesis, the model being tested, specifies a 9:3:3:1 ratio. Multiply 556 by each predicted fraction to obtain the expected counts. The table shows the calculation; O means observed and E means expected.

Seed categoryOE(O - E)2 / E, rounded
Round yellow315312.750.016
Round green108104.250.135
Wrinkled yellow101104.250.101
Wrinkled green3234.750.218

For example, round green contributes (108 - 104.25)2 / 104.25, approximately 0.135. Adding the four contributions gives chi-square approximately 0.47. Squaring prevents positive and negative differences from canceling.

Here the predicted proportions are fixed rather than estimated from the observations. There are four categories, so the test has 3 degrees of freedom: four counts can vary, but their total is fixed. At the conventional 5% significance level, the upper-tail critical value is 7.815.

Because 0.47 is below 7.815, we fail to reject the specified ratio. These counts are compatible with that prediction under the test's assumptions. The result does not prove that the genes are independent or establish the probability that the model is true.

The calculation uses counts, not percentages, because sample size affects how much random variation we expect. It also requires an appropriate sampling design and sufficiently large expected counts for the chi-square approximation. Here every expected count exceeds 5. A full analysis would also consider how the seeds were sampled and whether survival or other features of the experiment affected the counts.

Worth holding on to: A goodness-of-fit result evaluates a specified ratio. If the fit is poor, it does not by itself distinguish linkage from other explanations.

Common misconceptions

  • RrYy makes Rr and Yy gametes. Each gamete needs one allele from each gene. Its possibilities here are RY, Ry, rY, and ry.
  • Every two-gene problem gives 9:3:3:1. Changing a parent, dominance relationship, linkage pattern, or survival assumption can change the expected phenotype ratio.
  • Independent assortment and segregation are the same. Segregation concerns one allele pair; independent assortment concerns the relationship between different genes.
  • A small chi-square proves the explanation. A compatible result leaves the model plausible; it does not uniquely identify the biology that produced the counts.

Putting it together

When you meet several pairs of letters, separate the problem into genes. Write the gametes and state whether independence is justified. Use the resulting probabilities to predict genotypes, then translate those genotypes into phenotypes.

Key idea: The ratios are results of a model, not numbers to memorize in isolation. RrYy x RrYy gives 9:3:3:1 under the stated assumptions, while RrYy x rryy gives 1:1:1:1.

Sources

  1. Abbott, S., & Fairbanks, D. J. (2016). Experiments on plant hybrids by Gregor Mendel. Genetics, 204(2), 407-422. Primary report in translation, section on hybrids combining several differing characters; the 556-seed experiment. Genetics
  2. Clark, M. A., Douglas, M., & Choi, J. (2018). Biology 2e, section 12.3, Laws of inheritance. Independent assortment, probability and linkage. OpenStax. OpenStax
  3. National Institute of Standards and Technology. (n.d.). e-Handbook of Statistical Methods, section 1.3.5.15, Chi-square goodness-of-fit test. Formula, assumptions and degrees of freedom. NIST
  4. National Institute of Standards and Technology. (n.d.). e-Handbook of Statistical Methods, section 1.3.6.7.4, Critical values of the chi-square distribution. Upper-tail table, 3 degrees of freedom and probability 0.95. NIST table
Key terms
Dihybrid cross
A cross tracking two genes at once, such as RrYy x RrYy.
Law of independent assortment
Alleles of different genes are sorted into gametes independently when the genes are on different chromosomes.
Multiplication (product) rule
The probability of two independent events both occurring is the product of their separate probabilities.
Addition rule
The probability of either of two mutually exclusive events is the sum of their probabilities.
9:3:3:1 ratio
The expected phenotype ratio from two independently assorting genes with complete dominance in a double-heterozygote cross.
Gamete
A reproductive cell, such as an egg or sperm, carrying one chromosome set in the diploid inheritance model.

Beyond Simple Dominance

  • Distinguish incomplete dominance from codominance.
  • Explain multiple alleles using the ABO blood group system.
  • Define pleiotropy, epistasis, and polygenic inheritance.

The big picture

Cross a red snapdragon with a white snapdragon, and the offspring in a simple flower-color model are pink. Cross two of those pink plants, and red and white flowers reappear alongside pink ones. The appearance is intermediate, but the inherited alleles have not merged into a new pink allele.

To explain this result, keep the inheritance calculation and the phenotype calculation separate. The alleles still segregate. What changes is the relationship between the allele combination and the trait you observe.

Incomplete dominance: the heterozygote is intermediate

Incomplete dominance means that the heterozygote's phenotype falls between the phenotypes of the two homozygotes. We will keep the lesson's simple labels: RR for red snapdragons, rr for white, and Rr for pink. Here, the capital letter is a label, not a claim that red is completely dominant.

Each pink Rr plant produces R or r gametes with equal probability. Combining the gametes from Rr x Rr gives the same genotype probabilities as an ordinary monohybrid cross. Now give each genotype its own flower color:

GenotypeExpected fractionFlower color
RR1/4Red
Rr1/2Pink
rr1/4White

The phenotype ratio is 1 red : 2 pink : 1 white. It matches the genotype ratio because all three genotypes have different appearances in this model. There is no reason to combine RR and Rr into one class, as we did for completely dominant purple pea flowers.

Remember: The 1:2:1 genotype calculation has not changed. Incomplete dominance changes how you label the three resulting phenotypes.

Codominance: both effects are detectable

Codominance means that effects associated with both alleles can be detected in the heterozygote. Human blood type AB is an example: red blood cells display both A and B surface markers. They do not display one intermediate marker halfway between A and B.

Roan cattle provide a visible teaching example. A roan coat contains intermixed colored and white hairs. The hairs remain distinguishable, unlike the intermediate pink flower color in the snapdragon example. Look at what is being measured rather than assuming that any mixed appearance is incomplete dominance.

RelationshipWhat you observe in the heterozygote
Complete dominanceThe same phenotype category as one homozygote
Incomplete dominanceAn intermediate phenotype
CodominanceDetectable effects associated with both alleles

Multiple alleles: distinguish a person from a population

A diploid individual normally carries two allele copies at the gene being studied. Across a population, that gene may have many different versions. Multiple alleles means that more than two versions occur in the population, not that each person has all of them.

The introductory ABO model uses three common functional categories of alleles: IA, IB, and i. IA and IB are codominant with each other for the A and B markers, while each is dominant to i. Real ABO sequence variation is more extensive than these three classroom labels.

Genotype in the introductory modelABO phenotype
IAIA or IAiA
IBIB or IBiB
IAIBAB
iiO

Work the ABO cross before reading the answer

One parent is IAi and the other is IBi. The first can contribute IA or i. The second can contribute IB or i. Combine those gametes, with probability 1/2 for each allele from either parent.

Gamete from parent 2IA from parent 1i from parent 1
IBIAIB: ABIBi: B
iIAi: Aii: O

A, B, AB, and O each have probability 1/4 under this model. A family with four children is not guaranteed to include one of each type. The probabilities apply separately to each child.

The point: Predict inheritance from the allele pairs. A phenotype may leave more than one genotype possible.

What ABO alleles do inside a cell

An enzyme is a molecule, usually a protein, that speeds up a chemical reaction. The ABO gene provides instructions for an enzyme that modifies a sugar structure on the cell surface. It does not directly encode the finished surface marker.

The enzyme acts on a starting structure called the H antigen. An antigen is a structure the immune system can recognize. The A enzyme adds the sugar N-acetylgalactosamine; the B enzyme adds galactose. Those different final sugars help distinguish A from B.

An IAIB individual can make both forms of the enzyme. Their cells therefore display both kinds of marker. That is the biochemical reason for codominance in this example.

Common O alleles do not produce an active A- or B-modifying enzyme. A frequent O variant contains a one-base deletion in the DNA sequence. The deletion shifts how the protein-coding sequence is read, preventing the usual enzyme activity. This is a common mechanism, not a description of every O-associated variant.

The starting structure matters too. In the rare Bombay phenotype, variants affecting another gene prevent the usual H antigen from forming on red blood cells. Without that starting structure, the A and B enzymes cannot make their usual markers there. ABO genotype alone therefore does not explain every blood-typing result.

This example explains gene interaction; it is not a transfusion compatibility rule. Clinical laboratories use blood typing and compatibility testing because surface markers and immune reactions involve more than a classroom Punnett square.

Dominance comes from the biological effect

A variant that reduces a gene product's activity is called a loss-of-function variant. Some such variants are recessive because the other allele provides enough activity for the phenotype being measured. Calling all recessive alleles simply broken would hide both the range of molecular effects and their biological context.

For example, imagine an enzyme working in a pathway, a sequence of chemical reactions. Reducing the amount of that enzyme does not necessarily reduce the pathway's final output by the same fraction. Other steps may limit the rate. An intermediate enzyme amount can therefore produce an almost unchanged final phenotype.

A different mechanism is a dominant-negative effect. The altered product interferes with a functional product. If a protein works as an assembly of several parts, for example, an altered part can reduce the activity of an assembly that also contains functional parts.

These mechanisms explain why knowing that a variant changes protein function is not enough to assign dominance automatically. You need to know how its product acts alongside the other allele's product, and which phenotype you are measuring.

What matters here: Dominance is a relationship between alleles for a particular phenotype. Molecular activity, visible appearance, and effects on health do not always have the same dominance pattern.

One gene can affect several traits

Pleiotropy means that variation at one gene can influence more than one phenotypic feature. A protein may be used in several tissues, or a change in one process may have several downstream consequences.

Sickle cell disease illustrates the second route. The HBB gene supplies part of hemoglobin, the oxygen-carrying protein in red blood cells. In sickle cell disease, altered hemoglobin can change red-cell behavior. Cells may break down early or obstruct small blood vessels, contributing to anemia and pain.

Those different effects can follow from the same underlying genetic change. The example does not mean that carrying one altered HBB allele automatically causes sickle cell disease: the allele combination matters, and severity varies among people. Distinguish a gene's several effects from a prediction about one individual.

One gene can change another gene's visible effect

Epistasis is an interaction in which variation at one gene modifies or masks the phenotypic effect of variation at another. Dominance compares alleles at one gene; epistasis involves different genes.

Use a simplified model of Labrador coat color. At one gene, B and b distinguish black from brown pigment. At a second gene, E permits that pigment to appear in the hair, while ee produces yellow hair. Yellow hair contains pigment; it is not a complete absence of pigmentation.

Under this model, B_E_ is black, bbE_ is brown, and either B_ee or bbee is yellow. Read the second gene first: if the dog is ee, knowing B or b will not change the yellow hair category.

For BbEe x BbEe with independent assortment, begin with the four familiar genotype categories. B_E_ accounts for 9/16, bbE_ for 3/16, B_ee for 3/16, and bbee for 1/16. The last two look alike, so combine them: 3/16 + 1/16 = 4/16 yellow.

The resulting 9 black : 3 brown : 4 yellow ratio comes from regrouping genotypes by appearance. The alleles can still assort independently even though their effects on the phenotype interact.

Other interactions produce other groupings. If a dominant allele masks the effect of a second gene, combining the 9 and one 3 category can give 12:3:1. If two separate gene products are both needed for a result, only the double-dominant category may produce it, giving 9:7. Derive the grouping from the proposed mechanism instead of treating a ratio as proof of one explanation.

Many genes can influence one trait

Polygenic inheritance means that two or more genes influence a trait. Human height is an example. Many variants contribute to differences in height, and conditions such as nutrition also matter.

When many effects combine, a trait may vary across a broad range instead of falling into two categories. That does not mean every gene has an identical small effect or that every polygenic trait must follow a perfect bell-shaped distribution.

Separate two ideas that are easy to reverse. Pleiotropy follows one gene toward several effects. Polygenic inheritance follows several genes toward one trait. A gene can participate in both patterns: it may contribute to height while also influencing another feature.

Common misconceptions

  • Pink flowers require permanently blended alleles. Rr plants still transmit R or r. The intermediate phenotype does not erase the alleles.
  • Codominance means an average. Both allele-associated effects are detectable. Type AB has A and B markers, not an intermediate marker.
  • Multiple alleles means three copies in each person. The extra versions occur across the population. The ordinary diploid ABO model gives each person two copies.
  • Epistasis means independent assortment has stopped. The genes may still assort independently. Their products interact when a phenotype develops.
  • Genes determine a complex trait by themselves. Polygenic traits can also depend on environmental conditions. A genotype is not a complete prediction of an individual's outcome.

What to remember

First calculate which alleles can be inherited. Then ask how their products affect the measured trait: an intermediate phenotype, two detectable effects, or an interaction between genes. The same inheritance probabilities can lead to different phenotype ratios.

Key idea: Use incomplete dominance and codominance for relationships at one gene, multiple alleles for variation in a population, pleiotropy for one gene's several effects, epistasis for interactions between genes, and polygenic inheritance for several genes influencing one trait.

Sources

  1. Clark, M. A., Douglas, M., & Choi, J. (2018). Biology 2e, sections 12.2 and 12.3. Incomplete dominance, codominance, multiple alleles and epistasis. OpenStax. Section 12.2; Section 12.3
  2. Yamamoto, F., Clausen, H., White, T., Marken, J., & Hakomori, S. (1990). Molecular genetic basis of the histo-blood group ABO system. Nature, 345, 229-233. Nature
  3. Dean, L. (2005). The ABO blood group. In Blood groups and red cell antigens, chapter 5, Basic biochemistry and Molecular information. National Center for Biotechnology Information. NCBI Bookshelf
  4. Billiard, S., Castric, V., & Llaurens, V. (2021). The integrative biology of genetic dominance. Biological Reviews, 96(6), 2925-2942. Sections I and III. Biological Reviews
  5. Young, A., & Bellone, R. (2019, April 23). Coat color inheritance in the Labrador Retriever. University of California, Davis, School of Veterinary Medicine. UC Davis
  6. National Library of Medicine. (2024, March 14). Sickle cell disease. MedlinePlus Genetics; Description, Causes and Inheritance. MedlinePlus Genetics
  7. National Human Genome Research Institute. (n.d.). Polygenic trait. Talking Glossary of Genomic and Genetic Terms. NHGRI
Key terms
Incomplete dominance
The heterozygote has an intermediate phenotype, such as pink between red and white flowers; the alleles do not merge.
Codominance
Effects associated with both alleles are detectable in the heterozygote, as in blood type AB.
Multiple alleles
A gene with more than two versions present in a population, such as ABO.
Pleiotropy
One gene affecting several distinct phenotypic traits.
Epistasis
One gene masking or modifying the phenotypic effect of another gene.
Polygenic inheritance
Two or more genes influence a trait; environmental conditions may also affect its phenotype.

Module 2: Meiosis, Recombination, and Genetic Mapping

The cell division that shuffles chromosomes into gametes, how crossing over creates new allele combinations, and how recombination frequency lets us build gene maps.

Meiosis and the Chromosomal Basis of Inheritance

  • Contrast meiosis with mitosis in outcome and purpose.
  • Explain how meiosis physically carries out Mendel's two laws.
  • Define homologous chromosomes, haploid, and diploid.

The big picture

Mendel worked out his laws by counting peas, without ever seeing the machinery inside a cell. What was actually happening in there? This lesson reveals that machinery. It turns out that chromosomes, the packages of DNA inside cells, move during a special kind of cell division in exactly the way Mendel's factors must. Meiosis is the physical event that makes his abstract rules real.

By the end you will be able to connect a cross on paper to the dance of chromosomes in a dividing cell, and you will see why every sperm and egg is genetically unique. That link between the visible ratios and the invisible chromosomes is the heart of classical genetics.

The chromosome theory of inheritance

Decades after Mendel, biologists watching cells divide noticed that chromosomes behave precisely as Mendel's hereditary factors should. A chromosome is a single long molecule of DNA wound around proteins, carrying many genes in a fixed order. This observation led to the chromosome theory of inheritance: genes are located on chromosomes, and the movement of chromosomes during meiosis is what produces the inheritance patterns Mendel described. Genes are not free-floating; they ride on chromosomes, so tracking chromosomes tracks genes.

Diploid, haploid, and homologous pairs

Body cells, also called somatic cells, are diploid, meaning they carry two complete sets of chromosomes, one set from each parent. Diploid is often written 2n. In humans, 2n is 46 chromosomes, arranged as 23 pairs. The two members of each pair are called homologous chromosomes: a matched pair that carry the same genes in the same order, though they may carry different alleles of those genes. Think of homologous chromosomes as two editions of the same book, identical in chapter order but possibly differing in a few words.

Gametes are different. A gamete (egg or sperm) is haploid, carrying just one complete set of chromosomes, written n (23 in humans). This halving is essential: when a haploid sperm fertilizes a haploid egg, the two single sets combine to restore the diploid number, 46, in the offspring. Without halving, the chromosome number would double every generation.

Meiosis is the specialized cell division that halves the chromosome number, turning one diploid cell into haploid gametes. It is what makes sexual reproduction possible.

Key idea: Diploid cells carry two homologous sets (2n); meiosis halves this to haploid gametes (n) so fertilization can restore the diploid number.

Two divisions, four cells

So how does meiosis actually pull this off? Two consecutive divisions follow a single round of DNA replication. Because DNA is copied once but the cell divides twice, the chromosome number is halved.

  • Meiosis I is the reductional division. Homologous chromosomes pair up, then the two homologs of each pair are pulled to opposite ends of the cell and separated. This is the step that reduces the count from diploid to haploid, because each daughter cell now has only one chromosome from each original pair.
  • Meiosis II resembles an ordinary mitotic division. The sister chromatids of each chromosome (the two identical copies made during replication) separate from each other. No further reduction in chromosome number occurs here.

The end result is four haploid cells from one diploid starting cell. Contrast this with mitosis, the routine division that produces two genetically identical diploid cells for growth, repair, and asexual reproduction. Mitosis copies; meiosis both halves and shuffles.

FeatureMitosisMeiosis
DivisionsOneTwo
Daughter cellsTwoFour
Chromosome numberDiploid (unchanged)Haploid (halved)
Genetic resultIdentical to parentGenetically varied
PurposeGrowth and repairMaking gametes

Key idea: Meiosis is two divisions after one DNA replication, producing four varied haploid cells, while mitosis is one division producing two identical diploid cells.

How meiosis performs Mendel's laws

Meiosis is the physical machinery behind both of Mendel's laws, which is the deep reason the laws work.

The law of segregation is carried out in meiosis I. The two homologous chromosomes that carry the two alleles of a gene are pulled to opposite poles, so each resulting gamete receives only one of the two alleles. Segregation of alleles is simply the separation of homologous chromosomes.

The law of independent assortment arises because each homologous pair lines up and orients at random at the cell's midline during meiosis I, independently of every other pair. Which way one pair happens to face has no bearing on which way any other pair faces. This is why alleles of genes on different chromosomes are distributed independently.

Key idea: Segregation is the separation of homologs in meiosis I, and independent assortment is the random, independent orientation of each homologous pair.

Counting the variety meiosis creates

Just how much diversity does independent assortment alone generate? Enormous amounts. With each of the n homologous pairs able to orient in either of two ways, the number of chromosomally distinct gametes is 2 raised to the power n. Work a small case first: an organism with 3 pairs (2n = 6) can make 23 = 8 different gametes. For humans with 23 pairs, the figure is 223, which is 8,388,608, more than eight million, and that is before crossing over adds still more combinations. This is the deep source of the genetic uniqueness of every individual: no two gametes (except by rare chance) carry the same set of chromosomes.

Key idea: Independent assortment produces 2n chromosomally distinct gametes (over eight million in humans), the main reason siblings differ.

What holds chromosomes together, and what lets go

The reason meiosis I is reductional and meiosis II is not comes down to a single protein complex and the order in which it is cut.

When chromosomes are copied, the two identical sister chromatids are clamped together along their whole length by a ring-shaped protein called cohesin. During prophase I, homologous chromosomes find each other and are zipped together along their length by a protein scaffold, the synaptonemal complex. Crossovers form while they are zipped, and the resulting chiasmata physically tie the two homologs together, so each pair now behaves as a single unit on the spindle.

At anaphase I, cohesin is cut along the chromosome arms but deliberately protected at the centromeres by a guard protein. Cutting the arm cohesin releases the chiasmata, so the two homologs separate, but each one still holds its two sister chromatids together at the centromere. That is why the count halves at this step and why each daughter cell receives whole chromosomes rather than single chromatids. At anaphase II the centromeric protection is removed, the remaining cohesin is cut, and sisters finally separate.

This mechanism explains a striking clinical fact. Human oocytes enter prophase I before birth and then pause, sometimes for forty years, holding their cohesin the entire time. Cohesin is not replenished during that arrest, so it gradually deteriorates, chiasmata slip, and the risk of a chromosome missegregating rises with maternal age. The molecular story and the epidemiological one are the same story.

Key idea: Chiasmata plus cohesin hold each homologous pair together; cutting arm cohesin at anaphase I separates homologs while centromeric cohesin is protected until anaphase II, which is why meiosis I alone is reductional.

When separation fails: meiosis I versus meiosis II errors

Nondisjunction is the failure of chromosomes to separate, and the stage at which it happens leaves a distinguishable signature.

  • Nondisjunction in meiosis I. Both homologs travel to the same pole. Every one of the four resulting gametes is abnormal: two carry n + 1 chromosomes and, crucially, carry both the maternal and paternal homolog; two carry n - 1.
  • Nondisjunction in meiosis II. Only one secondary cell misdivides, so two gametes are normal, one carries n + 1 with two copies of the same homolog, and one carries n - 1.

Because the two cases differ in whether the extra chromosome came from one grandparent or from both, genetic markers can identify which division failed. Studies using that approach find that the great majority of human trisomies arise from errors in maternal meiosis I, which fits the cohesin story exactly.

Key idea: Meiosis I nondisjunction gives four abnormal gametes carrying both homologs, meiosis II nondisjunction gives two normal gametes and one carrying two copies of a single homolog, and markers can tell the two apart.

How much variety, all told

Independent assortment gives 223 = 8,388,608 chromosomally distinct gametes. Two further multipliers dwarf it.

  • Crossing over. A human meiosis makes roughly 50 crossovers in males and around 70 in females, scattered at effectively random positions. Every chromosome handed on is therefore a new patchwork of maternal and paternal segments rather than an intact grandparental chromosome, so the practical number of distinct gametes is astronomically larger than 223.
  • Fertilization. Any one of those gametes can meet any other. From assortment alone that is 8,388,608 x 8,388,608, which is about 7 x 1013 possible zygotes from a single couple, seven times more than there are cells in a human body.

This is the quantitative reason siblings resemble each other without being alike, and the reason a species stores its variation in combinations rather than in new mutations alone.

Key idea: Assortment gives 223 gametes, crossing over multiplies that enormously with about 50 to 70 exchanges per meiosis, and random fertilization squares the total to roughly 7 x 1013 possible zygotes per couple.

Where people get stuck

The first sticking point is losing track of what a chromosome is at each moment. After replication one chromosome consists of two sister chromatids, and it is still called one chromosome until the centromeres separate. Counting chromatids instead of centromeres is the usual source of wrong answers about chromosome number.

The second is assuming meiosis II halves the number again. It does not. It separates sisters, so four haploid cells result from two haploid cells, and the chromosome count stays at n throughout.

The third is picturing independent assortment as chromosomes choosing sides. Nothing chooses. Each bivalent attaches to whichever pole its kinetochores happen to face, and the independence of those orientations is simply the absence of any physical connection between different pairs.

Common misconceptions

  • Meiosis makes four cells, not two. Two divisions follow one replication, so the count doubles compared with mitosis while the chromosome number halves.
  • Homologous chromosomes are not identical. They match in gene order but can carry different alleles; that is exactly why offspring vary.
  • The chromosome number is halved in meiosis I, not meiosis II. Meiosis II separates sister chromatids and keeps the count haploid.
  • Sister chromatids are not homologous chromosomes. Sister chromatids are two identical copies of one chromosome; homologs are the maternal and paternal versions of a chromosome.

Recap

  • The chromosome theory places genes on chromosomes, whose meiotic movement explains Mendel's laws.
  • Somatic cells are diploid (2n, two homologous sets); gametes are haploid (n, one set), restoring 2n at fertilization.
  • Meiosis is two divisions after one replication, yielding four varied haploid cells; mitosis yields two identical diploid cells.
  • Meiosis I separates homologs (segregation) and orients each pair randomly (independent assortment).
  • Independent assortment alone makes 2n distinct gametes, over eight million in humans.

Sources

  1. Clark, M. A., Douglas, M., & Choi, J. (2018). Biology 2e (Ch. 11.1: The process of meiosis). OpenStax. openstax.org
  2. National Human Genome Research Institute. (n.d.). Meiosis. In Talking glossary of genomic and genetic terms. genome.gov
  3. Clark, M. A. (2008). Meiosis, genetic recombination, and sexual reproduction. Nature Education, 1(1), 208.
  4. Khan Academy. (n.d.). Meiosis (cellular and molecular biology). khanacademy.org
  5. Sutton, W. S. (1903). The chromosomes in heredity. The Biological Bulletin, 4(5), 231-250. doi.org/10.2307/1535741
  6. Ohkura, H. (2015). Meiosis: An overview of key differences from mitosis. Cold Spring Harbor Perspectives in Biology, 7(5), a015859. doi.org/10.1101/cshperspect.a015859
  7. Alberts, B., Johnson, A., Lewis, J., Raff, M., Roberts, K., & Walter, P. (2002). Meiosis. In Molecular biology of the cell (4th ed.). Garland Science. ncbi.nlm.nih.gov
Key terms
Diploid (2n)
Having two complete sets of chromosomes, one from each parent.
Haploid (n)
Having a single set of chromosomes, as in a gamete.
Homologous chromosomes
A matched pair carrying the same genes in the same order, one from each parent.
Meiosis
The two-division process that produces four haploid gametes from one diploid cell.
Meiosis I
The reductional division in which homologous chromosomes separate.
Chromosome theory of inheritance
The principle that genes reside on chromosomes whose meiotic movement explains Mendel's laws.

Crossing Over and Recombination

  • Describe crossing over and when it occurs.
  • Explain how recombination produces new allele combinations.
  • Define parental and recombinant offspring.

The big picture

Independent assortment shuffles whole chromosomes. But what about the alleles within a single chromosome? This lesson looks at that finer kind of mixing. During meiosis, paired chromosomes physically swap matching segments, creating combinations that neither parent chromosome had. This is crossing over, and it is both a major engine of genetic variety and, as you will see, the key that unlocks gene mapping.

The crucial insight is that the chance of a swap between two genes depends on how far apart they sit. That simple fact lets geneticists convert an offspring count into a physical distance along a chromosome, which is where the next lesson goes.

What crossing over is

During prophase of meiosis I, homologous chromosomes pair up so closely that they physically touch along their length. While paired, they exchange matching segments in a process called crossing over: the reciprocal swapping of corresponding pieces between homologous chromosomes. The points where the chromosomes cross and exchange are visible under a microscope as X-shaped structures called chiasmata (singular chiasma). Because the segments swapped are matching, no genes are gained or lost; only the alleles are rearranged.

Key idea: Crossing over is the reciprocal exchange of matching segments between paired homologous chromosomes in meiosis I, visible as chiasmata.

Parental and recombinant combinations

Consider a chromosome carrying alleles A and B together, paired with its homolog carrying a and b together. Before any crossover, gametes would receive either the AB combination or the ab combination, matching the original chromosomes. These original, unshuffled combinations are called parental. If a crossover occurs between the two genes, new combinations appear: one chromosome now carries A with b, and the other carries a with B. These new mixes are recombinant: allele combinations not present on either original chromosome.

Gametes carrying recombinant chromosomes produce recombinant offspring, while gametes with the original combinations produce parental offspring. Here is the full accounting for our example:

GameteTypeOrigin
ABParentalMatches an original chromosome
abParentalMatches the other original chromosome
AbRecombinantNew combination from a crossover
aBRecombinantNew combination from a crossover

Crossing over is therefore a second powerful source of the genetic variation that fuels evolution, working alongside independent assortment and the random union of gametes at fertilization. Independent assortment reshuffles whole chromosomes; crossing over reshuffles the alleles inside each one.

Key idea: A crossover between two genes converts parental allele combinations (AB, ab) into recombinant ones (Ab, aB), adding variation within a chromosome.

Distance controls how often genes recombine

Here is the insight that makes gene mapping possible. Crossovers happen at more or less random positions along a chromosome. The farther apart two genes lie, the more room there is between them for a crossover to fall, and so the more often they are separated into recombinant combinations. Two genes very close together are rarely separated, so they are almost always inherited together; two genes far apart are separated often.

We measure this with the recombination frequency: the fraction of offspring that are recombinant, calculated as the number of recombinant offspring divided by the total number of offspring. Because it rises with distance, recombination frequency is a direct measure of how far apart two genes are. Work an example: suppose a cross of 1000 offspring yields 430 AB, 420 ab, 75 Ab, and 75 aB.

The parental types (AB and ab) total 850 and the recombinant types (Ab and aB) total 150. The recombination frequency is 150 divided by 1000, which equals 0.15, or 15 percent. That 15 percent is the raw material the next lesson turns into a map.

Key idea: Recombination frequency equals recombinant offspring divided by total offspring, and it increases with the distance between two genes.

A picture that helps

Imagine two spots painted on a long rope that gets snipped once at a random point. Two spots near each other usually end up on the same piece after the snip; two spots at opposite ends almost always land on different pieces. Genes behave the same way along a chromosome: physical closeness translates directly into how often two alleles stay together. Close genes stay parental most of the time; distant genes become recombinant often.

Key idea: Like two marks on a randomly cut rope, close genes rarely separate while distant genes separate often, so recombination frequency reflects distance.

Computing a recombination frequency

Recombination frequency is measured with a test cross, because a homozygous recessive partner lets every offspring display the gamete it received from the parent under test.

Worked example. A fly heterozygous for two linked genes, with alleles arranged A B on one chromosome and a b on its homolog, is crossed to a doubly recessive fly. Among 1,000 offspring:

Offspring classNumberType
A B425Parental
a b415Parental
A b82Recombinant
a B78Recombinant
  • Recombinants total 82 + 78 = 160.
  • Recombination frequency = 160 / 1,000 = 0.16, or 16 percent.
  • The two genes therefore lie about 16 map units apart, and they are clearly linked, since unlinked genes would give four classes of roughly 250 each.

One detail explains the 50 percent ceiling. A crossover happens between two of the four chromatids in a bivalent, so a single exchange converts only half of the resulting gametes into recombinants. A crossover occurring in 32 percent of meioses therefore yields 16 percent recombinant gametes. Even if every meiosis had a crossover between two genes, only half the gametes would be recombinant, which is exactly why the frequency can approach but never exceed 50 percent.

Key idea: Recombination frequency is recombinants divided by total offspring, and because only two of four chromatids take part in each exchange, the value can approach but never exceed 50 percent.

What actually happens at the molecular level

Crossing over is not an accident of chromosomes tangling. It is a controlled reaction that the cell starts by deliberately breaking its own DNA.

  1. An enzyme called SPO11 cuts both strands of one chromatid, making a programmed double-strand break. A human meiosis makes roughly 200 to 300 of these.
  2. Nucleases chew back one strand on each side, leaving single-stranded tails.
  3. A tail invades the intact homologous chromosome, finds the matching sequence, and pairs with it. This is why crossing over is precise: the homolog itself is the template.
  4. DNA synthesis extends the invading strand, and the joined molecules can be resolved in two ways. One resolution swaps the flanking arms and produces a crossover; the other repairs the break without swapping arms and produces a non-crossover.

Only about 50 of those 200 to 300 breaks become crossovers in a human meiosis. The rest are repaired quietly. Nothing is lost either way, because the homolog supplies the missing information, which is why crossing over rearranges alleles without ever gaining or losing genes.

Key idea: Meiosis begins crossing over by deliberately cutting DNA with SPO11, then repairs each break using the homolog as template, and only a minority of the 200 to 300 breaks resolve as crossovers.

Crossovers are compulsory, and spaced out

Two rules govern where crossovers land, and both matter for the rest of the course.

First, every homologous pair needs at least one. The chiasma is the physical link that keeps a pair together on the meiosis I spindle, so a pair with no crossover has nothing holding it and is likely to segregate at random. The cell enforces an obligate crossover to prevent this. When it fails, aneuploidy follows, and the smallest human chromosomes, which have the least room for crossovers, are exactly the ones most often found in trisomies.

Second, crossovers avoid each other. One crossover suppresses the formation of another nearby, an effect called interference. The practical consequence is that double crossovers in a short interval are rarer than chance alone would predict, and the next lesson turns that shortfall into a number.

Key idea: Each bivalent must have at least one crossover to segregate properly, and interference spaces crossovers apart so double crossovers are rarer than independence would predict.

Proving that chromosomes really trade pieces

For two decades crossing over was inferred from breeding ratios alone, with no direct evidence that chromosomes physically exchange material. Harriet Creighton and Barbara McClintock closed that gap in maize in 1931 by finding a chromosome that was visibly unusual at both ends: it carried a dense knob at one end and an extra segment translocated from another chromosome at the other. Those two features acted as physical landmarks flanking two genes.

They then asked a simple question. Among the offspring, did the plants that were genetically recombinant also show a chromosome that had swapped its physical landmarks? They did. Genetic recombination and cytological exchange occurred together, in the same plants. Curt Stern reported the equivalent result in fruit flies the same year. Two independent organisms, one conclusion: the abstract shuffling of alleles is a visible exchange of chromosome segments.

Key idea: Creighton and McClintock used visible chromosome landmarks in maize to show that genetically recombinant offspring also carry physically exchanged chromosomes, tying the breeding ratios to real material.

Where people get stuck

The first sticking point is drawing a crossover between two sister chromatids. Do not: sisters are identical, so exchanging material between them changes nothing. A useful crossover happens between non-sister chromatids of homologous chromosomes.

The second is deciding which offspring classes are parental. The answer is not which alleles look dominant. It is which combinations were present on the parent's own chromosomes. In practice, in a test cross the two most numerous classes are the parental ones, and that is the standard way to identify them.

The third is confusing recombination with mutation. They are not the same: a recombinant chromosome carries no new alleles at all, only a new arrangement of ones that were already there.

Common misconceptions

  • Crossing over is between homologous chromosomes, not between sister chromatids or unrelated chromosomes. It happens in prophase of meiosis I.
  • Recombinant offspring are not mutants. Their alleles are unchanged; only the combination of alleles is new.
  • Recombination frequency does not exceed 50 percent. When genes are far apart or on different chromosomes, recombinants and parentals become equal, capping the frequency at 50 percent.
  • A crossover swaps matching segments, so genes are neither added nor lost. Only the pairing of alleles changes.

Recap

  • Crossing over is the reciprocal exchange of matching segments between homologous chromosomes in meiosis I, seen as chiasmata.
  • Parental combinations match the original chromosomes; recombinant combinations are new pairings produced by a crossover.
  • Crossing over adds variation within a chromosome, complementing independent assortment between chromosomes.
  • Recombination frequency is recombinant offspring over total offspring, and it rises with the distance between two genes.
  • Close genes rarely recombine and far genes recombine often, up to a ceiling of 50 percent.

Sources

  1. Clark, M. A., Douglas, M., & Choi, J. (2018). Biology 2e (Ch. 11.2: Sexual reproduction). OpenStax. openstax.org
  2. Miko, I. (2008). Thomas Hunt Morgan and sex linkage. Nature Education, 1(1), 143.
  3. National Human Genome Research Institute. (n.d.). Talking glossary of genomic and genetic terms (recombination, crossing over). genome.gov
  4. Khan Academy. (n.d.). Genetic linkage and recombination frequency (linkage mapping). khanacademy.org
  5. Creighton, H. B., & McClintock, B. (1931). A correlation of cytological and genetical crossing-over in Zea mays. Proceedings of the National Academy of Sciences, 17(8), 492-497. doi.org/10.1073/pnas.17.8.492
  6. Hunter, N. (2015). Meiotic recombination: The essence of heredity. Cold Spring Harbor Perspectives in Biology, a016618. doi.org/10.1101/cshperspect.a016618
  7. Brown, T. A. (2002). Mutation, repair and recombination. In Genomes (2nd ed.). Wiley-Liss. ncbi.nlm.nih.gov
Key terms
Crossing over
The exchange of matching segments between homologous chromosomes during meiosis I.
Recombination
The production of new allele combinations by crossing over or independent assortment.
Recombinant offspring
Offspring carrying a new allele combination not present in either parent chromosome.
Parental offspring
Offspring carrying the original, non-recombined allele combinations.
Chiasma
The X-shaped point where homologous chromosomes cross over and exchange segments.
Recombination frequency
The fraction of offspring that are recombinant, used to measure distance between genes.

Linkage and Genetic Mapping

  • Explain linkage and why linked genes violate the 9:3:3:1 ratio.
  • Convert recombination frequency into map units (centimorgans).
  • Construct a simple three-gene linear map from cross data.

The big picture

This lesson turns the idea from the last one into a ruler. If genes that sit close together tend to travel together, then measuring how often two genes get separated tells you how far apart they are. Do this for enough pairs of genes and you can draw a map of their order and spacing along a chromosome, without ever seeing the DNA itself.

You will learn the unit geneticists use for these distances, the simple rule that connects it to recombination frequency, and how to line up three genes into a single map from cross data. This is one of the most elegant pieces of reasoning in all of genetics.

Linkage: genes that travel together

Genes located on the same chromosome tend to be inherited together, a phenomenon called linkage. The full set of genes on one chromosome is a linkage group. Perfectly linked genes would always pass into gametes as a unit and never produce recombinants. In reality, crossing over separates linked genes some of the time, and the frequency of that separation is exactly the ruler we need. Linked genes therefore violate Mendel's 9:3:3:1 expectation, producing far more parental combinations than a dihybrid cross of unlinked genes would.

Key idea: Linked genes on the same chromosome are inherited together and deviate from independent assortment, but crossing over separates them at a measurable rate.

Map units and centimorgans

Geneticists define a map unit, also called a centimorgan (cM), as the distance between two genes that produces a 1 percent recombination frequency. The rule could not be simpler:

Map distance (in map units) = percent recombinant offspring.

So if two genes are separated in 12 percent of offspring, they lie 12 map units (12 cM) apart. If another pair is separated in 3 percent of offspring, they are 3 cM apart, and therefore closer together. The centimorgan is named after Thomas Hunt Morgan, whose lab pioneered this work with fruit flies.

This direct rule works well for genes that are reasonably close. For genes far apart, a double crossover (two crossovers between the same two genes) can restore the parental arrangement, so some recombination events go uncounted and the measured frequency underestimates the true distance. For that reason, long distances are built up by adding many short, accurate intervals rather than measuring the ends directly.

Key idea: One map unit (centimorgan) equals 1 percent recombination, so map distance equals the percent recombinant offspring for genes that are not too far apart.

Building a three-gene map

To order three genes, measure the recombination frequency between each pair and use the fact that distances add up along a line. Suppose three genes A, B, and C on one chromosome give:

Gene pairRecombination frequency
A and B8%
B and C12%
A and C20%

The trick is to find which two distances add up to the third. Here 8 (A to B) plus 12 (B to C) equals 20 (A to C), the largest distance. The gene that sits between the other two is the one whose two flanking distances sum to the total, so B lies between A and C. Place A at position 0, B at 8, and C at 20 map units. The resulting linear map is:

Linear genetic map placing gene A at 0, gene B at 8, and gene C at 20 centimorgans A B C 0 8 20

Notice that the A-to-C distance (20) is slightly less than the sum of the parts when double crossovers are common; in real data the outer distance often comes out a little short, which is itself a clue that the middle gene is being crossed over on both sides. For these clean teaching numbers, the parts add up exactly.

Key idea: The gene in the middle is the one whose two flanking recombination frequencies add up to the largest pairwise distance, which fixes the gene order and spacing.

When genes are effectively unlinked

Recombination frequency has a hard ceiling: 50 percent. Two genes so far apart that a crossover almost always occurs between them, or genes on entirely different chromosomes, produce recombinants and parentals in equal numbers, giving a recombination frequency of 50 percent. At that point the genes assort independently and appear unlinked, even if they happen to be on the same chromosome. This is why very distant genes on one chromosome can look just like genes on separate chromosomes.

Key idea: A recombination frequency of 50 percent is the maximum and means two genes assort independently, whether far apart on one chromosome or on different chromosomes.

The three-point test cross, worked in full

Pairwise measurements are fine, but here is something better: a single cross following three genes at once gives the order, both distances, and a measure of interference from one data set. It is the classic problem of a genetics course, and it is entirely mechanical once you know the sequence of steps.

Setup. A fly heterozygous at three linked genes, carrying + + + on one chromosome and a b c on its homolog, is test-crossed to a fly that is a b c / a b c. Among 1,000 offspring:

ClassNumber
+ + +383
a b c382
a + +63
+ b c60
+ + c52
a b +53
+ b +4
a + c3

Step 1: identify the parental classes. They are always the two most numerous, because no crossover is needed to make them. Here they are + + + (383) and a b c (382), totalling 765. This also confirms the arrangement of alleles on the parent's chromosomes.

Step 2: identify the double crossovers. They are always the two rarest, because two simultaneous exchanges are the least likely outcome. Here they are + b + (4) and a + c (3), totalling 7.

Step 3: find the middle gene. Compare a double-crossover class with the parental class it most resembles. Take + + + and + b +. Only the middle gene's allele has switched, and here that is b. So the order is a - b - c.

Step 4: compute each interval. A class is recombinant for an interval if the alleles flanking that interval no longer match the parental arrangement. Double crossovers are recombinant for both intervals, so they must be counted in each.

  • Interval a to b: the single crossovers in that region are a + + (63) and + b c (60), plus both double-crossover classes (7). Recombination frequency = (123 + 7) / 1,000 = 0.130, so 13.0 map units.
  • Interval b to c: the single crossovers there are + + c (52) and a b + (53), plus the doubles (7). Recombination frequency = (105 + 7) / 1,000 = 0.112, so 11.2 map units.

The map is therefore a --- 13.0 --- b --- 11.2 --- c, a total of 24.2 map units from a to c. Measuring a and c directly, ignoring b, would have given only (123 + 105) / 1,000 = 22.8 percent, because a double crossover leaves the outer two genes in their parental arrangement and so goes uncounted. That 1.4 percent gap is precisely the effect that makes long distances underestimate true map length.

Step 5: measure interference. If crossovers in the two intervals were independent, the expected frequency of doubles would be the product of the two single frequencies: 0.130 x 0.112 = 0.0146, which is 14.6 doubles among 1,000 offspring. Only 7 were observed.

  • The coefficient of coincidence is observed doubles divided by expected doubles: 7 / 14.6 = 0.48.
  • Interference is 1 minus that: 1 - 0.48 = 0.52.

An interference of 0.52 means that a crossover in one interval suppressed about half the expected crossovers in the neighboring interval. Interference near 1 means near-complete suppression, and 0 means the two intervals are independent. Values are usually positive in real organisms and fall toward zero as the intervals get further apart.

Key idea: In a three-point cross the parentals are the most frequent and the double crossovers the rarest; comparing them fixes the middle gene, each interval's frequency includes the doubles, and interference equals 1 minus observed doubles over expected doubles.

Where people get stuck

The first sticking point is forgetting to add the double crossovers into each interval. Leave them out and both distances come out too small, and the map no longer adds up.

The second is trying to identify the middle gene by looking at the map distances first. Do it from the double-crossover class instead: it is the only class that tells you the order directly, and it works even before any distance is computed.

The third is expecting genetic and physical distance to be proportional. They are not. Recombination is suppressed near centromeres and elevated at particular hotspots, so a region can be long in map units and short in base pairs, or the reverse. Genetic maps order genes reliably; only sequencing measures physical distance.

A fourth is assuming this method transfers directly to people. It cannot, because human matings are not designed test crosses and family sizes are small. Human linkage analysis instead collects many families, tracks marker variants alongside a trait, and scores the evidence statistically, reporting how much more likely the data are under linkage than under no linkage. The logic is the same; only the bookkeeping changes to suit data the geneticist did not get to design.

Common misconceptions

  • Map units measure recombination frequency, not physical length in nanometers. Regions with lots of crossing over look longer on a genetic map than their DNA length alone would suggest.
  • A recombination frequency cannot exceed 50 percent. If a calculation gives more, an error has been made.
  • Distances add only approximately over long stretches because double crossovers hide some events; short intervals are the most accurate.
  • Linkage does not mean two genes are always inherited together. It means they are inherited together more often than independent assortment would predict.

Recap

  • Linked genes sit on the same chromosome and are inherited together more than chance predicts, deviating from 9:3:3:1.
  • One map unit (centimorgan) equals 1 percent recombination, so map distance equals percent recombinant offspring.
  • Double crossovers make long distances underestimate the truth, so maps are built from short intervals.
  • The middle gene of three is the one whose flanking distances sum to the largest pairwise distance.
  • A recombination frequency of 50 percent is the maximum and means the genes are effectively unlinked.

Sources

  1. Clark, M. A., Douglas, M., & Choi, J. (2018). Biology 2e (Ch. 13.1: Chromosomal theory and genetic linkage). OpenStax. openstax.org
  2. Miko, I. (2008). Developing the chromosome theory. Nature Education, 1(1), 135.
  3. National Human Genome Research Institute. (n.d.). Centimorgan (cM). In Talking glossary of genomic and genetic terms. genome.gov
  4. Khan Academy. (n.d.). Linkage mapping and recombination frequency. khanacademy.org
  5. Sturtevant, A. H. (1913). The linear arrangement of six sex-linked factors in Drosophila, as shown by their mode of association. Journal of Experimental Zoology, 14(1), 43-59. doi.org/10.1002/jez.1400140104
  6. Lander, E. S., & Botstein, D. (1989). Mapping Mendelian factors underlying quantitative traits using RFLP linkage maps. Genetics, 121(1), 185-199. doi.org/10.1093/genetics/121.1.185
  7. National Human Genome Research Institute. (2020). Genetic mapping fact sheet. National Institutes of Health. genome.gov
Key terms
Linkage
The tendency of genes on the same chromosome to be inherited together.
Map unit / centimorgan (cM)
A unit of genetic distance equal to 1 percent recombination frequency.
Genetic map
A diagram of the linear order and relative distances of genes on a chromosome.
Double crossover
Two crossovers between the same two genes, which can restore the parental arrangement.
Recombination frequency
The percent of offspring that are recombinant, used directly as map distance.
Linkage group
A set of genes on the same chromosome that tend to be inherited together.

Module 3: Chromosomes and Chromosomal Disorders

How sex is determined, how genes on sex chromosomes are inherited, and how errors in chromosome number and structure cause genetic disorders.

Sex Determination and Sex-Linked Inheritance

  • Explain how the XY system determines sex in humans.
  • Predict inheritance patterns of X-linked recessive traits.
  • Explain why X-linked recessive traits are more common in males.

The big picture

Most chromosomes come in matched pairs. One special pair does not, and it decides biological sex while giving certain genes an unusual inheritance pattern along the way. This lesson explains how the XY system works and why disorders such as color blindness and hemophilia show up far more often in males than in females. The reasoning is a direct payoff of everything you know about dominant and recessive alleles, applied to a chromosome that males have only one copy of.

Once you can build a Punnett square for an X-linked gene, you can predict which sons and daughters are affected or carriers, and you can recognize the tell-tale pattern of an X-linked trait in a family tree.

How sex is determined

In humans, 22 of the 23 chromosome pairs are autosomes: any chromosome that is not a sex chromosome. The remaining pair are the sex chromosomes, the X and the Y, and their combination sets biological sex. In the human XY system, individuals with two X chromosomes (XX) typically develop as female, and those with one X and one Y (XY) typically develop as male.

The deciding factor is a single gene on the Y chromosome called SRY: when SRY is present it triggers male development, and when it is absent development follows the female pathway. Because the father contributes either an X or a Y while the mother always contributes an X, it is the father's gamete that determines the sex of the child.

Key idea: The XY system sets sex by the presence or absence of the Y chromosome's SRY gene, and the father's sperm (X or Y) determines a child's sex.

Genes on the X chromosome

The X chromosome is large and carries more than a thousand genes, most of which have nothing to do with sex, including genes for color vision and blood clotting. The Y chromosome is small and carries very few genes. A gene located on the X chromosome is called X-linked, and this location produces a distinctive inheritance pattern because of the mismatch between the sexes.

A female (XX) has two copies of every X-linked gene, so a recessive allele on one X can be masked by a dominant allele on the other, exactly like an autosomal gene. A male (XY) has only one X, so whatever allele he carries on it is expressed, whether dominant or recessive, because there is no second X to mask it. Males are said to be hemizygous for X-linked genes: having only one copy of a gene, so a single allele determines the phenotype.

Key idea: Females have two copies of X-linked genes and can mask a recessive allele, but males are hemizygous, so their single X-linked allele is always expressed.

Why males are affected more often

This asymmetry explains why X-linked recessive disorders, such as red-green color blindness and hemophilia, appear far more often in males. A male needs only one copy of the recessive allele to be affected, because he has a single X. A female needs two copies, one on each X, which is much rarer. A female with just one copy of the recessive allele does not show the trait; she is an unaffected carrier: a heterozygous individual who carries a recessive allele without showing the trait but can pass it on.

Work the classic cross. Let XA be the normal (dominant) allele and Xa the recessive disease allele. Cross a carrier mother (XAXa) with an unaffected father (XAY):

XA (from mother)Xa (from mother)
XA (father)XAXA daughter, unaffectedXAXa daughter, carrier
Y (father)XAY son, unaffectedXaY son, affected

Read the results by sex. Among the daughters, half are unaffected non-carriers (XAXA) and half are unaffected carriers (XAXa); none are affected, because the father gave every daughter his normal XA. Among the sons, half are unaffected (XAY) and half are affected (XaY). So a carrier mother and a normal father produce, on average, affected sons but no affected daughters.

Key idea: Crossing a carrier mother with a normal father gives about half the sons affected and no affected daughters, the classic X-linked recessive result.

Reading the pattern in families

X-linked recessive inheritance leaves a signature in a family tree. An affected son inherits his single X, and therefore the disease allele, from his mother, who is usually an unaffected carrier. The trait often appears to skip generations along the maternal line, surfacing in grandsons through carrier daughters. A key negative clue: an affected father cannot pass an X-linked recessive trait to his sons, because he gives sons his Y, not his X. Father-to-son transmission argues against X-linked inheritance.

Key idea: X-linked recessive traits pass from carrier mothers to affected sons and never directly from father to son, which is how the pattern is recognized.

Reciprocal crosses give different answers, and that is the test

For an autosomal gene it makes no difference which parent carries which allele. For an X-linked gene it makes all the difference, and that asymmetry is the cleanest diagnostic in classical genetics. Thomas Hunt Morgan found it in 1910 with a single white-eyed male fruit fly among thousands of red-eyed ones.

Cross 1: white-eyed male x homozygous red-eyed female. The father gives his X, carrying white, only to his daughters, along with his Y to his sons. Sons therefore get their single X from their mother and are all red-eyed; daughters get one white X from the father and one red X from the mother and are all red-eyed carriers. Every offspring is red-eyed. Let those F1 flies interbreed and the F2 shows the expected 3 red : 1 white overall, but with a twist that no autosomal gene could produce: every white-eyed fly is male.

Cross 2: the reciprocal. Now use a white-eyed female and a red-eyed male. Every son receives his only X from his white-eyed mother and is white-eyed. Every daughter receives the father's red X plus a white X and is red-eyed. The result is a complete reversal by sex, sometimes called criss-cross inheritance, because sons resemble their mother and daughters their father.

Compare the two crosses and the logic is inescapable. An autosomal gene gives identical results either way round. A gene that gives different results depending on which parent carried the allele must sit on a chromosome that the two sexes do not share equally.

Key idea: Reciprocal crosses give identical results for autosomal genes and opposite results for X-linked genes, which is how Morgan placed the white-eye gene on the X.

X inactivation makes every female a mosaic

A female has two X chromosomes and a male has one, yet both need roughly the same amount of X-encoded protein. The solution, proposed by Mary Lyon in 1961, is X inactivation: early in development each cell shuts down one of its two X chromosomes, condensing it into a dense body that is largely silent.

Three features of this process matter.

  • It is random. Each cell independently silences either the maternal or the paternal X.
  • It is early. In humans it happens when the embryo has only a few dozen cells.
  • It is inherited by descendants. Every cell descended from that early cell keeps the same X switched off.

The consequence is that a female heterozygous for any X-linked gene is not uniformly intermediate. She is a patchwork of clonal territories, some expressing one allele and some the other. Tortoiseshell and calico cats show this directly, since the gene for orange versus black coat color sits on the X, and the patches on the cat are the clones. The same principle explains why some female carriers of X-linked conditions have patchy signs, and why a carrier occasionally has symptoms if inactivation happened to fall unevenly and silenced the working copy in most of the relevant tissue.

Key idea: Random, early, clonally inherited X inactivation equalizes dosage between the sexes and makes every heterozygous female a mosaic of cell patches expressing one allele or the other.

Reading a pedigree step by step

A pedigree is a family diagram: squares for males, circles for females, filled symbols for affected individuals, horizontal lines for matings, and vertical lines to offspring. Reading one is a process of elimination, and the order of the questions matters.

  1. Does the trait skip generations? If unaffected parents have an affected child, the allele is recessive. If it appears in every generation and every affected person has an affected parent, it is probably dominant.
  2. Are the sexes affected equally? Roughly equal numbers point to an autosomal gene. A strong excess of affected males points to X-linked recessive.
  3. Look for father-to-son transmission. A father gives his son a Y, never his X, so a single clear case of an affected father with an affected son rules out X-linkage completely. This one observation is the most powerful line in pedigree analysis.
  4. Check the daughters of affected males. Under X-linked dominant inheritance an affected father passes the trait to all of his daughters and none of his sons, which is a distinctive signature.
  5. Consider the alternatives. Transmission from a mother to every one of her children, with no transmission from fathers, suggests a mitochondrial gene rather than a nuclear one.

Worked example. Two unaffected parents have four children: an affected son, an unaffected son, and two unaffected daughters. The mother's brother is also affected; no one else is. Work the questions in order. Unaffected parents with an affected child means recessive. All affected individuals are male, and the second affected person is on the mother's side, which fits transmission through carrier females. No father-to-son transmission appears anywhere. The best-supported mode is X-linked recessive, with the mother a carrier.

Be honest about the limits. This pedigree is also compatible with autosomal recessive inheritance, since small families produce lopsided sex ratios by chance alone. Pedigree analysis narrows the possibilities and ranks them; it rarely proves one mode from a handful of individuals, which is why molecular testing has largely replaced it for diagnosis while remaining essential for interpreting what a test result means for relatives.

Key idea: Work a pedigree in order - skipping generations for recessive, sex ratio for X-linkage, and father-to-son transmission as the decisive test - and state the conclusion as the best-supported mode rather than a proof.

Where people get stuck

The first sticking point is the phrase carrier. It means a heterozygote who does not show the trait, so it applies to females for X-linked recessive conditions and to either sex for autosomal recessive ones. A male cannot be a carrier of an X-linked recessive allele: with only one X, he either has it and shows it, or does not have it at all.

The second is expecting a Punnett square for an X-linked cross to look like an autosomal one. It will not: write the male's genotype with the Y included, as XaY rather than just a, or half the offspring classes will go missing.

The third is describing X-linked conditions as male diseases. That overstates it: females can be affected, either by inheriting two copies or through skewed X inactivation, and carrier females can have real clinical findings. The correct statement is that males are affected far more often, not that females are never affected.

Common misconceptions

  • The mother does not determine a child's sex. Because she always contributes an X, it is the father's X-or-Y sperm that decides.
  • X-linked is not the same as sex-limited. X-linked genes (like color vision) affect traits unrelated to sex; they simply sit on the X chromosome.
  • A carrier mother is not affected. She has a normal allele on her second X that masks the recessive one.
  • Fathers cannot pass X-linked recessive traits to sons. Sons get the father's Y, so an affected son's allele comes from the mother.

Recap

  • Autosomes are non-sex chromosomes; the X and Y sex chromosomes determine sex, with the Y's SRY gene triggering male development.
  • Males are hemizygous for X-linked genes, so a single recessive allele is expressed; females can mask it with a second X.
  • X-linked recessive disorders (color blindness, hemophilia) are more common in males for this reason.
  • A carrier mother crossed with a normal father yields about half affected sons and no affected daughters.
  • The trait passes from carrier mothers to sons and never father to son, its diagnostic pattern.

Sources

  1. Clark, M. A., Douglas, M., & Choi, J. (2018). Biology 2e (Ch. 13.2: Chromosomal basis of inherited disorders). OpenStax. openstax.org
  2. National Human Genome Research Institute. (n.d.). X-linked. In Talking glossary of genomic and genetic terms. genome.gov
  3. Chandley, A. C. (2008). Sex chromosomes and sex determination. Nature Education, 1(1).
  4. Khan Academy. (n.d.). Sex linkage and X-linked inheritance. khanacademy.org
  5. Morgan, T. H. (1910). Sex limited inheritance in Drosophila. Science, 32(812), 120-122. doi.org/10.1126/science.32.812.120
  6. Sinclair, A. H., Berta, P., Palmer, M. S., Hawkins, J. R., Griffiths, B. L., Smith, M. J., Foster, J. W., Frischauf, A.-M., Lovell-Badge, R., & Goodfellow, P. N. (1990). A gene from the human sex-determining region encodes a protein with homology to a conserved DNA-binding motif. Nature, 346(6281), 240-244. doi.org/10.1038/346240a0
  7. MedlinePlus Genetics. (2021). What are the different ways a genetic condition can be inherited? National Library of Medicine. medlineplus.gov
Key terms
Autosome
Any chromosome that is not a sex chromosome (chromosomes 1 through 22 in humans).
Sex chromosomes
The X and Y chromosomes, whose combination determines biological sex.
X-linked
Located on the X chromosome, giving a sex-dependent inheritance pattern.
Hemizygous
Having only one copy of a gene, as males do for X-linked genes.
Carrier
A heterozygous individual who carries a recessive allele without showing the trait.
SRY gene
The Y-chromosome gene that triggers male development.

Chromosomal Abnormalities and Genetic Disorders

  • Explain nondisjunction and how it leads to aneuploidy.
  • Describe common chromosomal disorders such as trisomy 21.
  • Distinguish changes in chromosome number from changes in structure.

The big picture

Meiosis is astonishingly precise. It is not, however, perfect. When chromosomes fail to divide correctly, or when pieces break and rejoin in the wrong place, the result can be a genetic disorder. This lesson sorts these errors into two clear families, changes in the number of chromosomes and changes in their structure, and shows how each one produces recognizable conditions such as Down syndrome.

The unifying idea is dosage: cells are finely tuned to have exactly two copies of each chromosome, so having one too many or one too few, or having genes rearranged, upsets the balance. Understanding these errors also explains how doctors detect them using a picture of a person's chromosomes.

Two families of error

Chromosomal errors fall into two broad categories. The first is a change in chromosome number, where a cell gains or loses whole chromosomes. The second is a change in chromosome structure, where chromosomes keep their number but pieces are lost, repeated, flipped, or moved. We take each in turn.

Nondisjunction and aneuploidy

The most common numerical error is nondisjunction: the failure of chromosomes to separate properly during meiosis. If a homologous pair fails to separate in meiosis I, or if sister chromatids fail to separate in meiosis II, some gametes end up with an extra chromosome and others with one too few. When such a gamete joins a normal gamete at fertilization, the offspring has an abnormal chromosome count, a condition called aneuploidy: having one or a few chromosomes more or fewer than the normal set. Having three copies of a particular chromosome is trisomy; having only one copy where there should be a pair is monosomy.

Key idea: Nondisjunction is the failed separation of chromosomes in meiosis, producing aneuploid gametes that lead to trisomy (three copies) or monosomy (one copy).

Common aneuploidy disorders

The best-known example is trisomy 21, or Down syndrome, in which a person has three copies of chromosome 21. Because chromosome 21 is small and carries relatively few genes, individuals with trisomy 21 survive and live full lives, though with characteristic physical features and some associated health considerations. This survivability is the exception rather than the rule: trisomies of larger, gene-rich chromosomes usually disrupt development so severely that the embryo does not survive, which is why most autosomal trisomies are never seen in liveborn children.

Aneuploidy of the sex chromosomes tends to be much better tolerated, because the Y carries few genes and cells naturally shut down extra X chromosomes. Examples include Turner syndrome (a single X, written 45,X, a monosomy) and Klinefelter syndrome (XXY, a trisomy of the sex chromosomes). The chance of nondisjunction rises with the age of the egg, which is why the frequency of trisomy 21 increases with maternal age.

Key idea: Trisomy 21 (Down syndrome) is survivable because chromosome 21 is small, while most other autosomal trisomies are not; sex-chromosome aneuploidies are generally better tolerated.

Changes in chromosome structure

Chromosomes can also break and rejoin incorrectly, altering their structure rather than their number. There are four main rearrangements:

  • In a deletion, a segment of a chromosome is lost.
  • In a duplication, a segment is repeated so that it appears twice.
  • In an inversion, a segment breaks out and reinserts backward, reversing the order of its genes.
  • In a translocation, a segment moves to a different, nonhomologous chromosome.

These rearrangements can disrupt a gene right at a break point, or change how nearby genes are regulated, and several are linked to specific cancers and inherited syndromes. The table summarizes both families of error side by side.

TypeWhat happens
NondisjunctionChromosomes fail to separate, giving abnormal counts
Trisomy / monosomyThree copies / one copy of a chromosome
DeletionA chromosome segment is lost
DuplicationA chromosome segment is repeated
InversionA segment is reversed in orientation
TranslocationA segment moves to a nonhomologous chromosome

Key idea: Structural changes keep chromosome number the same but rearrange the genetic material through deletion, duplication, inversion, or translocation.

Detecting chromosomal disorders

Both numerical and large structural changes can be seen by examining a karyotype: an organized display of all of an individual's chromosomes, arranged in order by size and shape. A karyotype instantly reveals an extra or missing chromosome (as in trisomy 21) and large rearrangements such as a translocation. What it cannot show is a single changed base in the DNA, which is far too small to appear at this scale; detecting those requires DNA sequencing, covered later. Karyotyping is a routine part of prenatal testing and cancer diagnosis.

Key idea: A karyotype displays all chromosomes by size, revealing whole-chromosome and large structural changes, but not single-base mutations.

Maternal age, in numbers

The association between maternal age and trisomy is one of the strongest and best-measured effects in human genetics, and it follows directly from the cohesin story in the meiosis lesson. Oocytes begin meiosis before birth and pause partway through, holding their chromosomes together with cohesin that is not replaced during the wait. The longer the wait, the more that glue degrades, and the more often a pair separates incorrectly.

The published age-specific chances of a live birth with trisomy 21 run roughly as follows:

Maternal ageApproximate chance at term
20about 1 in 1,500
30about 1 in 900
35about 1 in 350
40about 1 in 100
45about 1 in 30

Two further facts put this in context. Because most births are to younger people, the majority of children with Down syndrome are born to people under 35 even though the per-pregnancy chance is lower there. And chromosome errors at conception are far commoner than these figures suggest: roughly half of first-trimester pregnancy losses carry a chromosomal abnormality, so the numbers above describe what survives to term rather than what occurs.

Key idea: Age-related cohesin loss in long-arrested oocytes raises the chance of trisomy 21 at term from about 1 in 1,500 at age 20 to about 1 in 30 at 45, while about half of early pregnancy losses are chromosomally abnormal.

Three routes to trisomy 21, and why the distinction matters

Down syndrome is not a single genetic event. Three mechanisms produce it, and they carry different implications for a family.

  • Free trisomy 21 accounts for about 95 percent of cases and arises from nondisjunction in a single gamete. It is not inherited, and the chance of it recurring in a later pregnancy is only slightly above the age-related figure.
  • Robertsonian translocation accounts for roughly 3 to 4 percent. Here a copy of chromosome 21 is attached to another chromosome, most often 14. A parent can carry such a translocation in balanced form, with the normal amount of genetic material and no clinical features, yet produce unbalanced gametes. Recurrence risk for such a couple is substantially higher, which is why a karyotype is still ordered after a diagnosis even in the sequencing era.
  • Mosaicism accounts for 1 to 2 percent, where the error happens after fertilization so only some cell lines carry the extra chromosome.

Outcomes vary widely between individuals in all three groups. Down syndrome is associated with intellectual disability of variable degree and with a higher chance of certain heart and thyroid conditions, and average life expectancy has risen substantially over recent decades with better medical care. The genetics predicts the chromosome count; it does not predict an individual life.

Key idea: About 95 percent of trisomy 21 is free trisomy from nondisjunction, 3 to 4 percent involves a translocation that a balanced parent can transmit, and 1 to 2 percent is mosaic, so a karyotype still guides recurrence counseling.

Screening is not diagnosis: a worked predictive value

Cell-free DNA screening analyzes fragments of placental DNA in a pregnant person's blood. For trisomy 21 it detects about 99 percent of affected pregnancies with a false-positive rate near 0.1 percent, figures that sound conclusive. They are not, because what a positive result means depends on how common the condition is in that population. Work it twice.

Case 1: 100,000 pregnancies at age 25, where the chance is about 1 in 1,200.

  • Affected pregnancies: 100,000 / 1,200 = about 83. The test detects 99 percent, so 82 true positives.
  • Unaffected pregnancies: about 99,917. A 0.1 percent false-positive rate gives 100 false positives.
  • Positive predictive value = 82 / (82 + 100) = about 45 percent. A positive result is closer to a coin flip than to a diagnosis.

Case 2: 100,000 pregnancies at age 40, where the chance is about 1 in 100.

  • Affected: 1,000, of which 99 percent are detected, giving 990 true positives.
  • Unaffected: 99,000, giving 99 false positives.
  • Positive predictive value = 990 / (990 + 99) = about 91 percent.

Identical test, identical accuracy, and a positive result that means something very different in the two cases. This is why cell-free DNA testing is called a screen and why a positive result is followed by a diagnostic test such as chorionic villus sampling or amniocentesis, which examine fetal cells directly. The general lesson reaches well beyond prenatal testing: the predictive value of any test depends on the prevalence in the population being tested, not on the test alone.

Key idea: With 99 percent detection and 0.1 percent false positives, the positive predictive value of cell-free DNA screening for trisomy 21 is about 45 percent at age 25 and about 91 percent at age 40, so a positive screen requires diagnostic confirmation.

Where people get stuck

The first sticking point is thinking that a bigger chromosome means a milder trisomy. The opposite holds: trisomies are survivable only for the smallest, gene-poorest chromosomes such as 21, 18, and 13, because the extra dose of every gene on a large chromosome is not tolerated.

The second is treating a balanced translocation as harmless in every sense. It is not, quite: the carrier is usually completely healthy, since no genetic material is missing or extra, but their gametes can be unbalanced, so the consequence appears in the next generation rather than in them.

The third is reading a screening result as a diagnosis. It is not one: a screen sorts a population into higher and lower probability groups. Only a diagnostic test examines the fetal chromosomes themselves, and the arithmetic above shows how far apart those two things can be.

Common misconceptions

  • Down syndrome is a chromosome-number change (an extra chromosome 21), not a single-gene mutation.
  • Nondisjunction can happen in either meiosis I or meiosis II, and its likelihood rises with the age of the egg.
  • A translocation moves a segment between nonhomologous chromosomes and usually does not change the total chromosome count, so it is a structural change, not aneuploidy.
  • A karyotype cannot detect small mutations. It shows chromosome number and gross structure only.

Recap

  • Chromosomal errors are either changes in number or changes in structure.
  • Nondisjunction causes aneuploidy, producing trisomy (three copies) or monosomy (one copy).
  • Trisomy 21 causes Down syndrome and is survivable; most other autosomal trisomies are not, while sex-chromosome aneuploidies are better tolerated.
  • Structural changes include deletion, duplication, inversion, and translocation.
  • A karyotype reveals whole-chromosome and large structural changes but not single-base mutations.

Sources

  1. Clark, M. A., Douglas, M., & Choi, J. (2018). Biology 2e (Ch. 13.2: Chromosomal basis of inherited disorders). OpenStax. openstax.org
  2. National Human Genome Research Institute. (n.d.). Aneuploidy. In Talking glossary of genomic and genetic terms. genome.gov
  3. O'Connor, C. (2008). Chromosomal abnormalities: Aneuploidies. Nature Education, 1(1), 172.
  4. Khan Academy. (n.d.). Chromosomal mutations and nondisjunction (chromosomal basis of genetics). khanacademy.org
  5. Hassold, T., & Hunt, P. (2001). To err (meiotically) is human: The genesis of human aneuploidy. Nature Reviews Genetics, 2(4), 280-291. doi.org/10.1038/35066065
  6. Nagaoka, S. I., Hassold, T. J., & Hunt, P. A. (2012). Human aneuploidy: Mechanisms and new insights into an age-old problem. Nature Reviews Genetics, 13(7), 493-504. doi.org/10.1038/nrg3245
  7. MedlinePlus Genetics. (2021). What is noninvasive prenatal testing (NIPT) and what disorders can it screen for? National Library of Medicine. medlineplus.gov
Key terms
Nondisjunction
Failure of chromosomes to separate properly during meiosis.
Aneuploidy
Having an abnormal number of chromosomes, such as one extra or one missing.
Trisomy
The presence of three copies of a particular chromosome, as in trisomy 21.
Monosomy
The presence of only one copy of a chromosome that is normally paired.
Translocation
The movement of a chromosome segment to a nonhomologous chromosome.
Karyotype
An organized display of an individual's full set of chromosomes.

Module 4: DNA Structure, Replication, and Gene Expression

The molecular identity of the gene, how DNA copies itself, and how the information in DNA is transcribed and translated into proteins.

DNA Structure and Replication

  • Describe the double-helix structure of DNA and base pairing.
  • Explain semiconservative replication.
  • Name the major enzymes of DNA replication and their roles.

The big picture

Up to now a gene has been an abstract unit that follows rules. This lesson gives it a physical body. Genes are made of DNA, a long molecule whose elegant structure, discovered in 1953, immediately hinted at how it stores information and copies itself. Once you see the shape, the copying almost explains itself.

You will learn the parts of DNA, the strict base-pairing rule that lets one strand specify the other, and the team of enzymes that duplicates the entire genome every time a cell divides, with remarkable accuracy. This molecular foundation underlies replication, mutation, and everything in the rest of the course.

What genes are made of

For decades geneticists knew that genes ride on chromosomes, but what were genes actually made of, chemically? Mid-twentieth-century experiments settled the question: genes are made of DNA, deoxyribonucleic acid. In 1953, James Watson and Francis Crick, using the X-ray diffraction images produced by Rosalind Franklin, deduced its structure: a double helix, two strands wound around each other like a gently twisted ladder.

The structure of DNA

Each strand is a chain of building blocks called nucleotides. A nucleotide has three parts: a sugar (deoxyribose), a phosphate group, and one of four nitrogen-containing bases. The four bases are adenine (A), thymine (T), cytosine (C), and guanine (G). In the twisted ladder, the alternating sugars and phosphates form the two side rails (the backbone), while the bases point inward and pair up to form the rungs.

The pairing is strict and specific, a rule called complementary base pairing: A always pairs with T, and C always pairs with G. A pairs with T through two hydrogen bonds, and C pairs with G through three, which is why C-G pairs are a little more stable. The two strands run in opposite directions, described as antiparallel. The most important consequence of complementarity is that knowing the sequence of one strand automatically tells you the sequence of the other. If one strand reads A-T-G-C, its partner must read T-A-C-G.

Key idea: DNA is a double helix of nucleotides in which A pairs with T and C pairs with G, so each strand fully specifies its partner.

Reading a complementary strand

Because the strands are antiparallel and complementary, you can always reconstruct one strand from the other. Work an example. Suppose one strand reads, in the 5-prime to 3-prime direction, 5'-A T G C C G T A-3'. Apply the pairing rule base by base (A with T, T with A, G with C, C with G) and reverse the direction, because the partner runs the opposite way. The complementary strand is 3'-T A C G G C A T-5'. Each base determines its partner with no ambiguity, which is exactly what makes faithful copying possible.

Key idea: To write a complement, pair each base with its partner (A-T, C-G) along the antiparallel strand, and the result is fully determined by the original.

Semiconservative replication

Complementary base pairing immediately suggests how DNA copies itself. If the two strands unzip and separate, each old strand can serve as a template, a pattern for building a new complementary partner. The outcome is two DNA molecules, each made of one old strand and one brand-new strand. This mechanism is called semiconservative replication, because each daughter molecule conserves (keeps) half of the original. Matthew Meselson and Franklin Stahl confirmed it experimentally in 1958, in what has been called the most beautiful experiment in biology.

Key idea: Replication is semiconservative: the strands separate, each templates a new partner, and every daughter molecule keeps one original strand and gains one new one.

The enzymes of replication

Replication is carried out by a coordinated team of enzymes. Each has a specific job:

  • Helicase unwinds and separates the two strands, opening a Y-shaped region called the replication fork.
  • DNA polymerase reads each template strand and adds complementary nucleotides to build the new strand. It can only add nucleotides in one direction along a template.
  • DNA ligase stitches together the short pieces of new DNA that form on one of the two strands.

Because the strands are antiparallel, DNA polymerase can copy one new strand (the leading strand) smoothly and continuously, but must build the other (the lagging strand) in short segments that ligase then joins. DNA polymerase also proofreads its own work, backing up to remove a mismatched base before moving on. This proofreading keeps replication astonishingly accurate, which matters enormously: every time a human cell divides it must copy roughly three billion base pairs with only a handful of errors.

Key idea: Helicase unwinds the helix, DNA polymerase builds and proofreads the new strands, and ligase joins the fragments, together copying the genome with very high fidelity.

How the structure was actually deduced

The double helix was not guessed. It was assembled from three independent pieces of evidence, and knowing them makes the structure much harder to forget.

Chargaff's rules. Erwin Chargaff measured base composition across many species in the late 1940s and found something odd: the amount of adenine always equaled the amount of thymine, and guanine always equaled cytosine, even though the overall proportion of A plus T versus G plus C varied widely between organisms. Any correct model had to explain equality within pairs and variability between species.

X-ray diffraction. Rosalind Franklin and Raymond Gosling produced diffraction images of hydrated DNA fibers, of which the celebrated Photo 51 shows a clear X-shaped pattern. An X of that form is the diffraction signature of a helix. The spacing of the marks gave the numbers directly: a repeat every 3.4 nanometers along the fiber, individual bases stacked 0.34 nanometers apart, and a uniform width of about 2 nanometers.

The geometry that ties them together. A constant 2-nanometer width is only possible if every rung of the ladder is the same length. Purines (A and G) are two-ring bases and pyrimidines (C and T) are one-ring bases, so a purine must always pair with a pyrimidine: two purines would bulge and two pyrimidines would pinch. Combine that constraint with Chargaff's equalities and only one arrangement survives, A with T and G with C. Watson and Crick published that model in 1953, and the pairing rule immediately suggested how the molecule could be copied.

Key idea: Chargaff's base equalities, the helical X pattern and 2-nanometer width from X-ray diffraction, and the requirement that every rung be a purine paired with a pyrimidine together force the A-T and G-C pairing rule.

Not all base pairs are equal

An A-T pair is held by two hydrogen bonds and a G-C pair by three. That single difference has practical consequences that appear throughout molecular biology.

  • DNA rich in G and C takes more heat to separate into single strands, so its melting temperature is higher. Sequences from organisms living in hot springs tend to be GC-rich for exactly this reason.
  • Anyone designing a primer for a PCR reaction estimates its melting temperature from its base composition, because a primer that is too AT-rich will not stay bound at the working temperature.
  • Replication origins are typically AT-rich, since the helix has to be pried open there and the weaker pairing makes that easier.

One further structural point is easy to skim past and matters a great deal. The two strands run in opposite directions, described as antiparallel, so one runs 5 prime to 3 prime while its partner runs 3 prime to 5 prime. Because DNA polymerase can only add nucleotides to a 3 prime end, this single geometric fact forces one new strand to be built continuously and the other in fragments, which is where the leading and lagging strands come from.

Key idea: G-C pairs use three hydrogen bonds and A-T pairs two, so GC content sets melting temperature, and the antiparallel arrangement is what forces one strand to be copied discontinuously.

The fidelity budget

How accurate is replication, really? Not the product of one careful enzyme, but of three successive filters stacked together, and the numbers multiply.

  1. Base selection. The polymerase's active site fits correct base pairs better than incorrect ones. On its own this gives roughly one error in 105 bases.
  2. Proofreading. The same enzyme carries a second active site that clips off a mismatched nucleotide it has just added and tries again. This improves accuracy roughly a hundredfold, to about one error in 107.
  3. Mismatch repair. A separate system scans newly made DNA, spots the distortion a mismatch causes, and excises the wrong base from the new strand. This adds another hundred to thousandfold, giving a final rate near one error in 109 to 1010.

Worked example. A human cell copies about 6.4 x 109 base pairs each division. At a final error rate of 10-10 per base, the expected number of uncorrected errors is 6.4 x 109 x 10-10 = about 0.6 per cell division. Take away mismatch repair and set the rate at 10-7, and the same arithmetic gives about 640 new mutations per division, a thousandfold increase. That is why inherited defects in mismatch repair raise cancer risk so sharply.

Key idea: Base selection, proofreading, and mismatch repair multiply to about one error per 109 to 1010 bases, which is roughly one uncorrected error each time a human cell copies its 6.4 x 109 base pairs.

Where people get stuck

The first sticking point is the direction of synthesis. DNA polymerase always adds to a 3 prime end, without exception, so the new strand grows 5 prime to 3 prime while it reads its template 3 prime to 5 prime. Every question about leading and lagging strands resolves back to that one rule.

The second is expecting DNA polymerase to start a strand from nothing. It cannot. It only extends an existing 3 prime end, so a separate enzyme lays down a short RNA primer first, and that primer is later removed and replaced. The need for primers is also the root of the problem of copying the very ends of a linear chromosome.

The third is picturing a single replication fork travelling the length of a human chromosome. At about 50 nucleotides per second that would take weeks. Human cells instead open tens of thousands of origins at once and copy short stretches in parallel, which is how a genome of this size is duplicated in a few hours.

Common misconceptions

  • A does not pair with C or G. In DNA the only pairs are A-T and C-G; uracil replaces thymine only in RNA.
  • Replication is semiconservative, not conservative. Each new molecule contains one old strand and one new strand, not two new strands.
  • DNA polymerase does not start from nothing on a bare strand; it extends an existing primer, and it can only add nucleotides in one direction.
  • The two strands are antiparallel. They are not identical copies running the same way; they are complements running in opposite directions.

Recap

  • Genes are made of DNA, a double helix of nucleotides (sugar, phosphate, and a base).
  • Complementary base pairing (A-T, C-G) on antiparallel strands means each strand specifies the other.
  • You can write a complement by applying the pairing rule base by base.
  • Replication is semiconservative: each daughter molecule keeps one old strand and one new one.
  • Helicase, DNA polymerase (which proofreads), and ligase together copy the genome accurately.

Sources

  1. Clark, M. A., Douglas, M., & Choi, J. (2018). Biology 2e (Ch. 14: DNA structure and function). OpenStax. openstax.org
  2. National Human Genome Research Institute. (n.d.). Deoxyribonucleic acid (DNA). In Talking glossary of genomic and genetic terms. genome.gov
  3. Pray, L. A. (2008). Discovery of DNA structure and function: Watson and Crick. Nature Education, 1(1), 100.
  4. Khan Academy. (n.d.). DNA as the genetic material (DNA structure and replication). khanacademy.org
  5. Watson, J. D., & Crick, F. H. C. (1953). Molecular structure of nucleic acids: A structure for deoxyribose nucleic acid. Nature, 171(4356), 737-738. doi.org/10.1038/171737a0
  6. Meselson, M., & Stahl, F. W. (1958). The replication of DNA in Escherichia coli. Proceedings of the National Academy of Sciences, 44(7), 671-682. doi.org/10.1073/pnas.44.7.671
  7. Alberts, B., Johnson, A., Lewis, J., Raff, M., Roberts, K., & Walter, P. (2002). DNA replication mechanisms. In Molecular biology of the cell (4th ed.). Garland Science. ncbi.nlm.nih.gov
Key terms
Nucleotide
A DNA or RNA building block made of a sugar, a phosphate, and a nitrogenous base.
Complementary base pairing
The rule that A pairs with T and C pairs with G in DNA.
Double helix
The two-stranded, twisted-ladder structure of DNA.
Semiconservative replication
DNA copying in which each new molecule keeps one old strand and one new strand.
DNA polymerase
The enzyme that builds a new DNA strand by adding complementary nucleotides to a template.
Helicase
The enzyme that unwinds and separates the two DNA strands during replication.

Transcription and Translation

  • State the central dogma of molecular biology.
  • Describe transcription and the role of mRNA.
  • Explain how the genetic code is read during translation.

The big picture

DNA stores instructions, but it does not build anything itself. So how does a gene actually become a working protein? Through a two-step relay: first the gene is copied into a portable RNA message, then that message is read to assemble a chain of amino acids. This lesson follows that relay from start to finish and shows you how to translate a DNA sequence into the protein it encodes.

Getting comfortable with codons and the genetic code is what lets you predict exactly how a change in DNA will change a protein, which is the foundation for understanding mutations in the next module.

The central dogma

The overall flow of genetic information is captured in the central dogma of molecular biology: information moves from DNA to RNA to protein. The first step, copying DNA into RNA, is transcription. The second step, using that RNA to build a protein, is translation. DNA is the master archive that stays safe in the nucleus; RNA is the working copy that carries the message out to where proteins are made.

Key idea: The central dogma is DNA to RNA to protein, achieved by transcription (DNA to RNA) and translation (RNA to protein).

Transcription

Transcription copies the information in a gene from DNA into a molecule of messenger RNA (mRNA), the RNA copy that carries a gene's instructions to the ribosome. The enzyme RNA polymerase reads one strand of the DNA (the template strand) and builds a complementary RNA strand, following the same base-pairing logic as replication with one twist.

RNA differs from DNA in two ways: its sugar is ribose instead of deoxyribose, and it uses the base uracil (U) in place of thymine. So wherever the DNA template has an A, the new RNA gets a U rather than a T. For example, a DNA template reading A-C-G is transcribed into RNA as U-G-C. In eukaryotic cells the finished mRNA is processed and then travels out of the nucleus to the ribosomes in the cytoplasm, where translation happens.

Key idea: Transcription uses RNA polymerase to build an mRNA copy of a gene, pairing bases as in DNA but inserting uracil (U) wherever the template has adenine.

The genetic code

The mRNA is read in three-base words called codons. A codon is a sequence of three mRNA bases that specifies one amino acid or a stop signal. An amino acid is a building block of proteins; a protein is a chain of amino acids folded into a functional shape. The correspondence between codons and amino acids is the genetic code.

With four possible bases arranged in groups of three, there are 4 x 4 x 4 = 64 possible codons, far more than the 20 amino acids they need to specify. As a result the code is redundant (also called degenerate): most amino acids are specified by several different codons. Three special features are worth memorizing: the codon AUG signals the start of a protein and also codes for the amino acid methionine, and three codons (UAA, UAG, UGA) act as stop signals that end the protein. The code is also nearly universal, shared by almost all living things, which is powerful evidence of common ancestry.

Key idea: Codons are three-base words; 64 codons encode 20 amino acids plus start and stop signals, so the redundant genetic code has several codons per amino acid.

Translation

Translation is the synthesis of a protein from an mRNA sequence, and it takes place on the ribosome, the cellular machine that reads mRNA and links amino acids. The adapters that make translation work are molecules of transfer RNA (tRNA): each tRNA carries one specific amino acid and displays a three-base anticodon that pairs with the matching codon on the mRNA.

The ribosome moves along the mRNA one codon at a time. At each codon, the tRNA with the matching anticodon delivers its amino acid, and the ribosome links that amino acid to the growing chain. When a stop codon is reached, no tRNA matches it, so the finished protein is released and folds into the shape that determines its job. In short, the order of bases in DNA sets the order of codons in mRNA, which sets the order of amino acids in the protein, and that sequence determines what the protein does.

Key idea: During translation the ribosome reads mRNA codon by codon while tRNAs deliver matching amino acids, building a protein until a stop codon ends it.

A worked example: from gene to protein

Trace a tiny gene all the way to protein. Suppose the DNA template strand reads 3'-T A C G G A A T C-5'. Transcribe it into mRNA by complementary pairing (remembering U replaces T), reading the mRNA in the 5-prime to 3-prime direction: 5'-A U G C C U U A G-3'. Now split the mRNA into codons and look each one up in the genetic code:

CodonMeaning
AUGStart (methionine)
CCUProline
UAGStop

So this small gene codes for a chain beginning with methionine and proline, at which point the stop codon UAG ends translation. Notice how a single base change in the DNA would change one codon, and therefore possibly one amino acid, which is exactly how many mutations work.

Key idea: A DNA template is transcribed to mRNA, split into codons, and read into amino acids, so the base sequence directly dictates the protein sequence.

The codon table is organized, not arbitrary

Sixty-four codons for twenty amino acids leaves plenty of room for redundancy, and that redundancy is arranged in a way that softens the effect of mutation. Two patterns stand out.

First, the third base usually matters least. For eight amino acids, any of the four bases in the third position gives the same amino acid, and for most of the rest the third position sorts only into two groups. Second, codons that differ in the middle base usually specify amino acids with different chemistry, while codons that differ in the first base often specify chemically similar ones. The code has been arranged so that the commonest kinds of copying error do the least damage.

Worked example. Take the leucine codon CUU and count every single-base change that could occur, three at each of the three positions.

  • Third position: CUC, CUA, and CUG all still specify leucine. Three silent changes.
  • First position: AUU gives isoleucine, GUU gives valine, UUU gives phenylalanine. All three are missense, but all three are hydrophobic amino acids like leucine, so these are conservative substitutions likely to be tolerated in a protein.
  • Second position: CCU gives proline, CAU gives histidine, CGU gives arginine. All three are chemically very different from leucine, so these are the changes most likely to break the protein.

Of nine possible single-base changes, three are silent, three are conservative, and only three are likely to matter much. Run the same exercise on a codon of your choice and the pattern repeats. This is one reason a genome tolerates a background of mutation without falling apart.

Key idea: Third-position changes are usually silent and first-position changes usually conservative, so of the nine single-base changes to a typical codon only about a third are likely to alter protein function appreciably.

What happens to an mRNA before it is read

In bacteria a ribosome can start translating a message while it is still being transcribed. In eukaryotes transcription happens in the nucleus and translation in the cytoplasm, and in between the transcript is extensively processed. Three modifications turn a raw transcript into a working message.

  • A cap is added to the 5 prime end almost as soon as it emerges. It protects the end from enzymes and is the mark that the ribosome recognizes when it loads.
  • A poly-A tail of roughly 200 adenines is added to the 3 prime end. Its length is one of the things that sets how long the message survives in the cytoplasm, so it is a control point as well as a protection.
  • Splicing removes the internal non-coding stretches, the introns, and joins the coding exons together. A large assembly of RNA and protein recognizes the sequences marking each intron boundary and cuts precisely there.

Splicing is not merely tidying up. Most human genes with several exons can be spliced in more than one way, so a single gene can specify several different proteins depending on which exons are kept. This is a large part of the answer to a puzzle that surprised everyone when the human genome was sequenced: roughly 20,000 protein-coding genes support a proteome several times larger.

It also explains a class of disease mutations that looks harmless at first glance. A base change at an intron boundary need not alter any codon, yet it can cause an exon to be skipped or an intron to be retained, wrecking the protein. A substantial fraction of disease-causing point mutations act through splicing rather than by changing an amino acid directly.

Key idea: Eukaryotic transcripts gain a 5 prime cap and a poly-A tail and have introns spliced out, and because most genes can be spliced several ways, 20,000 genes yield many more proteins and splice-site mutations cause disease without changing a codon.

Where people get stuck

The first sticking point is which DNA strand is used. The template strand is read by RNA polymerase; the other strand, called the coding or sense strand, has the same sequence as the mRNA except with T where the RNA has U. If you are given the coding strand, you can write the mRNA straight off by swapping T for U: many exam mistakes come from complementing it unnecessarily.

The second is losing the reading frame. Watch for this: codons are counted from the start codon, not from the beginning of the mRNA, and there is nothing in the sequence itself that marks where a codon ends. That is exactly why an insertion or deletion of one or two bases is so destructive.

The third is imagining that a stop codon is read by a tRNA. It is not: there is no tRNA for a stop codon. A release factor protein recognizes it instead and triggers the ribosome to let go, which is why a nonsense mutation truncates a protein completely rather than substituting something odd at that position.

Common misconceptions

  • RNA uses uracil, not thymine. Where the DNA template has A, the RNA gets U.
  • A codon specifies one amino acid, not one gene or one protein. A gene contains many codons.
  • Transcription and translation are different steps. Transcription makes RNA from DNA; translation makes protein from RNA.
  • The redundancy of the code means some different codons make the same amino acid, which is why certain DNA changes have no effect on the protein.

Recap

  • The central dogma is DNA to RNA to protein.
  • Transcription uses RNA polymerase to copy a gene into mRNA, with U replacing T.
  • Codons are three-base words; the redundant genetic code maps 64 codons onto 20 amino acids plus start and stop.
  • Translation on the ribosome uses tRNA adapters to build a protein codon by codon until a stop codon.
  • A DNA sequence can be traced through mRNA and codons to the amino acid sequence of a protein.

Sources

  1. Clark, M. A., Douglas, M., & Choi, J. (2018). Biology 2e (Ch. 15: Genes and proteins). OpenStax. openstax.org
  2. National Human Genome Research Institute. (n.d.). Codon. In Talking glossary of genomic and genetic terms. genome.gov
  3. Clancy, S., & Brown, W. (2008). Translation: DNA to mRNA to protein. Nature Education, 1(1), 101.
  4. Khan Academy. (n.d.). Central dogma: Transcription and translation. khanacademy.org
  5. Nirenberg, M. W., & Matthaei, J. H. (1961). The dependence of cell-free protein synthesis in E. coli upon naturally occurring or synthetic polyribonucleotides. Proceedings of the National Academy of Sciences, 47(10), 1588-1602. doi.org/10.1073/pnas.47.10.1588
  6. Crick, F. H. C., Barnett, L., Brenner, S., & Watts-Tobin, R. J. (1961). General nature of the genetic code for proteins. Nature, 192(4809), 1227-1232. doi.org/10.1038/1921227a0
  7. Alberts, B., Johnson, A., Lewis, J., Raff, M., Roberts, K., & Walter, P. (2002). From DNA to RNA. In Molecular biology of the cell (4th ed.). Garland Science. ncbi.nlm.nih.gov
Key terms
Central dogma
The flow of genetic information from DNA to RNA to protein.
Transcription
Copying a gene's DNA sequence into messenger RNA.
Messenger RNA (mRNA)
The RNA copy of a gene that carries information to the ribosome.
Codon
A three-base sequence in mRNA that specifies one amino acid or a stop signal.
Translation
Building a protein from an mRNA sequence at the ribosome.
Transfer RNA (tRNA)
An adapter molecule that carries an amino acid and pairs its anticodon with an mRNA codon.

Gene Regulation

  • Explain why cells must regulate which genes are expressed.
  • Describe the operon model of prokaryotic gene control.
  • Summarize the main ways eukaryotes regulate gene expression.

The big picture

Every cell in your body carries the same complete set of genes, yet a neuron, a muscle cell, and a skin cell look and act nothing alike. This lesson explains the resolution to that puzzle: cells do not use all their genes at once. They switch genes on and off, and turn them up and down, so each cell expresses only the ones it needs. That selective control, not the gene list, is what makes a cell what it is.

You will see the cleanest example first, a simple bacterial switch, and then the many layers eukaryotes add to build far more complex bodies. Regulation also turns out to be central to disease, especially cancer.

Why regulation is necessary

Gene regulation is the control of which genes are expressed in a given cell, and how much. Why bother? Two reasons. First, it is efficient: a cell wastes energy if it makes proteins it does not need, so it keeps unused genes off. Second, it enables specialization and response: one genome can build hundreds of cell types, and a cell can react to its environment moment to moment, all by changing which genes are active. Development itself is a carefully choreographed program of switching genes on and off in the right cells at the right times.

Key idea: Because every cell shares the same genome, gene regulation, deciding which genes are on and how strongly, is what makes cells different and lets them respond to conditions.

The operon: regulation in bacteria

Bacteria offer the clearest example of a genetic switch. In bacteria, related genes are often grouped into an operon: a cluster of genes transcribed together under the control of a single switch. The textbook case is the lac operon of the bacterium E. coli, which holds the genes for digesting the sugar lactose. Its switch has two key parts working together: an operator, the stretch of DNA that acts as the switch, and a repressor, a protein that can sit on the operator to block transcription.

The logic is elegant. When lactose is absent, the repressor binds the operator and blocks RNA polymerase, so the lactose-digesting genes stay off and no energy is wasted. When lactose is present, it binds to the repressor and changes its shape, pulling it off the operator; transcription then proceeds and the enzymes are made. In short, the cell builds the lactose-digesting machinery only when there is lactose to digest.

ConditionRepressorOperon
Lactose absentBound to operatorOFF (no enzymes made)
Lactose presentReleased from operatorON (enzymes made)

Key idea: In the lac operon a repressor blocks transcription when lactose is absent and is released when lactose is present, so the cell makes lactose enzymes only when needed.

The lac operon has two switches, not one

The repressor is only half the story, and the missing half explains an experimental observation that the repressor alone cannot. A bacterium offered both glucose and lactose consumes the glucose first and ignores the lactose entirely, even though lactose is present and the repressor should therefore have let go. Something else is holding the operon down.

That something is a second, positive control. When glucose runs low the cell accumulates a signaling molecule called cyclic AMP, which binds an activator protein. The activator-cAMP complex then binds just upstream of the promoter and helps recruit RNA polymerase, raising transcription sharply. When glucose is plentiful, cyclic AMP stays low, the activator does not bind, and the promoter works only feebly even with the repressor gone.

Combining the two controls gives four situations, and only one of them produces strong expression.

GlucoseLactoseRepressorActivatorTranscription
PresentAbsentBound (blocks)Not boundOff
PresentPresentReleasedNot boundVery low
AbsentAbsentBound (blocks)BoundOff
AbsentPresentReleasedBoundHigh

Read the table as a logical statement: the operon runs strongly only when lactose is available and the preferred fuel is not. This is the molecular explanation for the two-step growth curve observed when a culture is given both sugars, and it is a general design principle. Combining a negative control that senses the substrate with a positive control that senses overall nutritional state lets a single promoter integrate two independent pieces of information.

Key idea: The lac operon combines negative control by a lactose-sensing repressor with positive control by a cyclic-AMP-dependent activator that senses glucose scarcity, so strong transcription requires lactose present and glucose absent.

Repressible operons run the logic backwards

The lac operon is inducible: normally off, switched on by the substrate it processes. Operons for biosynthetic pathways face the opposite problem, since the cell should make an amino acid continuously and stop only when there is already enough of it. Such an operon is repressible: normally on, switched off by the end product.

The tryptophan operon is the standard example. Its repressor protein cannot bind DNA by itself. When tryptophan is abundant, tryptophan molecules bind the repressor and change its shape so that it can now sit on the operator and shut the operon down. The end product of the pathway therefore switches off its own manufacture, a clean case of feedback control operating at the level of transcription rather than enzyme activity.

Key idea: Inducible operons such as lac are off until a substrate removes the repressor, while repressible operons such as trp are on until the end product activates the repressor.

Regulation in eukaryotes

Eukaryotic cells regulate genes at many more levels than bacteria, which is part of how they achieve their greater complexity. The main control points, roughly in the order information flows, are:

  • Chromatin structure. DNA is wound around proteins, and when it is packed tightly it is inaccessible and silent; loosening the packing exposes genes so they can be transcribed. Chemical tags added to the DNA or its packaging proteins are part of epigenetics, and they can switch genes on or off without changing the DNA sequence.
  • Transcription factors. These regulatory proteins bind near a gene and either help or block RNA polymerase. This is the main on-off decision for most eukaryotic genes.
  • RNA processing and stability. A single gene's RNA can be spliced in alternative ways to make several different proteins, and how long an mRNA survives affects how much protein is produced from it.
  • Translation and beyond. Cells can control how efficiently each mRNA is translated, and can modify, activate, or destroy proteins after they are made.

Key idea: Eukaryotes control genes at many stages, from chromatin packing and transcription factors to RNA processing and protein modification, giving fine, flexible control.

Epigenetics and inheritance of expression

Epigenetics refers to heritable changes in gene expression that do not alter the DNA sequence itself. Epigenetic marks, such as chemical tags on DNA, can silence or activate genes, and they can be copied when a cell divides, so a liver cell's daughters stay liver cells. Some epigenetic patterns respond to environment and experience, which helps explain how identical twins with the same DNA can differ over time. The DNA sequence is the same; what changes is which genes are read.

Key idea: Epigenetic marks change which genes are expressed without changing the DNA sequence, and they are passed on when cells divide.

When regulation fails

Because regulation decides which genes act, its failure is central to many diseases. Cancer is the clearest case: genes that control cell division are normally switched on only when new cells are needed, but mutations or faulty regulation can leave them stuck on, driving the uncontrolled division that defines a tumor. So the correct control of gene expression is not a minor detail; it is essential to health, development, and the very identity of every cell.

Key idea: Faulty gene regulation underlies many diseases, notably cancer, where genes that drive cell division are switched on when they should be off.

Chromatin as the first layer of eukaryotic control

Before a eukaryotic transcription factor can do anything, the DNA it needs must be physically accessible, and most of the genome most of the time is not. Roughly two meters of DNA are wound around histone proteins into nucleosomes and folded further, so the primary regulatory question in a eukaryotic cell is which regions are open.

Two chemical systems set that state. Adding acetyl groups to histone tails neutralizes their positive charge, weakens their grip on the negatively charged DNA, and loosens the packing; removing those groups tightens it again. Adding methyl groups to specific histone positions marks a region as active or repressed depending on exactly which residue is modified, so the same chemical modification can mean opposite things at different addresses. Separately, methylating the cytosine in clusters of CG dinucleotides at a promoter recruits proteins that shut the gene down and keep it down through cell division.

Once a region is open, the elements that control it need not be adjacent to it. Eukaryotic enhancers can sit tens or hundreds of thousands of bases away, on either side of a gene, and act by looping the intervening DNA so that the proteins bound to the enhancer are brought physically against the promoter. This is why a mutation far outside any coding sequence can abolish a gene's expression in one tissue while leaving it intact in another.

Key idea: Histone acetylation opens chromatin, histone and DNA methylation mark regions as active or silent, and distant enhancers reach their promoters by looping, so eukaryotic control begins with physical accessibility rather than with transcription factors alone.

Imprinting: when the parent of origin decides

For most genes it makes no difference which parent supplied a copy. For a small set, perhaps one or two hundred in humans, it makes all the difference, because one parental copy is marked during gamete formation and silenced. These imprinted genes are expressed from only one parental copy, so the individual is functionally hemizygous even though two copies are present.

The consequence appears clearly in a single region of chromosome 15. Losing the paternally contributed copy of that region produces one clinical picture, Prader-Willi syndrome, while losing the maternally contributed copy of the overlapping region produces a quite different one, Angelman syndrome. The deleted DNA can be essentially the same; what differs is which parent it came from, and therefore which genes were already silenced on the remaining copy. Both conditions vary considerably between individuals and both are managed with supportive care rather than cured.

Imprinting matters conceptually because it is the cleanest demonstration that a heritable epigenetic mark can be functionally decisive. The DNA sequence is unchanged; only its chemical annotation differs, and that annotation is erased and rewritten each generation as gametes are formed.

Key idea: Imprinted genes are silenced on one parental copy, so deletions in the same chromosome 15 region cause Prader-Willi or Angelman syndrome depending on which parent contributed the missing copy.

Where people get stuck

The first sticking point is treating regulation as a simple on-off switch. It rarely is: expression is quantitative, and the biologically important differences between cells are usually differences of degree rather than presence and absence, so the right question is normally how much rather than whether.

The second is confusing induction with mutation. They are not the same: an induced gene has not changed. A signal has merely allowed a pre-existing gene to be transcribed. Every liver cell and every neuron in a person carries the same genome, and their differences are differences of expression.

The third is overstating what epigenetics inherits. Marks are reliably passed to daughter cells during mitosis, which is well established. Transmission of an environmentally acquired mark through the germ line to a person's grandchildren is a much stronger claim, and in humans the evidence for it remains limited and contested.

Common misconceptions

  • Different cell types have the same genes, not different ones. They differ in which genes are expressed, not in the DNA they carry.
  • In the lac operon, lactose does not directly switch the genes on; it removes the repressor, which allows transcription.
  • Epigenetic changes do not alter the DNA sequence. They change how genes are read, and can be reversible.
  • Regulation is not only on or off. Cells also fine-tune how much of each protein is made.

Recap

  • Gene regulation controls which genes are expressed and how much, making cells different despite a shared genome.
  • In bacteria, an operon groups genes under one switch; the lac operon's repressor blocks transcription unless lactose is present.
  • Eukaryotes regulate at many levels: chromatin, transcription factors, RNA processing, and protein modification.
  • Epigenetic marks change gene expression without changing the DNA sequence and are heritable through cell division.
  • Failed regulation contributes to disease, especially cancer.

Sources

  1. Clark, M. A., Douglas, M., & Choi, J. (2018). Biology 2e (Ch. 16: Gene expression). OpenStax. openstax.org
  2. National Human Genome Research Institute. (n.d.). Epigenetics. In Talking glossary of genomic and genetic terms. genome.gov
  3. Ralston, A., & Shaw, K. (2008). Gene expression regulates cell differentiation. Nature Education, 1(1), 127.
  4. Khan Academy. (n.d.). Gene regulation (including the lac operon). khanacademy.org
  5. Jacob, F., & Monod, J. (1961). Genetic regulatory mechanisms in the synthesis of proteins. Journal of Molecular Biology, 3(3), 318-356. doi.org/10.1016/S0022-2836(61)80072-7
  6. Bird, A. (2007). Perceptions of epigenetics. Nature, 447(7143), 396-398. doi.org/10.1038/nature05913
  7. National Human Genome Research Institute. (2020). Epigenomics fact sheet. National Institutes of Health. genome.gov
Key terms
Gene regulation
The control of which genes are expressed, and how much, in a given cell.
Operon
A cluster of related genes transcribed together under one control region in prokaryotes.
Repressor
A protein that blocks transcription by binding to the operator.
Operator
The DNA control region where a repressor binds to switch an operon off.
Transcription factor
A protein that binds DNA to promote or block transcription of a gene.
Epigenetics
Heritable changes in gene expression that do not alter the DNA sequence itself.

Module 5: Mutation and DNA Repair

How the DNA sequence changes, the different kinds of mutations and their effects, and the systems that detect and fix damage.

Types of Mutations and Their Effects

  • Distinguish point mutations from frameshift mutations.
  • Classify substitutions as silent, missense, or nonsense.
  • Explain why mutations are both harmful and essential.

The big picture

A mutation is simply a change in the DNA sequence. But its consequences vary wildly, from nothing at all to a fatal disease. This lesson shows how to predict a mutation's effect by tracing it through the genetic code you learned in the last module. The key is to ask what happens to the codons downstream of the change.

You will also see the double nature of mutation: it is the source of the genetic variation that evolution depends on, and at the same time a cause of disease and cancer. Both sides come from the same molecular events.

What a mutation is

A mutation is any change in the DNA sequence of an organism. Mutations are the ultimate source of all genetic variation, and therefore of evolution, because they create the new alleles that natural selection can act on. Yet many mutations are harmful, and some cause serious disease. Classifying mutations by what they do to the DNA helps predict their effects on the resulting protein.

Point mutations and substitutions

A point mutation changes a single base in the DNA. The simplest kind is a substitution, in which one base is swapped for another (for example, an A replaced by a G). Because the genetic code is read in three-base codons, a single substitution changes just one codon, and that change has one of three possible outcomes depending on what the new codon means:

  • A silent mutation changes a codon but, thanks to the redundancy of the code, the new codon still specifies the same amino acid. The protein is unchanged. Example: GAA and GAG both code for glutamate, so a GAA to GAG change is silent.
  • A missense mutation changes the codon to one that specifies a different amino acid, altering the protein. Sickle-cell disease results from a single missense change that swaps one amino acid in a blood protein.
  • A nonsense mutation changes an amino-acid codon into a stop codon, cutting the protein short. The shortened protein is usually nonfunctional. Example: UAC (tyrosine) changing to UAA (stop).

Key idea: A substitution changes one codon, giving a silent (same amino acid), missense (different amino acid), or nonsense (premature stop) result.

Insertions, deletions, and the reading frame

Adding or removing bases can be far more disruptive than a substitution. The ribosome reads mRNA in non-overlapping groups of three, and the reading frame is the way the sequence is divided into those consecutive triplets. Inserting or deleting a number of bases that is not a multiple of three shifts the reading frame, so that every codon after the change is regrouped and misread. This is a frameshift mutation, and it usually produces a completely wrong, nonfunctional protein, often ending early at a newly created stop codon.

An English-sentence analogy makes the damage vivid. Read in three-letter words, a deletion garbles everything after the deleted letter:

SequenceRead in triplets
THE CAT ATE THE RAToriginal: every word makes sense
THE ATA TET HER ATafter deleting the first C: every group is scrambled

By contrast, inserting or deleting exactly three bases (or a multiple of three) adds or removes whole codons without shifting the frame, so the damage is usually limited to that spot rather than everything downstream.

Key idea: An insertion or deletion not divisible by three shifts the reading frame and misreads every downstream codon, which is why frameshifts are so damaging.

Where mutations come from

Mutations arise in two main ways. Some are spontaneous copying errors during DNA replication that escape proofreading. Others are caused by mutagens: agents such as ultraviolet light, ionizing radiation, and certain chemicals that damage DNA or cause mispairing. Mutations in body (somatic) cells are not passed to offspring but can cause cancer; mutations in the cells that make gametes can be inherited by the next generation.

Key idea: Mutations come from spontaneous replication errors and from mutagens like UV light, radiation, and chemicals.

Why mutations matter both ways

Mutation has two faces. Most individual mutations are neutral or harmful, and mutations in genes that control cell division are a root cause of cancer. Yet without mutation there would be no new alleles, no raw material for natural selection, and therefore no adaptation and no evolution. The same process that occasionally causes disease is also the wellspring of all biological diversity. Mutation is not simply an error to be eliminated; it is the source of the variation that life depends on.

Key idea: Mutation is both a cause of disease and the ultimate source of the genetic variation that makes evolution possible.

Working a substitution through to the protein

The categories only become concrete when the codons are written out. Take a short stretch of a coding strand reading TTC GAG CTG, which corresponds to the mRNA UUC GAG CUG and the amino acids phenylalanine, glutamate, leucine. Now change the middle codon three different ways.

  • GAG becomes GAA. Both specify glutamate, so the protein is unchanged. This is a silent substitution, and it happens most often at the third position for the reasons set out in the genetic code lesson.
  • GAG becomes GTG (mRNA GUG). Glutamate, which carries a negative charge, is replaced by valine, which is hydrophobic. This is a missense substitution, and it is a chemically drastic one.
  • GAG becomes TAG (mRNA UAG). That is a stop codon, so the ribosome releases the chain here and everything downstream is lost. This is a nonsense substitution.

The second of those is not a hypothetical. It is the change at codon 6 of the beta-globin gene that produces sickle hemoglobin. Replacing a charged surface residue with a hydrophobic one creates a sticky patch, and when the molecule gives up its oxygen those patches let hemoglobin molecules polymerize into long fibers that deform the red cell. Every clinical feature of sickle cell disease follows from that one substitution, which is also a compact demonstration of pleiotropy.

Key idea: Writing the codons out shows the three outcomes directly, and the GAG to GTG change replacing glutamate with valine at codon 6 of beta-globin is the single substitution underlying sickle cell disease.

Missense is a category, not a verdict

Describing a variant as missense says only that an amino acid has changed; it says almost nothing about consequence. Two considerations dominate.

The first is chemistry. Substituting one hydrophobic residue for another of similar size is usually tolerated, whereas exchanging a charged residue for a hydrophobic one, or introducing a proline into a helix, frequently is not. The second is position. A substitution in an enzyme's active site, or in the buried core that holds the fold together, is far more likely to matter than the same substitution on a flexible surface loop.

This is why laboratories reporting a genetic test often classify a finding as a variant of uncertain significance. The change is real and the sequencing is reliable, but the evidence linking that particular change to disease is insufficient. Large population databases have made this judgement more tractable, because a variant seen frequently in healthy people is unlikely to cause a severe early-onset condition. The honest reporting of uncertainty is a feature of good practice rather than a failure of the technology.

Key idea: The effect of a missense change depends on the chemical similarity of the substituted residues and on the position within the protein, which is why many variants are reported as being of uncertain significance.

Repeat expansions and anticipation

A distinct mutational mechanism involves short sequences repeated in tandem, which the replication machinery copies unreliably so that the number of copies can change between generations. Huntington's disease is the standard illustration. The relevant gene contains a run of CAG repeats, and the number of copies determines the outcome:

CAG repeatsConsequence
26 or fewerNot associated with the condition
27 to 35Intermediate; the person is not affected but the repeat may expand in offspring
36 to 39Reduced penetrance; some people develop the condition and some do not
40 or moreFull penetrance

Longer repeats are associated with earlier onset, and repeats tend to lengthen when transmitted, particularly through the father. A condition can therefore appear earlier and more severely in successive generations of a family, a pattern called anticipation. Fragile X syndrome involves an analogous CGG expansion in a different gene, with a premutation range that is itself associated with distinct health effects.

Two points deserve care here. The banding above is a summary of population data, not a prediction for an individual; onset age varies widely at any given repeat length. And because Huntington's disease is adult-onset and currently has no disease-modifying cure, predictive testing of people without symptoms is offered alongside genetic counseling, and many people at risk choose not to be tested. The genetics is only one part of the decision.

Key idea: Trinucleotide repeats can expand between generations, so repeat number predicts risk in bands rather than absolutely and produces anticipation, with earlier onset in successive generations.

Somatic and germline mutations are not the same problem

A mutation matters differently depending on which cells carry it. A germline mutation is present in the gametes and therefore in every cell of any resulting child, so it is heritable. A somatic mutation arises in a body cell during life, is passed only to that cell's descendants, and cannot be transmitted to offspring.

Somatic mutations accumulate steadily in every tissue with age, and while the great majority are harmless, a cell that happens to collect several in genes controlling division can escape normal restraint. Cancer is therefore best understood as a somatic genetic disease, which explains both why incidence rises steeply with age and why the same cancer type in two people can be driven by different mutations. An error occurring early in embryonic development lies between the two categories, producing mosaicism, in which one person carries two genetically distinct cell populations.

Key idea: Germline mutations are heritable and present in every cell, somatic mutations accumulate with age and drive cancer without being transmitted, and early embryonic errors produce mosaicism.

Where people get stuck

The first sticking point is the phrase a gene for a disease. Genes encode products, not outcomes: what is inherited is a variant that alters a product in a way that raises the chance of a condition under particular circumstances. Even for strongly determining variants, penetrance, severity, and age of onset vary, and for most common conditions the contribution of any single variant is small.

The second is assuming that a mutation must be harmful. It need not be: most substitutions are silent or inconsequential, some are beneficial, and the same allele can be advantageous in one environment and costly in another, as the sickle allele is with respect to malaria.

The third is equating mutation rate with risk. They are not the same: what determines whether a mutation appears in a lineage is the rate multiplied by the number of cell divisions and by how effectively repair operates, which is why exposure, age, and repair capacity all enter the picture alongside the intrinsic error rate.

Common misconceptions

  • Not all mutations change the protein. Silent mutations leave the amino acid sequence unchanged because the code is redundant.
  • A frameshift is usually far more damaging than a single substitution, because it garbles every codon downstream, not just one.
  • Mutations are not always bad. Many are neutral, and some are beneficial and drive adaptation.
  • Somatic mutations are not inherited. Only mutations in the cells that give rise to gametes can be passed to offspring.

Recap

  • A mutation is any change in DNA and is the ultimate source of genetic variation.
  • A substitution changes one codon, giving a silent, missense, or nonsense result.
  • An insertion or deletion not divisible by three causes a frameshift that misreads all downstream codons.
  • Mutations arise from replication errors and from mutagens such as UV light, radiation, and chemicals.
  • Mutation causes disease and cancer but also supplies the variation that evolution requires.

Sources

  1. Clark, M. A., Douglas, M., & Choi, J. (2018). Biology 2e (Ch. 14.6: DNA repair). OpenStax. openstax.org
  2. National Human Genome Research Institute. (n.d.). Mutation. In Talking glossary of genomic and genetic terms. genome.gov
  3. Clancy, S. (2008). Genetic mutation. Nature Education, 1(1), 187.
  4. Khan Academy. (n.d.). Types of mutations and their effects (DNA mutations). khanacademy.org
  5. Pauling, L., Itano, H. A., Singer, S. J., & Wells, I. C. (1949). Sickle cell anemia, a molecular disease. Science, 110(2865), 543-548. doi.org/10.1126/science.110.2865.543
  6. Ingram, V. M. (1957). Gene mutations in human haemoglobin: The chemical difference between normal and sickle cell haemoglobin. Nature, 180(4581), 326-328. doi.org/10.1038/180326a0
  7. MedlinePlus Genetics. (2021). What is a gene variant and how do variants occur? National Library of Medicine. medlineplus.gov
Key terms
Mutation
Any change in the DNA sequence of an organism.
Point mutation
A change affecting a single base, such as a substitution.
Silent mutation
A base change that still codes for the same amino acid, leaving the protein unchanged.
Missense mutation
A base change that swaps one amino acid for another in the protein.
Nonsense mutation
A base change that creates a premature stop codon, truncating the protein.
Frameshift mutation
An insertion or deletion that shifts the reading frame, garbling all downstream codons.

DNA Damage and Repair

  • Describe common sources of DNA damage.
  • Outline major DNA repair mechanisms.
  • Connect failed repair to cancer and inherited disorders.

The big picture

You might expect that the low mutation rate of living cells means DNA is rarely damaged. The truth is the opposite: DNA is under constant assault, and cells stay stable only because they run an elaborate set of repair systems that catch and fix damage almost as fast as it happens. This lesson surveys those repair systems and shows what goes wrong when they fail.

The payoff is a deeper understanding of both accuracy and disease. Repair is why replication is so faithful, and the breakdown of repair is a major route to cancer, which ties this lesson directly to the mutations you just studied.

DNA is constantly damaged

DNA is damaged by ultraviolet light from the sun, by ionizing radiation, by reactive chemicals in the environment and inside the cell, and by ordinary errors during replication. If this damage went uncorrected, it would accumulate as mutations and quickly become catastrophic for the cell and the organism. Cells survive because they invest heavily in DNA repair: a collection of systems that detect and correct damage or errors in DNA. The remarkably low mutation rate we observe is not because damage is rare, but because repair is so effective.

Key idea: DNA is damaged constantly, and the low mutation rate of cells reflects highly effective repair rather than a lack of damage.

Repair during and after replication

The first line of defense operates as DNA is being copied. It is proofreading by DNA polymerase: as the enzyme adds each new base, it checks the pairing and removes a mismatched base on the spot before continuing. Errors that slip past proofreading are caught by a second system, mismatch repair, which scans newly made DNA for mispaired bases, cuts out the incorrect stretch, and fills the gap correctly using the original strand as a guide. Together, proofreading and mismatch repair bring the replication error rate down to roughly one mistake per billion bases.

Key idea: Proofreading by DNA polymerase corrects errors during copying, and mismatch repair fixes those that slip through, giving very high replication accuracy.

Repairing damage from the environment

Damage caused by mutagens is handled by other systems. In excision repair, enzymes recognize a damaged or distorted section of DNA, cut out the bad stretch, and resynthesize it using the intact complementary strand as a template. This is how cells fix the damage caused by ultraviolet light, which fuses two adjacent bases together and kinks the helix. Because one strand is still intact and complementary, the correct sequence can always be rebuilt.

A harder problem is when both strands break at the same place, a double-strand break, since then there is no intact template on either side. Cells use specialized double-strand break repair pathways to rejoin the broken ends, though these are more error-prone than the copy-from-the-good-strand approach. The general principle holds: repair works best when at least one good strand remains to copy from.

Key idea: Excision repair cuts out environmental damage and rebuilds the DNA from the intact complementary strand, which is why an undamaged strand is so valuable.

When repair fails: disease and cancer

The importance of repair is clearest when it breaks down. People with the inherited disorder xeroderma pigmentosum cannot carry out excision repair of ultraviolet damage. As a result, their skin is extremely sensitive to sunlight and they develop skin cancers at a very high rate and young age, because unrepaired UV damage turns into mutations. More broadly, mutations in DNA repair genes are common in cancer, because a cell that cannot fix its DNA accumulates the additional mutations that drive a tumor.

This is why some repair genes are called guardians of the genome: they prevent the buildup of mutations that would otherwise unleash uncontrolled cell division. The overarching lesson is that maintaining the integrity of DNA is every bit as vital as copying it and reading it. A genome that cannot be protected is a genome that cannot be trusted.

Key idea: Failed DNA repair leads to disease, as in xeroderma pigmentosum, and drives cancer by letting mutations accumulate, so repair genes act as guardians of the genome.

The daily damage budget

The scale of spontaneous damage is easy to underestimate, and the published estimates make the point better than any adjective. In a single human cell on a single ordinary day, chemistry alone produces damage on the following order:

  • Depurination. The bond holding a purine to the sugar backbone hydrolyzes spontaneously, and roughly 10,000 adenines and guanines fall off per cell per day, leaving gaps with no base at all.
  • Deamination. Cytosine loses an amino group and becomes uracil, which would be read as thymine, at a rate of a few hundred events per cell per day. This is one reason cells maintain an enzyme whose sole job is recognizing uracil in DNA.
  • Oxidation. Reactive by-products of ordinary metabolism attack bases thousands of times per cell per day, with oxidized guanine the commonest product.
  • Ultraviolet light. A skin cell in bright sunlight can accumulate on the order of 100,000 covalent links between adjacent pyrimidines in an hour, and each of those distorts the helix enough to block replication.

Set that alongside the final mutation rate of roughly one error per 109 bases per division and the conclusion is unavoidable. Genomes are not stable because they are chemically inert; they are stable because damage is detected and reversed continuously, and repair is running at all times in every cell.

Key idea: Spontaneous chemistry inflicts tens of thousands of lesions per cell per day, so genome stability is an active achievement of repair rather than a property of the molecule.

Matching the pathway to the lesion

There is no general-purpose repair enzyme, because different lesions present different problems. Four systems handle most of the load, and each is defined by what it recognizes.

  • Base excision repair handles small chemical alterations such as a deaminated or oxidized base. A specialized enzyme recognizes the specific damaged base, flips it out of the helix, and cuts it off; the gap is then filled from the intact opposite strand.
  • Nucleotide excision repair handles bulky lesions that physically distort the helix, ultraviolet dimers foremost among them. Rather than recognizing a particular chemical group, it detects the distortion itself, excises a stretch of about two dozen nucleotides containing the damage, and resynthesizes it. A dedicated version of this pathway is coupled to transcription, so a lesion blocking an RNA polymerase is prioritized.
  • Mismatch repair corrects errors that escaped proofreading. Its interesting problem is deciding which of the two strands is wrong, since both bases are chemically normal; the system solves this by identifying the newly synthesized strand and correcting that one.
  • Double-strand break repair faces the most dangerous lesion, because both strands are cut and neither can serve as a template. Two routes exist. Non-homologous end joining simply trims and ligates the ends, which is fast, available at any point in the cell cycle, and frequently loses a few bases. Homologous recombination copies the missing information from the sister chromatid and is essentially error-free, but it is only possible after replication, when a sister chromatid exists.

Key idea: Base excision handles altered bases, nucleotide excision handles helix-distorting lesions, mismatch repair corrects replication errors on the new strand, and double-strand breaks are fixed either quickly and sloppily by end joining or accurately by recombination when a sister chromatid is available.

Inherited repair defects, and how medicine exploits them

Each pathway has a corresponding inherited condition, and the pattern of cancer risk in each one points directly back to the lesion that pathway handles.

People with xeroderma pigmentosum carry variants that impair nucleotide excision repair. They cannot remove ultraviolet damage efficiently, and their risk of skin cancer on sun-exposed skin is raised roughly a thousandfold, with tumors often appearing in childhood. Rigorous sun protection is the mainstay of management, and the specificity of the risk, confined largely to sunlight-exposed tissue, is itself evidence for what the pathway does.

Inherited defects in mismatch repair underlie Lynch syndrome, which raises the lifetime chance of colorectal and several other cancers. Tumors arising this way accumulate errors in short repeated sequences, a signature called microsatellite instability, and because such tumors carry an unusually large number of mutations they present many abnormal proteins to the immune system. That observation turned into treatment: mismatch-repair-deficient tumors respond notably well to immune checkpoint therapy, and this became one of the first approvals granted on the basis of a tumor's molecular feature rather than its organ of origin.

Variants in BRCA1 and BRCA2 impair homologous recombination and raise the chance of breast, ovarian, and several other cancers. Here the therapeutic logic is inverted. A cell that has lost homologous recombination depends heavily on a backup repair enzyme, so drugs inhibiting that backup kill the tumor cells while sparing normal cells that still have both pathways. This principle, called synthetic lethality, converted a repair defect from a purely bad prognosis into a specific vulnerability.

Key idea: Xeroderma pigmentosum, Lynch syndrome, and BRCA-related cancers each trace to one failed repair pathway, and the resulting dependence on remaining pathways is now the basis of checkpoint and synthetic-lethal therapies.

Where people get stuck

The first sticking point is treating damage and mutation as synonyms. They are not: damage is a chemical alteration that repair can usually reverse, while a mutation is a change in sequence that has become permanent because replication has copied it before repair reached it. Most damage never becomes mutation.

The second is assuming that all repair improves accuracy. It does not always: non-homologous end joining routinely loses or adds a few bases at the join, and cells accept that cost because an unrepaired double-strand break is worse than a small local error.

The third is expecting an inherited repair defect to cause cancer directly. It does not. It raises the rate at which mutations accumulate, so cancer becomes more likely and tends to occur earlier, which is a change in probability rather than a certainty.

Common misconceptions

  • Low mutation rates come from active repair, not from DNA being rarely damaged. Damage is frequent; correction is efficient.
  • Proofreading and mismatch repair are distinct. Proofreading acts during copying by the polymerase; mismatch repair scans the finished new strand afterward.
  • Excision repair needs an intact complementary strand to copy from, which is why double-strand breaks are harder to fix accurately.
  • A repair gene does not itself drive cell division. Its failure causes cancer indirectly, by allowing other mutations to pile up.

Recap

  • DNA is constantly damaged; the low mutation rate reflects effective repair.
  • Proofreading by DNA polymerase corrects errors during replication, and mismatch repair fixes those that escape.
  • Excision repair removes environmental damage such as UV lesions, rebuilding from the intact strand.
  • Double-strand breaks are harder to repair because no intact template remains nearby.
  • Failed repair causes disorders like xeroderma pigmentosum and contributes to cancer, so repair genes guard the genome.

Sources

  1. Clark, M. A., Douglas, M., & Choi, J. (2018). Biology 2e (Ch. 14.6: DNA repair). OpenStax. openstax.org
  2. National Human Genome Research Institute. (n.d.). Talking glossary of genomic and genetic terms (DNA repair, mutation). genome.gov
  3. Clancy, S. (2008). DNA damage and repair: Mechanisms for maintaining DNA integrity. Nature Education, 1(1), 103.
  4. Khan Academy. (n.d.). DNA proofreading and repair (DNA replication). khanacademy.org
  5. Lindahl, T. (1993). Instability and decay of the primary structure of DNA. Nature, 362(6422), 709-715. doi.org/10.1038/362709a0
  6. Sancar, A., Lindsey-Boltz, L. A., Ünsal-Kaçmaz, K., & Linn, S. (2004). Molecular mechanisms of mammalian DNA repair and the DNA damage checkpoints. Annual Review of Biochemistry, 73(1), 39-85. doi.org/10.1146/annurev.biochem.73.011303.073723
  7. Brown, T. A. (2002). Mutation, repair and recombination. In Genomes (2nd ed.). Wiley-Liss. ncbi.nlm.nih.gov
Key terms
DNA repair
Cellular systems that detect and correct damage or errors in DNA.
Proofreading
DNA polymerase's checking and correction of newly added bases during replication.
Mismatch repair
A system that finds and corrects mispaired bases in newly synthesized DNA.
Excision repair
Cutting out a damaged DNA section and resynthesizing it from the intact strand.
Mutagen
An agent such as UV light, radiation, or a chemical that causes DNA damage.
Xeroderma pigmentosum
An inherited disorder of failed UV-damage repair that causes extreme sun sensitivity and skin cancer.

Module 6: Population and Quantitative Genetics

How allele frequencies behave in whole populations, the Hardy-Weinberg equilibrium as a null model, the forces that change it, and the genetics of continuous traits.

The Hardy-Weinberg Principle

  • State the Hardy-Weinberg equations and their assumptions.
  • Calculate allele and genotype frequencies from population data.
  • Use Hardy-Weinberg as a null model for detecting evolution.

The big picture

So far you have followed genes through families, one cross at a time. This lesson zooms all the way out to whole populations and asks a different question: across thousands of individuals, what fraction carry each allele, and does that fraction stay steady or drift over the generations? Change in allele frequency over time is the genetic definition of evolution, so this is where genetics and evolution meet.

The centerpiece is a pair of simple equations that predict the genotype frequencies of a population that is not evolving. You will work a full numerical example with checked arithmetic, and you will see the surprising conclusion that most copies of a rare recessive allele are hidden in healthy carriers.

Thinking about whole populations

Population genetics is the study of allele and genotype frequencies in populations and how they change over time. Its central quantity is the allele frequency: the proportion of a particular allele among all copies of that gene in the population. If you imagine pooling every allele of a gene from every individual into one giant gene pool, the allele frequency is just the fraction of that pool made up of each version. The foundation of the whole field is the Hardy-Weinberg principle, a mathematical model that predicts the genotype frequencies expected in a population that is not evolving.

Key idea: Population genetics tracks allele frequencies in a gene pool, and the Hardy-Weinberg principle predicts the genotypes of a non-evolving population.

The two equations

Consider a gene with two alleles. Let p stand for the frequency of the dominant allele and q for the frequency of the recessive allele. Since these are the only two alleles, every copy in the pool is one or the other, so their frequencies must add up to 1:

p + q = 1

If mating is random, the chance of forming each genotype is found by combining allele frequencies the same way you would multiply probabilities, which is the expansion of (p + q) squared:

p2 + 2pq + q2 = 1

Each term is the expected frequency of one genotype: p2 is the frequency of homozygous dominant individuals, 2pq is the frequency of heterozygotes (there are two ways to get one of each allele, hence the 2), and q2 is the frequency of homozygous recessive individuals. A population whose genotype frequencies match these predictions is said to be in Hardy-Weinberg equilibrium.

Key idea: With two alleles, p + q = 1 for the alleles, and p2 + 2pq + q2 = 1 gives the frequencies of the homozygous dominant, heterozygous, and homozygous recessive genotypes.

A worked example, step by step

Suppose a recessive condition affects 1 in 400 people in a population. Only homozygous recessive individuals show the condition, so we start from q2 and work outward. Follow each step, and note that every number is checked at the end.

  1. The affected individuals are homozygous recessive, so q2 = 1/400 = 0.0025.
  2. Take the square root: q = the square root of 0.0025 = 0.05. The recessive allele frequency is 0.05.
  3. Use p + q = 1: p = 1 minus 0.05 = 0.95. The dominant allele frequency is 0.95.
  4. Homozygous dominant frequency: p2 = 0.95 x 0.95 = 0.9025, about 90 percent.
  5. Heterozygous carrier frequency: 2pq = 2 x 0.95 x 0.05 = 0.095, about 9.5 percent.

Now verify that the three genotype frequencies sum to 1, as they must: 0.9025 + 0.095 + 0.0025 = 1.0000. The arithmetic checks out exactly.

The striking result is this: although only 1 in 400 people (0.25 percent) show the condition, about 9.5 percent, or roughly 1 in 10 people, carry the recessive allele without knowing it. Carriers vastly outnumber affected individuals. This is exactly why recessive alleles persist in populations even when the recessive phenotype is rare; most copies of the allele are hidden safely in heterozygotes.

Key idea: From the frequency of a recessive phenotype (q2) you can compute q, then p, then all genotype frequencies, and hidden carriers usually greatly outnumber affected individuals.

Why the null model matters

Hardy-Weinberg equilibrium holds only under five idealized assumptions: no mutation, no migration (no gene flow), no natural selection, random mating, and a very large population (so no genetic drift). No real population meets all five perfectly. Sounds like a weakness, does it not? It is actually the whole point. The model's value is as a null hypothesis: a baseline expectation of what a non-evolving population would look like.

You compare a real population's measured genotype frequencies against the Hardy-Weinberg prediction. If they match, no evolutionary force is detectably acting on that gene. If they differ significantly, then one or more of the five assumptions is being violated, which tells you the population is evolving and points toward the responsible force. The equation is therefore a detector of evolution, not merely a description of stillness.

Key idea: Because no real population meets all five assumptions, Hardy-Weinberg serves as a null model: a mismatch between observed and predicted frequencies reveals that a population is evolving.

Running the calculation the other way

The worked example above started from a phenotype frequency and derived the allele frequencies. When every genotype is distinguishable, as it is for codominant markers, you can count genotypes directly and derive allele frequencies without assuming anything at all. Doing so is what makes the principle testable.

Worked example. A sample of 1,000 people is typed for a codominant blood group with alleles M and N, giving 360 MM, 480 MN, and 160 NN.

  • Count alleles. Each person carries two, so there are 2,000 alleles. The M count is 2 x 360 + 480 = 1,200, so p = 1,200 / 2,000 = 0.6, and q = 1 - 0.6 = 0.4.
  • Predict genotype counts from those frequencies. p2 = 0.36, giving 360 MM; 2pq = 2 x 0.6 x 0.4 = 0.48, giving 480 MN; q2 = 0.16, giving 160 NN.
  • The predicted numbers match the observed ones exactly, so this sample is in Hardy-Weinberg equilibrium at this locus.

Now a sample that is not. Suppose instead the counts were 400 MM, 400 MN, and 200 NN. The allele frequencies come out the same: p = (800 + 400) / 2,000 = 0.6 and q = 0.4. But the expected genotype counts are still 360, 480, and 160, so the observed and expected numbers now differ. Test the gap with chi-square:

  • MM: (400 - 360)2 / 360 = 1,600 / 360 = 4.44
  • MN: (400 - 480)2 / 480 = 6,400 / 480 = 13.33
  • NN: (200 - 160)2 / 160 = 1,600 / 160 = 10.00
  • chi-square = 4.44 + 13.33 + 10.00 = 27.8

Degrees of freedom here are 1, not 2, because one allele frequency was estimated from the same data. The critical value at the 5 percent level with 1 degree of freedom is 3.841, so 27.8 rejects the null decisively. The direction of the deviation is informative: heterozygotes are scarce and both homozygote classes are inflated, which is the signature of inbreeding or of pooling two populations with different allele frequencies into one sample.

Key idea: Counting alleles from genotype counts gives p and q without assumptions, and comparing observed with expected genotype numbers by chi-square, with one degree of freedom for a two-allele locus, tests whether the population is at equilibrium.

Hardy-Weinberg on the X chromosome

The standard equations assume every individual carries two copies, which fails for X-linked genes in males. Because a male has a single X, his phenotype is determined by his one allele, so the frequency of affected males is simply q. Females carry two copies, so the frequency of affected females is q2.

Worked example. The common form of red-green color vision deficiency has an allele frequency of roughly q = 0.08 in populations of European ancestry.

  • Affected males: q = 0.08, or 8 percent.
  • Affected females: q2 = 0.0064, or 0.64 percent.
  • Carrier females: 2pq = 2 x 0.92 x 0.08 = 0.147, so about 15 percent of women carry one copy without being affected.

The ratio of affected males to affected females is q / q2 = 1 / q, which here is 12.5 to 1. That single expression captures the whole reason X-linked recessive conditions are commoner in males, and it also predicts that the rarer the allele, the more lopsided the ratio becomes. For an allele at frequency 0.001 the ratio would be 1,000 to 1, which is why some X-linked conditions are almost never seen in females.

Key idea: For X-linked genes the frequency of affected males is q and of affected females q2, so the male-to-female ratio is 1/q and grows as the allele becomes rarer.

Where people get stuck

The first sticking point is squaring the wrong quantity. The frequency of the recessive phenotype is q2, so recovering the allele frequency requires taking the square root: taking the square root of the affected count instead of the affected frequency is the commonest arithmetic error in this topic.

The second is expecting heterozygotes to be rare when the recessive allele is rare. The opposite holds: when q is small, 2pq is far larger than q2, since the ratio of carriers to affected individuals is 2p/q, which grows as q falls. For q = 0.01 there are roughly 198 carriers for every affected person, which is why recessive alleles persist in populations even under strong selection.

The third is reading equilibrium as evidence that nothing is happening. It is not: a population can sit at Hardy-Weinberg proportions at one locus while evolving rapidly at others, and a single generation of random mating restores the proportions even after a disturbance, so the test is a snapshot of one gene rather than a verdict on a population.

Common misconceptions

  • q2, not q, equals the frequency of the recessive phenotype. To get the allele frequency q you must take the square root.
  • Dominant alleles do not automatically become more common. Allele frequencies stay constant under Hardy-Weinberg regardless of dominance.
  • The 2 in 2pq is not optional. Heterozygotes can form in two ways (allele A from either parent), so their frequency is doubled.
  • Real populations are rarely in perfect equilibrium. The model is a comparison baseline, not a claim that populations never change.

Recap

  • Population genetics studies allele frequencies; evolution is change in those frequencies over time.
  • For two alleles, p + q = 1 and p2 + 2pq + q2 = 1 give allele and genotype frequencies.
  • From a recessive phenotype frequency q2, take the square root for q, subtract for p, and compute p2 and 2pq; the three genotypes sum to 1.
  • Carriers (2pq) usually far outnumber affected individuals (q2), so recessive alleles persist even when rare.
  • Hardy-Weinberg is a null model; deviation from its prediction signals that a population is evolving.

Sources

  1. Clark, M. A., Douglas, M., & Choi, J. (2018). Biology 2e (Ch. 19.2: Population genetics). OpenStax. openstax.org
  2. National Human Genome Research Institute. (n.d.). Talking glossary of genomic and genetic terms (allele frequency, population genetics). genome.gov
  3. Edwards, A. W. F. (2008). G. H. Hardy (1908) and Hardy-Weinberg equilibrium. Nature Education, 1(1).
  4. Khan Academy. (n.d.). The Hardy-Weinberg equation and its assumptions. khanacademy.org
  5. Hardy, G. H. (1908). Mendelian proportions in a mixed population. Science, 28(706), 49-50. doi.org/10.1126/science.28.706.49
  6. Crow, J. F. (1999). Hardy, Weinberg and language impediments. Genetics, 152(3), 821-825. doi.org/10.1093/genetics/152.3.821
  7. Clark, M. A., Douglas, M., & Choi, J. (2018). Biology 2e (Section 19.2: Population genetics). OpenStax. openstax.org
Key terms
Population genetics
The study of allele and genotype frequencies in populations and how they change.
Allele frequency
The proportion of a particular allele among all copies of a gene in a population.
Hardy-Weinberg principle
A model predicting genotype frequencies (p-squared, 2pq, q-squared) in a non-evolving population.
Hardy-Weinberg equilibrium
The state in which allele and genotype frequencies stay constant across generations.
Null hypothesis
A baseline expectation (here, no evolution) against which real data are compared.
Genetic drift
Random change in allele frequencies, strongest in small populations.

Forces That Change Allele Frequencies

  • Identify the five forces that alter allele frequencies.
  • Distinguish genetic drift from natural selection.
  • Explain the founder effect and bottleneck effect.

The big picture

The Hardy-Weinberg model told you what a population looks like when nothing is changing it. This lesson is the flip side: it names the forces that break each of the model's five assumptions and so cause allele frequencies to change, which is exactly what evolution is. If you understand what each broken assumption does, you understand the basic engine of evolution.

The most important distinction here is between natural selection, which is directional and builds adaptation, and genetic drift, which is random and matters most in small populations. Getting that contrast clear is the single most useful thing in this lesson.

The five evolutionary forces

Hardy-Weinberg equilibrium requires five conditions: no mutation, no gene flow, no natural selection, random mating, and a very large population. Break any one and allele frequencies change, meaning the population evolves. Each broken assumption corresponds to a real evolutionary force:

  • Mutation introduces brand-new alleles into a population. It is the ultimate source of all genetic variation, but by itself it changes allele frequencies very slowly, because mutations are rare per generation.
  • Gene flow (also called migration) is the transfer of alleles between populations as individuals or their gametes move from one place to another. Gene flow tends to make separate populations more genetically similar to each other.
  • Natural selection is the differential survival and reproduction of different genotypes. It is the only force that consistently produces adaptation: a trait that increases fitness, that is, an organism's reproductive success. Selection increases the frequency of alleles that help their carriers survive and reproduce.
  • Genetic drift is random change in allele frequencies due to chance alone. It is strongest in small populations, and it can eliminate or fix an allele regardless of whether that allele is helpful or harmful.
  • Non-random mating occurs when individuals choose mates based on genotype or relatedness. It changes genotype proportions (for example, increasing homozygotes when relatives mate) even if it does not by itself change the underlying allele frequencies.

Key idea: Mutation, gene flow, natural selection, genetic drift, and non-random mating are the five forces; each breaks a Hardy-Weinberg assumption and can change a population genetically.

Selection versus drift

Two of these forces deserve a careful contrast, because students most often confuse them. Natural selection is not random. It systematically favors alleles that raise fitness, so over generations it produces organisms well suited to their environment. If a darker coat hides a mouse from predators, dark mice survive and reproduce more, and the dark allele rises predictably.

Genetic drift, by contrast, is entirely random. Which alleles happen to increase or decrease is a matter of chance, like the luck of which few individuals happen to leave offspring. In a large population, these chance fluctuations average out and drift is negligible, so selection dominates. In a small population, chance looms large, and drift can overwhelm selection, spreading even a harmful allele or wiping out a beneficial one. The bigger the population, the weaker the drift.

Key idea: Natural selection is directional and builds adaptation, while genetic drift is random and matters most in small populations, where it can override selection.

Two dramatic cases of drift

Drift is especially powerful in two named situations, both involving a small number of individuals.

In the founder effect, a small group breaks away from a larger population to start a new one, carrying only a random, unrepresentative sample of the original gene pool. Just by chance, some alleles will be over-represented and others missing, so the new population's allele frequencies can differ sharply from the source. Human populations founded by a handful of settlers often show unusually high frequencies of otherwise rare alleles for this reason.

In the bottleneck effect, a population is drastically reduced in size by a disaster such as disease, hunting, or habitat loss, and the survivors carry only a chance sample of the original variation. When the population later recovers, it rebuilds from that reduced sample. Both effects reduce genetic diversity and can leave a lasting genetic signature, which is one reason small and endangered populations are genetically fragile and vulnerable.

Key idea: The founder effect (a small group starting a new population) and the bottleneck effect (a population sharply reduced) are both forms of drift that reduce genetic diversity.

How fast can selection actually remove an allele?

Selection is often described as powerful, and against a common allele it is. Against a rare recessive one it is remarkably slow, and the arithmetic explains a fact that puzzles many students: why severe recessive conditions persist despite generations of selection.

Consider the strongest possible case, an allele that is completely lethal when homozygous and has no effect in heterozygotes. Under that assumption the allele frequency after t generations is qt = q0 / (1 + t q0).

Worked example. Start at q0 = 0.01, so one person in 10,000 is affected.

  • After 10 generations: q = 0.01 / (1 + 10 x 0.01) = 0.01 / 1.1 = 0.0091.
  • To halve the frequency to 0.005 requires t = (1/0.005) - (1/0.01) = 200 - 100 = 100 generations, roughly 2,500 years in humans.

The reason is arithmetic rather than biological. When q is 0.01, the fraction of copies of the allele sitting in homozygotes, where selection can see them, is only q2 / (q2 + 2pq), which is about 0.5 percent. Essentially every copy is hidden in a heterozygote and is invisible to selection. The rarer the allele becomes, the better it hides, so removal slows down exactly as it proceeds.

Key idea: With complete selection against a recessive homozygote, qt = q0/(1 + t q0), so halving a frequency of 0.01 takes 100 generations because almost every copy is sheltered in heterozygotes.

Mutation-selection balance

If selection removes alleles and mutation creates them, a frequency exists at which the two rates cancel. For a recessive allele under selection coefficient s, that equilibrium is q = the square root of (mu / s), where mu is the mutation rate to that allele per generation.

Worked example. Take a typical per-gene mutation rate of mu = 10-6 and a lethal recessive allele with s = 1.

  • q = square root of (10-6 / 1) = 10-3, so the allele sits at about one copy in a thousand.
  • The frequency of affected individuals is q2 = 10-6, or one in a million.
  • Carriers are 2pq, which is about 0.002, so roughly one person in 500 carries it. Carriers outnumber affected individuals by about 2,000 to 1.

For a dominant allele the equilibrium is different, q = mu / s, because every copy is exposed to selection immediately. That difference explains why severe dominant conditions are typically rare and often arise as new mutations in an unaffected family, while severe recessive conditions can reach appreciable carrier frequencies.

Key idea: Mutation and selection settle at q = square root of (mu/s) for a recessive allele and q = mu/s for a dominant one, which is why recessive conditions accumulate carriers and severe dominant ones usually appear as new mutations.

When the heterozygote wins: balancing selection

Selection does not always drive an allele to fixation or loss. If the heterozygote has the highest fitness, both alleles are maintained indefinitely, a situation called heterozygote advantage or overdominance. The sickle cell allele in regions with endemic malaria is the best-documented human example: homozygotes for the sickle allele have sickle cell disease, while heterozygotes have substantial protection against severe malaria and so out-survive both homozygous classes.

The equilibrium frequency follows from the two selection coefficients. Writing s1 for the disadvantage of the normal homozygote and s2 for that of the sickle homozygote, the sickle allele settles at q = s1 / (s1 + s2).

Worked example. Suppose malaria imposes a 10 percent fitness cost on the normal homozygote, so s1 = 0.1, and sickle cell disease historically imposed a near-complete cost, s2 = 1.0.

  • q = 0.1 / (0.1 + 1.0) = 0.1 / 1.1 = about 0.09, or 9 percent.

That predicted figure is close to the sickle allele frequencies actually observed in populations from historically high-malaria regions, and it also predicts what should happen where malaria is controlled: with s1 falling toward zero, the equilibrium frequency falls with it. The general lesson is that an allele's fitness is a property of an environment, not of the allele.

Key idea: Heterozygote advantage maintains both alleles at q = s1/(s1 + s2), which predicts a sickle allele frequency near 9 percent under historical malaria pressure and a decline where malaria is controlled.

How strong is drift? Effective population size

Drift is often described qualitatively as mattering more in small populations. The relationship can be made exact. The variance in allele frequency introduced by drift in one generation is pq / (2Ne), where Ne is the effective population size, the size of an idealized population that would drift at the observed rate.

Worked example. Take an allele at p = q = 0.5, so pq = 0.25.

  • With Ne = 1,000: variance = 0.25 / 2,000 = 0.000125, so the standard deviation of the change in one generation is about 0.011.
  • With Ne = 10: variance = 0.25 / 20 = 0.0125, and the standard deviation is about 0.11, ten times larger.

Two consequences follow. A brand-new neutral allele in a diploid population has a probability of eventually reaching fixation of just 1/(2N), simply because it is one copy among 2N. And effective population size is usually far smaller than the census count, because unequal family sizes, unequal sex ratios, and past bottlenecks all reduce it. A species that looks numerous can therefore be drifting like a much smaller one, which is a central concern in conservation genetics.

Key idea: Drift produces a per-generation variance of pq/(2Ne), so halving effective size roughly increases the standard deviation of change by 40 percent, and a new neutral allele fixes with probability 1/(2N).

Where people get stuck

The first sticking point is treating drift and selection as alternatives. They are not: they operate simultaneously, and which one dominates for a given allele depends on whether the selection coefficient is large or small compared with 1/(2Ne). An allele that is effectively neutral in a small population can be strongly selected in a large one.

The second is expecting selection to eliminate deleterious alleles entirely. It will not: mutation regenerates them continuously, so what selection achieves is an equilibrium frequency rather than removal.

The third is describing a bottleneck as reducing only the number of individuals. Its lasting genetic effect runs deeper: the loss of alleles, particularly rare ones, and that loss is not restored when numbers recover, because population growth multiplies the survivors rather than reintroducing what was lost.

Common misconceptions

  • Only natural selection reliably produces adaptation. Drift, gene flow, and mutation change frequencies but do not systematically improve fit to the environment.
  • Genetic drift is random, not goal-directed. It can spread harmful alleles or remove helpful ones purely by chance.
  • The founder and bottleneck effects are both drift, not selection, because which alleles survive is a matter of chance, not fitness.
  • Mutation alone changes allele frequencies very slowly; it supplies variation that the other forces then act on.

Recap

  • Five forces change allele frequencies: mutation, gene flow, natural selection, genetic drift, and non-random mating.
  • Natural selection is the only force that consistently produces adaptation by favoring fitness-raising alleles.
  • Genetic drift is random change, strongest in small populations, where it can overwhelm selection.
  • The founder effect and bottleneck effect are forms of drift that reduce genetic diversity.
  • Each force corresponds to breaking one Hardy-Weinberg assumption.

Sources

  1. Clark, M. A., Douglas, M., & Choi, J. (2018). Biology 2e (Ch. 19.3: Adaptive evolution). OpenStax. openstax.org
  2. National Human Genome Research Institute. (n.d.). Genetic drift. In Talking glossary of genomic and genetic terms. genome.gov
  3. Andrews, C. A. (2010). Natural selection, genetic drift, and gene flow do not act in isolation in natural populations. Nature Education Knowledge, 3(10), 5.
  4. Khan Academy. (n.d.). Mechanisms of evolution and genetic drift (evolution and natural selection). khanacademy.org
  5. Kimura, M. (1968). Evolutionary rate at the molecular level. Nature, 217(5129), 624-626. doi.org/10.1038/217624a0
  6. Lewontin, R. C., & Hubby, J. L. (1966). A molecular approach to the study of genic heterozygosity in natural populations. II. Amount of variation and degree of heterozygosity in natural populations of Drosophila pseudoobscura. Genetics, 54(2), 595-609. doi.org/10.1093/genetics/54.2.595
  7. University of California Museum of Paleontology. (n.d.). Mechanisms: The processes of evolution. Understanding Evolution. evolution.berkeley.edu
Key terms
Gene flow
The transfer of alleles between populations through migration of individuals or gametes.
Natural selection
Differential survival and reproduction of genotypes, the only force that produces adaptation.
Genetic drift
Random change in allele frequencies, strongest in small populations.
Founder effect
Reduced, skewed genetic variation when a small group founds a new population.
Bottleneck effect
Loss of genetic variation when a population is sharply reduced in size.
Adaptation
A trait that increases fitness, produced by natural selection.

Quantitative and Complex Traits

  • Explain why polygenic traits vary continuously.
  • Define heritability and interpret it correctly.
  • Distinguish the roles of genes and environment in complex traits.

The big picture

Mendel's traits fell into neat categories. But the traits people usually care about, height, weight, blood pressure, crop yield, do not. They vary smoothly across a whole range, with most individuals in the middle. This lesson explains why: such traits are shaped by many genes at once and by the environment. It closes the loop between the simple gene-by-gene genetics you learned early on and the messy reality of complex human traits and diseases.

You will also meet one of the most useful and most misunderstood ideas in all of genetics, heritability, and learn to state carefully what it does and does not mean.

Why some traits vary continuously

Many important traits do not fall into a few clean classes but vary smoothly across a range. These are quantitative traits: traits that vary continuously and can be measured on a scale, such as height or weight. The branch of the field that studies them is quantitative genetics. Quantitative traits behave differently from Mendel's clear-cut characters for two combined reasons: they are polygenic, and they are strongly influenced by the environment.

A polygenic trait is one influenced by many genes, each contributing a small additive effect (a small amount that sums with the others). A trait controlled by a single gene with two alleles produces only a few possible phenotypes. But when dozens of genes each nudge the trait up or down a little, the number of possible combinations becomes enormous, and the phenotypes blend into a smooth continuous variation: a range of values rather than a few discrete categories. Add the influence of environment, such as nutrition affecting height, and the outcome is the familiar bell-shaped distribution, with most individuals near the average and fewer toward the extremes.

Key idea: Quantitative traits vary continuously because many genes each add a small effect and the environment contributes too, blending phenotypes into a bell-shaped range.

A worked way to picture it

Imagine a simplified trait set by just three genes, each with an add-a-unit allele and an add-nothing allele. An individual could carry anywhere from 0 to 6 add-a-unit alleles, and the most common counts (around 3) can be reached by many different allele combinations, while the extremes (0 or 6) can be reached only one way each.

So middle values are common and extremes are rare, producing a peaked distribution, even from just three genes. Real quantitative traits involve far more genes plus environment, which smooths the steps into a continuous curve. This is why height does not come in a few discrete heights but in a smooth spread.

Key idea: Because middle phenotype values can be produced in many more ways than extreme values, polygenic traits naturally form a peaked, bell-shaped distribution.

Heritability

Since both genes and environment shape quantitative traits, a natural question is how much of the variation in a trait, within a particular population, is due to genetic differences among individuals. That proportion is called heritability. A heritability near 1 means most of the variation in that population is genetic; a heritability near 0 means most of it is environmental. Heritability is important in agriculture and animal breeding, where it predicts how quickly a trait can be improved by selective breeding: high-heritability traits respond fast to selection, low-heritability traits slowly.

Key idea: Heritability is the proportion of trait variation in a population that is due to genetic differences, and it predicts how well a trait will respond to selection.

What heritability does not mean

Heritability is one of the most misunderstood ideas in genetics, so three cautions are worth stating plainly:

  • Heritability describes variation within a population, not any single individual. A heritability of 0.8 for height does not mean 80 percent of your own height is genetic; that statement is meaningless for one person.
  • Heritability is specific to a particular population in a particular environment. Change the environment, and the value can change. A trait can have high heritability in one setting and low heritability in another.
  • High heritability does not mean a trait cannot be changed by the environment, and it says nothing about the causes of differences between groups. Group differences can be entirely environmental even for a highly heritable trait.

Key idea: Heritability applies to variation within one population in one environment; it does not partition a single person's trait and does not explain differences between groups.

Complex traits and modern genetics

With those cautions in mind, quantitative genetics bridges the gene-by-gene view of classical genetics and the reality of complex traits: traits shaped by the combined action of many genes and the environment. Most common diseases, such as diabetes and heart disease, are complex traits, influenced by many genetic variants of small effect plus lifestyle and environment. Modern genome-wide studies try to find those many variants, and quantitative genetics is the foundation for predicting disease risk and for improving crops and livestock. The simple rules you learned for single genes still operate; they are just summed over many genes at once.

Key idea: Complex traits, including most common diseases, arise from many small-effect genes plus environment, and quantitative genetics underlies efforts to predict risk and improve crops.

Splitting the variance

Quantitative genetics does not analyze individuals; it analyzes variation. The total observed variation in a trait within a population is written VP, for phenotypic variance, and the whole field rests on dividing it into components.

The first division separates genetic from environmental sources: VP = VG + VE. The genetic part then divides further. VA, the additive variance, comes from allele effects that simply add up and are therefore transmitted predictably from parent to offspring. The remainder comes from dominance interactions between alleles at one locus and from interactions between loci, neither of which survives the reshuffling of meiosis reliably.

That distinction produces two different heritabilities, and confusing them is a common source of error.

  • Broad-sense heritability, H2 = VG / VP, is the fraction of variation attributable to genetic differences of any kind.
  • Narrow-sense heritability, h2 = VA / VP, is the fraction attributable to additive effects alone. This is the one that predicts resemblance between parents and offspring, and therefore the one that matters for breeding and for response to selection.

Key idea: Phenotypic variance splits into genetic and environmental parts, the genetic part into additive and non-additive components, and narrow-sense heritability is the additive fraction that predicts parent-offspring resemblance.

The breeder's equation, worked

Narrow-sense heritability earns its keep in a single equation that predicts how much a population will change under selection: R = h2 x S. Here S is the selection differential, the difference between the mean of the selected parents and the mean of the whole population, and R is the response, the difference between the offspring mean and the original population mean.

Worked example. A population of grain plants yields a mean of 100 grams per plant. A breeder selects as parents only those plants averaging 120 grams, and the narrow-sense heritability of yield in this population is 0.4.

  • Selection differential S = 120 - 100 = 20 grams.
  • Response R = 0.4 x 20 = 8 grams.
  • The offspring generation should therefore average about 108 grams, not 120. The parents' advantage is only partly heritable, and the rest of it was environmental or non-additive.

Two extensions follow immediately. Rearranging gives h2 = R / S, so running the experiment and measuring the response is itself a way to estimate heritability, a quantity known as realized heritability. And repeating selection generation after generation depletes the additive variance it feeds on, so response slows and eventually plateaus. That plateau is a routine observation in long-running selection experiments and in commercial breeding programs alike.

Key idea: R = h2 x S predicts the response to selection, so a selection differential of 20 grams at h2 = 0.4 yields an 8-gram gain, and repeated selection exhausts additive variance and plateaus.

Finding the genes: association studies and polygenic scores

How do you find the actual genes involved? Modern work tries to identify the individual variants underlying quantitative variation. A genome-wide association study genotypes hundreds of thousands to millions of variants in very large samples and tests each one for association with a trait. Applied to height in samples of millions of people, this approach has identified more than ten thousand associated variants. Almost every one has a tiny individual effect, on the order of a fraction of a millimeter, and together they account for a substantial but incomplete share of the heritable variation.

Adding those effects together for one person gives a polygenic score. Such scores can rank a population meaningfully, and at the extremes of the distribution they carry real information about average risk. Three limitations bound what they mean.

  • They predict distributions, not individuals. A person in the top few percent of a score for a common disease still has a modest absolute probability of developing it, and someone at the bottom is not exempt.
  • They travel badly between populations. Because the large majority of participants in association studies have been of European ancestry, and because the correlations between measured markers and causal variants differ across ancestries, scores lose a substantial fraction of their accuracy when applied to people of other ancestries. Deploying them clinically without addressing this would widen existing health inequities rather than narrow them.
  • They describe association, not mechanism. A variant reliably associated with a trait may sit in a regulatory region affecting a gene some distance away, and identifying the causal gene is separate work.

Key idea: Association studies find thousands of tiny-effect variants per complex trait, and polygenic scores built from them rank risk within a population but predict poorly for individuals and transfer unreliably across ancestries.

Where people get stuck

The first and most consequential sticking point is reading heritability as inevitability. Phenylketonuria makes the point unanswerably: the underlying variation is genetic and the condition is highly heritable, yet detecting it at birth and restricting one amino acid in the diet prevents the intellectual disability entirely. A heritable trait can be completely modifiable, because heritability measures the sources of variation under current conditions and says nothing about what a new intervention would do.

The second is applying a heritability estimate outside the population and environment where it was measured. If environmental conditions become more uniform, environmental variance falls, and heritability rises even though nothing genetic has changed. The number is a property of a population at a time, not of a trait.

The third is the most serious error in this whole area: inferring that because a trait is heritable within groups, differences between groups must be genetic. The inference is invalid. Within-group heritability can be high while an entire between-group difference is environmental, and the standard demonstration is two genetically identical batches of seed grown in rich and poor soil, where height is highly heritable within each tray and the difference between trays is entirely due to the soil. Any claim about group differences requires evidence of a completely different kind.

Common misconceptions

  • Continuous variation is not a departure from Mendelian genetics; it is the summed result of many Mendelian genes plus environment.
  • Heritability is not the fraction of an individual's trait that is genetic. It describes variation across a population.
  • High heritability does not mean the environment is irrelevant or that the trait is fixed.
  • Heritability says nothing about why groups differ; between-group differences can be purely environmental.

Recap

  • Quantitative traits vary continuously because they are polygenic and environmentally influenced.
  • Many genes with small additive effects, plus environment, produce a bell-shaped distribution.
  • Heritability is the proportion of trait variation in a population that is genetic, and it predicts response to selection.
  • Heritability applies to a population in an environment, not to individuals or between-group differences.
  • Complex traits, including most common diseases, combine many small-effect genes with the environment.

Sources

  1. Clark, M. A., Douglas, M., & Choi, J. (2018). Biology 2e (Ch. 12.3: Laws of inheritance, polygenic and continuously varying traits). OpenStax. openstax.org
  2. National Human Genome Research Institute. (n.d.). Polygenic trait. In Talking glossary of genomic and genetic terms. genome.gov
  3. National Human Genome Research Institute. (n.d.). Talking glossary of genomic and genetic terms (heritability and complex disease). genome.gov
  4. Khan Academy. (n.d.). Variations on Mendelian genetics (polygenic inheritance and continuous variation). khanacademy.org
  5. Fisher, R. A. (1919). The correlation between relatives on the supposition of Mendelian inheritance. Transactions of the Royal Society of Edinburgh, 52(2), 399-433. doi.org/10.1017/S0080456800012163
  6. Visscher, P. M., Hill, W. G., & Wray, N. R. (2008). Heritability in the genomics era: Concepts and misconceptions. Nature Reviews Genetics, 9(4), 255-266. doi.org/10.1038/nrg2322
  7. Martin, A. R., Kanai, M., Kamatani, Y., Okada, Y., Neale, B. M., & Daly, M. J. (2019). Clinical use of current polygenic risk scores may exacerbate health disparities. Nature Genetics, 51(4), 584-591. ncbi.nlm.nih.gov
Key terms
Quantitative trait
A trait that varies continuously, such as height or weight.
Polygenic trait
A trait influenced by many genes, each adding a small effect.
Continuous variation
A smooth range of phenotypes rather than a few discrete categories.
Heritability
The proportion of trait variation in a population that is due to genetic differences.
Additive effect
The small, summing contribution of each gene to a polygenic trait.
Complex trait
A trait shaped by the combined action of many genes and the environment.

Module 7: Biotechnology and Genomics

The laboratory tools that let us copy, read, and edit DNA, including PCR, sequencing, and CRISPR, and what reading whole genomes reveals.

Tools of Genetic Engineering

  • Explain how restriction enzymes and plasmids enable cloning.
  • Describe how PCR amplifies a specific DNA sequence.
  • Outline how gel electrophoresis separates DNA fragments.

The big picture

Modern genetics is not only about understanding DNA. It is about handling it directly. This lesson introduces the core laboratory toolkit that lets scientists cut DNA at chosen spots, copy it, move it between organisms, and sort it by size. These few techniques underlie everything from producing life-saving insulin in bacteria to DNA fingerprinting in a forensics lab.

Each tool does one specific job, and they combine like a workshop's saw, copier, and sorting tray. Once you know what each does, you can follow how a gene gets moved from a human into a bacterium and mass-produced.

Cutting and pasting DNA

Biotechnology is the use of biological tools and organisms to solve practical problems. Its foundational cutting tools are restriction enzymes: proteins, originally discovered in bacteria, that cut DNA at a specific short recognition sequence. Because a given restriction enzyme always cuts at the same sequence, scientists can reliably snip out a gene of interest at predictable points. Many restriction enzymes cut in a staggered way that leaves short single-stranded overhangs called sticky ends, which readily pair with a matching overhang on another piece of DNA.

To paste two cut pieces together, DNA ligase (the same enzyme that seals fragments in replication) joins their backbones. A common destination for a cut gene is a plasmid: a small circular DNA molecule from bacteria that replicates on its own, separate from the main chromosome. A gene inserted into a plasmid and put back into bacteria is copied every time the bacteria divide, a process called cloning (making many identical copies of a gene or organism). This is exactly how bacteria are engineered to mass-produce human insulin for people with diabetes.

Key idea: Restriction enzymes cut DNA at specific sequences, ligase pastes pieces together, and a plasmid carries a gene into bacteria, which clone it every time they divide.

PCR: copying DNA in a tube

Often researchers need many copies of one specific stretch of DNA, for example to study it or to test for its presence. The polymerase chain reaction (PCR) does exactly that, amplifying a chosen DNA sequence millions of times in a few hours, right in a test tube. PCR repeats a simple three-step cycle:

  1. Heat separates (denatures) the two DNA strands.
  2. Cooling lets short primers bind to the target sequence, marking where copying will start.
  3. A heat-stable DNA polymerase extends the primers, building new complementary strands.

Each cycle doubles the amount of the target region, so after n cycles you have roughly 2 to the power n copies. That is why PCR is described as exponential: 10 cycles give about a thousandfold increase, 20 cycles about a millionfold. PCR is the workhorse behind DNA fingerprinting, diagnostic tests for infections and genetic conditions, and countless research applications.

Key idea: PCR amplifies a specific DNA sequence exponentially by repeating denature, primer-binding, and extension cycles, doubling the target each cycle.

Sorting DNA by size

To analyze DNA fragments, scientists separate them by size using gel electrophoresis: a method that uses an electric field to pull DNA through a gel, sorting fragments by length. A DNA sample is loaded into a slot at one end of a gel, and an electric field is applied across it. Because DNA is negatively charged (from its phosphate backbone), it moves toward the positive electrode. The gel acts like a molecular sieve: smaller fragments slip through its mesh faster and travel farther, while larger fragments lag behind.

The result is a pattern of bands, sorted from large (near the loading slot) to small (far from it). So if you cut a sample and get fragments of 500, 1500, and 3000 base pairs, the 500-base-pair fragment travels farthest and the 3000-base-pair fragment stays closest to the start. Gel electrophoresis is used to check the results of a restriction cut or a PCR reaction and to compare DNA samples, as in forensic identification and paternity testing.

Key idea: In gel electrophoresis, negatively charged DNA moves toward the positive electrode and smaller fragments travel farther, sorting fragments by size into visible bands.

Putting the tools together

These tools combine into a standard workflow for genetic engineering. To make bacteria produce a human protein such as insulin: use a restriction enzyme to cut out the human insulin gene and to open a plasmid at a matching site; use ligase to paste the gene into the plasmid; introduce the recombinant plasmid into bacteria; and let the bacteria clone and express the gene, producing human insulin. PCR can supply many copies of the gene to start with, and gel electrophoresis can confirm at each step that the right fragments are present. The toolkit is modular, and mastering what each piece does lets you follow almost any modern genetics protocol.

Key idea: Restriction enzymes, ligase, plasmids, PCR, and gel electrophoresis combine into a workflow that can move a human gene into bacteria and mass-produce its protein.

Restriction enzymes: why the sites look the way they do

Restriction enzymes exist because bacteria need a defense against bacteriophages, and they cut foreign DNA at short specific sequences while a partner enzyme methylates the cell's own copies of those sequences to protect them. Two structural features of their recognition sites have practical consequences.

First, most sites are palindromic, reading the same 5 prime to 3 prime on both strands. EcoRI recognizes GAATTC, whose complement read in the same direction is also GAATTC. This symmetry exists because the enzyme works as a pair of identical subunits, each engaging one strand.

Second, many enzymes cut the two strands at offset positions, leaving short single-stranded overhangs. EcoRI cuts between G and A on both strands, so every fragment ends in a four-base AATT overhang. Any two fragments cut with the same enzyme therefore have complementary sticky ends that pair spontaneously, which is exactly what makes it possible to join human DNA to a bacterial plasmid. Other enzymes cut both strands at the same position and leave blunt ends, which will join to anything but do so far less efficiently.

Worked example. How often should a given site occur by chance? A six-base site has one particular sequence out of 46 = 4,096 possibilities, so on random DNA it appears roughly every 4,096 base pairs. Cutting a 48,000 base-pair phage genome with such an enzyme should therefore give about 48,000 / 4,096 = 12 fragments. A four-base cutter, by the same reasoning, cuts every 44 = 256 bases and would produce nearly 190 fragments from the same DNA. Choosing the recognition-site length is choosing the average fragment size.

Key idea: Restriction sites are palindromes because the enzymes are symmetric dimers, offset cuts give complementary sticky ends that make cloning possible, and a site of n bases occurs about every 4n base pairs.

PCR by the numbers

Each PCR cycle doubles the target, so after n cycles the amount is multiplied by 2n. The consequences of that exponent are worth calculating rather than asserting.

  • After 20 cycles: 220 = about 1 million-fold.
  • After 30 cycles: 230 = about 1.07 billion-fold. A single starting molecule becomes enough DNA to see on a gel.
  • Real reactions are not perfectly efficient. At 90 percent efficiency each cycle multiplies by 1.9 rather than 2, and 1.930 is about 2.4 x 108, roughly four times less than the ideal. Reactions also plateau once reagents run low, which is why simply adding more cycles does not keep the exponential going.

Three practical points follow from the chemistry. The polymerase must survive repeated heating to about 95 degrees Celsius, which is why the enzyme originally came from a hot-spring bacterium. The annealing temperature is set a few degrees below the primers' melting temperature, which in turn depends on their length and G-C content, so primer design is really thermodynamics. And because a single contaminating molecule is amplified as faithfully as the intended one, contamination control is the dominant practical concern in any PCR laboratory.

Quantitative PCR turns this into a measurement by watching fluorescence rise during the reaction and recording the cycle at which the signal crosses a threshold. Since each cycle doubles, a sample that crosses one cycle earlier started with twice as much target, and about 3.3 cycles corresponds to a tenfold difference, because 23.3 is approximately 10.

Key idea: PCR multiplies the target by 2n, giving roughly a billionfold gain in 30 cycles, real efficiency below 100 percent reduces that severalfold, and in quantitative PCR one cycle equals a twofold difference in starting amount.

Finding the cells that worked

A transformation is inefficient: only a small minority of bacteria take up a plasmid, and only some of those plasmids contain the intended insert. Two filters are therefore built into every cloning plasmid.

Selection removes cells that took up nothing. The plasmid carries an antibiotic resistance gene, and plating on that antibiotic kills every untransformed cell, so every colony that grows carries a plasmid.

Screening then distinguishes plasmids that received an insert from those that simply re-closed on themselves. In the classic version, the cloning site sits inside a gene for an enzyme that turns a colorless dye blue. A successful insert interrupts that gene, so colonies with an insert are white and colonies without one are blue, and the correct ones can be picked by eye.

Key idea: Antibiotic selection identifies cells that took up any plasmid, and a screen such as blue-white color identifies the subset whose plasmid actually carries the insert.

Where people get stuck

The first sticking point is expecting a bacterium to express a human gene straight from genomic DNA. It cannot: bacteria cannot splice, so introns would be translated as nonsense. The standard solution is to start from the messenger RNA and copy it back into DNA with reverse transcriptase, which yields an intron-free version of the gene.

The second is confusing PCR with cloning. They are not the same: PCR copies a defined region in a tube and needs to know the flanking sequences in advance, while cloning inserts DNA into a living cell that then replicates it indefinitely. They answer different questions, and modern work usually uses both.

The third is reading a gel backwards. Do not: small fragments migrate furthest, so the bands nearest the bottom are the smallest, and DNA always runs toward the positive electrode because its phosphate backbone is negatively charged at working pH.

Common misconceptions

  • Restriction enzymes cut DNA; they do not copy it. PCR is what makes copies.
  • In gel electrophoresis, smaller fragments travel farther, not larger ones, because they move through the gel mesh more easily.
  • A plasmid is a small circular DNA that replicates independently; it is not part of the main bacterial chromosome.
  • PCR growth is exponential, not linear. Each cycle doubles the target, so copies rise very fast.

Recap

  • Restriction enzymes cut DNA at specific sequences, and ligase joins cut pieces together.
  • Plasmids carry genes into bacteria, which clone the gene as they divide.
  • PCR amplifies a chosen DNA sequence exponentially through repeated heating and copying cycles.
  • Gel electrophoresis sorts DNA fragments by size, with smaller fragments moving farther toward the positive electrode.
  • Together these tools enable feats such as engineering bacteria to produce human insulin.

Sources

  1. Clark, M. A., Douglas, M., & Choi, J. (2018). Biology 2e (Ch. 17: Biotechnology and genomics, cloning, PCR, and gel electrophoresis). OpenStax. openstax.org
  2. National Human Genome Research Institute. (n.d.). Polymerase chain reaction (PCR). In Talking glossary of genomic and genetic terms. genome.gov
  3. National Human Genome Research Institute. (n.d.). Talking glossary of genomic and genetic terms (restriction enzyme, plasmid, and cloning). genome.gov
  4. Khan Academy. (n.d.). Biotechnology: DNA cloning and PCR. khanacademy.org
  5. Saiki, R. K., Gelfand, D. H., Stoffel, S., Scharf, S. J., Higuchi, R., Horn, G. T., Mullis, K. B., & Erlich, H. A. (1988). Primer-directed enzymatic amplification of DNA with a thermostable DNA polymerase. Science, 239(4839), 487-491. doi.org/10.1126/science.2448875
  6. Cohen, S. N., Chang, A. C. Y., Boyer, H. W., & Helling, R. B. (1973). Construction of biologically functional bacterial plasmids in vitro. Proceedings of the National Academy of Sciences, 70(11), 3240-3244. doi.org/10.1073/pnas.70.11.3240
  7. National Human Genome Research Institute. (2020). Polymerase chain reaction (PCR) fact sheet. National Institutes of Health. genome.gov
Key terms
Biotechnology
The use of biological tools and organisms to solve practical problems.
Restriction enzyme
A protein that cuts DNA at a specific recognition sequence.
Plasmid
A small circular DNA molecule from bacteria used to carry genes in cloning.
Cloning
Making many identical copies of a gene or organism.
Polymerase chain reaction (PCR)
A technique that amplifies a specific DNA sequence millions of times.
Gel electrophoresis
A method that separates DNA fragments by size using an electric field.

DNA Sequencing, CRISPR, and Genomics

  • Explain what DNA sequencing determines and why it matters.
  • Describe how CRISPR-Cas9 edits genes precisely.
  • Define genomics and give examples of what genomes reveal.

The big picture

This final lesson reaches the current frontier of genetics: reading and rewriting whole genomes. Two technologies did most of the work of getting us here: fast DNA sequencing, which lets us read the entire genetic instruction set, and CRISPR gene editing, which lets us change it at a chosen spot. Together they have transformed biology and medicine in the last two decades.

The course opened with Mendel counting peas and closes with the ability to read and edit the complete DNA of any organism. Understanding these tools, and the questions they raise, is essential for making sense of modern genetics and the medical and ethical debates around it.

Reading the genome

DNA sequencing determines the exact order of the bases (A, T, C, and G) in a stretch of DNA. Knowing the sequence is the starting point for almost everything else: identifying genes, spotting mutations, and comparing organisms. The landmark effort was the Human Genome Project, an international collaboration completed in 2003, which produced the first reference sequence of a human genome, roughly three billion base pairs. That first genome took over a decade and enormous resources.

Since then, the cost and time of sequencing have fallen dramatically, so that a human genome can now be sequenced quickly and inexpensively, in a matter of days rather than years. This flood of affordable sequence data gave rise to genomics: the study of whole genomes and all of an organism's genes together, rather than one gene at a time. Genomics is what you get when reading DNA becomes cheap enough to read all of it.

Key idea: DNA sequencing reads the exact order of bases; the Human Genome Project first sequenced the human genome in 2003, and falling costs since then launched genomics, the study of whole genomes.

Editing the genome with CRISPR

CRISPR-Cas9 is a precise, programmable gene-editing tool adapted from a bacterial immune system that bacteria use to cut up invading viruses. It has two working parts:

  • A guide RNA: a short RNA molecule designed to match a chosen target sequence in the genome. It acts like a homing address.
  • The Cas9 protein: an enzyme that cuts DNA. It acts like the scissors.

The guide RNA leads Cas9 to the exact matching site in the genome, where Cas9 makes a cut in the DNA. The cell's own repair machinery then heals the break, and scientists exploit that repair step to make an edit: they can disable a gene by letting the repair introduce errors, or insert a new sequence at the cut site. The reason CRISPR spread through biology with astonishing speed is that it is cheap, precise, and easy to reprogram: to aim at a new target you simply design a new guide RNA, leaving Cas9 unchanged. Editing DNA went from difficult and specialized to something almost any lab can do.

Key idea: CRISPR-Cas9 edits genes by using a programmable guide RNA to lead the Cas9 cutting enzyme to a chosen sequence, and it is easy to retarget by changing only the guide RNA.

Promise and ethical questions

CRISPR is being developed to treat genetic diseases, and early therapies for conditions such as sickle-cell disease have shown real success by editing a patient's own cells. But the same power raises serious ethical questions, especially about editing human embryos or reproductive cells in ways that would be inherited by all future generations, a step widely regarded as crossing a line that current science should not cross. The technology's ease of use makes these questions urgent rather than hypothetical.

Key idea: CRISPR promises cures for genetic diseases like sickle-cell disease, but editing inheritable embryo or germline DNA raises serious ethical concerns.

What genomes reveal

Reading genomes at scale has reshaped our understanding of life in several ways:

  • Comparing genomes across species confirms common ancestry and reveals which genes are shared and conserved, turning evolution into something readable in DNA.
  • Within our own species, genomics links particular DNA variants to disease risk and to differences in how patients respond to drugs. This is the basis of personalized medicine: tailoring medical treatment and prevention to an individual's genetic makeup.
  • Genomics also studies the collective genomes of whole microbial communities, such as the human microbiome, revealing the vast unseen genetics of the organisms that live in and on us.

This course began with Mendel deducing hidden factors from counting peas, and it ends with the ability to read and edit the complete instruction set of any organism. That shift, from inferring genes to reading and rewriting them, places genetics at the center of twenty-first-century biology and medicine.

Key idea: Genomics confirms evolutionary relationships, links DNA variants to disease and drug response (personalized medicine), and reads the genomes of whole microbial communities.

Three generations of sequencing

The method has changed three times, and each change altered what questions were askable.

Sanger sequencing works by including a small proportion of modified nucleotides that terminate the growing chain wherever they are incorporated. The result is a nested set of fragments of every possible length, which are separated by size and read off base by base. It is accurate and produces reads of several hundred bases, and it remains the method of choice for checking a single short region. It was also the method behind the Human Genome Project, which took thirteen years and on the order of a few billion dollars for one composite reference sequence, declared essentially complete in 2003.

Next-generation sequencing replaced one reaction at a time with hundreds of millions in parallel on a single surface. The individual reads are short, typically a few hundred bases, and are assembled computationally by overlap. Cost per genome has fallen by orders of magnitude since 2007, far faster than computing costs fell over the same period, which is why sequencing moved from a national project to a routine laboratory service.

Long-read sequencing addresses what short reads cannot do. Repetitive regions longer than a read cannot be placed unambiguously, so the 2003 reference left several percent of the genome unresolved, largely around centromeres and other repeats. Reads tens of thousands of bases long closed those gaps, and a genuinely complete, gap-free human genome sequence was published in 2022, almost two decades after the draft.

Key idea: Sanger sequencing gave accurate short reads and the 2003 reference genome, massively parallel short reads collapsed the cost, and long reads finally resolved the repetitive regions to produce a gap-free human genome in 2022.

How CRISPR-Cas9 finds and cuts its target

The system was borrowed from bacteria, where it functions as an adaptive immune mechanism that stores fragments of past viral invaders and uses them to recognize the same virus again. Three components do the work in the laboratory version.

  1. A guide RNA of about twenty bases is designed to match the intended target sequence. Retargeting the system to a different gene means synthesizing a different guide, which is why the technique spread so quickly.
  2. The Cas9 protein carries that guide, scans the genome, and pairs the guide with matching DNA.
  3. A short adjacent motif, called a PAM and consisting of two guanines for the most widely used Cas9, must sit immediately next to the target. Without it Cas9 will not cut, which both restricts where edits can be made and prevents the enzyme from attacking the bacterial CRISPR array itself.

Cas9 then cuts both DNA strands a few bases from the PAM, and what happens next is done by the cell, not by the enzyme. If the break is sealed by end joining, small insertions or deletions are frequently introduced, which shifts the reading frame and knocks the gene out. That is easy and reliable. If a donor template is supplied and the cell uses homologous recombination instead, the sequence can be rewritten precisely, but this route is much less efficient and works mainly in dividing cells. Newer tools avoid the double-strand break entirely: base editors chemically convert one base into another in place, and prime editors write a short specified sequence directly.

Key idea: A twenty-base guide RNA plus an adjacent PAM directs Cas9 to cut, after which cellular end joining produces knockouts easily while precise rewriting by recombination remains inefficient, motivating base and prime editors.

What CRISPR does not yet do reliably

The technique is genuinely transformative and it is also frequently overstated. The honest limitations are these.

  • Off-target cutting. A guide can tolerate mismatches and cut at unintended sites. Screening methods now detect these, and higher-fidelity Cas9 variants reduce them, but they are not eliminated.
  • Unwanted on-target damage. Repair of the intended break can produce large deletions and complex rearrangements around the site, effects that standard short-range checks can miss entirely.
  • Delivery. Getting the editing machinery into the right cells in a living body remains the central practical obstacle, and it is the reason the first approved therapies edit cells outside the body and return them.
  • Efficiency and mosaicism. Not every targeted cell is edited, and editing an embryo can produce an individual whose cells are a mixture of edited and unedited, which is difficult to detect in advance and impossible to correct afterwards.

Against that, the clinical achievement is real. A therapy for sickle cell disease and beta thalassemia approved in late 2023 removes a patient's own blood stem cells, edits a regulatory sequence so that the cells resume making fetal hemoglobin, and returns them. It sidesteps delivery by working in a dish, and it edits body cells only, so the change is not passed to the patient's children.

Key idea: Off-target cuts, large on-target rearrangements, delivery into living tissue, and mosaicism remain unsolved, which is why the first approved CRISPR therapy edits a patient's own cells outside the body.

Germline editing: a question that is not settled

The distinction that carries the ethical weight is between somatic and germline editing. A somatic edit affects only the treated person and is, in principle, an ordinary medical intervention subject to ordinary standards of consent and evidence. A germline edit, made in an embryo, egg, or sperm, would be present in every cell of the resulting person and in their descendants.

The debate became concrete in 2018, when a researcher announced the birth of children from embryos he had edited. The response from the scientific community was near-uniform condemnation on grounds of inadequate safety data, questionable consent, and no genuine medical necessity, since safer alternatives existed; he was subsequently convicted in his own jurisdiction. National academies and international bodies called for a moratorium on clinical germline use, and in 2021 the World Health Organization published a governance framework recommending registries, oversight mechanisms, and international coordination, while not endorsing heritable editing at present.

The arguments on each side are worth stating plainly rather than resolving.

  • In favour. Some couples cannot have an unaffected genetically related child by any current means. Preventing a severe heritable condition permanently, rather than treating each generation, is a coherent medical goal.
  • Against. The people most affected cannot consent. Off-target and mosaicism risks are inherited along with the intended change. Almost every case put forward can already be addressed by testing embryos and selecting an unaffected one. And the line between preventing disease and selecting for preferred traits is difficult to define and easier still to move, with equity consequences if access follows wealth.

Disability rights scholars have also argued that framing certain conditions as defects to be eliminated reflects a judgement about which lives are worth living, and that this judgement should not be made implicitly by technical decisions. Whatever position a reader reaches, the responsible summary is that this is an open and contested question of governance and values, not one that the biology settles.

Key idea: Somatic edits affect one consenting patient while germline edits are inherited, and after the condemned 2018 case the international position is oversight and restraint rather than settled approval, with serious arguments on both sides.

Where people get stuck

The first sticking point is treating a sequenced genome as an interpreted one. They are not the same thing: reading the bases is now cheap and fast, but knowing what a given variant does is neither, and most variants found in any individual genome have no established significance.

The second is imagining CRISPR as a word processor for DNA. It is not: cutting is easy and reliable, but replacing is neither, and the cell rather than the enzyme decides what the final sequence looks like.

The third is assuming that a genetic therapy for a condition makes the condition a solved problem. It does not: cost, access, delivery, and long-term follow-up all remain, and a therapy approved in one country may be unavailable in the places where the condition is most common.

Common misconceptions

  • Sequencing reads DNA; CRISPR edits it. They are different tools for different jobs.
  • In CRISPR, the guide RNA provides the targeting and Cas9 does the cutting; to hit a new target you change the guide RNA, not Cas9.
  • The Human Genome Project did not invent CRISPR or discover DNA's structure; it produced the first reference human genome sequence.
  • Genomics studies whole genomes, not single genes in isolation.

Recap

  • DNA sequencing reads the exact order of bases; the Human Genome Project first sequenced the human genome in 2003.
  • Falling sequencing costs created genomics, the study of whole genomes.
  • CRISPR-Cas9 edits genes using a programmable guide RNA to direct the Cas9 cutting enzyme.
  • CRISPR offers cures for genetic diseases but raises ethical concerns about heritable editing.
  • Genomics confirms common ancestry, enables personalized medicine, and studies microbial communities.

Sources

  1. Clark, M. A., Douglas, M., & Choi, J. (2018). Biology 2e (Ch. 17: Biotechnology and genomics, sequencing and applications). OpenStax. openstax.org
  2. National Human Genome Research Institute. (n.d.). The Human Genome Project [Fact sheet]. genome.gov
  3. National Human Genome Research Institute. (n.d.). CRISPR. In Talking glossary of genomic and genetic terms. genome.gov
  4. Khan Academy. (n.d.). Biotechnology: DNA sequencing and genome editing. khanacademy.org
  5. Sanger, F., Nicklen, S., & Coulson, A. R. (1977). DNA sequencing with chain-terminating inhibitors. Proceedings of the National Academy of Sciences, 74(12), 5463-5467. doi.org/10.1073/pnas.74.12.5463
  6. Jinek, M., Chylinski, K., Fonfara, I., Hauer, M., Doudna, J. A., & Charpentier, E. (2012). A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity. Science, 337(6096), 816-821. doi.org/10.1126/science.1225829
  7. World Health Organization. (2021). Human genome editing: Recommendations. who.int
  8. Kosicki, M., Tomberg, K., & Bradley, A. (2018). Repair of double-strand breaks induced by CRISPR-Cas9 leads to large deletions and complex rearrangements. Nature Biotechnology, 36(8), 765-771. ncbi.nlm.nih.gov
Key terms
DNA sequencing
Determining the exact order of bases in a DNA molecule.
Genomics
The study of whole genomes and all of an organism's genes together.
Human Genome Project
The international effort, completed in 2003, that first sequenced a reference human genome.
CRISPR-Cas9
A precise, programmable gene-editing tool adapted from a bacterial defense system.
Guide RNA
The RNA that directs Cas9 to a specific matching DNA sequence to be cut.
Personalized medicine
Tailoring medical treatment to an individual's genetic makeup.

Open the interactive version with quizzes and progress →