⚧ Gender & Women’s Studies · Undergraduate · WGS 340

Gender, Sexuality & Society

A college-level course on gender and sexuality that leads with measurement. It separates three concepts that ordinary speech runs together: sex, a set of biological characteristics; gender, a social system of roles, expectations and identity; and sexuality, which research measures along three independent dimensions of attraction, behaviour and identity. From there the course works through the…

Start the interactive course (quizzes, progress, videos) →

Free forever. No sign-up, no ads. 15 lessons. The full lesson text is below so you can read it right here.

Module 1: Concepts and Bodies

Three concepts that ordinary speech treats as one, the biology of sex difference reported as effect sizes rather than adjectives, and the evidence on how children become gendered.

Sex, Gender, Sexuality: Three Concepts, Three Literatures

  • Distinguish sex, gender and sexuality as separate constructs and state how each is operationalized in research.
  • Explain why sexuality requires three independent measures and what happens when only one is collected.
  • Identify which construct a given study is actually measuring, and name the measurement choice its authors made.

A word that had to be borrowed

In 1955 a psychologist at Johns Hopkins named John Money published a paper about children born with reproductive anatomy that did not sort cleanly into the two boxes on a birth certificate. He had a vocabulary problem. The word sex was already doing four jobs at once: it named chromosomes, it named genitals, it named an act, and it named the social category a newborn was sorted into and then treated as for life. Money needed to write about the fourth thing on its own, so he borrowed a term that grammarians had used for centuries to sort nouns, and wrote about gender role: the whole set of things a person does that present them to others as a boy or a man, a girl or a woman.

Thirteen years later Robert Stoller, a psychiatrist who ran a gender identity clinic at UCLA, published Sex and Gender (1968) and added the second half of the pair: gender identity, a person's internal sense of being male, female, both or neither. In 1972 the British sociologist Ann Oakley put the whole distinction into general circulation with Sex, Gender and Society, and it moved out of the clinic and into the social sciences, where it has been argued over ever since.

You already met that distinction in WGS 2010. This lesson does something narrower and more useful: it treats sex, gender and sexuality as three separate measurement problems, because that is where most confusion in this field actually starts. When a headline says that men and women differ in some trait, the first question is not whether that is true. The first question is which of the three things the study measured, and how.

Key idea: Sex, gender and sexuality are not three words for one thing; they are three constructs with different definitions, different instruments and different research literatures, and a study can only be evaluated once you know which one it collected.

Sex: what the word refers to when researchers use it

In biology, sex refers to a cluster of characteristics organized around two gamete types: large, few, immobile ova and small, many, mobile sperm. Around that anchor sit chromosomes, gonads, hormone profiles, internal ducts, external genitalia and the secondary characteristics that appear at puberty. In typical development these line up. Roughly six to seven weeks after conception, the SRY gene on the Y chromosome switches the undifferentiated gonad toward testis development; without that signal the gonad develops as an ovary, and the hormonal cascade follows from there.

The cluster does not always line up, and the exceptions are informative rather than exotic. In complete androgen insensitivity syndrome, a fetus with XY chromosomes and functioning testes cannot respond to androgens, and develops typical external female anatomy; the condition is often discovered only when menstruation does not begin. In congenital adrenal hyperplasia, an XX fetus is exposed to elevated androgens and may be born with virilized external genitalia. Klinefelter syndrome (47,XXY) occurs in something on the order of one in 500 to one in 1,000 births of babies assigned male; Turner syndrome (45,X) in roughly one in 2,500 births of babies assigned female.

How common is all this? Here is a lesson in how definitions manufacture numbers. In 2000 Anne Fausto-Sterling and colleagues published an estimate in the American Journal of Human Biology putting the frequency of bodies that differ from a strict developmental ideal at about 1.7 percent of births. In 2002 Leonard Sax replied in the Journal of Sex Research that the 1.7 percent figure was dominated by conditions such as late-onset congenital adrenal hyperplasia and Klinefelter syndrome, which do not produce genital ambiguity at birth. Restricting the count to cases where the sex of a newborn is not immediately apparent, Sax arrived at roughly 0.018 percent, about one in 5,500. Both numbers are arithmetically correct. They answer different questions. Whenever you meet a prevalence figure in this course, ask what the denominator and the inclusion criteria were before you ask whether the number is big.

One practical warning. In almost every dataset you will ever use, sex is a single binary variable, usually recorded once, usually by someone other than the person it describes. That is a measurement decision with consequences, and it is invisible in the resulting tables.

The point: Sex is a cluster of biological characteristics that usually but not always co-vary, and published prevalence estimates for atypical development range across two orders of magnitude purely because of what each author counted.

Gender: an identity, an interaction, and a structure

Gender research operates at three levels at once, and papers that seem to contradict each other are often working at different ones.

At the individual level, gender is identity and expression: what a person understands themselves to be, and how they present. Current survey practice, endorsed in a 2022 National Academies report on measuring sex, gender identity and sexual orientation, is the two-step question: ask sex assigned at birth, then ask current gender identity, and compare. Asking one question with mixed categories produces data you cannot disaggregate later.

At the interactional level, gender is something people do to and for each other. Candace West and Don Zimmerman argued in a 1987 article in Gender and Society that gender is not a property you carry into a room but an accomplishment produced in the room, moment by moment, under the constant possibility of being held accountable for doing it wrong. Their example is unglamorous and convincing: the endless small decisions about voice, gesture, deference and who reaches for the check are not expressions of an inner state so much as an ongoing performance that others assess. Judith Butler's Gender Trouble (1990) pushed further, arguing that the repeated acts do not express a prior gendered essence but constitute the appearance of one. Butler has spent thirty years correcting the most common misreading of that claim, which is that gender is a costume you select each morning; the argument is closer to the opposite, that the repetition is compelled.

At the structural level, gender is a property of institutions. Joan Acker's 1990 analysis of gendered organizations pointed out that the abstract worker assumed by most job design is a person with no pregnancy, no school pickup and no elder care, which is not a neutral assumption. Structural measures include occupational segregation indexes, the share of legislative seats held by women, and composite instruments such as the World Economic Forum's Global Gender Gap Index.

What matters here: Gender is measured at three different levels, individual identity, interactional accomplishment and institutional structure, and a finding at one level does not automatically transfer to another.

Sexuality: why one question is never enough

Sexuality research settled decades ago on three dimensions that have to be measured separately because they do not coincide in the same people.

  • Attraction. Who a person is drawn to, sexually or romantically.
  • Behaviour. Who a person has actually had sexual contact with, over a stated period.
  • Identity. The label a person applies to themselves: straight, gay, lesbian, bisexual, queer, asexual, or none of these.

Edward Laumann and colleagues built this three-part structure into the 1992 National Health and Social Life Survey, and the American national surveys that followed kept it. The reason is that the three produce different numbers from the same respondents. In the National Survey of Family Growth for 2011 to 2013, about 17 percent of women aged 18 to 44 reported some same-sex sexual contact in their lifetime, while under 7 percent identified as lesbian or bisexual. For men, both figures were lower and the gap between them was smaller. Public health surveillance uses the awkward phrase men who have sex with men precisely because identity does not predict behaviour reliably enough to base an HIV prevention program on.

Identity figures also move fast, which tells you something about identity rather than about desire. Gallup's tracking, which asks a single identity question, put the share of American adults identifying as LGBTQ+ at around 7.6 percent in its 2024 release, with roughly one in five adults in Generation Z doing so, against low single digits among the oldest cohorts. Attraction and behaviour estimates have not moved anywhere near as sharply over the same period. The most defensible reading is that the willingness to apply a label has changed faster than the underlying distribution of desire.

A note on the most famous number in the field. Alfred Kinsey's volumes of 1948 and 1953 reported that a substantial minority of men had had some same-sex experience, and his seven-point scale treated orientation as a continuum rather than a pair of boxes. The claim that Kinsey found ten percent of men to be homosexual is a compression of a much narrower finding about a specific behavioural window, applied to a sample that was never nationally representative. Kinsey's real contribution was the continuum and the sheer scale of the interviewing, not that one figure.

Remember: Attraction, behaviour and identity are three separate measurements, they diverge substantially in the same population, and a survey that collects only identity will systematically undercount same-sex behaviour.

The three side by side

ConstructRefers toUsually measured byFields that study it
SexGametes, chromosomes, gonads, hormones, anatomyA single recorded binary variable; occasionally karyotype or hormone assayEndocrinology, genetics, epidemiology, sports science
GenderIdentity, expression, roles, and the structure of institutionsTwo-step identity items, expression scales, segregation and representation indexesSociology, psychology, economics, political science, anthropology
SexualityAttraction, behaviour, identityThree separate self-report items, ideally self-administeredPublic health, demography, psychology, sociology

Read the third column across and you can see why interdisciplinary arguments in this field go badly. An endocrinologist and a sociologist using the word gender are often not disagreeing about the world; they are reporting different instruments.

Where the three come apart

Three concrete cases will fix the distinctions better than another definition.

A woman with complete androgen insensitivity syndrome has XY chromosomes and internal testes. She was assigned female at birth on the evidence available, was raised as a girl, identifies as a woman, and may be attracted to men. Sex, at the chromosomal level, points one way; every other element points another. Whether you call her sex female depends entirely on which element of the cluster you have decided the word tracks, which is a definitional choice, not a discovery.

A man in a rural county reports on a confidential survey that he has had sex with men in the past year and describes himself, without hesitation, as straight. He is not lying on either item. Identity and behaviour are separate variables, and a health department that plans outreach from the identity item will not find him.

A trans woman describes herself as a lesbian. Gender identity and sexual orientation are orthogonal axes; knowing where someone sits on one tells you nothing about the other. Confusing the two is the single most common error in undergraduate writing on this material, and it is the reason the acronym keeps a T that stands for something categorically different from the L, G and B.

There is a fourth case worth naming: mode of administration changes answers. Studies comparing face-to-face interviewing with audio computer-assisted self-interviewing consistently find higher reports of stigmatized behaviour when a machine, rather than a person, hears the answer. That is not a curiosity. It means that for sensitive items, part of what you are measuring is the interview.

The upshot: Sex, gender identity and sexual orientation vary independently in real people, so a study that collapses them into one variable has destroyed information it cannot recover.

Common misconceptions

  • Sex is biological and gender is social, so they never interact. The distinction is about what each word denotes, not a claim that bodies and social life are sealed off from each other; hormones shape behaviour and social conditions shape hormones.
  • Intersex conditions are so rare they can be ignored, or so common that sex is not bimodal. Both claims come from the same literature; the range from 0.018 percent to 1.7 percent is produced entirely by inclusion criteria.
  • Gender performativity means gender is a free daily choice. Butler's argument is that repetition under social compulsion produces the appearance of an essence, which is close to the opposite of voluntarism.
  • Kinsey found that ten percent of people are gay. He reported a narrower behavioural finding on a non-representative sample, and his lasting contribution was the continuum scale.
  • Gender identity and sexual orientation are two versions of the same thing. They are independent axes; a person's gender does not predict who they are attracted to.

The short version

  • Money borrowed gender from grammar in 1955, Stoller added gender identity in 1968, and Oakley carried the pair into sociology in 1972.
  • Sex names a cluster of biological features that usually co-vary; prevalence estimates for atypical development range from about 0.018 percent to about 1.7 percent depending purely on what is counted.
  • Gender is measured at three levels: identity, interaction and institutional structure. The two-step question is the current standard for the first.
  • Sexuality requires three separate items, since attraction, behaviour and identity diverge; about 17 percent of US women aged 18 to 44 reported same-sex contact against under 7 percent identifying as lesbian or bisexual.
  • The three constructs vary independently in real people, and the mode of interviewing itself changes the answers to sensitive questions.

Sources

  1. National Academies of Sciences, Engineering, and Medicine. (2022). Measuring sex, gender identity, and sexual orientation. The National Academies Press. nationalacademies.org
  2. West, C., and Zimmerman, D. H. (1987). Doing gender. Gender and Society, 1(2), 125-151. SAGE Journals
  3. National Center for Health Statistics. (n.d.). National Survey of Family Growth. Centers for Disease Control and Prevention. cdc.gov
  4. Gallup. (2024). LGBTQ+ identification in the United States. news.gallup.com
  5. Encyclopaedia Britannica. (n.d.). Gender identity. britannica.com
Key terms
Gender role
Money's 1955 term for the full set of behaviours by which a person presents to others as a boy or man, girl or woman.
Gender identity
A person's internal sense of being male, female, both or neither, named as a separate construct by Stoller in 1968.
Two-step question
The survey method of asking sex assigned at birth and current gender identity separately so the two can be compared.
Doing gender
West and Zimmerman's argument that gender is an accomplishment produced in interaction under accountability, not a property carried into it.
Differences of sex development
Congenital conditions in which chromosomal, gonadal or anatomical sex characteristics do not follow the typical pattern.
Three dimensions of sexuality
Attraction, behaviour and identity, measured separately because they diverge within the same population.
Men who have sex with men
A behavioural public health category used because sexual identity does not reliably predict sexual behaviour.
Mode effect
The change in survey answers produced by how a question is administered, such as higher reporting of stigmatized behaviour under self-administration.

The Biology, Honestly: Effect Sizes, Overlap, and Open Questions

  • Compute and interpret Cohen's d, and convert an effect size into overlap and into a plain-language probability.
  • State which measured sex differences are large, which are small, and which are contested, with the evidence for each.
  • Explain why a replicated difference does not by itself establish a cause, and name the evidence that would bear on causation.

Seven million test scores

In 2008 Janet Hyde and four colleagues did something unglamorous and decisive. The No Child Left Behind Act had obliged American states to test every child in mathematics, which meant that for the first time the scores of essentially an entire population existed in one place. Hyde's team pulled the results for grades 2 through 11 from ten states, more than seven million individual scores, and calculated how far apart boys and girls actually were. The answer, published in Science, was that the standardized difference sat between about zero and 0.06 in every grade and every state. On a scale where 0.20 counts as a small effect, that is indistinguishable from nothing.

What makes this the right place to start is not the finding but the unit. Hyde reported a number called Cohen's d, and until you can read one you cannot evaluate any claim in this field, in either direction. People who want to argue that the sexes are basically the same and people who want to argue that they differ profoundly are usually citing the same literature. What separates them is which effect sizes they mention.

Why this matters: Nearly every dispute about sex differences is a dispute about magnitude, not existence, and magnitude has a standard unit that you can learn in ten minutes.

How to read an effect size

Cohen's d is the difference between two group means divided by the pooled standard deviation of the two groups. Dividing by the spread is the whole trick: it converts a difference measured in centimetres, or test points, or newtons of grip force, into a common currency, so that a difference in height and a difference in vocabulary can be compared.

Work one example all the way through. In American adults, mean male height is roughly 175 cm and mean female height roughly 162 cm, with a standard deviation of about 7 cm in each group. The difference of 13 cm divided by 7 gives d of about 1.9. Jacob Cohen's conventional labels call 0.2 small, 0.5 medium and 0.8 large, so 1.9 is off the top of the usual scale. Height is about as separated as human sex differences get.

Now translate that into something you can picture. Two conversions do the work. The first is overlap: the proportion of the two distributions that occupy the same range of values. The second is the plain-language probability that a randomly chosen member of the higher group exceeds a randomly chosen member of the lower group.

Cohen's dLabelDistribution overlapChance the drawn man exceeds the drawn woman
0.10Trivial96 percent53 percent
0.20Small92 percent56 percent
0.50Medium80 percent64 percent
0.80Large69 percent71 percent
1.00Very large62 percent76 percent
1.90Height34 percent91 percent

Look at the bottom row and hold onto it, because it recalibrates everything else. Height is one of the largest sex differences in the human body, and even there, if you pick one man and one woman at random, roughly one pair in eleven has the woman taller. That is what a very large sex difference looks like. Now look at the row for d of 0.20, where 92 percent of the two distributions occupy the same range, and note that many published claims about how men and women differ are describing effects that size or smaller.

In short: An effect size of 0.2 leaves 92 percent overlap; even the height difference, at 1.9, leaves a woman taller in about one randomly drawn pair in eleven.

What turns out to be small

In 2005 Hyde reviewed 46 meta-analyses covering the psychological literature on sex differences and reported that 78 percent of the effect sizes were 0.35 or below. She called this the gender similarities hypothesis, and it has held up reasonably well since. In the small-or-negligible column sit mathematics performance, reading comprehension, most measures of verbal ability, leadership effectiveness, self-esteem (around 0.21, favouring boys and men, and larger in adolescence than at any other age), most measures of moral reasoning, and the great majority of personality facets when measured one at a time.

Two of these deserve a note. The verbal advantage often attributed to girls is real in specific subskills, notably spelling and writing quality, where the effect can reach 0.3 to 0.5, but is close to zero for reading comprehension and vocabulary. And the mathematics result matters especially because the belief in a large male advantage persists among teachers and parents in the face of population-level data showing almost none.

What turns out to be large

The similarities hypothesis is not a claim that everything is the same, and the exceptions cluster in recognizable places.

  • Physical performance after puberty. Throwing velocity and distance produce some of the largest effect sizes in the psychological literature, above 1.5 and in some samples above 2.0. Grip strength, sprint speed and upper-body power are similarly separated. These effects are near zero before puberty and open sharply after it.
  • Vocational interests. Su, Rounds and Armstrong's 2009 meta-analysis of over half a million respondents found a difference of about 0.93 on a things-versus-people dimension: men on average more oriented to working with objects and systems, women to working with people. This is one of the largest and most replicated psychological differences known, and it is central to arguments about occupational segregation you will meet in Module 3.
  • Some sexual attitudes and behaviours. Meta-analyses find differences near 0.8 to 1.0 for frequency of masturbation and for permissive attitudes toward casual sex, while differences in most other sexual attitudes, including attitudes toward same-sex relationships, are small.
  • Physical aggression. John Archer's meta-analyses put direct physical aggression around 0.5 to 0.6, with the gap widest at the extreme tail: homicide perpetration is overwhelmingly male in every society with reliable records.
  • Three-dimensional mental rotation. Around 0.5 to 0.9 depending on the test and the timing conditions, the largest of the spatial abilities and the only cognitive measure in this list. Training reduces it substantially without eliminating it.

There is a second statistical fact that matters more than most of these means. In many traits, male variance is slightly larger than female variance, with variance ratios typically between 1.05 and 1.20. A ratio that modest is invisible in the middle of the distribution and consequential at the extremes, because small differences in spread produce larger differences in the ratio of the sexes several standard deviations out. This is one reason arguments about representation at the very top of a field cannot be settled by looking at average scores. It is also not universal: the variance ratio differs by country and by trait, which is itself a clue that it is not fixed.

Bottom line: The large differences concentrate in post-pubertal physical performance, vocational interests, some sexual behaviours, physical aggression and 3D rotation; almost everything else measured in psychology falls below 0.35.

The brain, and an honest argument about it

Here the literature is genuinely divided, and it is worth watching two competent camps disagree.

The first position: sex differences in the brain are real, measurable and clinically relevant. Adult male brains average roughly 10 percent larger in total volume, much of which tracks body size. Structural imaging finds reliable average differences in the relative size of particular regions and in the ratio of grey to white matter. Studies using the UK Biobank, with samples in the thousands, report these differences with narrow confidence intervals. Researchers in this camp, Larry Cahill prominent among them, argue that ignoring these differences has real costs in neurology and psychiatry, given that conditions from depression to Parkinson disease to autism have markedly different sex ratios and treatment responses.

The second position: the differences are real on average and do not sort brains into two types. Daphna Joel and colleagues, in a 2015 paper in the Proceedings of the National Academy of Sciences, took the regions showing the largest sex differences and asked how many individual brains were consistently at the male end or consistently at the female end across all of them. In their samples, only a small percentage were internally consistent; the overwhelming majority were mosaics of features more common in males and features more common in females. A 2021 review by Lise Eliot and colleagues went through the imaging literature region by region and argued that once brain size is controlled and multiple comparisons are handled properly, few differences survive replication.

Where does that leave you? With a distinction worth carrying. Group-level average differences in brain structure exist and are measurable. A reliable classification of individual brains into male and female types does not follow from them, any more than the height data license a rule that tall people are men. Both statements are supported by the same imaging datasets. When a popular article slides from the first to the second, it has made an inferential leap the data do not support.

The core of it: Average sex differences in brain structure are documented; the claim that individual brains fall into two coherent types is a separate claim, and the mosaic analyses tell against it.

What the biology does not establish

Suppose a difference is large and replicated. You still do not know why it exists, and this is where most public argument goes wrong. Four lines of evidence bear on causation, and none of them is decisive alone.

Prenatal hormones. Girls with congenital adrenal hyperplasia, exposed to elevated androgens before birth, show on average more interest in toys typically preferred by boys, in studies conducted since the 1990s. The samples are small, the girls know they have a medical condition, and parents know too, so the finding is suggestive rather than clean.

Non-human primates. Studies of vervet monkeys (2002) and rhesus macaques (2008) reported that males spent more time with wheeled toys and females with plush toys, in animals with no exposure to human gender socialisation. The samples were small and the toy categories have been criticized as arbitrary. Suggestive, contested, not settled.

Cross-cultural variation. If a difference were purely biological you would expect it to be roughly constant across societies. Many are not. The mathematics gap varies substantially by country and has closed in several. Conversely, some differences are larger in more gender-equal countries: Stoet and Geary reported in 2018 that the sex gap in graduation from science and engineering degrees is wider in countries scoring higher on gender equality indexes, a pattern often called the gender-equality paradox. That finding has been challenged on how the outcome variable was constructed, and the replication argument is live. Either way, cross-national variation shows the differences are not fixed.

Change over time. The mathematics gap in the United States narrowed to near zero within about three decades, which is far too fast for a genetic explanation and is the single strongest piece of evidence that measured differences respond to opportunity.

The honest summary is that biology and social environment both contribute, that their relative contribution differs by trait, and that for most traits nobody can currently put a number on the split. Beware anyone who can.

Common misconceptions

  • A statistically significant difference is a big difference. With seven million test scores, differences of 0.01 reach significance. Significance is about sampling error; the effect size is about magnitude.
  • Male and female brains are distinct types. Average regional differences are documented, but the mosaic analyses find few individual brains that are internally consistent across the most sex-differentiated regions.
  • The gender similarities hypothesis says there are no differences. Hyde's claim was that 78 percent of measured effects are 0.35 or below, and she has always named the exceptions, including throwing distance and some sexual behaviours.
  • Larger differences in more equal countries prove the differences are innate. The gender-equality paradox is a real reported pattern whose measurement is disputed, and several explanations for it, including economic security and cultural expectations about self-expression, do not require innateness.
  • If a trait is biological, it cannot change. The near closure of the United States mathematics gap in about thirty years happened far too fast for genetic change and is the clearest counterexample.

Where this leaves us

  • Cohen's d is a difference in means divided by the pooled standard deviation; 0.2 is small, 0.5 medium, 0.8 large, and adult height is about 1.9.
  • Even at 1.9, roughly one randomly drawn man-woman pair in eleven has the woman taller, which is the calibration to keep for every smaller effect.
  • Hyde's review of 46 meta-analyses found 78 percent of psychological sex differences at or below 0.35, including mathematics, reading comprehension and leadership effectiveness.
  • The large exceptions are post-pubertal physical performance, the things-versus-people interest dimension at about 0.93, some sexual behaviours, physical aggression and 3D mental rotation.
  • Average brain differences are real; the inference to two brain types is not supported, and causal attribution for behavioural differences remains genuinely open.

Sources

  1. Hyde, J. S. (2005). The gender similarities hypothesis. American Psychologist, 60(6), 581-592. American Psychological Association
  2. Hyde, J. S., Lindberg, S. M., Linn, M. C., Ellis, A. B., and Williams, C. C. (2008). Gender similarities characterize math performance. Science, 321(5888), 494-495. Science
  3. Joel, D., et al. (2015). Sex beyond the genitalia: The human brain mosaic. PNAS, 112(50), 15468-15473. PNAS
  4. Su, R., Rounds, J., and Armstrong, P. I. (2009). Men and things, women and people: A meta-analysis of sex differences in interests. Psychological Bulletin, 135(6), 859-884. American Psychological Association
  5. Office of Research on Women's Health. (n.d.). Sex as a biological variable. National Institutes of Health. orwh.od.nih.gov
Key terms
Cohen's d
A standardized effect size: the difference between two group means divided by the pooled standard deviation.
Distribution overlap
The proportion of two distributions occupying the same range of values; 92 percent at d of 0.2 and 34 percent at d of 1.9.
Gender similarities hypothesis
Hyde's summary that 78 percent of measured psychological sex differences are 0.35 or smaller.
Variance ratio
The ratio of male to female variance in a trait, typically 1.05 to 1.20, which matters chiefly at the tails of a distribution.
Things-people dimension
The vocational interest axis on which the sexes differ by about 0.93, the largest replicated psychological difference.
Brain mosaic
Joel and colleagues' finding that most individual brains combine features more common in males with features more common in females.
Gender-equality paradox
The reported pattern that sex gaps in science degree completion are wider in countries scoring higher on gender equality, whose measurement is disputed.
Congenital adrenal hyperplasia
A condition involving prenatal androgen exposure in XX fetuses, used in research on hormonal influences on behaviour.

Becoming Gendered: Socialisation from the Delivery Room Onward

  • Trace the developmental sequence by which children acquire gender categories, preferences and rigidity.
  • Weigh the evidence on how much parents, peers, teachers and children themselves contribute to gender typing.
  • Explain why children are best modelled as active seekers of gender information rather than passive recipients.

The same baby, three names

In 1975 three psychologists, Carol Seavey, Phyllis Katz and Sue Rosenberg Zalk, sat adult volunteers in a room with a three-month-old infant dressed in a yellow jumpsuit. Some were told the baby was a girl. Some were told it was a boy. Some were told nothing. Three toys were within reach: a rubber ring, a small football, and a doll. The adults who thought they were holding a girl reached for the doll. The adults who thought they were holding a boy reached for the football. The adults given no label spent noticeably more time examining the baby, apparently trying to work out which it was, and several asked outright. It was the same infant every time.

A year earlier, Jeffrey Rubin and colleagues had interviewed thirty sets of parents within twenty-four hours of a birth. Hospital records showed no meaningful differences in the newborns' length, weight or Apgar scores. The parents of daughters described their babies as softer, finer-featured and more delicate. The parents of sons described theirs as firmer, larger and better coordinated. Fathers, who in 1974 had typically observed the infant only through nursery glass, produced the most sharply differentiated descriptions.

Gender socialisation, in other words, is already running on day one, and the adults doing it are not aware that they are doing it. This lesson traces what happens over the following six years, and then asks the harder question of how much of the outcome the adults actually cause.

Key idea: Differential treatment by gender begins before an infant has any behaviour to respond to, and the adults producing it sincerely report that they are responding to the child.

The developmental sequence

Children do not receive gender in one lump. They assemble it, and the assembly follows a fairly predictable order across societies studied.

AgeWhat appearsHow it is tested
6 to 12 monthsDiscrimination of male and female faces and voicesHabituation: infants look longer at a face from a new category
18 to 24 monthsFirst sex-typed toy preferences; some gender labellingFree-play observation; pointing tasks
2 to 3 yearsGender identity: reliable self-labelling as a boy or a girlDirect question, sorting photographs
3 to 5 yearsGender stability over time; strong same-sex peer preferenceAsking whether a boy will grow into a man or a woman
5 to 7 yearsGender constancy across appearance; peak stereotype rigidityAsking whether a boy in a dress is still a boy
7 to 10 yearsRigidity relaxes; stereotypes become known but negotiableJudgments about whether violations are wrong or merely unusual

Lawrence Kohlberg proposed the identity, stability and constancy sequence in 1966, and it has survived remarkably well. The step that surprises students is the last one in the middle rows. Between about five and seven, children are typically at their most rigid: this is the age at which a child will insist that a woman cannot be a firefighter, sometimes while their own mother is a firefighter. The rigidity is not learned prejudice so much as a cognitive stage. Children who have just worked out that a category is permanent tend to over-apply it, and the same over-application shows up in their reasoning about animals and objects.

The peer sorting is equally striking and much less discussed. Eleanor Maccoby's observational work found that by about age four and a half, children were spending roughly three times as much play time with same-sex as other-sex peers, and that by about six and a half the ratio had climbed to something like eleven to one. Nobody organizes this. It happens in playgrounds where adults are actively encouraging mixed play, and it happens in a wide range of societies. Maccoby argued that the two segregated peer groups then develop different interaction styles, more turn-taking and reciprocal talk in one, more direct commands and rough-and-tumble in the other, and that children spend years practising in these separate cultures.

So what?: Gender develops in a sequence, peaks in rigidity around ages five to seven for cognitive rather than attitudinal reasons, and involves powerful voluntary peer segregation that adults do not create and cannot easily undo.

Three theories, and what each explains

Social learning theory, in the tradition of Albert Bandura, holds that children acquire gendered behaviour through observation, imitation and differential reinforcement. It explains a great deal about why children in different societies acquire different content. Its weakness is timing: children often become more stereotyped than the adults around them, which reinforcement alone struggles to explain.

Cognitive-developmental theory, Kohlberg's, holds that once children categorize themselves, they actively seek out information about what their category does and are motivated to conform to it. This explains the over-rigid five-year-old with the firefighter mother. Its weakness is that sex-typed toy preferences appear well before full gender constancy.

Gender schema theory, developed by Sandra Bem and by Carol Martin and Charles Halverson in the early 1980s, splits the difference. Children build mental schemas of what goes with each gender, and then use those schemas to filter attention and memory. Show a child a photograph of a woman using a hammer, and a week later a substantial share of children will recall a man. This is not a failure to look; it is memory reconstructing the image toward the schema. Schema theory explains both the content-specificity that social learning captures and the self-driven quality that Kohlberg captured, and it is the framework most current research uses.

Martin and Ruble later described the young child as a gender detective: actively hunting for cues about which things belong to which category, and treating adult behaviour as evidence rather than instruction. That reframing matters for the next section.

How much do parents actually do?

Here is where the honest answer is more interesting than the intuitive one. Hugh Lytton and David Romney published a meta-analysis in 1991 covering 172 studies of how parents treat sons and daughters differently. Across most of what they measured, including warmth, responsiveness, restrictiveness, discipline and encouragement of achievement, the differences were small and often not distinguishable from zero. One area stood out clearly: encouragement of sex-typed activities. Parents buy different toys, praise different play, and steer children toward different activities, and they do this substantially.

So the popular picture of parents systematically socializing children into different personalities is not well supported, while the narrower claim about activities and interests is. That narrower channel may matter more than it sounds, since interests drive practice, practice drives skill, and skill drives later choices. The spatial-rotation difference discussed in the previous lesson is a plausible candidate: boys receive more construction toys, spend more hours manipulating objects in three dimensions, and mental rotation is trainable.

Two other adult channels are documented. First, expectations. Studies following American families found that parents estimated sons' mathematics ability higher than daughters' at identical school grades, and that those estimates predicted the children's own self-assessments better than the grades did. Second, language. Adults use more emotion words and more references to sadness with girls, and more references to anger with boys, in ordinary storybook reading. Neither of these is a deliberate curriculum.

The largest source of enforcement, though, is not adults. It is other children, and the enforcement is asymmetric. Cross-culturally, boys who play with toys coded as feminine face sharper peer sanction than girls who play with toys coded as masculine. A girl in a superhero costume is unremarkable in most Western primary schools; a boy in a princess dress is not. That asymmetry runs through the whole course, and Lesson 4 takes it up directly.

Worth holding on to: Parents differ mainly in the activities they encourage rather than in warmth or discipline, adult expectations shape children's self-assessments independent of grades, and peers do most of the day-to-day enforcement, disproportionately against boys.

What varies, and what that tells you

If gender content were biologically fixed you would expect the same tasks to be coded the same way everywhere. They are not. Market trading is women's work in much of West Africa and men's work in much of South Asia. Agricultural fieldwork is heavily female in parts of sub-Saharan Africa and heavily male in parts of Latin America. Computer programming was coded as clerical and largely female work in the 1950s and 1960s in Britain and the United States, and the share of American computer science bachelor's degrees going to women peaked around 1984 near 37 percent before falling sharply, which is not a trajectory any biological account predicts.

What does look near-universal is the existence of a gender division of some kind, along with the greater involvement of women in direct infant care and of men in activities involving physical risk. The lesson to draw is specific: the fact of division is close to universal, the content of the division is highly variable, and confusing the two produces most of the bad arguments in this area.

Common misconceptions

  • Children are blank slates that adults write gender onto. Children actively seek gender information and frequently become more rigid than the adults around them, which passive-transmission accounts cannot explain.
  • A five-year-old insisting women cannot be doctors has learned sexism at home. Peak rigidity around ages five to seven is a documented cognitive stage that appears even in children of counter-stereotypical parents.
  • Parents treat sons and daughters completely differently. The meta-analytic evidence finds small differences in warmth and discipline, and a clear difference in one channel: encouragement of sex-typed activities.
  • Same-sex play preference is imposed by adults. It emerges strongly in settings where adults actively encourage mixed play, reaching roughly eleven to one by age six and a half.
  • Gender-neutral parenting has been shown to eliminate gender typing. No study has demonstrated that, partly because no child is raised outside a wider society, and honest researchers say so.

Putting it together

  • The Baby X studies and newborn-description research show differential treatment beginning before any behaviour exists to respond to.
  • Development runs from face discrimination in infancy to self-labelling around age two or three, stability, then constancy and peak rigidity at five to seven, after which stereotypes loosen.
  • Gender schema theory explains most of the evidence: children build schemas, then filter attention and even memory through them.
  • Parents differ mostly in the activities they encourage; peers enforce more, and enforce boys' boundaries harder than girls'.
  • The existence of a gender division of labour is near-universal; its content varies enormously, as the rise and fall of women in computing shows.

Sources

  1. Seavey, C. A., Katz, P. A., and Zalk, S. R. (1975). Baby X: The effect of gender labels on adult responses to infants. Sex Roles, 1(2), 103-109. Springer
  2. Lytton, H., and Romney, D. M. (1991). Parents' differential socialization of boys and girls: A meta-analysis. Psychological Bulletin, 109(2), 267-296. American Psychological Association
  3. Maccoby, E. E. (1998). The two sexes: Growing up apart, coming together. Harvard University Press. Harvard University Press
  4. Martin, C. L., and Ruble, D. N. (2004). Children's search for gender cues. Current Directions in Psychological Science, 13(2), 67-70. SAGE Journals
  5. Encyclopaedia Britannica. (n.d.). Gender role. britannica.com
Key terms
Baby X studies
Experiments in which adults respond to the same infant differently depending on the gender label they are given.
Gender constancy
The understanding, typically consolidated between five and seven, that gender persists despite changes in appearance.
Peak rigidity
The developmental window around ages five to seven when children apply gender rules most inflexibly, for cognitive rather than attitudinal reasons.
Gender schema
A mental structure organizing what belongs to each gender, which then filters what a child attends to and remembers.
Gender detective
Martin and Ruble's description of the child as an active seeker of gender cues rather than a passive recipient of instruction.
Sex segregation in play
Children's voluntary preference for same-sex playmates, reaching roughly eleven to one by about age six and a half.
Differential encouragement
The one channel in which parents reliably differ by child's sex: which activities and toys they promote.
Asymmetric enforcement
The pattern by which boys face sharper peer sanction for gender-atypical behaviour than girls do.

Module 2: Masculinities and Desire

Men and masculinity as objects of study rather than the unmarked default, and sexuality measured along its three dimensions, including what the research on the origins of orientation can and cannot support.

Masculinities: Connell's Framework and the Measured Facts of Men's Lives

  • Explain Connell's distinction between hegemonic, complicit, subordinated and marginalized masculinities, and identify the two common misreadings of the term hegemonic.
  • Summarize the documented sex differences in mortality, incarceration and educational attainment, with sources.
  • Weigh norm-based and structural explanations for male excess mortality and state what evidence would separate them.

Nine in ten

In 2022 the Bureau of Labor Statistics counted 5,486 fatal work injuries in the United States. More than nine in ten of the people killed were men. That figure does not appear because men are more careless. It appears because roofing, logging, commercial fishing, refuse collection, structural iron work and long-haul trucking are between eighty and ninety-eight percent male, and those are the occupations that kill people. The number is a fact about the division of labour, expressed in deaths.

For most of the twentieth century, men were the unmarked category in social research. Studies of workers meant studies of men; studies of gender meant studies of women. The field that emerged in the 1980s to study men as gendered did two things at once. It asked what masculinity demands of the men who live under it, and it insisted that masculinity is not one thing. This lesson takes both questions seriously and then puts numbers on the outcomes.

Key idea: Studying men as gendered subjects means asking what masculinity requires and what it costs, without treating men either as the default human or as a single undifferentiated group.

Connell's four positions

Raewyn Connell's Masculinities (1995) is the field's organizing text, and its central move is pluralization. In any society at any moment there are multiple masculinities standing in relation to one another and to women, and they are not equal in status.

  • Hegemonic masculinity is the pattern that currently legitimates the gender order: the version most honoured, the one used as the standard against which other men are measured. In the contemporary United States it involves some combination of physical competence, heterosexual success, emotional control, occupational achievement and independence.
  • Complicit masculinity describes the large majority of men who do not embody the hegemonic pattern, and may not try to, but who benefit from the general arrangement without directly enforcing it.
  • Subordinated masculinity describes patterns actively positioned below, historically and most sharply gay masculinity, along with men marked as effeminate.
  • Marginalized masculinity describes patterns whose authority is limited by race, class or immigration status, where a man may exemplify hegemonic traits and still be denied the standing they normally confer.

Two misreadings are worth heading off, because both are extremely common in student writing. First, hegemonic does not mean statistically typical. Connell's point is that very few men actually meet the standard, which is part of how it works: it operates as an aspiration that most men fail. Second, hegemonic masculinity is not a synonym for toxic masculinity. Connell has been explicit that the concept describes a position in a structure, not a personality defect, and she has objected to readings that turn a relational analysis into a list of bad traits.

The point: Hegemonic masculinity names the most legitimated pattern in a hierarchy of masculinities, not the most common one and not a character flaw.

The rules, as boys describe them

In 1976 Robert Brannon compressed American masculine norms into four rules that have proved uncomfortably durable: no sissy stuff, be a big wheel, be a sturdy oak, and give them hell. Avoid anything coded feminine. Achieve status. Show no vulnerability. Take risks and be aggressive when required.

Since 2003 the standard research instrument has been James Mahalik's Conformity to Masculine Norms Inventory, which scores respondents on distinct subscales including emotional control, self-reliance, risk-taking, primacy of work, dominance, playboy, power over women, violence, winning and heterosexual self-presentation. Treating these separately turns out to matter enormously.

Y. Joel Wong and colleagues published a meta-analysis in 2017 covering 78 samples and more than 19,000 respondents. Overall conformity to masculine norms was associated with poorer mental health and with lower likelihood of seeking psychological help. But the association was carried by specific subscales. Self-reliance, playboy and power over women were consistently associated with worse outcomes. Primacy of work was essentially unrelated. That result is the whole reason to use a multi-scale instrument: the finding is not that masculinity is bad for men, it is that particular norms, above all the injunction against needing anyone, predict measurable harm.

What the numbers show

Outcome (United States, recent years)Male share or gapSource
Deaths by suicideAbout four in five; male rate roughly four times femaleNational Center for Health Statistics
Fatal work injuriesMore than 90 percentBureau of Labor Statistics
Homicide victimizationRoughly three in fourFederal Bureau of Investigation and BJS
State and federal prisonersAround 93 percentBureau of Justice Statistics
Life expectancy at birthMen roughly five to six years lowerNational Center for Health Statistics
Bachelor's degrees conferredMen about 42 percent, a minority since the early 1980sNational Center for Education Statistics
Referral to special education and school suspensionBoys roughly two to three times girlsDepartment of Education civil rights data

Put that table next to the table you will meet in Module 3, on earnings, authority and unpaid care, and you have the shape of the honest picture: on measures of income, wealth, senior authority and freedom from sexual violence, the disparities run one way; on mortality, incarceration and educational attainment, they run the other. Any account that reports only one column is describing half of a distribution.

Bottom line: Men die earlier, are killed at work and by violence more often, are imprisoned far more, and now earn a minority of degrees, while holding most senior positions and higher lifetime earnings; both columns come from the same statistical agencies.

Two explanations, and what would separate them

Take suicide, where the pattern is sharp enough to test. In the United States men die by suicide roughly four times as often as women. Women report attempting suicide more often than men, in survey after survey. This combination has been called the gender paradox of suicide, and there are two serious explanations.

The norms account says that masculine socialisation restricts help-seeking and emotional disclosure. Men are less likely to have a primary care physician, less likely to attend a mental health appointment, and more likely to be diagnosed late. Wong's meta-analytic finding on self-reliance supports this: the norm most strongly associated with poor outcomes is precisely the one prohibiting the request for help.

The means account says the difference is largely about method lethality. In the United States, firearms are used in roughly half of male suicides and a substantially smaller share of female ones, and firearm attempts are lethal in the large majority of cases while overdoses, the most common method in female attempts, are not. On this account, comparable levels of distress and intent produce different death rates because of what is within reach.

What would separate them? Comparative data. If lethality were the whole story, the male-to-female ratio should track firearm availability across countries, and it partly does: the ratio is far higher in the United States than in many countries with tight firearm regulation. If norms were the whole story, the ratio should track measures of masculine norm endorsement across cultures. Neither prediction holds cleanly. China was for years the major exception, with female suicide rates exceeding male, associated with pesticide access in rural areas, and rates fell substantially as urbanization and pesticide regulation changed. That exception is instructive: it shows means mattering enormously, without showing norms to be irrelevant. The defensible conclusion is that both operate, and that their relative weight varies by country.

Where masculinity has already moved

Norms are not static, and the clearest evidence is fatherhood. American fathers' time on child care roughly tripled between 1965 and the mid-2010s, from around two and a half hours a week to around eight, according to time-use data compiled by Pew Research. Housework roughly doubled. Both remain below mothers' hours, and the trend line is unmistakable.

Policy moves it too, and the natural experiments are clean. Iceland's 2000 parental leave reform allocated three months to the mother, three months to the father and three months to share, with the individual months non-transferable. Father uptake went from marginal to near-universal within a few years. When Quebec introduced its own parental insurance plan in 2006 with a father-specific benefit, the share of eligible fathers taking leave rose from roughly one in five to more than four in five, while the rest of Canada, under the older transferable scheme, barely moved. A norm that looked like a stable feature of male identity turned out to be highly responsive to whether the leave was labelled as his and lost if unused.

Worth holding on to: When leave is non-transferable and individually assigned, fathers' uptake changes by tens of percentage points within a few years, which tells you the norm was never fixed.

Common misconceptions

  • Hegemonic masculinity means the way most men behave. It names the most legitimated pattern, which most men do not meet; the gap between standard and reality is part of the mechanism.
  • Hegemonic and toxic masculinity are the same concept. Connell's term describes a structural position; toxic masculinity is a later popular coinage naming a set of harmful behaviours, and she has objected to the conflation.
  • Masculinity as such damages men's mental health. The meta-analytic evidence points at specific norms, above all self-reliance, playboy and power over women, while others show no association.
  • Men's higher suicide death rate means men are more suicidal. Women report more attempts; the death-rate gap involves method lethality as well as help-seeking, and the balance varies by country.
  • Men's overrepresentation in workplace deaths shows greater recklessness. It mostly reflects occupational sorting into the specific industries where fatal injuries occur.

What to carry forward

  • Connell's framework treats masculinities as plural and hierarchical: hegemonic, complicit, subordinated and marginalized.
  • Brannon's four rules and Mahalik's multi-scale inventory operationalize masculine norms; Wong's 2017 meta-analysis shows the harm is carried by particular subscales, not by conformity in general.
  • Men account for over 90 percent of fatal work injuries, about four in five suicides, around 93 percent of prisoners, and a minority of bachelor's degrees.
  • The suicide gap involves both restricted help-seeking and method lethality, and cross-national comparison shows neither alone explains it.
  • Fathers' care time has roughly tripled since 1965, and non-transferable leave in Iceland and Quebec moved paternal uptake dramatically within a few years.

Sources

  1. Connell, R. W. (1995). Masculinities. University of California Press. University of California Press
  2. Bureau of Labor Statistics. (n.d.). Injuries, illnesses, and fatalities. U.S. Department of Labor. bls.gov
  3. National Center for Health Statistics. (n.d.). FastStats: Suicide and self-harm injury. Centers for Disease Control and Prevention. cdc.gov
  4. Wong, Y. J., Ho, M.-H. R., Wang, S.-Y., and Miller, I. S. K. (2017). Meta-analyses of the relationship between conformity to masculine norms and mental health-related outcomes. Journal of Counseling Psychology, 64(1), 80-93. American Psychological Association
  5. Pew Research Center. (n.d.). Parenting and family life research. pewresearch.org
Key terms
Hegemonic masculinity
Connell's term for the currently most legitimated pattern of masculinity, which most men do not embody but which sets the standard.
Complicit masculinity
The position of men who do not enact the hegemonic pattern but benefit from the gender arrangement it sustains.
Marginalized masculinity
Masculinity whose authority is limited by race, class or immigration status even when hegemonic traits are present.
Conformity to Masculine Norms Inventory
Mahalik's instrument scoring distinct masculine norms separately, including self-reliance, emotional control and playboy.
Gender paradox of suicide
The pattern of higher reported attempts among women and higher death rates among men.
Method lethality
The probability that a given suicide method results in death, a major contributor to sex differences in suicide mortality.
Non-transferable leave
Parental leave months reserved for a specific parent and forfeited if unused, the design feature associated with large increases in fathers' uptake.
Occupational sorting
The distribution of workers across industries, which accounts for most of the male share of fatal work injuries.

Sexuality: Scripts, Practice, and What a Survey Can Reach

  • Explain why national surveys produce arithmetically impossible partner counts and what that reveals about self-report.
  • Apply sexual script theory at the cultural, interpersonal and intrapsychic levels to an ordinary encounter.
  • State which findings about sexual behaviour replicate across national surveys and which are artefacts of measurement.

An impossible number

Britain's third National Survey of Sexual Attitudes and Lifestyles, fielded between 2010 and 2012 with more than 15,000 respondents, reported that men had a mean of about twelve lifetime opposite-sex partners and women about eight. American surveys find the same shape. So do surveys in France, Australia and Norway.

Now do the arithmetic. In a closed population, every opposite-sex partnership adds exactly one to a man's count and one to a woman's count. The totals must be equal, and with roughly equal numbers of men and women, the means must be equal too. They never are. Men consistently report between a third and twice as many partners as women.

Something is wrong, and working out what is wrong teaches more about sexuality research than any list of findings. Four candidate explanations exist, and you can test them.

Key idea: The partner-count discrepancy is a mathematical impossibility that appears in every national survey, which makes it the cleanest available window onto how self-report about sex actually behaves.

Four suspects, and the evidence on each

Suspect one: the population is not closed. Men might be counting partners outside the survey frame, including partners abroad or people paid for sex, who are underrepresented in household samples. This is real and contributes something, but modelling suggests it cannot account for a gap of this size in countries where commercial sex is relatively uncommon.

Suspect two: different counting rules. Ask people how they arrived at their number and men are more likely to estimate, women more likely to enumerate. Estimation rounds upward and rounds to salient numbers, which is why survey responses cluster on multiples of five and ten. This is well documented and contributes substantially.

Suspect three: different definitions of what counts. In a 1999 study published in JAMA, Stephanie Sanders and June Reinisch asked American college students whether they would say they had had sex if the encounter involved oral-genital contact. Roughly 60 percent said they would not. If men and women apply different thresholds for what enters the count, the totals diverge without anyone lying.

Suspect four: motivated reporting. Here is the experiment that settles the question of whether social pressure is operating. In 2003 Michele Alexander and Terri Fisher assigned undergraduates to three conditions. In one, respondents believed their answers might be seen by the researcher. In another, answers were anonymous. In the third, the respondent was connected to what they were told was a functioning polygraph, a technique called the bogus pipeline. Women's reported number of partners was lowest in the exposure condition, higher when anonymous, and highest of all when they believed they could be caught lying. Men's numbers moved in the opposite direction. Under the bogus pipeline, the sex difference shrank sharply.

The lesson generalizes beyond partner counts. When you read any statistic about sexual behaviour, ask what the respondent thought the consequence of answering honestly would be. That question is not cynicism; it is measurement.

The upshot: The partner gap is produced by a combination of sampling frames, estimation versus enumeration, differing definitions of sex, and social desirability, and the bogus pipeline experiment shows the last of these is real and gendered.

Sexual scripts

In 1973 John Gagnon and William Simon proposed that sexual conduct is scripted in the way that social conduct generally is: people do not simply act on drives, they follow learned sequences about who does what, in what order, meaning what. They distinguished three levels, and the distinction is genuinely useful.

  • Cultural scenarios are the society-wide narratives: what a date is, what counts as seduction, who is expected to initiate, what a first time is supposed to feel like.
  • Interpersonal scripts are what two specific people negotiate, mostly without discussing it, in a given encounter.
  • Intrapsychic scripts are the internal fantasies and self-narratives through which a person interprets their own desire.

Script theory does real explanatory work on the sexual double standard. If the cultural scenario assigns men the role of pursuer and treats accumulation as achievement, while assigning women the role of gatekeeper and treating accumulation as damage, then the same behaviour carries opposite meanings depending on who performs it. That asymmetry is measurable: vignette experiments in which the same described behaviour is attributed to a male or female character reliably produce different judgments of the character, though the size of the effect has fallen in Western samples since the 1980s and varies considerably by the specific behaviour described.

Scripts also explain something the biological literature cannot: why the coital sequence in Western societies is so standardized, and why it ends where it ends. That question leads directly to the next section.

The orgasm gap and what it rules out

In a 2018 analysis of more than 50,000 American adults, David Frederick and colleagues asked how often respondents usually orgasmed with a familiar partner. The pattern is worth reading carefully.

GroupUsually orgasm (approximate)
Heterosexual men95 percent
Gay men89 percent
Bisexual men88 percent
Lesbian women86 percent
Bisexual women66 percent
Heterosexual women65 percent

Look at the fourth and sixth rows together. Lesbian women report orgasm at rates roughly twenty points higher than heterosexual women. Whatever explains the heterosexual gap therefore cannot be a fact about female anatomy or female physiology, because the same anatomy is present in both rows. The candidate explanations that survive are about practice and script: the range of activities included in an encounter, how the encounter is understood to end, and how comfortably partners communicate. Studies of what actually happens in encounters support this, finding that the activities most associated with women's orgasm are more frequent in same-sex encounters.

This is a good example of a general method. When a group difference could have a biological or a social explanation, look for a comparison in which the biology is held constant and the social arrangement varies.

What matters here: The twenty-point difference between lesbian and heterosexual women rules out anatomy as the explanation for the heterosexual orgasm gap and points to practice and script instead.

Desire, and a model that changed clinical practice

The linear model of sexual response inherited from Masters and Johnson ran desire, arousal, plateau, orgasm, resolution, with desire first. In 2000 the Canadian clinician Rosemary Basson proposed an alternative circular model in which many people, disproportionately women in long-term relationships, begin an encounter from emotional intimacy or willingness rather than spontaneous desire, and experience desire after arousal begins. The distinction between spontaneous and responsive desire has since become standard in sex therapy.

The clinical consequence was immediate. Under the linear model, a woman who did not experience spontaneous desire met criteria for a disorder. Under the responsive model, she may simply have a different and equally common pattern. This is a rare case where a conceptual revision measurably changed who was diagnosed.

Asexuality sits nearby and is genuinely distinct. Analysing the 1994 British survey data, Anthony Bogaert found that roughly one percent of respondents reported never having felt sexual attraction to anyone. Asexuality is not the same as low desire, celibacy or a disorder, and the current diagnostic manuals explicitly exclude people who identify as asexual from the relevant diagnoses.

Common misconceptions

  • Men really do have more partners than women. In a closed population that is arithmetically impossible; the reported gap is a measurement artefact with several identified sources.
  • Anonymous surveys eliminate social desirability bias. The bogus pipeline experiments show reported numbers still shift when respondents believe dishonesty could be detected.
  • The orgasm gap reflects female anatomy. Lesbian women report rates about twenty points higher than heterosexual women, which the anatomical account cannot accommodate.
  • Low spontaneous desire is a dysfunction. Responsive desire, in which desire follows arousal rather than preceding it, is a common and normal pattern that Basson's model made visible.
  • Asexuality is just a low sex drive. It is defined by absence of sexual attraction, is reported by around one percent of respondents, and is distinguished from disorder in current diagnostic criteria.

Summing up

  • Every national survey produces a partner-count gap that cannot be true, and diagnosing it teaches you how sexual self-report works.
  • Four mechanisms contribute: sampling frames, estimation versus enumeration, differing definitions of sex, and motivated reporting demonstrated by the bogus pipeline.
  • Gagnon and Simon's scripts operate at cultural, interpersonal and intrapsychic levels, and explain the double standard better than drive-based accounts.
  • Frederick's 2018 data show a roughly thirty-point heterosexual orgasm gap and a twenty-point difference between lesbian and heterosexual women, which points to practice rather than anatomy.
  • Basson's responsive desire model changed clinical thresholds, and asexuality, reported by about one percent, is a separate category from low desire.

Sources

  1. National Survey of Sexual Attitudes and Lifestyles. (n.d.). Natsal study overview. natsal.ac.uk
  2. Frederick, D. A., St. John, H. K., Garcia, J. R., and Lloyd, E. A. (2018). Differences in orgasm frequency among gay, lesbian, bisexual, and heterosexual men and women in a US national sample. Archives of Sexual Behavior, 47(1), 273-288. Springer
  3. The Kinsey Institute. (n.d.). Research and data on human sexuality. Indiana University. kinseyinstitute.org
  4. Society for the Scientific Study of Sexuality. (n.d.). About the field. sexscience.org
  5. National Center for Health Statistics. (n.d.). National Survey of Family Growth: Key statistics. Centers for Disease Control and Prevention. cdc.gov
Key terms
Bogus pipeline
An experimental technique in which respondents believe a device can detect dishonesty, used to estimate social desirability bias.
Sexual script
Gagnon and Simon's concept of learned sequences governing sexual conduct at cultural, interpersonal and intrapsychic levels.
Cultural scenario
The society-wide narrative about how a sexual encounter is supposed to proceed and what its parts mean.
Sexual double standard
The pattern by which identical sexual behaviour is evaluated differently depending on the actor's gender.
Orgasm gap
The roughly thirty-point difference in usual orgasm frequency between heterosexual men and heterosexual women.
Responsive desire
Basson's model in which desire follows arousal and intimacy rather than arising spontaneously beforehand.
Asexuality
The absence of sexual attraction, reported by about one percent of respondents and distinguished from low desire or disorder.
Enumeration versus estimation
Two counting strategies for reporting partner numbers, the second of which rounds upward and produces clustering at multiples of five.

Where Orientation Comes From: Reading the Evidence

  • Summarize what genome-wide association and twin studies establish about the heritability of same-sex sexual behaviour.
  • Explain the fraternal birth order effect and the design feature that makes it evidence for a prenatal mechanism.
  • Evaluate claims about the origins of orientation, including claims about fluidity and about attempts to change it.

Half a million genomes

In August 2019 a team led by Andrea Ganna published in Science the largest study of its kind: a genome-wide association analysis of 477,522 people from the UK Biobank and the customer database of a consumer genetics company, testing whether reported same-sex sexual behaviour was associated with variation anywhere in the genome.

They found five loci reaching genome-wide significance. Together, those five explained less than one percent of variation in the behaviour. Taking all measured common variants together, the analysis attributed somewhere between about 8 and 25 percent of variation to genetics in aggregate. The authors stated in the paper, and repeated in every interview, that the results cannot be used to predict anyone's orientation.

That is a strange-looking result if you expect a study to either find a gay gene or find nothing. It is exactly what geneticists expect for a complex human trait: many variants of tiny effect, substantial aggregate heritability, and no predictive power at the individual level. Height works the same way, and nobody doubts height is heritable. This lesson works through what the whole body of evidence supports, and, just as importantly, what it does not.

Key idea: The largest genetic study of same-sex sexual behaviour found no gene for orientation, moderate aggregate heritability, and no individual predictive power, which is the standard signature of a complex polygenic trait.

Twins, and a lesson about sampling

Twin studies estimate heritability by comparing identical and fraternal pairs. In 1991 Michael Bailey and Richard Pillard reported that among identical twin pairs where one man was gay, about 52 percent of the co-twins were also gay, against about 22 percent for fraternal twins. The finding was widely reported as strong evidence for a genetic basis.

Then the sampling problem surfaced. Bailey and Pillard had recruited through advertisements in gay publications, which oversamples pairs where both twins are gay and out, because a man with a gay twin is more likely to see the advertisement and respond. Later population-based studies, drawing from national twin registries rather than volunteers, produced lower numbers. A Swedish registry study by Niklas Långström and colleagues in 2010, covering more than 3,800 twin pairs, estimated heritability at roughly 34 to 39 percent for men and 18 to 19 percent for women, with most of the remaining variance attributed to non-shared environment rather than to family environment.

Two conclusions follow, and both matter. First, orientation is meaningfully heritable but not determined. Second, and this is the transferable lesson, the sampling frame can move an estimate by more than the effect being measured. Whenever a striking behavioural genetic finding appears, ask how the participants were recruited before you ask what the number means.

The point: Volunteer-recruited twin studies put concordance far higher than population registries do; the registry estimates of roughly 34 to 39 percent heritability in men are the defensible figures.

The fraternal birth order effect

Of all the findings in this area, one is unusually robust. Ray Blanchard and colleagues, working from the 1990s onward, documented that a man's odds of being homosexual rise with the number of older biological brothers born to the same mother, by roughly a third with each one. The effect has replicated across many samples in several countries.

The obvious objection is social: perhaps growing up with older brothers changes a boy's development. Anthony Bogaert tested that directly in 2006 by comparing men raised with older brothers to whom they were not biologically related, and men with biological older brothers they had not been raised with. The effect tracked the biological older brothers, not the ones in the household. A boy raised with three older stepbrothers showed no elevated likelihood; a boy raised apart from an older biological brother did.

That design points at something happening during gestation rather than during childhood. The leading hypothesis is maternal immunization: with each male pregnancy, the mother's immune system may develop an increasing response to proteins expressed by the male fetus, altering aspects of prenatal development. A 2018 study in the Proceedings of the National Academy of Sciences found higher levels of antibodies to a Y-linked protein in mothers of gay sons, particularly those with older sons, which is a direct test of the mechanism rather than of the pattern. The hypothesis is not confirmed, and the effect explains only a modest share of cases in the population, perhaps around 15 to 30 percent of gay men in some estimates. It is nevertheless the clearest evidence available for a prenatal biological contribution.

What has been ruled out

Some explanations have been tested repeatedly and have failed.

  • Parenting style. The mid-century clinical claims about distant fathers and overbearing mothers came from psychiatric patient samples with no comparison group. Studies with proper controls have not supported them.
  • Recruitment or exposure. Children raised by same-sex parents do not show elevated rates of same-sex orientation at levels that would support a transmission account, and the large-scale studies of their outcomes generally find no meaningful differences in psychological adjustment once family stability and income are accounted for. The one high-profile study reporting otherwise, published by Mark Regnerus in 2012, was criticized on the grounds that its comparison groups conflated family structure with family disruption, and its own journal published an internal audit of the review process.
  • Choice. No study has documented anyone changing orientation deliberately, and self-reports of choice are rare even among people with strong incentives to report it.

What has not been ruled out is a substantial role for environment in a broad sense. The registry twin studies attribute most variance to non-shared environment, which is a category that includes prenatal conditions, differential experiences, and, importantly, measurement error. Non-shared environment is a residual, not an explanation.

Fluidity, and why it is not the opposite of innateness

Lisa Diamond followed a sample of about 80 non-heterosexual women for more than a decade, interviewing them repeatedly, and found that a substantial share changed how they described their orientation at least once, often more than once, and that changes ran in several directions rather than converging on any endpoint. Larger longitudinal datasets have found the same pattern, with identity change more common among women than men and more common in adolescence and early adulthood than later.

People frequently read fluidity as undermining biological accounts. It does not, for a simple reason: heritability and stability are different properties. Body weight is highly heritable and changes throughout life. What Diamond's work establishes is that orientation is better modelled as a distribution of attractions that can shift, particularly for women, than as a fixed slot. It also implies that a survey asking about identity at one moment is taking a snapshot of something that moves.

Remember: Fluidity and heritability are compatible; a trait can be substantially heritable and still change within a person's lifetime.

Attempts to change orientation

The evidence here is unusually clear for a contested subject. The American Psychological Association's 2009 task force reviewed the peer-reviewed literature on sexual orientation change efforts and concluded that studies claiming success were methodologically weak, typically lacking control groups, relying on self-report from participants with strong incentives, and following participants only briefly. The task force found no reliable evidence of orientation change and documented reports of harm including depression, anxiety and suicidality.

The most striking single episode belongs to the researcher who supplied the best-known supporting study. Robert Spitzer, the psychiatrist who had led the removal of homosexuality from the diagnostic manual in 1973, published an interview study in 2001 reporting that some highly motivated participants described change. In 2012 he wrote to the journal that had published it retracting his interpretation, on the grounds that he had no way to verify the self-reports, and apologized to people who had undergone such treatment on the strength of his paper. It is a rare and instructive example of a scientist publicly withdrawing a finding that had become politically useful to others.

One clarification worth keeping. Rights arguments do not rest on immutability. Religion is not immutable and is protected everywhere; the legal and moral case against discrimination has never actually depended on establishing that a trait cannot change.

Common misconceptions

  • Science found a gay gene. The 2019 study found five loci explaining under one percent of variance and stated explicitly that its results cannot predict any individual's orientation.
  • Twin studies show orientation is about half genetic. The 52 percent figure came from a volunteer sample recruited through gay publications; population registries put heritability nearer 34 to 39 percent in men and lower in women.
  • The older brother effect is about being raised with brothers. It tracks biological older brothers regardless of upbringing, which is what makes it evidence for a prenatal mechanism.
  • Fluidity disproves biological influence. Heritability and lifelong stability are separate properties, and many highly heritable traits change over time.
  • Conversion therapy is unproven rather than disproven. The reviewed literature shows no reliable evidence of change and documented harm, and the best-known supporting study was retracted by its own author.

The takeaway

  • Ganna's 2019 study of 477,522 genomes found a polygenic signal with no predictive power, the standard picture for a complex trait.
  • Population-registry twin studies put heritability around 34 to 39 percent in men and 18 to 19 percent in women, well below the earlier volunteer-sample estimates.
  • The fraternal birth order effect raises odds by about a third per older biological brother and tracks biology rather than household, supporting a prenatal mechanism.
  • Parenting style, exposure and deliberate choice have been tested and are not supported; non-shared environment remains a large residual rather than an explanation.
  • Orientation can shift, particularly among women, and no intervention has been shown to change it, with the best-known supporting study retracted by its author in 2012.

Sources

  1. Ganna, A., et al. (2019). Large-scale GWAS reveals insights into the genetic architecture of same-sex sexual behavior. Science, 365(6456). Science
  2. American Psychological Association. (2009). Report of the task force on appropriate therapeutic responses to sexual orientation. apa.org
  3. Bogaert, A. F., et al. (2018). Male homosexuality and maternal immune responsivity to the Y-linked protein NLGN4Y. PNAS, 115(2), 302-306. PNAS
  4. Diamond, L. M. (2008). Sexual fluidity: Understanding women's love and desire. Harvard University Press. Harvard University Press
  5. Williams Institute. (n.d.). Research on LGBT populations. UCLA School of Law. williamsinstitute.law.ucla.edu
Key terms
Genome-wide association study
An analysis testing millions of genetic variants across a large sample for association with a trait or behaviour.
Polygenic trait
A trait influenced by many variants of individually tiny effect, producing aggregate heritability without individual predictive power.
Heritability
The share of variation in a trait within a population attributable to genetic variation; a property of populations, not individuals.
Fraternal birth order effect
The finding that a man's odds of being homosexual rise by roughly a third with each older biological brother.
Maternal immune hypothesis
The proposal that maternal immune responses to male-specific proteins accumulate across male pregnancies and affect prenatal development.
Non-shared environment
The residual variance category in behavioural genetics covering everything not genetic and not shared by siblings, including measurement error.
Sexual fluidity
Diamond's finding that attractions and identity labels can shift over time, more commonly among women.
Sexual orientation change efforts
Interventions claiming to alter orientation, for which reviewed evidence shows no reliable change and documented harm.

Module 3: Work, Money, and Care

The pay gap taken apart into its measured components, the earnings penalty that arrives with a first child, the arithmetic of unpaid care, and how gender operates inside medicine.

The Pay Gap, Decomposed

  • Distinguish the unadjusted and adjusted pay gaps and explain what each control adds and removes.
  • Reproduce the main components of a standard decomposition and state the size of the unexplained residual.
  • Explain the child penalty and the greedy work hypothesis, and evaluate what the audit study evidence supports.

Why pharmacists are nearly equal

American pharmacy has one of the smallest gender earnings gaps of any well-paid profession. In 1970 it had one of the larger ones. Nothing about the chemistry changed. What changed was the organization of the work, and Claudia Goldin and Lawrence Katz reconstructed the sequence in detail.

In 1970 a large share of pharmacists owned or worked in independent stores. The owner knew the customers, kept irregular hours, and could not be replaced by a colleague on short notice. By the 2000s most pharmacists worked for chains, hospitals or mail-order operations, where records are centralized, procedures are standardized, and one licensed pharmacist can pick up exactly where another left off. Two consequences followed. Part-time work stopped carrying a wage penalty per hour, because a pharmacist working twenty hours is genuinely half of a pharmacist working forty. And the premium for being continuously available collapsed, because there was nothing left that only you could do.

Hold that example in mind. It contains, in miniature, the argument that Goldin was awarded the 2023 Nobel Memorial Prize in Economic Sciences for building: that in rich countries today, most of what remains of the pay gap is produced not by employers paying two people differently for identical work, but by the way jobs are structured around time.

Key idea: Pharmacy went from a large gender pay gap to a small one because the work became substitutable and pay became linear in hours, without any change in the underlying skill.

Two numbers, and why they differ

Every discussion of the pay gap involves two figures that are constantly confused.

The unadjusted or raw gap compares all women to all men. In the United States, women working full time earn somewhere around 82 to 84 cents for each dollar earned by men, depending on whether you use the weekly earnings series from the Bureau of Labor Statistics or the annual series from the Census Bureau. Include part-time workers and the ratio drops further, because part-time work is disproportionately female and pays less per hour.

The adjusted gap compares women and men who match on a specified list of characteristics: education, experience, occupation, industry, hours, region, union status. Depending on the control set, the adjusted figure typically lands between about 92 and 98 cents.

People treat the second number as the true gap and the first as propaganda, or the first as the true gap and the second as an apologist's trick. Both readings misunderstand what a control does. Controlling for occupation answers the question: among a man and a woman who are both pharmacists, how much do they differ? It deliberately sets aside the question of why fewer women are surgeons and more are pediatricians. If that sorting is itself the product of steering, harassment, expectations or scheduling incompatible with caregiving, then controlling for occupation removes a large part of the phenomenon from view.

The right way to hold both is this. The unadjusted gap measures the difference in what women and men take home. The adjusted gap measures how much of that difference remains once you have made two workers look alike on paper. Each answers a real question and neither answers the other's.

The point: The unadjusted gap and the adjusted gap answer different questions, and the controls that shrink the number also conceal the mechanisms that produce it.

Taking the gap apart

Francine Blau and Lawrence Kahn published the standard decomposition in the Journal of Economic Literature in 2017, using American panel data from 1980 and 2010. The shape of their result, in rounded terms:

ComponentDirection and rough contributionWhat changed since 1980
EducationNow slightly favours women; reduces the gapReversed sign as women overtook men in degree attainment
Work experienceExplains a modest share, on the order of a seventhShrank sharply as women's continuous employment rose
Occupation and industryThe single largest measured componentGrew in relative importance as other components shrank
Union status, region, raceSmallLittle change
Unexplained residualRoughly a third or more of the total gapFell in absolute size but rose as a share

Read the third row and the fifth row together and you have the modern picture. The classic explanation, that women were less educated and less experienced, has largely dissolved; American women now hold the majority of bachelor's and master's degrees. What remains is concentrated in where people work and in a residual nobody can fully account for.

The penalty that arrives with a child

The sharpest evidence on when the gap opens comes from Denmark, where administrative records track every individual's earnings year by year. Henrik Kleven, Camille Landais and Jakob Sogaard lined up men's and women's earnings relative to the birth of a first child and plotted them.

Before the birth, the two lines run in parallel. At the birth, women's earnings drop sharply. Men's do not move at all. Women's earnings partially recover but never rejoin the counterfactual path, leaving a long-run shortfall of roughly twenty percent that is still visible a decade later. The authors' striking summary is that in recent Danish cohorts, essentially the entire remaining gender earnings gap can be attributed to what happens after the first child arrives.

The same method applied across countries produces a revealing spread. Long-run child penalties are smallest in Denmark and Sweden, in the low twenties as a percentage, larger in the United Kingdom and the United States, in the thirties and forties, and largest in Germany and Austria, above fifty percent. Those differences do not track biology; they track parental leave design, childcare availability, school hours, and prevailing beliefs about maternal employment. Kleven and colleagues also found that the size of a woman's penalty correlates with the penalty her own mother experienced, which points at transmitted norms rather than at a purely institutional story.

Why this matters: In administrative data, the gender earnings gap in recent cohorts opens almost entirely at the birth of a first child, and its size across countries varies by a factor of more than two according to policy and norms.

Greedy work

Goldin's explanation for why the child penalty converts into a lasting pay gap turns on the shape of the pay-hours relationship. In many high-paid occupations, pay is not proportional to hours. Someone who works sixty hours earns more than one and a half times someone who works forty, because clients want continuity, deals do not pause, and the person who is always reachable becomes the person who is trusted. Goldin calls these greedy jobs.

Now put a couple in that world. If both partners work greedy jobs, someone must take the call at seven in the evening. The household maximizes joint income by having one partner go all in and the other take the flexible role. Which partner does which is not determined by the economics, but it is heavily determined by convention, and in most couples the woman takes the flexible role. The pay gap then follows from the convexity of the pay schedule, without any individual employer treating two identical workers differently.

The prediction this generates is testable and has held up: the gender pay gap is largest in law, finance, consulting and corporate management, where hours are greedy, and smallest in pharmacy, veterinary medicine, optometry and much of technology, where a shift can be handed over. That is the pharmacy story again, generalized.

What is in the residual

The unexplained portion is not a measurement of discrimination. It is a measurement of what the model failed to capture, which includes discrimination but also unmeasured productivity differences, negotiation, job attributes not in the data, and error. To learn about discrimination directly you need experiments.

The best-known are correspondence studies, in which matched applications differing only in a name or a detail are sent to real employers. Shelley Correll, Stephen Benard and In Paik ran a version in 2007 in which otherwise identical resumes differed by whether the applicant mentioned membership in a parent-teacher association. Mothers were roughly half as likely to be called back as non-mothers, and evaluators recommended a starting salary about eleven thousand dollars lower. Fathers, in the same design, were not penalized and in some measures did better. A 2012 experiment by Corinne Moss-Racusin and colleagues sent science faculty an identical application for a lab manager position under a male or female name; the male applicant was rated more competent and offered a higher salary, by both male and female faculty.

A note on the most cited study of all. Goldin and Cecilia Rouse reported in 2000 that the introduction of blind auditions, with musicians playing behind a screen, increased women's success in orchestra hiring. The finding entered the popular canon. It has since been re-examined, and several statisticians have argued that the key estimates are imprecise and not clearly distinguishable from chance given the sample sizes involved. The correspondence-study evidence is much stronger than the audition evidence, and if you are going to cite one, cite the experiments.

Common misconceptions

  • The unadjusted gap means women are paid less for the same work. It compares all women to all men and includes differences in occupation, hours and experience.
  • The adjusted gap is the real gap. Controlling for occupation removes exactly the sorting that much of the research is trying to explain.
  • Women's lower education explains the gap. That reversed decades ago; the education term now works in women's favour in American decompositions.
  • The unexplained residual measures discrimination. It measures what the model missed, which includes discrimination alongside unmeasured job characteristics and error.
  • The blind audition study proved orchestral hiring was biased. Its estimates have been challenged as imprecise; the correspondence studies are the stronger evidence for hiring discrimination.

Pulling it together

  • Pharmacy's gap shrank when the work became substitutable and pay became linear in hours, which is the model case for Goldin's argument.
  • The unadjusted United States gap sits around 82 to 84 cents; adjusted figures land around 92 to 98 depending on controls, and the controls themselves are analytically loaded.
  • Blau and Kahn's decomposition leaves occupation and industry as the largest measured component and roughly a third or more unexplained.
  • Danish administrative data show earnings diverging at the first birth and never reconverging, with long-run penalties from the low twenties to above fifty percent across countries.
  • Correspondence experiments show a motherhood penalty of about half the callback rate and a fatherhood bonus, and identical applications under male names rated more competent.

Sources

  1. Blau, F. D., and Kahn, L. M. (2017). The gender wage gap: Extent, trends, and explanations. Journal of Economic Literature, 55(3), 789-865. American Economic Association
  2. Kleven, H., Landais, C., and Sogaard, J. E. (2019). Children and gender inequality: Evidence from Denmark. American Economic Journal: Applied Economics, 11(4), 181-209. American Economic Association
  3. The Nobel Prize. (2023). The Sveriges Riksbank Prize in Economic Sciences in Memory of Alfred Nobel 2023: Claudia Goldin. nobelprize.org
  4. Bureau of Labor Statistics. (n.d.). Current Population Survey: Earnings data. U.S. Department of Labor. bls.gov
  5. Correll, S. J., Benard, S., and Paik, I. (2007). Getting a job: Is there a motherhood penalty? American Journal of Sociology, 112(5), 1297-1339. University of Chicago Press
Key terms
Unadjusted pay gap
The ratio of all women's earnings to all men's, without controlling for occupation, hours or experience.
Adjusted pay gap
The residual difference in earnings between women and men matched on a specified list of characteristics.
Decomposition
A statistical procedure attributing shares of an observed gap to measured characteristics and to an unexplained remainder.
Child penalty
The persistent fall in a woman's earnings relative to trend that begins at the birth of a first child.
Greedy job
An occupation in which pay rises more than proportionally with hours, so continuous availability is rewarded disproportionately.
Correspondence study
An experiment sending matched applications differing in one attribute to real employers to measure differential treatment.
Motherhood penalty
The measured reduction in callbacks and recommended salary for otherwise identical applicants signalled as mothers.
Occupational sorting
The unequal distribution of women and men across occupations, the largest measured component of the modern pay gap.

Unpaid Care: The Arithmetic of the Second Shift

  • Explain how time-use diaries collect data and which kinds of care work they systematically undercount.
  • State the size of the unpaid work gap in national and global data and convert it into an economic magnitude.
  • Compare relative-resources, time-availability and gender-display explanations for the division of household labour.

An extra month of days

Between 1980 and 1988 the sociologist Arlie Hochschild and her research partner Anne Machung followed fifty two-earner couples with young children, interviewing them repeatedly and observing many of them at home. The couples described themselves, almost uniformly, as sharing. The time diaries told a different story. Adding housework and childcare to paid employment, Hochschild calculated that the women in her sample worked roughly fifteen hours a week more than the men, which she rendered in the line that made the book famous: an extra month of twenty-four-hour days each year.

The Second Shift (1989) also documented something subtler than the hours, which Hochschild called the leisure gap and the family myth. Couples who divided labour very unequally often maintained an elaborate shared account in which they did not. One husband was described as sharing because he handled the downstairs, meaning the garage and the dog, while his wife handled the upstairs, meaning the cooking, cleaning, laundry and children.

Fifty couples is not a national estimate. What Hochschild provided was the hypothesis. This lesson provides the measurement.

Key idea: Hochschild's fieldwork produced both the central finding, that women's total work exceeds men's once unpaid hours are counted, and the mechanism that hides it, which is an account of sharing that the couple both believes.

How you measure something nobody bills for

The instrument is the time-use diary. In the American Time Use Survey, run by the Bureau of Labor Statistics since 2003, an interviewer walks a randomly selected respondent through the previous twenty-four hours, from four in the morning to four in the morning, asking what they were doing, how long it lasted, where they were and who was present. Around nine thousand people are interviewed each year. Diaries beat direct estimation badly: asked how many hours a week they do housework, people produce numbers that do not sum to a plausible week.

Three limits are worth knowing before you read any time-use table.

  • Primary activity coding. Each minute is assigned to one main activity. A parent cooking dinner while supervising a toddler is coded as cooking. The survey does collect a separate measure of time with a child in the respondent's care, and that measure is far larger than the childcare figure people usually quote.
  • Availability is not activity. Being the parent who cannot leave the house because the baby is asleep does not appear as work in any column, though it constrains the day completely.
  • Cognitive labour is invisible. Noticing that the shoes no longer fit, holding the vaccination schedule in mind, and remembering that the gift needs buying occupy no minutes a diary can capture.

That last one has been measured another way. Allison Daminger's 2019 study in the American Sociological Review interviewed couples about the cognitive dimension of household labour and separated it into four steps: anticipating a need, identifying options, deciding, and monitoring the result. Decision-making was the step couples shared most equally. Anticipation and monitoring, the invisible bookends, fell overwhelmingly to women. That structure explains a common household argument in which one partner truthfully says they help whenever asked and the other truthfully says that having to ask is the work.

The numbers

American Time Use Survey results in recent years, in round terms: on an average day roughly 85 percent of women and roughly 70 percent of men do some household activity, and on the days they do it, women spend around two and a half hours against around two for men. For care of household children under six, women's time runs to roughly twice men's. The gap has narrowed substantially since the 1960s, driven both by men doing more and by women doing considerably less housework than their mothers did, as standards changed and appliances and purchased services absorbed part of the load.

The comparison that most surprises students is total work. Add paid and unpaid hours together and, in several rich countries, women's and men's totals come out reasonably close, with women's slightly higher in most. The composition differs sharply while the sum does not. That finding is real and it is not the end of the story, for three reasons: the sum is much less equal in most low- and middle-income countries; unpaid hours do not generate wages, pensions or promotion; and the flexibility demanded by unpaid work shapes which paid jobs are available in the first place, which is exactly the child penalty from the previous lesson.

Globally the International Labour Organization's 2018 assessment put the numbers at a scale worth memorizing. Women perform about 76 percent of all unpaid care hours worldwide, roughly 3.2 times men's contribution. Unpaid care work amounts to something on the order of 16 billion hours every day. Valued at an hourly minimum wage, that work would represent around nine percent of global economic output, which is more than most countries spend on health care.

In short: Women perform roughly three quarters of the world's unpaid care hours; valued at minimum wage that work would equal about nine percent of global output, and none of it accrues wages or pension credit.

Three explanations, and a result that broke one of them

Sociologists have tested three accounts of how couples divide domestic work.

AccountPredictionHow it fares
Time availabilityWhoever has fewer paid hours does more houseworkPartly supported; explains some variation but not the level of the gap
Relative resourcesThe partner earning less has less bargaining power and does moreSupported up to the point where earnings are equal
Gender displayHousework performs gender, so norm violations get compensated elsewhereProposed to explain what happens past equal earnings; contested

The interesting case is what happens when a wife out-earns her husband. Julie Brines reported in 1994 in the American Journal of Sociology that the expected downward slope stopped and, in some specifications, reversed: as wives moved from equal earnings to primary earner, their housework hours went up rather than down. The interpretation offered was gender display, that couples deviating from the breadwinner norm in earnings compensated by exaggerating conventional roles at home.

That finding became a textbook staple, and then it was challenged. Sanjiv Gupta argued in 2007 that the apparent reversal was an artefact of modelling housework against the wife's share of household income rather than her own absolute earnings. Using absolute earnings, the relationship is straightforward and monotonic: the more a woman earns, the less housework she does, largely because she buys some of it out. Later work has gone both ways, and the honest position is that the compensating-display effect is small at best and may be a specification artefact, while the substitution effect of a woman's own earnings is robust.

This is a good example of how an appealing sociological story can survive for a decade on a modelling choice. When you meet a counterintuitive curve, ask what is on the horizontal axis.

The upshot: A woman's own absolute earnings reliably reduce her housework hours; the striking claim that out-earning a husband increases them rests on a contested modelling choice.

What actually moves the numbers

Three levers have measurable effects. Non-transferable parental leave, discussed in Lesson 4, changes the division of infant care and has persistent effects on later involvement in several national studies. Affordable childcare changes maternal employment sharply; Quebec's low-fee childcare programme, introduced in 1997, is associated with a large and durable rise in maternal labour force participation. And the marketization of care shifts work rather than eliminating it: cleaning, cooking and eldercare purchased on the market are performed overwhelmingly by women, disproportionately migrant women, who often leave their own children in the care of relatives in another country. Hochschild later named this pattern the global care chain, and it is the reason a policy that looks like a solution in one country can be a redistribution of the same problem across borders.

Common misconceptions

  • Time-use diaries capture all care work. They code one primary activity per minute, so supervision, availability and cognitive labour largely disappear from the headline figures.
  • Men and women in rich countries now do equal amounts of unpaid work. The gap has narrowed substantially, but women still do roughly twice as much care of young children in American data.
  • Because total work hours are close to equal, the division does not matter. Unpaid hours generate no wages, pension credits or promotion, and they constrain which paid jobs are reachable.
  • Women who out-earn their husbands do more housework to compensate. That reversal depends on modelling relative rather than absolute income and does not survive the alternative specification cleanly.
  • Buying domestic services solves the problem. It transfers the work to other women, frequently migrants, whose own care obligations are then met by someone further down the chain.

What you now know

  • Hochschild's fifty-couple study produced both the second shift finding and the family myth that conceals it.
  • Time-use diaries beat direct estimation but code one activity per minute, which hides supervision, availability and the cognitive labour Daminger measured separately.
  • American women do household activities more often and for longer, and roughly twice as much care of young children; totals of paid plus unpaid work are closer than the components suggest.
  • Globally, women perform about 76 percent of unpaid care hours, some 16 billion hours a day, worth roughly nine percent of global output at minimum wage.
  • Relative-resources and time-availability accounts explain much of the variation; the gender-display reversal at high female earnings is contested and probably a specification artefact.

Sources

  1. Bureau of Labor Statistics. (n.d.). American Time Use Survey. U.S. Department of Labor. bls.gov
  2. International Labour Organization. (2018). Care work and care jobs for the future of decent work. ilo.org
  3. Daminger, A. (2019). The cognitive dimension of household labor. American Sociological Review, 84(4), 609-633. SAGE Journals
  4. Brines, J. (1994). Economic dependency, gender, and the division of labor at home. American Journal of Sociology, 100(3), 652-688. University of Chicago Press
  5. UN Women. (n.d.). Redistribute unpaid work. unwomen.org
Key terms
Second shift
Hochschild's term for the unpaid domestic work performed after paid employment, disproportionately by women.
Family myth
A shared account by which a couple describes an unequal division of labour as equal sharing.
Time-use diary
An instrument recording the previous twenty-four hours minute by minute, which outperforms direct estimation of weekly hours.
Primary activity coding
The convention of assigning each minute to one main activity, which conceals supervision and simultaneous care.
Cognitive labour
The anticipating, identifying, deciding and monitoring involved in running a household, distributed unequally at the anticipation and monitoring stages.
Relative resources
The account holding that the lower-earning partner has less bargaining power and therefore performs more household work.
Gender display
The contested proposal that couples deviating from earnings norms compensate by exaggerating conventional roles at home.
Global care chain
The transfer of care work across borders, in which migrant women perform paid care while relatives care for their own children.

Gender in Medicine: Dosing, Diagnosis, and Who Gets Studied

  • Trace the regulatory history that excluded women from clinical trials and the reforms that reversed it.
  • Distinguish differences in drug handling, disease presentation and clinical treatment, with examples of each.
  • Assess evidence on diagnostic delay and on men's lower engagement with health services.

A sleeping pill, twenty-one years late

In January 2013 the Food and Drug Administration announced that the recommended bedtime dose of zolpidem, the insomnia drug sold as Ambien, should be cut in half for women. Driving simulator studies and blood measurements had shown that women clear the drug more slowly than men, so that eight hours after a standard dose, a substantial share still had enough in their blood to impair driving. The drug had been on the American market since 1992.

Nothing was hidden. The pharmacokinetic difference was in the data. What was missing was a regulatory system that treated sex as a variable worth analysing before approval rather than after two decades of accident reports. This lesson is about how medicine acquired that blind spot, how it has been partially corrected, and what remains.

Key idea: The zolpidem correction took twenty-one years not because the evidence was unavailable but because nothing in the approval process required anyone to look for it.

How the exclusion happened

In 1977 the Food and Drug Administration issued a guideline recommending that women of childbearing potential be excluded from early-phase drug trials. The reasoning was not casual. Thalidomide, prescribed for morning sickness between 1957 and 1961, had caused severe limb malformations in thousands of infants, and diethylstilbestrol, given to pregnant women into the 1970s, produced cancers in their daughters years later. Regulators concluded that the safest course was to keep women out of early testing.

The unintended consequence took twenty years to become visible. Excluding women from early trials meant that dosing, side effects and interactions were characterized in male bodies, and the results were then applied to everyone. A 2001 report from the United States General Accounting Office examined ten prescription drugs withdrawn from the market since January 1997 and found that eight of them posed greater health risks for women than for men. Some of those risks would have been detectable earlier in a trial population that included women in sufficient numbers.

There was a further problem inside the studies that did include women: many analysed results without breaking them down by sex, so a differential effect could be present in the dataset and absent from the paper.

The corrections

Two reforms matter most. The NIH Revitalization Act of 1993 required that women and members of minority groups be included in National Institutes of Health funded clinical research, and that phase III trials be designed to allow analysis of whether interventions affect them differently. In 2016 the NIH went further with its policy on sex as a biological variable, requiring applicants to account for sex in vertebrate animal and human studies, not merely in the human phase. That second reform matters because the male-default habit began in the laboratory: for decades, researchers preferred male rodents on the grounds that the oestrous cycle introduced variability, a rationale that has since been challenged by analyses finding female rodents no more variable than males on most measures.

Progress is real and partial. Women are now well represented in many trial areas overall, while remaining underrepresented in some cardiovascular and oncology trials relative to disease burden, and sex-stratified analysis is still not universally reported.

Why this matters: The 1993 Act fixed enrolment and the 2016 policy attacked the male-default habit at its origin in animal research, but sex-stratified reporting is still inconsistent.

Three different things that get called sex differences in medicine

TypeWhat it isExample
PharmacokineticThe body handles the drug differentlySlower zolpidem clearance in women, driving the 2013 dose change
PathophysiologicalThe disease itself differs in frequency or presentationRoughly four in five autoimmune disease patients are women; myocardial infarction presents without classic chest pain more often in women
ClinicalThe same presentation is treated differently by cliniciansWomen presenting to emergency departments with acute abdominal pain wait longer and are less likely to receive opioid analgesia

Keeping these apart is the analytic core of the lesson. A pharmacokinetic difference calls for a different dose. A pathophysiological difference calls for different diagnostic criteria and different research. A clinical difference calls for changing what clinicians do. Treating all three as one problem produces bad policy, because the fixes are not interchangeable.

Cardiovascular disease illustrates all three at once. It is the leading cause of death for women in the United States, a fact that survey after survey finds women themselves underestimate. Women having heart attacks more often present with fatigue, nausea, jaw or back pain rather than crushing chest pain, which delays recognition by patients and clinicians alike. Women have historically been underrepresented in cardiovascular trials relative to their share of the disease burden. And a 2018 study in the Proceedings of the National Academy of Sciences, analysing two decades of Florida emergency admissions, found that female heart attack patients had higher mortality when treated by male physicians than by female physicians, with the difference narrowing for male physicians who had treated more female patients previously. That result does not identify the mechanism, and it has generated an active methodological debate, but it is the kind of finding that only appears when someone thinks to stratify.

Diagnostic delay

Endometriosis affects something on the order of one in ten women of reproductive age. The interval between first symptoms and diagnosis is commonly reported at seven to ten years across national studies. Multiple factors contribute: the definitive diagnosis has historically required laparoscopic surgery, symptoms overlap with common gastrointestinal complaints, and severe menstrual pain is frequently normalized by patients, families and clinicians together. Autoimmune conditions show a similar pattern, with patient surveys reporting several years and several physicians before a correct diagnosis.

Pain treatment has been studied experimentally and observationally, and the direction is consistent. Studies of emergency department care find that women with equivalent presentations wait longer for analgesia and are less likely to receive opioids. Vignette studies, in which clinicians assess identical case descriptions differing only in the patient's sex, find that women's pain is more often attributed to psychological factors. The effect sizes are modest, and the studies vary in quality, but the consistency of direction across designs is what makes the finding credible.

The other side of the ledger

Men die roughly five to six years earlier than women in the United States and lead in most causes of death. They are less likely to have a usual source of care, less likely to have visited a physician in the past year, and more likely to present at a late stage. Some of this is the self-reliance norm measured in Lesson 4; some is structural, since routine contact with the health system through contraception, cervical screening and pregnancy care gives many women a relationship with a clinician that many men simply never form.

Research attention is uneven in both directions. Male contraception has produced no new approved method in decades, with trials repeatedly halted over side effects that are comparable to those long tolerated in female methods, an asymmetry worth noticing. Osteoporosis and breast cancer in men are underdiagnosed partly because both are coded as women's diseases. The point is not to run a competition over who is worse served. It is that a medical system organized around a default patient serves everyone who is not that patient worse, and the default has not always been the same one.

One number belongs here, because it shows how sex and race interact. American maternal mortality rose to 32.9 deaths per 100,000 live births in 2021, during the pandemic, and fell to 22.3 in 2022. Throughout, the rate for Black women has run roughly two and a half to three times the rate for white women, a disparity that persists after adjusting for education and income. Any account of gender and health that stops at the average has missed the largest variation in the data.

Bottom line: A medical system built around a default patient underserves everyone else, and the disparities within the category of women, particularly by race, are larger than the average difference between women and men.

Common misconceptions

  • Women were excluded from trials out of indifference. The 1977 guideline was a response to thalidomide and diethylstilbestrol; the harm came from a protective rule with unexamined consequences.
  • Including women in trials solved the problem. Enrolment improved after 1993, but results are still not always analysed and reported separately by sex.
  • Sex differences in medicine are all biological. Pharmacokinetic, pathophysiological and clinical differences are three separate categories with three different remedies.
  • Heart disease is primarily a men's problem. It is the leading cause of death for American women, and atypical presentation contributes to delayed recognition.
  • Men's worse mortality is purely behavioural. Behaviour matters, and so does the fact that many women have routine clinical contact through reproductive care that men never establish.

Looking back

  • The 2013 zolpidem dose change corrected a twenty-one-year-old error that the approval process had never been designed to catch.
  • A 1977 FDA guideline excluded women of childbearing potential from early trials; a 2001 GAO review found eight of ten recently withdrawn drugs posed greater risks to women.
  • The 1993 NIH Revitalization Act mandated inclusion, and the 2016 sex as a biological variable policy extended the requirement to animal research.
  • Pharmacokinetic, pathophysiological and clinical differences require different remedies; cardiovascular disease shows all three.
  • Endometriosis diagnosis commonly takes seven to ten years, women receive less analgesia for equivalent pain, men engage with care less, and Black maternal mortality runs at roughly two and a half to three times the white rate.

Sources

  1. Office of Research on Women's Health. (n.d.). NIH policy on sex as a biological variable. National Institutes of Health. orwh.od.nih.gov
  2. U.S. Food and Drug Administration. (2013). Drug safety communication: Risk of next-morning impairment after use of insomnia drugs. fda.gov
  3. U.S. General Accounting Office. (2001). Drug safety: Most drugs withdrawn in recent years had greater health risks for women (GAO-01-286R). gao.gov
  4. Greenwood, B. N., Carnahan, S., and Huang, L. (2018). Patient-physician gender concordance and increased mortality among female heart attack patients. PNAS, 115(34), 8569-8574. PNAS
  5. National Center for Health Statistics. (n.d.). Maternal mortality rates in the United States. Centers for Disease Control and Prevention. cdc.gov
Key terms
Pharmacokinetics
How the body absorbs, distributes, metabolizes and clears a drug, which can differ systematically by sex.
NIH Revitalization Act
The 1993 law requiring inclusion of women and minority participants in federally funded clinical research.
Sex as a biological variable
The 2016 NIH policy requiring applicants to account for sex in vertebrate animal as well as human research.
Sex-stratified analysis
Reporting results separately for female and male participants rather than pooling them.
Atypical presentation
Symptoms that depart from the textbook picture of a condition, as in myocardial infarction without classic chest pain.
Diagnostic delay
The interval between first symptoms and correct diagnosis, commonly seven to ten years for endometriosis.
Male-default research
The practice of characterizing physiology and dosing in male subjects and generalizing the findings to everyone.
Usual source of care
A regular clinician or practice a patient returns to, which men are measurably less likely to have.

Module 4: Harm and Law

Why two federal surveys of the same country differ sevenfold on sexual violence, what the intimate partner violence literature actually shows, and the statutes and cases that govern sex discrimination in American law.

Counting Violence: Why the Numbers Disagree

  • Explain why crime surveys and public health surveys produce sharply different estimates of sexual violence.
  • Distinguish intimate terrorism from situational couple violence and state the evidence for the distinction.
  • Describe attrition through the criminal justice process and the definitional history that hid male victims.

One country, one year, a factor of seven

For the year 2010, two agencies of the United States federal government published estimates of how many rapes and sexual assaults had occurred. The Bureau of Justice Statistics, using the National Crime Victimization Survey, put the figure at roughly 188,000. The Centers for Disease Control and Prevention, using the National Intimate Partner and Sexual Violence Survey, put the number of completed or attempted rapes against women alone at roughly 1.27 million.

Neither agency was lying, incompetent, or working from a different population. Both surveys were well funded, professionally administered and probability sampled. The gap of nearly sevenfold comes from the instruments, and pulling it apart is the most useful thing anyone can learn about violence statistics, because the same lesson applies to every number you will ever encounter on this subject.

Key idea: Two rigorous federal surveys of the same country in the same year differed by a factor of about seven on sexual violence, and the difference is produced by design choices rather than by error.

Four design choices that move the number

Design choiceNational Crime Victimization SurveyNational Intimate Partner and Sexual Violence Survey
FramingPresented as a survey about crimePresented as a survey about health and relationships
Question styleScreening questions about attacks and unwanted sexual activityBehaviourally specific questions naming acts explicitly
DefinitionTracks criminal categoriesIncludes incidents where the respondent was unable to consent because of alcohol or drugs
Sample and modeHousehold panel, ages 12 and up, other household members may be presentTelephone survey of adults, lower response rate

Framing is the largest single factor. In a survey introduced as being about crime, a respondent must first classify what happened to her as a crime before it becomes reportable. A woman who was assaulted by a boyfriend, who did not go to the police and does not think of the event in criminal terms, may answer no to a crime screener and yes to a question asking whether anyone has ever had sexual contact with her when she was too intoxicated to consent. Both answers are honest.

Behavioural specificity works the same way. Asking whether someone has been raped requires the respondent to apply a legal label; asking whether anyone has ever used physical force to make them have sexual intercourse asks about an event. Studies that vary only the question wording reliably find the second form producing higher estimates.

Every design choice has a cost. Behaviourally specific questions on a health survey will capture events that no prosecutor could charge and that some respondents themselves would not call assault. Crime surveys with a stable panel design produce excellent year-over-year trends but systematically undercount. Neither instrument is the true one. The right question is which is appropriate for the purpose: allocating police resources, or estimating the public health burden.

What matters here: Crime framing lowers estimates because the respondent must first classify the event as a crime; behaviourally specific health framing raises them because it asks about acts rather than labels.

What the instruments agree on

Focusing on the discrepancy obscures a substantial consensus. Across surveys, methodologies and countries, four findings recur.

  • Victimization is concentrated in adolescence and early adulthood, with a large share of first victimizations occurring before age 25 and a substantial share before 18.
  • Most perpetrators are known to the victim. The stranger in the alley accounts for a minority of cases in every dataset.
  • Reporting to police is low. American estimates typically put the share of sexual assaults reported at around one in four or fewer.
  • Rates of intimate partner violence are high enough globally to be a leading cause of injury among women of reproductive age. The World Health Organization estimates that roughly 27 percent of ever-partnered women aged 15 to 49 have experienced physical or sexual violence from a partner.

The symmetry debate, and how it was resolved

Here is a genuine scientific dispute with a genuine resolution, which makes it worth working through.

Beginning in the 1970s, family sociologists using the Conflict Tactics Scale asked both partners how often each had performed specific acts during disagreements: pushed, slapped, threw something, hit with a fist. Across many samples, the counts came out close to symmetric. Women reported perpetrating physical acts at rates similar to men. Murray Straus, who developed the instrument, argued for decades that this finding was suppressed for political reasons.

Meanwhile, researchers working from police reports, emergency departments, and shelters found overwhelming asymmetry. Women were three to four times more likely to be injured. In the United States, roughly a third of female homicide victims are killed by an intimate partner, against a small fraction of male victims. Fear, control and repeated victimization ran heavily in one direction.

Michael Johnson proposed the resolution: the two literatures were sampling different phenomena. He distinguished situational couple violence, arising from specific conflicts that escalate, generally lower in severity, rarely involving a pattern of control, and roughly symmetric in perpetration, from intimate terrorism, in which physical violence is one instrument within a broader pattern of coercive control involving isolation, surveillance, economic restriction and threat, which escalates over time and is overwhelmingly perpetrated by men against women. General population surveys, which reach mostly ordinary households, capture mainly the first. Shelters, courts and hospitals see mainly the second. Both literatures were measuring what their sampling frames contained.

The practical payoff is large. If you treat all partner violence as one phenomenon, you get interventions that fail: couples counselling is reasonable for situational conflict and dangerous where coercive control is present, because it presumes two parties who can safely speak. Several jurisdictions have since written coercive control into criminal law, with England and Wales doing so in the Serious Crime Act 2015 and Scotland going further in the Domestic Abuse (Scotland) Act 2018.

The upshot: Partner violence is not one thing; situational couple violence is roughly symmetric and lower in severity, while intimate terrorism is built around coercive control, escalates, and is strongly asymmetric.

What happens after a report

Prevalence is only half the picture. Attrition through the criminal process is severe and measurable. American estimates compiled from federal data suggest that of every thousand sexual assaults, roughly 310 are reported to police, around 50 lead to an arrest, about 28 to a felony conviction, and about 25 to incarceration. England and Wales publish comparable figures, and the charge rate as a proportion of recorded rape offences fell to about one to two percent in the early 2020s before rising modestly under a national action plan.

The reasons are structural rather than mysterious: most cases turn on consent between people who know each other, physical evidence rarely settles that question, delay degrades what evidence exists, and complainants withdraw at high rates, often citing the process itself. Any policy proposal in this area has to specify which stage of this funnel it is trying to change.

Male victims and a definition that erased them

Until 2012 the Federal Bureau of Investigation defined rape, for national reporting purposes, as the carnal knowledge of a female forcibly and against her will. A man could not be recorded as a rape victim in the national crime statistics, whatever had happened to him. The definition, in place since 1927, was replaced in January 2012 with a gender-neutral formulation covering penetration of any person without consent.

Measurement caught up more slowly than the definition. The CDC survey distinguishes being penetrated from being made to penetrate someone else, and in several survey years the twelve-month prevalence of men reporting being made to penetrate was comparable to the twelve-month prevalence of women reporting rape, a finding highlighted by Lara Stemple and Ilan Meyer in the American Journal of Public Health in 2014. Lifetime figures remain very different, with women reporting far higher rates. The methodological point stands regardless of how the categories are combined: a category that does not exist in an instrument cannot appear in its results, and for eighty-five years the national instrument had no place to record a male victim.

Common misconceptions

  • Conflicting statistics mean the researchers are biased. The main federal estimates diverge because of framing, question wording and definitions, all documented in the survey methodology.
  • Behaviourally specific questions inflate the numbers dishonestly. They measure acts rather than labels, which is a defensible choice with a known cost: they include events no prosecutor could charge.
  • Partner violence is symmetric between the sexes. Act counts in general population surveys are close to symmetric; injury, fear, homicide and coercive control are not, and Johnson's typology explains why both findings appear.
  • Low conviction rates mean most reports are false. Attrition is concentrated at reporting, evidence and withdrawal stages, and estimated false report rates in the research literature are in the low single digits to around eight percent.
  • Male victimization is negligible. The national crime definition excluded male victims entirely until 2012, and survey categories that capture it were introduced only recently.

What to remember

  • For 2010 the crime survey estimated about 188,000 rapes and sexual assaults while the health survey estimated about 1.27 million against women, a gap produced by design rather than error.
  • Framing, behavioural specificity, definitions of incapacitation, and sampling mode are the four levers that move violence estimates.
  • Both traditions agree that victimization concentrates in youth, that most perpetrators are known, and that police reporting is low.
  • Johnson's distinction between situational couple violence and intimate terrorism reconciles the symmetry dispute and has changed both intervention design and criminal law.
  • Of every thousand American sexual assaults, roughly 310 are reported and about 25 result in incarceration, and the FBI definition excluded male victims until 2012.

Sources

  1. Bureau of Justice Statistics. (n.d.). National Crime Victimization Survey. Office of Justice Programs. bjs.ojp.gov
  2. World Health Organization. (2021). Violence against women (fact sheet). who.int
  3. Federal Bureau of Investigation. (n.d.). Uniform Crime Reporting: Rape definition. ucr.fbi.gov
  4. Stemple, L., and Meyer, I. H. (2014). The sexual victimization of men in America. American Journal of Public Health, 104(6), e19-e26. American Public Health Association
  5. RAINN. (n.d.). The criminal justice system: Statistics. rainn.org
Key terms
Crime framing
Presenting a survey as being about crime, which requires respondents to classify an event as criminal before reporting it.
Behaviourally specific questions
Items naming acts explicitly rather than using legal labels, which reliably produce higher prevalence estimates.
Conflict Tactics Scale
Straus's instrument counting specific acts during disagreements, which yields near-symmetric perpetration rates in general samples.
Situational couple violence
Violence arising from escalating specific conflicts, generally lower in severity and roughly symmetric between partners.
Intimate terrorism
Violence embedded in a broader pattern of coercive control, which escalates and is strongly asymmetric by sex.
Coercive control
A pattern of isolation, surveillance, economic restriction and threat, criminalized in England and Wales in 2015 and Scotland in 2018.
Attrition
The loss of cases at each stage between offence and conviction, from non-reporting through withdrawal and evidentiary failure.
Made to penetrate
The survey category capturing victimization of men that was absent from national crime definitions before 2012.

Law and Rights: From Reed to Bostock

  • Explain the tiers of constitutional scrutiny and identify which applies to sex classifications and why.
  • Summarize the coverage of Title VII, the Pregnancy Discrimination Act and Title IX, and the leading cases interpreting each.
  • Apply the relevant standard to a concrete employment or education fact pattern.

An estate in Idaho

When Richard Reed died in 1967 without a will, his adoptive parents, Sally and Cecil Reed, who were separated, each applied to administer his small estate. The Idaho probate code provided that where two applicants were otherwise equally entitled, males must be preferred to females. The probate court, following the statute, appointed Cecil.

Sally Reed appealed, and in 1971 the Supreme Court held unanimously in Reed v. Reed that the Idaho provision violated the Equal Protection Clause of the Fourteenth Amendment. The brief for Reed listed as an author a Rutgers law professor named Ruth Bader Ginsburg. It was the first time in the amendment's hundred-and-three-year history that the Court struck down a law for discriminating on the basis of sex.

That is a startlingly late date, and it tells you something structural. The constitutional architecture governing sex discrimination in the United States is barely fifty years old, was built case by case rather than by any single amendment, and rests on statutes that were in several instances passed for reasons unrelated to the outcomes they produced.

Key idea: American sex discrimination law dates from 1971, was assembled case by case, and is statutory as much as constitutional.

How courts test a sex-based rule

American equal protection analysis sorts government classifications into tiers, each with a different burden.

StandardApplies toGovernment must showLeading case
Rational basisMost classifications, including age and wealthAny conceivable legitimate purposeApplied by default
Intermediate scrutinySex, and legitimacy of birthAn important objective and a substantially related meansCraig v. Boren (1976)
Strict scrutinyRace, national origin, fundamental rightsA compelling interest and the least restrictive meansApplied to racial classifications

Sex sits in the middle tier, and it got there by an odd route. In Frontiero v. Richardson (1973), involving an Air Force lieutenant denied the dependent benefits a male officer would have received automatically for a spouse, four justices wanted to treat sex like race and apply strict scrutiny. They were one vote short. Three years later, in a case about an Oklahoma statute allowing women to buy low-alcohol beer at 18 while men had to wait until 21, the Court settled on the intermediate standard in Craig v. Boren.

The standard sharpened in 1996. United States v. Virginia concerned the Virginia Military Institute, which admitted only men and proposed a separate women's leadership programme at another college as a remedy. Writing for the Court, Justice Ginsburg held that a sex classification requires an exceedingly persuasive justification, that the justification must be genuine rather than invented for litigation, and that it must not rest on overbroad generalizations about the talents or preferences of either sex. VMI's own evidence, that some women could meet its standards, defeated its case.

The point: Sex classifications face intermediate scrutiny, tightened by the VMI requirement of an exceedingly persuasive and non-post-hoc justification that does not rest on generalizations about either sex.

The workplace: Title VII and what grew from it

Title VII of the Civil Rights Act of 1964 prohibits employment discrimination because of race, colour, religion, national origin and sex. The inclusion of sex has a peculiar history: it was added by amendment on the House floor by Howard Smith of Virginia, an opponent of the bill, and passed. Whatever the intent, the word was in the statute, and the litigation that followed built most of American employment gender law.

  • Phillips v. Martin Marietta (1971) established that a rule applied to one sex only, there a refusal to hire mothers of preschool children while hiring fathers of preschool children, is sex discrimination even though it does not exclude all women.
  • General Electric v. Gilbert (1976) held that excluding pregnancy from a disability plan was not sex discrimination. Congress disagreed and overrode it with the Pregnancy Discrimination Act of 1978, which is a useful reminder that statutes can reverse judicial interpretations.
  • Meritor Savings Bank v. Vinson (1986) established that a hostile work environment, not merely a lost job or promotion, can violate Title VII.
  • Price Waterhouse v. Hopkins (1989) held that penalizing an employee for failing to conform to sex stereotypes is sex discrimination. Ann Hopkins had brought in more business than any other partnership candidate in her year and was told to walk, talk and dress more femininely and to wear makeup.
  • Oncale v. Sundowner Offshore Services (1998) held unanimously that same-sex harassment is actionable, and that the statute's protections do not depend on the sexes of the parties.
  • Ledbetter v. Goodyear (2007) held that the filing clock for a discriminatory pay decision runs from the decision itself, which Congress reversed with the Lilly Ledbetter Fair Pay Act of 2009, restarting the clock with each affected paycheck.
  • Bostock v. Clayton County (2020) held six to three that firing an employee for being gay or transgender is discrimination because of sex under Title VII, on the reasoning that you cannot penalize a man for attraction to men without taking his sex into account.

Two statutes sit alongside. The Equal Pay Act of 1963 addresses unequal pay for substantially equal work in the same establishment, and unlike Title VII it does not require proof of intent, though it allows a defence for any factor other than sex. The Pregnant Workers Fairness Act, enacted in 2022 and effective in 2023, requires reasonable accommodation for pregnancy-related limitations, closing a gap that Young v. UPS (2015) had only partially addressed.

Education: thirty-seven words

Title IX of the Education Amendments of 1972 is a single sentence: no person in the United States shall, on the basis of sex, be excluded from participation in, be denied the benefits of, or be subjected to discrimination under any education program or activity receiving federal financial assistance.

Nothing in it mentions sports. Athletics became the most visible application because a 1979 policy interpretation set out a three-part test under which an institution complies if it meets any one of three conditions: participation opportunities substantially proportionate to enrolment, a history and continuing practice of expanding opportunities for the underrepresented sex, or full and effective accommodation of that sex's interests and abilities. Meeting any one suffices, which is why describing the test as a quota misstates it.

The measured effect is large. Girls' participation in American high school sports rose from roughly 294,000 in the 1971-72 school year to more than three million in recent years. Women went from a small minority of law and medical students to roughly half or more.

The courts then extended the statute well beyond athletics. Cannon v. University of Chicago (1979) established that individuals can sue under it. Franklin v. Gwinnett County Public Schools (1992) allowed monetary damages, opening the way to sexual harassment claims. Davis v. Monroe County Board of Education (1999) held that a school can be liable for student-on-student harassment where it is deliberately indifferent to known severe and pervasive conduct. Federal enforcement guidance on campus sexual misconduct has swung substantially with successive administrations, in 2011, 2020 and 2024, with continuing litigation, which is why any specific procedural claim in this area needs a current date attached.

Remember: Title IX contains thirty-seven words and no mention of athletics; the three-part test requires only one of three routes to compliance, and the statute's reach into harassment came through the courts.

Outside the United States

The Convention on the Elimination of All Forms of Discrimination against Women, adopted by the UN General Assembly in 1979 and in force from 1981, has been ratified by the large majority of states, with the United States among the few that signed but never ratified. CEDAW obliges parties to eliminate discrimination in law and practice and to report periodically to a monitoring committee, though many states entered reservations to specific articles, most commonly on family law. In 2019 the International Labour Organization adopted Convention 190, the first international treaty specifically addressing violence and harassment in the world of work.

Working a fact pattern

Try one. A public university requires all students in its nursing programme to complete clinical rotations, and denies a request from a pregnant student to reschedule two rotations, telling her she may withdraw and reapply next year. What are the questions?

First, is this a covered entity? A public university receiving federal funds is covered by Title IX, and as a state actor is also subject to equal protection analysis. Second, what is the claim? Title IX regulations have long treated discrimination on the basis of pregnancy as sex discrimination in education, requiring that pregnancy be treated like other temporary conditions. Third, what is the comparator? If a student with a broken leg would be permitted to reschedule, the differential treatment is the violation, and identifying that comparator is usually where a case is won or lost. Fourth, what remedy is available? Cannon supplies the private right of action and Franklin supplies damages. Notice that the analysis turned on a comparison, not on a general claim of unfairness. That is what legal reasoning in this area looks like.

Common misconceptions

  • Title IX is a sports law. Its text does not mention athletics; the statute governs all education programmes receiving federal funds, including admissions, harassment and pregnancy.
  • The three-part test is a quota. Proportionality is one of three alternative routes to compliance, and an institution needs to satisfy only one.
  • Sex classifications get strict scrutiny. Frontiero fell one vote short in 1973; sex classifications receive intermediate scrutiny, sharpened by the VMI standard.
  • The Equal Pay Act and Title VII cover the same ground. The Equal Pay Act addresses equal work in one establishment and requires no proof of intent; Title VII is broader and reaches hiring, promotion and harassment.
  • Bostock settled all questions about gender identity in law. It construed Title VII's employment protections; questions in education, health coverage, and public accommodation continue to be litigated under other statutes.

Recap

  • Reed v. Reed (1971) was the first decision striking a law for sex discrimination, and the constitutional framework is younger than most students expect.
  • Sex classifications face intermediate scrutiny from Craig v. Boren, tightened by the exceedingly persuasive justification standard in the VMI case.
  • Title VII generated the law of sex stereotyping, hostile environment, same-sex harassment and, in Bostock, protection for sexual orientation and gender identity, with Congress twice reversing the Court by statute.
  • Title IX is thirty-seven words that never mention sports, and its three-part athletics test offers three alternative routes to compliance.
  • CEDAW binds most of the world with widespread reservations, and the United States signed but never ratified it.

Sources

  1. Legal Information Institute. (n.d.). Equal protection and sex discrimination. Cornell Law School. law.cornell.edu
  2. U.S. Department of Justice, Civil Rights Division. (n.d.). Title IX of the Education Amendments of 1972. justice.gov
  3. U.S. Equal Employment Opportunity Commission. (n.d.). Title VII of the Civil Rights Act of 1964. eeoc.gov
  4. Oyez. (n.d.). Case summaries: Reed v. Reed, Craig v. Boren, United States v. Virginia, Bostock v. Clayton County. oyez.org
  5. Office of the High Commissioner for Human Rights. (n.d.). Convention on the Elimination of All Forms of Discrimination against Women. United Nations. ohchr.org
Key terms
Intermediate scrutiny
The standard applied to sex classifications, requiring an important objective and means substantially related to it.
Exceedingly persuasive justification
The VMI formulation requiring that a sex classification's rationale be genuine and not invented in response to litigation.
Sex-plus discrimination
Treating a subset of one sex differently, as in refusing to hire mothers of small children while hiring fathers of small children.
Hostile work environment
Harassment severe or pervasive enough to alter conditions of employment, actionable under Title VII since Meritor.
Sex stereotyping
Penalizing an employee for failing to conform to expectations about how their sex should behave, held unlawful in Price Waterhouse.
Three-part test
The Title IX athletics compliance framework offering proportionality, continuing expansion, or full accommodation of interests as alternatives.
Deliberate indifference
The standard for institutional liability for known student-on-student harassment established in Davis.
CEDAW
The 1979 UN convention on eliminating discrimination against women, widely ratified with reservations and never ratified by the United States.

Module 5: Culture and Difference

What content analysis can and cannot establish about media representation, the small number of causal estimates that survive scrutiny, and gender systems outside the Euro-American frame described in their own terms.

Counting the Screen: Representation and What It Does

  • Explain how content analysis operationalizes a construct such as sexualization, and state what a count of characters can and cannot support.
  • Report the measured representation figures for film, news and reference works, with their sources and time spans.
  • Distinguish self-report, experimental and quasi-experimental evidence on media effects, and evaluate a contested estimate.

A three-line rule from a 1985 comic strip

In 1985 Alison Bechdel drew a half page for her series Dykes to Watch Out For in which one woman explains to another why they will not be seeing a film that evening. She has a rule. She goes only to films that have at least two women in them, who talk to each other, about something other than a man. The last one that qualified, she says, was Alien, on the strength of a brief exchange between two women about the monster. Bechdel has credited the rule to her friend Liz Wallace and has spent the decades since explaining that it was a joke in a comic strip, not a research instrument.

It became one anyway. Film festivals adopted it, Swedish cinemas put it on posters in 2013, and a volunteer database now scores thousands of titles against it. The reason a throwaway gag travelled that far is worth pausing on, because it explains the whole field this lesson covers. The rule converts a vague dissatisfaction, that films seem to treat women as accessories, into a criterion that two strangers can apply to the same film and mostly agree on. Agreement between independent observers is the entire currency of content analysis. Without it you have opinion; with it you have data.

You should also see immediately what the rule cannot do. Gravity fails it: one woman, alone in orbit for most of the running time. A film in which two women discuss shoe shopping for ninety seconds and are then killed passes. The Bechdel test is a floor placed absurdly low, and the fact that a large minority of major films still trip over it is the only claim it was ever designed to support.

So what?: A representation measure is only as good as the question it was built to answer, and the first thing to ask of any such measure is what it would fail to notice.

How you actually count a film

The most sustained count of American film is run by Stacy L. Smith and the Annenberg Inclusion Initiative at the University of Southern California, which has coded the 100 top-grossing films of each year since 2007. Its method is worth walking through, because every representation figure you will ever read was produced by something like it.

First, the unit of analysis. Annenberg codes the independent speaking character: anyone who says one or more words on screen or is addressed by name. That choice determines everything downstream. Count named characters instead and the numbers shift; count screen minutes instead and they shift again. Second, the sampling frame: the year's 100 highest-grossing domestic releases. That is a frame with a built-in bias, since it captures the industry's commercial centre and misses independent and international film entirely, which is a limitation the reports state rather than hide.

Third, and this is where the real work sits, the codebook. Suppose you want to measure whether a character is sexualized. You cannot ask coders to record whether a character seems sexualized, because then you have measured the coders. So the construct is broken into observable indicators: is the character shown in clothing that exposes the chest, midriff or upper thigh; is there partial or full nudity; does another character make a verbal reference to their physical attractiveness. Each is a yes or no judgement a stranger can make from the footage. Fourth, reliability. Two coders work independently on an overlapping subset, and the agreement between them is reported as a statistic. Low agreement means the codebook is vague, and the honest response is to rewrite the codebook rather than publish the number.

What this method cannot reach is meaning. It records that a character appeared in revealing clothing; it cannot record whether the film treated that as her strength, her trap, or a joke at the audience's expense. Content analysis counts what is on the screen. Interpreting it is a separate job with separate methods.

The counts, held still

Here are the better-established figures, each with the instrument that produced it. Read the middle column as an estimate with a method attached, not as a fact of nature.

What was countedThe measured figureSource
Girls and women among speaking characters, 100 top-grossing US filmsAbout 34 percent in 2019, and never outside the low thirties in any year from 2007 onwardAnnenberg Inclusion Initiative
Top-grossing films with a girl or woman as lead or co-leadFewer than half in 2019, which was nonetheless the highest count the series had recordedAnnenberg Inclusion Initiative
Women among the directors of those 100 filmsRoughly one in ten in 2019, itself a record, against low single digits in most earlier years of the seriesAnnenberg Inclusion Initiative
Sexualized clothing indicators by age bandGirls aged 13 to 20 coded at rates close to those for women aged 21 to 39Annenberg Inclusion Initiative
Women among the people heard, read about or seen in news25 percent in 2020, up from 17 percent when the count began in 1995Global Media Monitoring Project
Women among biographies on English WikipediaUnder a fifth, up from roughly 15 percent in the mid 2010sWikipedia editor project tracking

Read down the table and two patterns separate. The behind-the-camera figures and the news figures have moved, slowly and from a low base: eight percentage points in news over twenty-five years, a directing share that went from almost nothing to about one in ten. The on-screen speaking-character share has barely moved at all across thirteen years of counting. That is a strange combination, and it is the most interesting thing on the page. It suggests the constraint is not simply who is hired but something about the shape of the stories themselves, in which the number of characters who need to speak is set by a genre template that has not changed.

The news row rewards a second look. The Global Media Monitoring Project, run every five years since 1995 by volunteers who code a single day of news in over a hundred countries, is the longest-running comparison of its kind. Its 2020 round found women were a quarter of the people appearing in newspaper, television and radio news. Twenty-five years of movement produced a figure still a long way from the population share, and the project's own breakdowns show the gap is widest in politics and economics coverage and narrowest in health and social stories.

What matters here: Representation has moved measurably behind the camera and in news sourcing while the on-screen speaking share has been close to flat since 2007, which means the two are not driven by the same mechanism.

From counting to consequences

Counting is the easy half. The hard question is whether any of it does anything, and the evidence on that comes in three tiers of very different quality.

Tier one, self-report. A 2018 survey commissioned by the Geena Davis Institute asked women working in science and engineering about Dana Scully, the FBI agent and physician in The X-Files. A majority of respondents who were familiar with the character said she had been a role model. That is worth something and it is not evidence of causation. People asked whether a character influenced them have no access to the counterfactual version of themselves who never watched, and the sample was of women already in those fields.

Tier two, experiments. Barbara Fredrickson and Tomi-Ann Roberts proposed objectification theory in 1997: that living in a culture which routinely appraises women's bodies leads women to take an observer's view of themselves, and that this costs attention. In 1998 Fredrickson and colleagues tested it directly. Participants were asked to try on and evaluate a garment alone in a dressing room, either a swimsuit or a sweater, then complete a maths test while still wearing it. Women in the swimsuit condition performed worse on the maths test than women in the sweater condition. Men showed no such difference. The manipulation is small, the setting is artificial, and the mechanism it demonstrates is real: self-observation consumes cognitive resources, and it was triggered by a mirror and a garment.

Tier three, natural experiments at scale. The strongest media-effects evidence comes from cases where something arrived in some places before others for reasons unrelated to the outcome. Robert Jensen and Emily Oster studied the staggered introduction of cable television to villages in rural India in the early 2000s. In villages that got cable, reported acceptance of wife beating fell, reported son preference fell, and girls' school enrolment rose, with the authors noting the size of the attitude shifts was comparable to what several years of additional education would produce. Eliana La Ferrara and colleagues did something similar with the spread of the Rede Globo signal across Brazil, where the arrival of a broadcaster whose telenovelas featured small families was followed by measurably lower fertility, with the largest effects among women closest in age to the main characters.

The dispute over one show, and what would settle it

In 2015 Melissa Kearney and Phillip Levine published an estimate in the American Economic Review that MTV's 16 and Pregnant, which premiered in June 2009, caused a reduction in teen births of about 4 percent in the eighteen months after it aired, which they calculated as roughly a third of the total decline in teen childbearing over that period. Their design used variation in MTV viewership across media markets before the show existed, on the logic that places already watching more MTV would be more exposed to the new programme, and they supported it with Google search and Twitter data showing spikes in contraception-related terms after episodes.

David Jaeger, Theodore Joyce and Robert Kaestner replied that the result did not survive. Their objection was that markets with high MTV viewership already had teen birth rates falling faster than other markets before the show premiered, so the estimated effect could be picking up a pre-existing divergence rather than anything the programme did. When they allowed for market-specific trends, the effect shrank toward nothing.

Notice what this argument is and is not about. Nobody disputes the correlation. The disagreement is entirely about whether the comparison markets were on the same trajectory before treatment, which is the assumption every difference-in-differences design rests on and the assumption that is never directly testable. What would settle it is more pre-period data, or a setting where exposure was assigned rather than chosen. Both teams behaved well: one published a design and its data, the other tested the design's weakest joint. That is what a working literature looks like from the inside, and it is why single headline findings about media effects deserve less confidence than the count data in the table above.

Common misconceptions

That a film passing the Bechdel test is a film with good gender politics. The test asks three yes or no questions about ninety seconds of dialogue. Films with contemptuous portrayals of women pass routinely; films with a single formidable woman alone in the story fail. It measures one specific absence at the population level and was never meant to grade individual films.

That if media shape attitudes, attitudes must be downstream of media. The credible causal estimates in this lesson are modest, specific and mostly come from settings where a medium arrived somewhere it had not existed before. None of them supports the idea that portrayals are the main driver of what people believe about gender, and the flat speaking-character share alongside decades of large attitude change is evidence against it.

That content analysis is subjective because a human does the coding. A human also reads a thermometer. What makes a measurement usable is not the absence of a person but a written rule specific enough that two people applying it independently land in the same place, and a published statistic telling you how often they did.

Putting it together

Representation research earns its keep by being boring and repeatable: a defined unit, a stated frame, a codebook that turns a fuzzy idea into observable indicators, and a reliability figure that lets you judge the whole thing. The resulting counts are steady enough to be useful, and they say that the on-screen speaking share in top-grossing American film has sat near a third for well over a decade while the share of women directing those films and appearing in news has climbed slowly from a very low base.

The effects literature is thinner and should be handled accordingly. The strong results come from natural experiments in which broadcast reached some places before others, and even the best known of the American estimates is contested on grounds that go to the core of the method. The upshot: be confident about what is on the screen, be careful about what it does, and always ask what the instrument was built to notice.

Sources

  1. Smith, S. L., Choueiti, M., and Pieper, K. Inequality in Popular Films series. USC Annenberg Inclusion Initiative. USC Annenberg
  2. World Association for Christian Communication. (2021). Who Makes the News? Global Media Monitoring Project 2020. GMMP
  3. Geena Davis Institute on Gender in Media. Research reports on gender in family entertainment. Geena Davis Institute
  4. Fredrickson, B. L., Roberts, T.-A., Noll, S. M., Quinn, D. M., and Twenge, J. M. (1998). That swimsuit becomes you: Sex differences in self-objectification, restrained eating, and math performance. Journal of Personality and Social Psychology, 75(1), 269-284.
  5. Jensen, R., and Oster, E. (2009). The power of TV: Cable television and women's status in India. Quarterly Journal of Economics, 124(3), 1057-1094.
  6. Kearney, M. S., and Levine, P. B. (2015). Media influences on social outcomes: The impact of MTV's 16 and Pregnant on teen childbearing. American Economic Review, 105(12), 3597-3632; and the reply by Jaeger, D. A., Joyce, T. J., and Kaestner, R. (2020). Journal of Business and Economic Statistics, 38(1).
Key terms
Content analysis
A method that codes media artefacts against a written codebook so that independent observers produce comparable counts.
Unit of analysis
The thing being counted, such as an independent speaking character, chosen before coding and determining every figure that follows.
Codebook
The written rules translating a construct such as sexualization into observable yes or no indicators a coder can apply.
Intercoder reliability
A published statistic reporting how often two independent coders assigned the same code to the same material.
Bechdel test
A three-part screening question from a 1985 comic strip: two women, who talk to each other, about something other than a man.
Objectification theory
Fredrickson and Roberts's 1997 account in which habitual appraisal of women's bodies produces self-observation that consumes attention.
Parallel trends assumption
The requirement in a difference-in-differences design that treated and comparison groups were moving alike before treatment; untestable in principle and the usual point of attack.
Global Media Monitoring Project
A volunteer study coding one day of news across more than a hundred countries every five years since 1995.

Gender Systems Outside the Euro-American Frame

  • Describe hijra, two-spirit, fa'afafine and burrnesha social categories in the terms of the societies that produced them.
  • Explain how colonial law shaped the legal position of these categories, and identify the statutes involved.
  • State three specific errors made when non-Western categories are used as evidence in Western arguments, and avoid them in your own writing.

A courtroom in New Delhi, 15 April 2014

Justices K. S. Radhakrishnan and A. K. Sikri of the Supreme Court of India delivered judgment in National Legal Services Authority v. Union of India. The Court held that transgender persons have a constitutional right to be recognised in the gender they identify with, including as a third gender, and directed central and state governments to treat them as socially and educationally backward classes for the purposes of reservations in education and public employment. Much of the judgment is history: pages on hijras in the epics and in Mughal courts, and on what changed under British administration.

Five years later Parliament passed the Transgender Persons (Protection of Rights) Act 2019, which requires a person seeking recognition to apply to a District Magistrate for a certificate. Trans and hijra activists argued in the streets and in petitions that a magistrate's certificate is the opposite of the self-identification NALSA had recognised. That sequence, a court granting recognition and a legislature attaching a gatekeeper to it, is a good corrective to any story in which legal change moves in one direction.

This lesson looks at four gender systems that developed outside the Euro-American frame. The discipline it asks of you is harder than it sounds: describe each in the terms of the society that produced it, before asking what it means for any argument you already hold.

Remember: These categories are not variants of a Western model, and the fastest way to misread them is to arrive with the question of what they prove.

Hijra: lineage, livelihood, and a law written in 1871

Hijras in India, Pakistan and Bangladesh are usually assigned male at birth, live in feminine social roles, and belong to households organised around a guru and her chelas, disciples who join the household and take its lineage name. The structure is a kinship system, with inheritance, obligation and discipline, and people move between households at real social cost.

The classic livelihood is badhai: arriving at a household after a birth or a wedding to sing, dance and bless the family, and to be paid for it. The transaction rests on a claimed power to confer fertility and, if refused, to withhold or curse. Serena Nanda's fieldwork in Neither Man nor Woman (1990) established much of the ethnographic record in English. Gayatri Reddy's With Respect to Sex (2005), based on work in Hyderabad, pushed against the tendency to read hijra life through sexuality alone: her informants organised their world around izzat, respect, and located themselves through religion, class, kinship and neighbourhood as much as through gender.

Now the legal history, which matters because it recurs across the region. Section 377 of the Indian Penal Code, drafted under Thomas Macaulay and enacted in 1861, criminalised carnal intercourse against the order of nature; the Supreme Court read it down as to consenting adults in Navtej Singh Johar v. Union of India on 6 September 2018. Part II of the Criminal Tribes Act 1871 went further and specifically targeted eunuchs, requiring local registers of hijras and criminalising their appearance in public in female dress and their performance for payment. Neither statute was indigenous. Much of the legal apparatus that made hijra life precarious was written in English by colonial administrators, and versions of it outlived the empire across South Asia, Africa and the Caribbean.

Do not read any of this as a happier alternative. The NALSA judgment itself catalogued exclusion from schooling, employment and health care, and the livelihoods available to many hijras narrow to badhai, begging and sex work. Recognition of a category and decent treatment of the people in it are separate things.

Two-spirit: an English term coined in Winnipeg in 1990

At the third annual intertribal Native American and First Nations gay and lesbian gathering, held near Winnipeg in 1990, participants proposed the English term two-spirit for community use. It was intended to replace berdache, a word anthropologists had used for a century, which entered English through French from a Persian and Arabic root meaning a kept boy. The colonial insult was in the vocabulary itself.

Two-spirit is therefore an English umbrella, thirty-five years old, adopted by Indigenous people in North America for their own purposes. It is not a translation of any single nation's word, and treating it as one repeats the flattening it was coined to correct. The specific categories differ: Diné nádleehí, a term glossed roughly as one who changes or transforms; Lakota winkte; Zuni lhamana. Their roles differed too, across ceremony, craft, mediation and warfare, and by no means every nation had such a category.

One life is documented in unusual detail. We'wha, a Zuni lhamana who lived from about 1849 to 1896, was a weaver and potter of great standing at Zuni Pueblo. In 1886 We'wha spent roughly six months in Washington, hosted by the anthropologist Matilda Coxe Stevenson, demonstrated weaving, and was received by President Grover Cleveland. Will Roscoe's The Zuni Man-Woman (1991) reconstructs the record. The interesting part is not that Washington society was charmed but that almost everyone who met We'wha assumed they were meeting a woman, which tells you the category had no slot waiting for it in the vocabulary of the visited society.

The core of it: Two-spirit is a modern pan-Indigenous English term with a datable origin, standing in for many distinct nation-specific categories that it does not translate.

Samoa: fa'afafine, and a hypothesis someone actually tested

Fa'afafine, from the Samoan fa'a, in the manner of, and fafine, woman, are assigned male at birth and live in feminine social roles within their families and villages. The category is ordinary in the sense that matters: families expect that some sons will be fa'afafine, and there is a recognised place in household labour and obligation.

Samoa is also the site of one of the few places where an evolutionary hypothesis about same-sex attraction has been given a real test. The kin selection hypothesis proposes that a trait reducing an individual's own reproduction can persist if it raises the reproductive success of close relatives, since relatives carry copies of the same genes. Paul Vasey and Doug VanderLaan measured avuncular tendencies, the willingness to invest time and resources in nieces and nephews, and reported that fa'afafine scored higher than either Samoan women or gynephilic Samoan men, and that the elevation was specific to their own kin rather than to children generally. That specificity is what makes the finding interesting, because a generalised fondness for children would not distinguish kin selection from simple nurturance.

Be precise about what this establishes. It shows a measured difference in reported willingness to invest, in one society, on an instrument the researchers designed. It does not close the fitness accounting: nobody has shown that the additional investment translates into enough extra surviving nieces and nephews to pay for the forgone direct reproduction, and comparable elevations have not been found in every society tested. Related roles exist nearby, including fakaleiti in Tonga and mahu in Hawai'i and Tahiti, and they are not interchangeable with one another.

Northern Albania: the Kanun and the burrnesha

In the highlands of northern Albania, Kosovo and Montenegro, customary law known as the Kanun of Leke Dukagjini governed households for centuries and was written down by the priest Shtjefen Gjecovi, published in 1933 after his death. Its provisions on blood feud, property and household headship all assumed a male head. A household that lost its men, whether to feud or to emigration, faced ruin.

The Kanun contained an exit. A woman could swear before village elders to remain celibate for life, and thereafter live as a man: wear men's clothing, carry a weapon, drink and smoke with men, speak in the assembly, head the household and inherit. The oath was public, permanent and honoured. These are the burrnesha, sworn virgins. Antonia Young's Women Who Become Men (2000) collects interviews with a number of them.

What matters analytically is the trigger. This is a role generated by a specific property and kinship system under specific pressure, and several burrnesha interviewed in the twentieth century said plainly that they took the oath because a household needed a head or a betrothal needed escaping. Read that as a claim about identity resembling modern transgender identity and you will get the case wrong; read it as pure social utility and you will miss those who describe the life as the one they wanted. The population is now very small, likely dozens rather than hundreds, and falling, because the conditions that produced the role have gone.

Four systems, side by side

CategoryWhereWhat organises itWhat it is not
HijraIndia, Pakistan, BangladeshGuru and chela lineage households; ritual performance at births and weddingsNot a private identity: it is entry into a household with obligations
Two-spiritNorth America, pan-IndigenousAn English umbrella coined in 1990 for many distinct nation-specific rolesNot an ancient word, and not a single category
Fa'afafineSamoaRecognised feminine social role for people assigned male, embedded in family labourNot equivalent to a Western gay or trans identity
BurrneshaNorthern Albania and neighbouring highlandsSworn celibacy under the Kanun, allowing a woman to take a male household roleNot a third gender: the burrnesha lives as a man

Read the last column across and the family resemblance dissolves. Two of these are not third categories at all. One is a household you join rather than a self you discover. One is an umbrella term whose whole purpose is to refuse a single description. What they share is only that the Euro-American two-box model does not accommodate them, which is a fact about the model.

Recognition in law, and the trouble with league tables

Legal recognition has moved fastest in places rarely cited in these debates. Nepal's Supreme Court ordered recognition of a third gender in the 2007 Sunil Babu Pant case. Pakistan's Supreme Court did so in 2009, followed by the Transgender Persons Act 2018. Bangladesh recognised hijra as a distinct category in 2013. Argentina's Gender Identity Law of 2012 went further than anything in Europe or North America at the time by making legal gender change a matter of self-declaration with no medical or judicial requirement, a model later adopted by Malta in 2015 and Germany in its 2024 self-determination law.

You will also meet composite indices: the UNDP's Gender Inequality Index and the World Economic Forum's Global Gender Gap Index, whose 2024 report put the global gap at roughly 68 percent closed and estimated well over a century to parity at the observed rate of change. Treat these as summaries rather than measurements. A composite index is a weighted sum of components chosen by its authors, and a country can move several ranks by a change in one component, which is why the same country often sits in different positions in the two indices in the same year. Use them to generate questions, then go to the component data to answer them.

Common misconceptions

That the existence of these categories settles a Western argument. It is used in both directions: as proof that gender is culturally constructed, and as proof that all societies simply have ways of accommodating same-sex attraction. Neither inference works, because the categories are heterogeneous in exactly the ways that would have to be uniform for the argument to run. Two of the four in the table are not third genders at all.

That two-spirit is an ancient Indigenous word. It is English, proposed at a gathering near Winnipeg in 1990, and was adopted precisely because the older anthropological term carried a slur in its etymology.

That non-Western societies are uniformly more accepting. Recognition and welfare come apart. India recognises a third gender in constitutional law while the NALSA judgment itself documents exclusion from schooling, work and health care, and much of the criminal law that made hijra life dangerous was drafted in London.

What to carry forward

Four systems, four different organising principles: a lineage household with a ritual livelihood, a modern umbrella term standing over many older roles, a recognised feminine role embedded in family labour, and a sworn oath that solved a problem in customary property law. Colonial statute is a common thread in their legal treatment, from Section 377 of the 1861 Indian Penal Code to Part II of the Criminal Tribes Act 1871.

In short: describe first, in the society's own terms and with its own history attached; only then ask what, if anything, the case bears on. The order matters, because a category recruited into an argument before it has been described will be described to suit the argument.

Sources

  1. Supreme Court of India. (2014). National Legal Services Authority v. Union of India. Case overview
  2. Two-spirit: origin of the term and its use across nations. Wikipedia
  3. United Nations Development Programme. Gender Inequality Index, Human Development Reports. UNDP
  4. Nanda, S. (1999). Neither Man nor Woman: The Hijras of India (2nd ed.). Wadsworth.
  5. Reddy, G. (2005). With Respect to Sex: Negotiating Hijra Identity in South India. University of Chicago Press.
  6. Roscoe, W. (1991). The Zuni Man-Woman. University of New Mexico Press; and Young, A. (2000). Women Who Become Men: Albanian Sworn Virgins. Berg.
  7. Vasey, P. L., and VanderLaan, D. P. (2010). An adaptive cognitive dissociation between willingness to help kin and nonkin in Samoan fa'afafine. Psychological Science, 21(2).
Key terms
Hijra
A South Asian category, usually assigned male at birth, living in feminine roles within guru and chela lineage households with a ritual performance livelihood.
Badhai
The hijra practice of performing blessings at births and weddings for payment, resting on a claimed power to confer or withhold fertility.
Two-spirit
An English umbrella term proposed at a 1990 intertribal gathering near Winnipeg, replacing the anthropological term berdache.
Lhamana
The Zuni category to which We'wha belonged, involving work and dress associated with women alongside other roles.
Fa'afafine
Samoan category, from fa'a, in the manner of, and fafine, woman: assigned male at birth and living in a recognised feminine social role.
Kin selection hypothesis
The proposal that a trait lowering an individual's own reproduction can persist if it raises the reproductive success of close relatives.
Kanun
The customary law of the northern Albanian highlands, written down by Shtjefen Gjecovi and published in 1933, under which the burrnesha oath was made.
Burrnesha
A person assigned female who swore lifelong celibacy before elders and thereafter lived as a man, heading a household and inheriting.
Criminal Tribes Act 1871
Colonial Indian legislation whose Part II required registers of eunuchs and criminalised hijra dress and performance in public.
Composite index
A weighted summary of chosen components, such as the Gender Inequality Index, useful for generating questions rather than for measurement.

Module 6: The Live Disputes

What is actually known about trans and nonbinary populations and where the numbers come from, followed by four arguments that remain unsettled, each given through the evidence its strongest advocates cite.

Trans and Nonbinary Lives: What the Data Support

  • Explain why national estimates of the transgender population differ, and identify the definitional and sampling choices that produce the differences.
  • Distinguish what a large convenience sample can and cannot establish, using the US Transgender Survey as the worked case.
  • State the minority stress model's central prediction and name evidence that bears on it.

Two questions added to a federal survey in July 2021

On 21 July 2021 the Census Bureau launched Phase 3.2 of the Household Pulse Survey with two new items: one asking sex assigned at birth, one asking current gender identity. It was the first Census Bureau household survey to collect gender identity at national scale. The survey itself was an emergency instrument, built quickly in 2020 to track pandemic effects, which is why an addition of this kind could be made in months rather than the years a decennial instrument requires.

Before that, almost everything known about the American transgender population came from three thin sources: state-level modules attached to the Behavioral Risk Factor Surveillance System, a gender identity question added to the Youth Risk Behavior Survey in some jurisdictions, and community surveys recruited online. This lesson is about what those instruments can support. The 2022 National Academies report you met in the first lesson made the operational recommendation that now governs the field: ask two questions, not one, so that the answers can be compared and disaggregated later.

Why this matters: Every disputed claim in the next lesson rests on population numbers, and population numbers are the product of a question, a mode and a frame, all chosen recently and still changing.

How many people, and why the published figures differ

Two respectable estimates, published within weeks of each other in 2022, look inconsistent until you read their definitions.

The Williams Institute at UCLA School of Law, drawing on the state BRFSS modules for adults and the Youth Risk Behavior Survey for teenagers, estimated that about 1.6 million people aged 13 and over in the United States identify as transgender, roughly 0.6 percent of that population. Its striking finding was an age gradient: the estimated share among 13 to 17 year olds was more than double the share among adults aged 25 to 64.

Pew Research Center, surveying adults, reported that 1.6 percent identify as transgender or nonbinary, and that among adults under 30 the figure was 5.1 percent. Same country, same year, figures that differ by a factor of nearly three at the headline level.

Nothing is wrong with either. Pew's category includes nonbinary respondents; the Williams estimate is of transgender identification specifically. Pew surveyed adults; Williams included teenagers. The instruments differ, the modes differ, and the age ranges differ. This is the lesson from intersex prevalence in Module 1 arriving in a new place: before asking whether a number is big, ask what its denominator was and who was counted inside it.

The age gradient is where the interesting argument sits, and there are three live explanations. The first is disclosure: younger cohorts are more willing to report an identity on a survey, exactly as the Gallup identity series suggests for sexual orientation. The second is a genuine cohort difference in identification. The third is measurement: question wording and survey mode behave differently with adolescents than with adults, and the youth estimates lean on a school-administered instrument. What would separate these is a panel that follows the same cohort for two decades, and nobody has one yet, because the questions are barely a decade old.

The largest survey there is, and what it cannot tell you

The 2015 US Transgender Survey collected 27,715 responses, recruited online through several hundred community organisations. The 2022 wave collected 92,329. These are the largest datasets on trans people in existence, and they are convenience samples. There is no sampling frame, so there is no way to weight the respondents back to a population, which means every figure describes the people who answered rather than trans Americans in general. People reachable through community organisations are not a random draw.

Held to that standard, the 2015 results are still worth stating. Thirty-nine percent of respondents showed serious psychological distress in the past month on the Kessler-6 screen, against about 5 percent in the comparable general population sample. Forty percent reported a suicide attempt at some point in their lives, against 4.6 percent nationally. Thirty percent had experienced homelessness at some point. Around three in ten of those who had held a job in the past year reported being fired, denied a promotion, or otherwise mistreated at work because of their gender identity.

Are those magnitudes right? Probably not precisely, and the bias could run in either direction: distressed people may be more motivated to answer a community survey, or the most isolated people may never see it. But the direction is corroborated by probability samples. State BRFSS modules and the Youth Risk Behavior Survey, which do have sampling frames, also find substantially elevated distress and suicidality among transgender respondents. When a convenience sample and a probability sample point the same way, believe the sign and hold the magnitude loosely.

The 2022 wave's early findings run the other way on one measure worth noting: among respondents who had undergone any form of transition, the great majority reported being more satisfied with their lives than before. That too is a convenience sample, and it too should be read as direction rather than as a rate.

The point: A large convenience sample gives you rich detail about respondents and no licence to state a population rate; its value is in what it can be checked against.

Minority stress, and what would falsify it

The dominant explanatory model comes from Ilan Meyer's 2003 review in Psychological Bulletin. Minority stress theory says that the elevated rates of distress in stigmatised populations are produced by identifiable stressors rather than by anything intrinsic to the identity: distal stressors, meaning discrimination, rejection and violence that happen to a person; and proximal ones, meaning the expectation of rejection, the labour of concealment, and internalised stigma.

The model earns its place by making predictions that could fail. If distress is caused by stressors, then it should track the stressors: gaps should be smaller where families are supportive, where schools intervene, and where the law protects. In the US Transgender Survey, respondents whose immediate families were supportive reported far lower psychological distress and far lower rates of suicide attempts than those whose families rejected them, with differences on some measures of tens of percentage points. In a natural experiment on the adjacent population, Julia Raifman and colleagues reported in JAMA Pediatrics in 2017 that state-level legalisation of same-sex marriage was followed by a reduction in adolescent suicide attempts, with the effect concentrated among sexual minority students. That is a policy change acting on a population's stress environment rather than on any individual's identity, which is exactly the shape minority stress predicts.

What the model does not do is licence a leap from correlation to any specific intervention's effect. Supportive families differ from unsupportive ones in many ways beyond support. Keep the model as an organising account of where to look, not as an estimate.

An instrument that created a category, and one that broke

Nonbinary identification is the clearest case of an obvious point: a survey that offers only two boxes measures a nonbinary population of exactly zero. The category becomes visible when an instrument makes room for it, which is why nonbinary estimates begin in the 2010s and why the age profile of those estimates partly reflects who was being surveyed with which form.

The reverse failure is instructive too. England and Wales asked a voluntary gender identity question in the 2021 census, the first time it had been asked. Analysts subsequently found that some respondents with lower English proficiency appeared to have misread the question, inflating counts in exactly the areas where language support was weakest, and in 2024 the UK statistics regulator withdrew the accredited official statistics designation from those estimates. No fraud, no bad faith, and a large national dataset degraded by the wording of one question. When the next contested statistic reaches you, that is the failure mode to check for first.

Common misconceptions

That rising numbers mean something is spreading. Identification figures are measurements of willingness to report on an instrument that mostly did not exist twenty years ago. In Module 1 you saw the same pattern for sexual orientation: identity figures moved much faster than attraction and behaviour figures over the same period. Rising counts are consistent with rising disclosure, with genuine cohort change, and with better instruments, and the available data do not currently separate them.

That the survey shows 40 percent of trans people are suicidal. Three errors in one sentence. The figure is a lifetime attempt rate, not a current state; it comes from a convenience sample with no population weighting; and the comparison figure of 4.6 percent is also lifetime. Cite it as what it is: a rate among respondents to a community survey, corroborated in direction by probability samples.

That better data would settle the political questions. Better data will settle some empirical questions in the next lesson. It will not settle questions about how to trade off competing claims, which are arguments about values that use evidence rather than arguments that evidence can conclude.

Where this leaves us

National measurement of this population is roughly a decade old and improving quickly. The credible estimates run from about 0.6 percent of Americans aged 13 and over identifying as transgender to about 1.6 percent of adults identifying as transgender or nonbinary, and the difference between those figures is definitional rather than factual. The largest datasets are convenience samples that describe their respondents well and their population not at all, and their central findings on distress and discrimination are corroborated in direction by probability samples that have sampling frames.

Worth holding on to: ask which instrument, which definition, which frame, and which year, and most apparent contradictions between two published numbers dissolve before you have to adjudicate anything.

Sources

  1. Herman, J. L., Flores, A. R., and O'Neill, K. K. (2022). How Many Adults and Youth Identify as Transgender in the United States? Williams Institute, UCLA School of Law. Williams Institute
  2. Brown, A. (2022). About 5% of young adults in the U.S. say their gender is different from their sex assigned at birth. Pew Research Center. Pew Research Center
  3. National Center for Transgender Equality. US Transgender Survey, 2015 and 2022 waves. USTS
  4. US Census Bureau. Household Pulse Survey, Phase 3.2 onward. Census Bureau
  5. Meyer, I. H. (2003). Prejudice, social stress, and mental health in lesbian, gay, and bisexual populations. Psychological Bulletin, 129(5), 674-697.
  6. Raifman, J., Moscoe, E., Austin, S. B., and McConnell, M. (2017). Difference-in-differences analysis of the association between state same-sex marriage policies and adolescent suicide attempts. JAMA Pediatrics, 171(4), 350-356.
Key terms
Two-step question
Asking sex assigned at birth and current gender identity as separate items so responses can be compared and disaggregated.
Convenience sample
A sample with no sampling frame, recruited through available channels; it describes respondents and supports no population rate.
Sampling frame
The defined list from which a probability sample is drawn, and the thing that makes weighting to a population possible.
BRFSS
The Behavioral Risk Factor Surveillance System, a state-based probability survey whose optional modules supplied early transgender adult estimates.
Minority stress
Meyer's model attributing elevated distress in stigmatised groups to distal stressors such as discrimination and proximal ones such as concealment.
Distal and proximal stressors
Events that happen to a person versus internal processes such as expected rejection and internalised stigma.
Kessler-6
A six-item screen for serious psychological distress in the past month, which allows comparison against general population benchmarks.
Accredited official statistics
A UK designation of statistical quality, withdrawn from the 2021 census gender identity estimates in 2024 after question comprehension problems emerged.

Four Arguments That Are Not Over

  • State the strongest version of each side in the disputes over youth gender medicine, sport eligibility, single-sex provision and data collection.
  • Separate the empirical question inside each dispute from the question about values, and name what evidence would move the empirical part.
  • Distinguish transgender athletes from athletes with differences of sex development, and explain why conflating them corrupts the argument.

A report published on 10 April 2024

Hilary Cass, a paediatrician and former president of the Royal College of Paediatrics and Child Health, published the final report of her independent review of gender identity services for children and young people in England. NHS England had commissioned it in 2020 after concerns about the country's single specialist clinic. The review commissioned systematic reviews of the evidence from the University of York, and its findings became the most argued-over document in this field anywhere in the world within a week of publication.

This lesson takes four disputes that are genuinely open and works each the same way: the strongest case on one side, the strongest case on the other, the evidence each rests on, and what finding would actually move it. Where a dispute contains a question about values rather than about facts, the lesson says so instead of pretending a study will resolve it.

One: medical treatment of adolescents

The case for restriction. The York systematic reviews appraised the published studies of puberty suppression and of masculinising and feminising hormones in under-18s and found that the great majority scored poorly on standard quality appraisal: small samples, short follow-up, high loss to follow-up, and outcomes measured with instruments that varied between studies. The review also documented a changed case mix. Referrals to the English service rose steeply through the 2010s, and the referred population shifted toward adolescents registered female at birth presenting in the teenage years, with high rates of co-occurring autism and mental health difficulties, a group not well represented in the older Dutch cohort studies on which practice was founded. If the population has changed, older follow-up data may not transfer. NHS England ended routine prescription of puberty blockers outside a research protocol in March 2024 and committed to running a clinical trial. Sweden's National Board of Health and Welfare in 2022 and Finland's COHERE guidance in 2020 had already restricted hormonal treatment of minors to research settings or exceptional cases, prioritising psychosocial support.

The case against restriction. The Endocrine Society's clinical practice guideline, the World Professional Association for Transgender Health's Standards of Care version 8 (2022), and the American Academy of Pediatrics hold that individualised care after assessment, including suppression and hormones for some adolescents, is appropriate. Critics of the York appraisals, including an analysis published in 2024 by a group of Yale-affiliated clinicians and lawyers, argued that downgrading studies for lacking randomisation and blinding applies a standard that cannot be met here: you cannot blind an adolescent to whether puberty is proceeding, and the ethics of a placebo arm are contested. Applied consistently, they argued, the same appraisal would strip the evidence base from large areas of paediatric medicine that nobody proposes to stop. They further point to follow-up cohorts reporting low rates of regret and discontinuation, and to the fact that restriction is itself an intervention with outcomes that ought to be measured.

What would move it. Prospective cohorts with pre-registered outcomes, follow-up measured in years rather than months, comparison against the closest realistic alternative rather than against nothing, and consistent instruments. England's decision to run a trial rather than impose a permanent prohibition is, whatever one thinks of the pause, an attempt to generate exactly that.

American law took a different route. In United States v. Skrmetti, decided in June 2025, the Supreme Court upheld a Tennessee statute barring puberty blockers and hormones for gender transition in minors. The majority reasoned that the law classified on age and on medical indication rather than on sex, so only rational basis review applied, and under that standard the legislature's judgement about a contested medical question stood. The dissent argued the classification is a sex classification on the face of the statute, because whether a drug may be prescribed depends on the sex of the patient receiving it: testosterone for one purpose is lawful, testosterone for another is not. Return to Module 4 and you will see the whole argument turns on the tier-of-scrutiny machinery you learned there.

Bottom line: the disagreement about adolescent medicine is not mainly about whether the evidence is thin, which nearly everyone concedes; it is about what a clinician should do while the evidence is thin, and that is a question about risk tolerance under uncertainty.

Two: eligibility in women's sport

Start by separating two populations that headlines routinely merge. Transgender women are people registered male at birth who identify and live as women. Athletes with differences of sex development, such as Caster Semenya, are women, raised as women, with a variation in sex development that in some cases raises circulating testosterone. These are different biological situations, they are governed by different regulations, and an argument that moves between them is not an argument.

The case for restrictive eligibility. At elite level, the performance gap between men and women is around 10 to 13 percent in running and swimming events and substantially larger in strength and power events. Emma Hilton and Tommy Lundberg's 2021 review in Sports Medicine assembled the evidence on what testosterone suppression does to that advantage and reported that lean body mass, muscle cross-sectional area and strength decline modestly, on the order of a few percent, over the first year of suppression. If the gap is 10 percent and suppression removes something closer to 5 percent of muscle mass, a substantial part of the advantage conferred by male puberty persists, including skeletal dimensions that do not revert at all. World Athletics moved in March 2023 to exclude athletes who had experienced male puberty from female world-ranking competition, and World Aquatics adopted a comparable policy in 2022 alongside an open category.

The case for inclusive eligibility. The International Olympic Committee's 2021 Framework on Fairness, Inclusion and Non-Discrimination holds that no athlete should be presumed to have an unfair advantage on the basis of gender identity or sex variations, and that eligibility should be set sport by sport on evidence specific to that sport, since the physiology that matters in shooting is not the physiology that matters in weightlifting. The numbers of affected athletes are very small. And the history of enforcement is genuinely bad: compulsory sex testing of female competitors ran from 1968 to 1998 and was abandoned because it produced false accusations and no benefit, while the modern regulations required Semenya to suppress her natural testosterone in order to compete. She lost at the Court of Arbitration for Sport in 2019; the Grand Chamber of the European Court of Human Rights ruled in 2025 that Switzerland had failed to subject her case to the rigorous judicial review it required, a holding about procedure that did not decide whether the regulations themselves are lawful. Enforcement also falls unevenly: the women pulled aside for scrutiny have disproportionately been women from the Global South and women whose appearance departs from expectation.

What would move it. Sport-specific measurement of retained advantage after defined suppression periods, since the current evidence is thin outside a few sports. Data on whether open categories attract entries. And clarity from federations about which value they are optimising, because a rule cannot maximise inclusion and minimise residual advantage at once, and choosing between them is not something a physiology paper can do for you.

Three: single-sex provision

In April 2025 the UK Supreme Court decided For Women Scotland v The Scottish Ministers, holding that the words man, woman and sex in the Equality Act 2010 refer to biological sex. The consequence is that single-sex services and associations may be defined on that basis under the Act's exceptions.

The case for sex-based provision. Some services exist because of a category defined by risk and by what users can predict about who else will be present: rape crisis centres, prison accommodation, hospital wards, changing rooms, intimate searches. Providers need a rule that front-line staff can administer without adjudicating anyone's inner life, and a self-identification rule leaves them with no rule at all. Proponents emphasise that the argument is about provision and safeguarding rather than about whether anyone's identity is real.

The case against. A blanket rule sweeps in people who present no risk whatever, and it imposes measurable harm on them, including exclusion from health care and from services after victimisation. The offending data invoked in these debates are drawn from very small numbers: the transgender prison population in England and Wales is counted in the hundreds, and conclusions built on a few dozen cases are fragile in both directions. And a rule stated as biological sex is enforced in practice by staff looking at people, which means the scrutiny lands on any woman who looks atypical.

What would move it. Linked administrative data with adequate numbers, and evaluation of whether specific policies achieve their stated aims. What no data will settle is how to weigh a small probability of serious harm to one group against a certain harm to another, which is a question about values that people can answer differently while reading the same table.

What matters here: in this dispute the empirical component is unusually small and the numbers unusually thin, so an argument conducted entirely in statistics is generally an argument in disguise.

Four: what the state should record

The last dispute is the quietest and probably the most consequential, because it determines what anybody can measure about the other three.

Sex and gender identity are different variables that answer different questions. Cervical screening pathways and drug dosing need natal sex; the analysis of discrimination in employment needs gender identity. Collect only one and the other becomes unanswerable, retrospectively and permanently, because you cannot go back and ask. That is the argument for the two-step approach the National Academies recommended in 2022, and it is a strong one.

The arguments for caution are real too. As you saw in the previous lesson, adding an identity question to the England and Wales census in 2021 produced estimates that the UK statistics regulator later stripped of their accredited status because of comprehension problems. Asking every respondent for their sex at birth is intrusive for a large majority who will never be analysed on that variable. And small populations in general-purpose surveys produce estimates with confidence intervals wide enough to support almost any headline, which is how the same dataset ends up cited by both sides.

None of these is a reason not to measure. They are reasons to cognitively test questions before fielding them, to publish the wording, and to report uncertainty honestly rather than as a footnote.

How to argue about any of these

Three habits, and they will serve you outside this subject too.

  1. Split the question. Almost every dispute here contains an empirical part that evidence can move and a values part that it cannot. Youth medicine contains both: what happens to these adolescents over ten years is empirical; how much uncertainty is acceptable before intervening is not.
  2. State your defeater. Before you argue, write down the finding that would change your position. If you cannot write one, you are not holding a position about the world.
  3. Check the population. Ask whether both sides are describing the same people. Transgender athletes and athletes with differences of sex development, prison populations and general populations, one clinic's referrals and a national cohort: most apparent contradictions in this field are two accurate statements about two different groups.

Common misconceptions

That these four are equally open. They are not. The physiology of retained advantage after suppression is measured, if incompletely and in few sports. The long-run outcomes of adolescent treatment are genuinely thin, and both sides say so. The provision dispute is mostly not empirical at all. Treating all four as equally unsettled is as wrong as treating any of them as closed.

That a professional body's position settles a scientific question. A guideline is evidence about the collective judgement of a field, which is worth something. It is not a substitute for the studies, and in this area several national bodies reading the same literature have reached different conclusions, which by itself tells you the literature does not compel one answer.

That the loudest positions are the actual poles. The published disagreement between the Cass Review and its critics is narrower than the public argument about it: both sides accept the evidence base is weak, and both accept some adolescents benefit from some care. They differ on what follows.

The short version

Four disputes, four different structures. Adolescent medicine turns on how to act under acknowledged uncertainty, and the strongest evidence in play is systematic reviews reporting that the underlying studies are weak. Sport turns on a measured residual advantage after suppression, on the order of a few percent of muscle loss against a performance gap around 10 to 13 percent, and on a choice between values that no measurement can make. Single-sex provision has the thinnest data and the largest values component. Data collection determines whether any of the others can be settled at all.

Across the whole course the same move has kept working. Ask which construct was measured, ask what the denominator was, ask what the instrument could not see, and ask what would change your mind. The upshot: that habit is the transferable part of this subject, and it will still be useful when every specific number in these fifteen lessons has been superseded.

Sources

  1. Cass, H. (2024). Independent Review of Gender Identity Services for Children and Young People: Final Report. The Cass Review
  2. World Professional Association for Transgender Health. (2022). Standards of Care for the Health of Transgender and Gender Diverse People, Version 8. WPATH
  3. International Olympic Committee. (2021). Framework on Fairness, Inclusion and Non-Discrimination on the Basis of Gender Identity and Sex Variations. IOC
  4. Supreme Court of the United States. (2025). United States v. Skrmetti. Supreme Court of the United States
  5. UK Supreme Court. (2025). For Women Scotland Ltd v The Scottish Ministers. UK Supreme Court
  6. Hilton, E. N., and Lundberg, T. R. (2021). Transgender women in the female category of sport: Perspectives on testosterone suppression and performance advantage. Sports Medicine, 51, 199-214.
  7. National Academies of Sciences, Engineering, and Medicine. (2022). Measuring Sex, Gender Identity, and Sexual Orientation. National Academies Press.
Key terms
Systematic review
A structured search and quality appraisal of all studies meeting stated criteria, which reports the strength of a literature rather than adding new data.
Case mix
The composition of a patient population; a change in case mix is a reason older follow-up studies may not transfer to current patients.
Rational basis review
The most deferential tier of constitutional scrutiny, applied in Skrmetti on the majority's holding that the statute classified by age and medical use rather than by sex.
Difference of sex development
A variation in the development of chromosomes, gonads or anatomy; athletes with a DSD are not transgender and are governed by separate regulations.
Retained advantage
The portion of performance advantage from male puberty that persists after a defined period of testosterone suppression.
Open category
A competition class open to all entrants regardless of sex or gender, adopted by some federations alongside a protected female category.
Defeater
The specific finding that would change your position, written down in advance; a position without one is not a claim about the world.
Cognitive testing
Pre-fielding research on how respondents actually understand a survey question, absent or inadequate in several of the measurement failures in this course.

Open the interactive version with quizzes and progress →