🗣️ Linguistics · Undergraduate · LING 320

Sociolinguistics: Language in Society

A complete, college-level course in sociolinguistics, the study of language as social behavior. It begins with the discipline's founding move: replacing opinions about good and bad speech with counted evidence about how real people actually talk. You will learn what a linguistic variable is, how sociolinguists collect data despite the observer's paradox, and how William Labov's studies on Martha's…

Start the interactive course (quizzes, progress, videos) →

Free forever. No sign-up, no ads. 16 lessons. The full lesson text is below so you can read it right here.

Module 1: Foundations: Language as Social Behavior

What sociolinguistics studies and how it studies it: the descriptive stance argued from evidence rather than asserted, the linguistic variable as the field's basic unit of measurement, and the practical craft of collecting speech data honestly and ethically.

What Sociolinguistics Studies: Description, Prescription, and the Evidence

  • Define sociolinguistics and explain how it differs from the structural linguistics you met in an introductory course.
  • Distinguish the claim that a form is socially stigmatized from the claim that it is linguistically defective, and explain why only the first survives contact with data.
  • Marshal three independent lines of evidence for the descriptive stance, and state honestly what descriptivism does and does not imply about standard-language prestige.

The big picture

Somewhere today, a hiring manager will glance at a resume, hear a voice on the phone, and decide the candidate is not quite right. A teacher will circle a sentence in red that is perfectly grammatical in the dialect the child speaks at home. A juror will find one witness more credible than another, and will not be able to say exactly why. None of these people think of themselves as making a linguistic judgment. All of them are. Sociolinguistics is the field that studies what is actually happening in moments like these, and it does so by treating language as what it has always been: social behavior, performed by people who belong to communities and who are read by other people as belonging to them.

An introductory linguistics course teaches you that language is a structured system. You learn that speech is built from a finite inventory of sounds, that words are assembled from morphemes, and that sentences are built by rules operating over hierarchical structure. That course, in effect, studies the system as if one idealized speaker used it uniformly. This course studies what happens when you put the system back into a room full of actual people. The moment you do, one fact leaps out: nobody speaks uniformly. Speakers vary, and their variation is not noise. It is patterned, and it is patterned along social lines you can measure.

That is the field's founding insight and its whole research program. Where an earlier generation of linguists treated the messiness of real speech as a nuisance to be idealized away, sociolinguists discovered that the messiness is itself lawful. If you count carefully, you find that a speaker's pronunciation of a single vowel correlates with their income, their neighborhood, their age, who they are talking to, and how much attention they are paying. Those correlations are stable enough to predict. Language variation, it turns out, is one of the most sensitive instruments we have for reading social structure.

Key idea: Sociolinguistics studies language as social behavior, and it rests on the discovery that variation in speech is not random noise but a lawful, measurable reflection of social structure.

What the field actually covers

The territory is broad, and it helps to see the map before walking it. At one end sits what is often called variationist sociolinguistics, or micro-sociolinguistics: the quantitative study of how small linguistic features, a vowel here, a consonant there, a verb ending, distribute across social groups and across situations. This is the tradition that this course spends its first two modules on, because its methods are the field's backbone.

At the other end sits the sociology of language, or macro-sociolinguistics: the study of whole languages in whole societies. Which language do people use in parliament, in church, at home, on television? What happens when two languages share a country? Why do communities abandon a language within three generations, and what does it take to bring one back? Those questions occupy Module 5.

Between the two sits the study of language in interaction: how politeness works, how conversations are organized, what happens when a speaker of one variety is judged by a speaker of another in a courtroom or a rental office. That is Module 6. The unifying thread across all three is that language does two jobs at once. It carries propositional content, the thing you are saying, and it carries social information, the signal of who you are and how you are positioning yourself right now. Structural linguistics studies the first job well. Sociolinguistics studies the second.

Key idea: The field spans the quantitative study of small variables, the study of whole languages in societies, and the study of language in face to face interaction, unified by the recognition that every utterance carries social meaning alongside its content.

Two very different claims that people constantly confuse

Now to the argument that organizes everything else in this course. When someone says a form is bad English, they are almost always running together two claims that need to be pulled apart, because one of them is true and the other is false.

The first claim is social: this form is stigmatized. Using it in certain settings will cost you. People will judge you, and some of those people control jobs, grades, and loans. That claim is straightforwardly true, and this course takes it seriously rather than waving it away. The second claim is linguistic: this form is defective, illogical, lazy, or the result of a failure to learn rules. That claim is testable, and when tested against data it fails, every time, in every variety anyone has bothered to study carefully.

The confusion between the two is what linguists call the difference between description and prescription. A descriptive rule states what speakers of a variety actually do. A prescriptive rule states what someone thinks they ought to do. Sociolinguistics is descriptive, not because linguists are indifferent to social consequences, but because you cannot study the social consequences of a form until you have accurately described what the form is. You would not trust a doctor who diagnosed diseases by whether the symptoms sounded distasteful. The rest of this lesson lays out the evidence, because the descriptive stance deserves to be argued rather than announced.

Key idea: The claim that a form is stigmatized and the claim that a form is defective are separate claims; the first is often true and consequential, the second has never survived empirical testing.

Evidence one: stigmatized forms are rule-governed

The strongest evidence is the simplest. Take any feature that gets condemned as sloppy, and look for the rule. You will find one, and it will usually be more intricate than the standard alternative.

Consider the sentence He be working. Speakers of standard varieties often hear this as a failed attempt at He is working, produced by someone who has not learned to conjugate. It is not. In African American English, the uninflected be marks habitual aspect: it means he works regularly, as a matter of routine. He working, without be, means he is working right now. These are different sentences with different meanings, and speakers of the variety do not confuse them. Module 4 works through the system in detail, but note what this already shows. The variety makes a grammatical distinction that Standard English cannot make in a single verb form at all. Standard English has to reach for an adverb: he usually works. On this point the stigmatized variety is not simpler. It is richer.

Or take the double negative, as in I did not see nothing. It is condemned in English on the theory that two negatives make a positive, a rule borrowed from logic and applied to grammar by analogy in the eighteenth century. But language is not arithmetic. Standard French requires two negative elements in a clause, ne and pas. Standard Spanish requires no and nada together in the sentence no vi nada, literally I did not see nothing, and every educated Spanish speaker produces it. Negative concord is ordinary grammar across the world's languages, and it was ordinary in English too: Chaucer used multiple negatives freely. Nothing about the construction is illogical. What changed was its social standing.

This generalization holds up wherever anyone has checked. Every human variety that has been described in detail, including every variety whose speakers are told they speak badly, turns out to be a complete rule-governed system that its speakers use with precision and that outsiders systematically misread. The Linguistic Society of America states this as the professional consensus of the field, not as a courtesy.

Key idea: Stigmatized forms consistently turn out to encode systematic distinctions, sometimes finer ones than the standard variety makes, which is exactly the opposite of what a deficiency account predicts.

Evidence two: prescriptive rules have datable inventors

The second line of evidence is historical. If prescriptive rules tracked logic or clarity, they would not have birthdays. Many of them do.

The rule against splitting an infinitive, as in to boldly go, was essentially unknown before the nineteenth century. Its rationale came from Latin, where an infinitive is a single word and therefore cannot be split; English infinitives are two words, so the analogy never applied. The rule against ending a sentence with a preposition traces largely to the poet John Dryden in 1672, who disliked the construction and revised his own earlier work to remove it. Robert Lowth's influential grammar of 1762 recorded a preference in the same direction, though Lowth himself acknowledged that the stranded preposition suited the familiar style of English. Neither rule describes English. Both describe the taste of particular men whose opinions found an audience.

The table below sets a few well known prescriptions beside what the record actually shows.

Prescriptive ruleStated rationaleWhat the historical record shows
Never split an infinitiveInfinitives are indivisibleTrue of Latin, not of English; the rule appears only in the 1800s
Never end a sentence with a prepositionElegance and Latin word orderTraced to Dryden's personal taste in 1672; stranding is native English
Two negatives make a positiveLogicNegative concord is standard in French, Spanish, and older English
Singular they is an errorNumber agreementAttested since the 1300s, used by Chaucer, Shakespeare, and Austen
Ain't is not a wordIt is not standardA regular contraction that lost prestige in the 1800s, formerly used across classes

Notice the pattern. In each case the rule postdates the construction, often by centuries, and the rationale offered turns out on inspection to be about something other than English grammar. A rule that has to be taught, that speakers violate constantly without communication breaking down, and that has a documented origin in one person's preference is not a fact about the language. It is a fact about the history of taste.

Key idea: Many famous prescriptive rules can be dated to specific writers and specific centuries, long after the constructions they condemn were in ordinary use, which is not how facts about a language behave.

Evidence three: the same feature flips prestige across communities

The third line of evidence is the decisive one, because it isolates the social variable while holding the linguistic one constant. If some pronunciations were intrinsically better, prestige would track features. It does not. It tracks who uses them.

Consider the pronunciation of r after a vowel, in words like car and fourth. In most of England, dropping that r is the prestige pattern: it is a feature of Received Pronunciation, the accent long associated with the BBC and elite schooling. In New York City, dropping the same r has been a marker of lower social status, and pronouncing it is the prestige pattern. One sound, two cities, opposite social values. No property of the sound itself can explain that, because the sound is the same in both places. Only the social facts differ.

The same reversal happens across time in one place. In New York, r-lessness was itself the prestige norm in the nineteenth century and lost that status in the twentieth. If a feature can be prestigious and then stigmatized in a single community within living memory, without changing acoustically at all, then prestige is plainly a property of the community rather than of the sound. Labov built an entire experimental design around this fact, and Module 2 walks through it.

Key idea: A single linguistic feature can be prestigious in one community and stigmatized in another, or in the same community at different times, which shows that prestige attaches to speakers rather than to sounds.

What descriptivism does not claim

Here is where careless versions of this argument go wrong, and where students are right to push back. Saying that all varieties are linguistically equal is not the same as saying that all varieties are socially equal. They are not, and pretending otherwise would be a disservice to anyone who has to navigate a job interview.

Standard-language prestige is real, and its consequences are measurable. Later lessons cover audit studies in which the same person, speaking in different varieties, receives different responses when inquiring about apartments. Students who write in a stigmatized variety on standardized tests lose points. Witnesses whose speech is unfamiliar to jurors are believed less. These are facts about the world, and no amount of correct linguistic analysis makes them go away. A course that told you variation carries no cost would be lying to you.

What the descriptive stance does claim is narrower and firmer. It claims that the cost is imposed socially rather than caused linguistically, and therefore that the remedy is a question of policy, education, and justice rather than of repair. It also claims that mistaking social stigma for cognitive deficit produces predictable harm: children misdiagnosed with language disorders, students taught that their home speech is broken, and speakers penalized for something no one can defend as an actual defect. Distinguishing the two claims is not a moral flourish added to the science. It is the difference between an accurate account and an inaccurate one, and getting it right is the reason the field exists.

Key idea: Descriptivism holds that varieties are linguistically equal but not socially equal, which is why the costs speakers face are best understood as social facts requiring social remedies rather than linguistic defects requiring correction.

Common misconceptions

  • Sociolinguistics is just linguistics with opinions attached. The field is quantitative and empirical. Its central claims come from counted data, controlled experiments, and replications across dozens of communities, not from advocacy.
  • Descriptivists think correctness does not exist. Descriptivists think correctness is relative to a variety and a setting. Every variety has forms its speakers reject as ungrammatical; the question is whose norms are being applied and with what authority.
  • Standard English is the original, and dialects are what happened to it. The reverse is closer to true. Standard English descends from one regional variety that gained institutional power, and it has been changing the whole time, as Module 3 documents.
  • If all varieties are equal, teaching the standard is oppressive. Most sociolinguists argue the opposite: because prestige is real, students deserve full command of the standard, taught as an addition rather than as a replacement or a correction.
  • Casual speech is degraded formal speech. Casual and formal styles are different registers with different rules, and every speaker commands several. Neither is a damaged version of the other.

Recap

  • Sociolinguistics studies language as social behavior, founded on the discovery that variation is patterned and measurable rather than random.
  • The field spans quantitative variation studies, the sociology of whole languages, and the analysis of language in interaction.
  • The claim that a form is stigmatized differs sharply from the claim that it is defective; only the first holds up.
  • Stigmatized varieties are rule-governed and sometimes encode finer distinctions than the standard, as habitual be shows.
  • Famous prescriptive rules often have datable inventors and rationales borrowed from Latin or from logic rather than from English.
  • The same feature can carry opposite prestige in different communities, which locates prestige in society rather than in sound.
  • Linguistic equality does not mean social equality, and the real consequences of standard-language prestige are part of the field's subject matter.

Sources

  1. Linguistic Society of America. (n.d.). What is sociolinguistics? linguisticsociety.org
  2. Linguistic Society of America. (n.d.). What is linguistics? linguisticsociety.org
  3. Britannica. (n.d.). Sociolinguistics. Encyclopaedia Britannica. britannica.com
  4. Labov, W. (1963). The social motivation of a sound change. Word, 19(3), 273-309. doi.org/10.1080/00437956.1963.11659799
  5. MacNeil, R., & Cran, W. (2005). Do you speak American? PBS. pbs.org
Key terms
Sociolinguistics
The study of language as social behavior, including how linguistic variation correlates with social factors and what social work language does.
Descriptive rule
A statement of what speakers of a variety actually do, arrived at by observing and counting real usage.
Prescriptive rule
A statement of what someone believes speakers ought to do, typically enforced by institutions and reflecting the prestige of particular groups.
Variety
A neutral cover term for any distinguishable form of a language, used to avoid the ranking implied by words like dialect or accent.
Stigmatized form
A linguistic feature that carries social penalty in certain settings, regardless of whether it is structurally regular.
Habitual be
The uninflected be of African American English, which marks an action as recurring or customary rather than ongoing at the moment of speech.
Negative concord
The grammatical requirement or option for more than one negative element in a single clause, standard in many languages including French and Spanish.
Standard language ideology
The widely held belief that one variety is the correct or natural form of a language and that others are deviations from it.

The Linguistic Variable: Making Variation Countable

  • Define the linguistic variable and its variants, and explain why the concept is what turned impressions about accent into measurable science.
  • Apply the envelope of variation and the principle of accountability to decide which tokens in a recording actually count.
  • Distinguish internal from external conditioning factors, and place a variable on Labov's scale of indicator, marker, and stereotype.

The big picture

Suppose you tell me that working class New Yorkers drop their r sounds. I ask a fair question: how do you know, and how much? You could point to someone you heard. I could point to a working class New Yorker who pronounced every r in the sentence. Now what? We are stuck trading anecdotes, which is exactly where the study of accent sat for most of its history.

The move that broke the deadlock was deceptively small. Instead of asking whether a group uses a form, ask how often, out of all the chances they had. That single reframing turns a shouting match into an arithmetic problem. It also fits the data far better, because real speakers are almost never categorical. The same New Yorker pronounces the r in some words and not others, more in careful speech and less in casual speech, more when reading a word list and less when telling a story about a childhood fight. The pattern lives in the proportions, not in the presence or absence of the feature.

The unit that makes this possible is the linguistic variable. It is the central tool of the field, and this lesson builds it carefully, because almost everything in Modules 2 through 4 is an application of it. If you understand what a variable is, what counts as a token of it, and how to keep yourself honest while counting, you can read the primary literature of sociolinguistics without needing anyone to interpret it for you.

Key idea: The linguistic variable reframes questions about who speaks how as questions about relative frequency, which is what makes variation measurable rather than merely arguable.

Variables and variants

A linguistic variable is a point in the grammar where speakers have more than one way of doing the same thing. The alternatives are called variants. By convention, variables are written in parentheses to distinguish them from phonemes in slashes or sounds in brackets. So the New York r is the variable (r), with two variants: r pronounced and r absent.

The classic example, and the one worked hardest in the literature, is (ing): the ending of words like walking, singing, and something. It has two common variants. One ends in the velar nasal [ŋ], the other in the alveolar nasal [n]. People describe the second as dropping the g, which is a spelling-based misdescription worth correcting early. Nothing is dropped. The speaker produces a nasal consonant in either case; the tongue simply makes contact at the alveolar ridge rather than at the velum. That is a substitution of one sound for another, not a deletion, and the alveolar variant is in fact the older form historically. Calling it lazy gets both the phonetics and the history wrong.

Variables are not limited to pronunciation. Morphosyntactic variables include was and were leveling, as in we was there against we were there. Discourse variables include the quotative system, where a speaker introduces reported speech with said, went, or be like: and then she was like, no way. Lexical variables include soda against pop against coke for a carbonated drink. Each is a place where speakers choose, and each choice carries social information.

Key idea: A linguistic variable is a site of structured choice with two or more variants, written in parentheses, and it may be phonological, morphosyntactic, discourse-level, or lexical.

The envelope of variation

Before you can count, you have to decide what counts. This is the envelope of variation: the precise set of contexts in which the variable could occur, so that each occurrence is a genuine choice between the variants and not something else in disguise.

Take (ing) again. If you simply searched a transcript for the letters ing, you would collect words like sing, king, thing, and ring. Those must be excluded. No English speaker says sin for sing, so those tokens are not choices at all. Including them would flood your data with a variant that was never in competition. The envelope for (ing) is unstressed word-final -ing, which appears in progressive participles like running, in gerunds like the running of the race, in some nouns like ceiling and morning, and in the pronoun-like something and nothing.

Even inside that envelope, the grammar matters. Researchers going back to John Fischer's 1958 study of children in a New England village found that grammatical category conditions the choice: verbal forms like running and hitting favor the alveolar variant, while nominal forms like ceiling and morning favor the velar one. That is a fact about the linguistic environment, and if you do not track it you will mistake a grammatical effect for a social one. If one speaker happens to talk about buildings and another about actions, their raw percentages will differ for reasons that have nothing to do with class.

Key idea: The envelope of variation defines exactly which tokens represent a real choice between variants, and drawing it wrongly imports linguistic effects into your social conclusions.

The principle of accountability

Here is the discipline that separates sociolinguistics from impressionistic reporting. Labov's principle of accountability requires that you report every case where the variable could have occurred, including the cases where the variant you are interested in did not appear.

This sounds obvious and is constantly violated in casual reasoning. Human attention is drawn to the salient variant. If you are listening for r-lessness, you will notice every dropped r and glide past dozens of pronounced ones. Your impression will be that the speaker is r-less, when the count might show forty percent. The principle says: build the denominator first. Every token inside the envelope goes into the total, and the number you report is the count of one variant over that total.

Work an example. You record two speakers for twenty minutes each and extract every token inside the (ing) envelope.

SpeakerAlveolar [n] tokensVelar [ŋ] tokensTotal in envelopePercent alveolar
Speaker A, casual conversation62188078
Speaker A, reading a passage9314023
Speaker B, casual conversation21598026
Speaker B, reading a passage238405

Read what this table tells you. Speaker A uses the alveolar variant more than Speaker B in both styles, so the two speakers differ. But both speakers shift the same direction when reading, so both are responding to the situation. Neither speaker is categorical anywhere: even Speaker B, the most standard user here, produces two alveolar tokens while reading. That combination of stable between-speaker differences and parallel within-speaker shifting is the signature finding of the whole field, and you can only see it because you counted the denominator.

Key idea: The principle of accountability requires counting every context where the variable could have appeared, which converts selective impressions into a rate that can be compared across speakers and styles.

Internal and external conditioning

Every variable is pushed around by two families of factors, and a competent analysis separates them.

Internal, or linguistic, factors come from inside the grammar. For (ing), grammatical category is one. For consonant cluster reduction, as in tes for test, the following sound matters enormously: reduction is far more likely before a consonant, as in tes case, than before a vowel, as in test area. For was and were leveling, whether the subject is singular or plural, and whether the clause is affirmative or negative, both make a difference.

External, or social, factors come from outside: socioeconomic class, age, gender, ethnicity, region, network, and the speech situation itself. These are the factors Module 2 is about.

Because both families operate at once, sociolinguists use multivariate statistical models to estimate each factor's contribution while holding the others constant. Historically this was done with a program called Varbrul, later Goldvarb, and now typically with mixed-effects logistic regression in R. The output is a set of factor weights, numbers between zero and one, where values above 0.5 indicate that a factor favors the variant and values below 0.5 indicate that it disfavors it. You do not need the mathematics to read the literature. You need to know that when a paper reports that following consonant favors deletion at 0.71 while following vowel disfavors it at 0.32, it is telling you the size and direction of a linguistic effect after the social effects have been accounted for.

Key idea: Internal linguistic factors and external social factors condition the same variable simultaneously, and multivariate models estimate each one's effect while holding the others constant.

Indicators, markers, and stereotypes

Not all variables sit at the same level of social awareness, and Labov proposed a useful three-way scale.

An indicator varies across social groups but not across styles within a speaker. Speakers are essentially unaware of it. If you find that one neighborhood pronounces a vowel slightly differently from another, but nobody shifts when reading aloud and nobody comments on it, you have an indicator.

A marker varies across social groups and also across styles. Speakers are not consciously aware of it, but they are responsive to it, shifting toward the prestige variant as formality rises. Both (ing) and New York (r) are markers, which is why the table above showed both between-speaker and within-speaker patterning.

A stereotype is a variable that has risen fully into public consciousness. It gets named, imitated, joked about, and corrected. The New York pronunciation of words like third and bird, often written in dialect spelling as toid and boid, became a stereotype. So did the Boston treatment of r in park the car. Stereotypes are socially loud but linguistically unreliable: because they are under conscious control, speakers suppress them, and popular imitations exaggerate them well beyond what any speaker actually produces. When a feature becomes a stereotype it often begins to recede from real speech even as its caricature persists.

Key idea: Indicators pattern by group only, markers pattern by group and style, and stereotypes have entered public awareness and are therefore consciously suppressed or caricatured.

An honest complication: sameness of meaning

The definition of a variable includes the phrase two ways of saying the same thing, and that phrase does real work. For phonological variables it is uncontroversial: walking and walkin refer to the same activity, so the two variants are semantically equivalent and the choice between them can only be social.

Above the level of sound, this gets contested. Beatriz Lavandera raised the objection sharply in 1978: if two syntactic constructions differ even slightly in meaning or emphasis, then a speaker who uses one more often may simply be saying different things, not saying the same thing differently. Is the passive really equivalent to the active? Does he was like carry the same nuance as he said? The field's working response has been to require functional comparability rather than perfect synonymy, to define envelopes narrowly enough that plausible meaning differences are controlled, and to report the criteria openly so others can challenge them. It is worth knowing that this is a live methodological question rather than a settled one, because a course that presents every method as airtight has taught you to read uncritically.

Key idea: Semantic equivalence between variants is straightforward for sound but genuinely debatable for syntax and discourse, and careful analysts define narrow envelopes and state their criteria rather than assuming the problem away.

Common misconceptions

  • Saying walkin means dropping the g. Both variants have a nasal consonant; only its place of articulation differs. The alveolar variant is also historically older than the velar one.
  • A speaker either has a feature or does not. Almost no sociolinguistic variable is categorical for a speaker. Rates, not presence and absence, are the data.
  • You count only the interesting variant. Without the denominator of all possible contexts, a rate cannot be computed and comparisons across speakers are meaningless.
  • Statistical models tell you what causes what. Factor weights describe association strength within a modeled dataset. Causal claims still require argument, replication, and often experiment.
  • Stereotyped features are the most common ones. Stereotypes are the most noticed, which is different. Conscious awareness usually drives a feature down in real speech while its caricature grows louder.

Recap

  • A linguistic variable is a site of structured choice, written in parentheses, with two or more variants.
  • Variables exist at every level: phonological, morphosyntactic, discourse, and lexical.
  • The envelope of variation specifies exactly which tokens represent a genuine choice and must be defined before counting.
  • The principle of accountability requires counting all contexts where the variable could occur, not just the salient ones.
  • Internal linguistic factors and external social factors act at once and are separated with multivariate models reporting factor weights.
  • Indicators pattern by group, markers by group and style, and stereotypes have entered conscious public awareness.
  • Semantic equivalence between variants is clean in phonology and genuinely contested above it, a live methodological debate.

Sources

  1. Fischer, J. L. (1958). Social influences on the choice of a linguistic variant. Word, 14(1), 47-56. doi.org/10.1080/00437956.1958.11659655
  2. Labov, W. (1963). The social motivation of a sound change. Word, 19(3), 273-309. doi.org/10.1080/00437956.1963.11659799
  3. Lavandera, B. R. (1978). Where does the sociolinguistic variable stop? Language in Society, 7(2), 171-182. doi.org/10.1017/S0047404500005510
  4. Wikipedia contributors. (n.d.). Variation (linguistics). en.wikipedia.org
  5. Linguistic Society of America. (n.d.). What is sociolinguistics? linguisticsociety.org
Key terms
Linguistic variable
A point in the grammar where speakers choose among two or more variants that do the same linguistic work; written in parentheses, as in (ing).
Variant
One of the alternative realizations of a linguistic variable, such as the alveolar or velar nasal in walking.
Envelope of variation
The precisely defined set of contexts in which a variable could occur, so that every token counted represents a real choice.
Principle of accountability
Labov's requirement to report every context where the variable could have appeared, including non-occurrences, so that rates are meaningful.
Internal factor
A conditioning influence from within the grammar, such as the following segment or the grammatical category of the word.
Factor weight
A number between 0 and 1 from a multivariate model, above 0.5 when a factor favors the variant and below 0.5 when it disfavors it.
Marker
A variable that patterns both across social groups and across speaking styles, showing that speakers respond to it without full conscious awareness.
Stereotype
A variable that has entered explicit public awareness, becoming subject to comment, imitation, and conscious suppression.

Doing Sociolinguistics: The Observer's Paradox, Interviews, Corpora, and Ethics

  • State the observer's paradox and describe the specific techniques the sociolinguistic interview uses to work around it.
  • Compare interview, rapid anonymous survey, ethnographic, and corpus methods, and say what each is good for and what it cannot deliver.
  • Explain the ethical obligations of recording human speech, including consent, anonymization, and the principle of returning value to the community studied.

The big picture

Everything in the last lesson assumed you have data. This lesson is about where the data comes from, and it opens with a problem that has no clean solution.

The speech sociolinguists most want is the speech people use when they are not thinking about speech: the unmonitored, everyday vernacular of a person talking to their own people about their own concerns. That is where the community's grammar shows most clearly, because attention to form pulls speakers toward the standard and blurs exactly the patterns under study. But you cannot analyze speech you have not recorded, and the act of recording makes people think about how they sound. Labov named this the observer's paradox: to observe how people talk when they are not being observed, you must observe them.

There is no trick that dissolves this. What the field has instead is a toolkit of partial workarounds, each with known strengths and known distortions, plus a professional habit of stating clearly which method produced which number. Learning the toolkit is what lets you read a study and know how much weight its findings can bear.

Key idea: The observer's paradox is the structural obstacle of the field: the vernacular is the target, observation disturbs it, and every method is a managed compromise rather than a solution.

The sociolinguistic interview

The workhorse method is a long recorded conversation, typically an hour or more, structured but not rigid. It is called an interview, but the goal is nearly the opposite of a job interview. The interviewer is trying to stop being the center of attention.

The interview is organized into modules, loose topic clusters the fieldworker can move between: childhood games, the neighborhood and how it has changed, school, work, dating and marriage, fights, and beliefs about luck or fate. These topics are chosen because they invite extended narrative, and narrative is the key. When a person tells a story about something that actually happened to them, they become absorbed in the telling. Emotional involvement competes with self-monitoring, and the vernacular surfaces.

The most famous single instrument is the danger of death question, usually phrased along the lines of: were you ever in a situation where you thought you were going to die? Speakers who answer it typically produce their most involved, least monitored speech of the whole session. Labov reported that the shift is often audible within a sentence or two. Modern fieldworkers use it more cautiously, and often replace it, since it can reopen genuine trauma and many ethics boards now expect gentler prompts such as a close call, a moment of embarrassment, or a story about a beloved pet.

Alongside conversation, the interview usually includes tasks that deliberately raise attention to speech: reading a connected passage, reading a word list, and reading minimal pairs such as source and sauce or pin and pen. These produce progressively more careful speech. The result is a style continuum running from casual conversation through careful conversation, reading passage, word list, and minimal pairs. That continuum is a research instrument in itself, since the slope of a speaker's shift across it tells you how socially charged a variable is.

Key idea: The sociolinguistic interview uses narrative modules to lower self-monitoring and reading tasks to raise it deliberately, producing a style continuum from casual speech to minimal pairs.

Who is holding the microphone

A fact often left out of methods summaries: the interviewer is part of the data. Speakers accommodate to whoever they are talking to, which is the subject of the audience design lesson later in Module 2, and that accommodation does not switch off during research.

The effect is well documented. Studies of African American speakers have repeatedly found that the same person produces measurably different rates of vernacular features depending on the interviewer's race, familiarity, and speech. Community insiders, or fieldworkers who have spent months in a neighborhood before recording anything, obtain different and generally more vernacular data than strangers with clipboards. This is not a flaw to be embarrassed about; it is an effect to be measured, reported, and where possible designed around, for instance by having community members conduct interviews or by comparing insider and outsider recordings of the same speakers.

Key idea: Interviewer identity systematically shifts the speech recorded, so who collects the data is a methodological variable that must be reported rather than assumed away.

Other ways in

The interview is not the only route, and its limits motivated several alternatives.

The rapid and anonymous survey elicits a single targeted token from many strangers in a public place, in seconds, with no recording equipment and no sense of being studied. Labov's department store study, which the next lesson works through in full, is the classic case. Its strength is that the observer's paradox is nearly eliminated. Its weakness is that you get one or two tokens per person and almost no information about who they are.

Ethnographic and participant observation methods go the opposite way. The researcher spends months or years inside a community, learning its internal social categories rather than importing external ones. Penelope Eckert's multi-year study in a Detroit-area high school is the model: instead of sorting students by parental income, she discovered the locally meaningful categories the students themselves used, the school-oriented Jocks and the school-alienated Burnouts, and found that these predicted vowel patterns better than conventional class measures did. The cost is time and scale.

Corpus methods use large collections of recorded and transcribed speech, letting researchers search millions of words for a variable that would take years to collect by hand. The Santa Barbara Corpus of Spoken American English, the spoken portion of the British National Corpus, and the Corpus of Contemporary American English are widely used. Corpora give statistical power and permanence, and they let later researchers test earlier claims. What they usually cannot give is fine social detail about each speaker or the acoustic quality needed for precise phonetic measurement.

Two more approaches recur throughout this course. Experimental methods, especially the matched guise technique covered in Module 3, measure listeners' attitudes rather than speakers' production. And large crowd-sourced dialect surveys, of which the Harvard Dialect Survey is the best known, gather self-reported usage from tens of thousands of people, trading precision for extraordinary geographic coverage.

MethodBest forMain limitation
Sociolinguistic interviewRich speech plus detailed social data on each speakerObserver's paradox; slow; one interviewer's effect on all data
Rapid anonymous surveyMany speakers fast, minimal observation effectOne or two tokens; little social information
EthnographyLocally meaningful social categoriesYears of fieldwork; small numbers; hard to replicate
Corpus analysisStatistical power; reanalysis by othersThin speaker metadata; audio may be unsuitable for acoustics
Matched guise experimentListener attitudes, isolated from other cuesMeasures reactions in a lab, not behavior in the world
Crowd-sourced surveyGeographic breadth at very large scaleSelf-report, which diverges from actual usage

Key idea: Each method trades depth against breadth and naturalness against control, so strong claims in sociolinguistics usually rest on convergence across several methods rather than on any single study.

Sampling and social categories

Who you record determines what you can conclude. Early survey work aimed at random samples of a population, which is statistically ideal and practically brutal, since a truly random sample of a city requires contacting people who have no interest in talking to you.

Most sociolinguistic work therefore uses a judgment sample, also called a quota sample: the researcher decides in advance which social cells matter, for instance three age groups by two genders by three class levels, and fills each cell with several speakers. The claim is not that the sample mirrors the city's proportions but that it covers the relevant categories well enough to compare them.

Measuring social class is its own problem. Studies have used occupation alone, or composite indices combining occupation, education, income, and housing. Every choice embeds assumptions, and the assumptions are visible in the results: the Eckert example above shows that externally imposed class categories can be less predictive than locally salient identities. Good practice is to state exactly how the categories were built so readers can judge whether the social variable measures what it claims to.

Key idea: Sociolinguists generally use judgment samples that fill defined social cells rather than random samples, and the way social categories are constructed shapes what the study can find.

Ethics: recording people is not a neutral act

Speech data is personal. A recording carries a person's voice, their opinions, and often their most vulnerable stories, since the interview technique is designed to elicit exactly those. That places real obligations on the researcher, and modern practice codifies them.

Informed consent comes first. Participants must know they are being recorded, what the recording is for, who will hear it, how long it will be kept, and that they may stop or withdraw at any point without penalty. Covert recording, which would neatly solve the observer's paradox, is rejected in professional practice for precisely this reason. The three principles laid out in the Belmont Report, respect for persons, beneficence, and justice, underlie the institutional review board review that university research must pass before a single recording is made.

Anonymization follows. Names, employers, schools, and identifying details are replaced with pseudonyms in transcripts, and audio is stored securely. This is harder than it sounds, because a voice is itself an identifier and because a small community can be re-identified from a few local details.

Beyond compliance sits a further obligation many sociolinguists take seriously: research on a community should return something to it. Walt Wolfram has argued for a principle of linguistic gratuity, holding that researchers who gain from a community's speech owe it active advocacy in return, whether through dialect awareness materials for local schools, museum exhibits, testimony, or teacher training. Community-based participatory approaches go further and involve community members in setting research questions, conducting interviews, and controlling archives. In work with Indigenous communities, this now often extends to data sovereignty: the community, not the university, holds authority over recordings of its language and decides how they may be used. These commitments are not decoration. Given how often linguistic research has been used to characterize communities that had no say in the characterization, they are part of doing the work correctly.

Key idea: Ethical sociolinguistics requires informed consent, careful anonymization, and increasingly a positive obligation to return value to the community studied and to respect its authority over its own linguistic data.

Common misconceptions

  • The observer's paradox can be solved with a hidden microphone. Covert recording violates informed consent and is not accepted practice, so the paradox is managed rather than eliminated.
  • An interview is a questionnaire. The sociolinguistic interview aims at extended narrative and often succeeds when the interviewer talks least; a question-and-answer format produces careful, atypical speech.
  • Reading tasks contaminate the data. They are included on purpose. The contrast between reading styles and conversation is the measurement, not an error.
  • A bigger sample is always better. A thousand thin tokens with no social information answer different questions than twenty deeply documented speakers, and neither replaces the other.
  • Ethics review is bureaucratic overhead. Consent, anonymization, and reciprocity respond to a documented history of communities being studied, described, and stigmatized without recourse.

Recap

  • The observer's paradox: the vernacular is the target, but observing speech changes it.
  • The sociolinguistic interview uses narrative modules to lower monitoring and reading tasks to raise it, yielding a style continuum.
  • Interviewer identity measurably changes the speech recorded and must be reported as part of the method.
  • Rapid anonymous surveys, ethnography, corpora, experiments, and crowd-sourced surveys each trade naturalness, depth, and scale differently.
  • Judgment samples filling defined social cells are the norm, and how social categories are constructed shapes the findings.
  • Consent, anonymization, and secure storage are required, and covert recording is not an acceptable workaround.
  • Many sociolinguists accept a further obligation to return value to communities and to respect community authority over linguistic data.

Sources

  1. Wikipedia contributors. (n.d.). William Labov. en.wikipedia.org
  2. National Commission for the Protection of Human Subjects of Biomedical and Behavioral Research. (1979). The Belmont Report. U.S. Department of Health and Human Services. hhs.gov
  3. Du Bois, J. W., Chafe, W. L., Meyer, C., Thompson, S. A., Englebretson, R., & Martey, N. (2000-2005). Santa Barbara Corpus of Spoken American English. University of California, Santa Barbara. linguistics.ucsb.edu
  4. Davies, M. (n.d.). Corpus of Contemporary American English (COCA). english-corpora.org
  5. Linguistic Society of America. (n.d.). What is sociolinguistics? linguisticsociety.org
Key terms
Observer's paradox
Labov's formulation of the field's central methodological obstacle: the aim is to record how people speak when not observed, which requires observing them.
Vernacular
The relatively unmonitored everyday speech a person uses with intimates, treated as the most systematic form of a speaker's grammar.
Sociolinguistic interview
A long, loosely structured recorded conversation organized into topic modules designed to elicit extended narrative and reduce self-monitoring.
Style continuum
The graded series from casual conversation through careful speech, reading passage, word list, and minimal pairs, reflecting increasing attention to speech.
Rapid and anonymous survey
A method that elicits a single targeted token from many strangers in public, minimizing observation effects at the cost of social detail.
Judgment sample
A sampling design that fills predefined social cells with speakers rather than drawing randomly from the population.
Informed consent
A participant's knowing, voluntary agreement to be recorded after being told the purpose, audience, storage, and right to withdraw.
Linguistic gratuity
Wolfram's principle that researchers who benefit from a community's language owe that community active advocacy and returned value.

Module 2: Variation and Its Social Correlates

The empirical core of the field: two founding studies worked through as method, then the social dimensions that variation tracks, including class, network, ethnicity, gender, age, and the audience a speaker is designing for.

Labov's Landmark Studies: Martha's Vineyard, the Department Stores, and Apparent Time

  • Reconstruct the design, findings, and interpretation of the Martha's Vineyard and New York department store studies.
  • Explain what each study contributed methodologically: local social meaning in the first, rapid stratified sampling and style shift in the second.
  • Apply the apparent-time construct to infer change in progress, and identify age grading as its principal confound.

The big picture

Two studies, completed within about two years of each other by the same graduate student, established how sociolinguistics would be done. One was on a small island off Massachusetts, the other in three department stores in Manhattan. Neither took long. Both were designed with a precision worth studying in itself, because in each case the design does the intellectual work.

Read them not as history but as templates. The first shows that a linguistic variant can carry local social meaning that has nothing to do with prestige as usually understood. The second shows how to get a stratified sample and a style shift out of strangers in a few hours. Together they answer the question the previous lessons raised: how do you get from a definition of a variable to a finding you can defend?

Key idea: The Martha's Vineyard and department store studies are the field's founding demonstrations that variation can be sampled, counted, and interpreted with a design as tight as any experiment.

Martha's Vineyard: identity, not prestige

In 1961 Labov went to Martha's Vineyard, an island whose year-round population of a few thousand swelled enormously each summer with mainland vacationers. The island's traditional fishing economy was in decline, and land was being bought up by outsiders. That social situation is the whole point of the study.

The variables were the first elements of two diphthongs: (ay), the vowel of right and life, and (aw), the vowel of house and out. Most American speakers begin these with a low vowel. Some Vineyarders began them with a centralized vowel, closer to the sound in but, giving pronunciations that a mainlander might hear as roight and hoose. This centralization was an old New England feature that had been receding for generations.

Labov recorded island speakers and computed a centralization index for each. Then he cross-tabulated it against age, occupation, geography, and, crucially, a factor he had to construct himself: each speaker's orientation toward the island, judged from what they said about staying or leaving. The results are worth stating precisely.

ComparisonPattern in centralization
By ageHighest in the 31 to 45 group, lower among both the oldest and the youngest speakers
By region of the islandHighest up-island, especially in the fishing village of Chilmark
By occupationHighest among fishermen, the group most identified with traditional island life
By orientation to the islandHighest among speakers who intended to stay, lowest among those planning to leave

Now interpret it. If centralization were simply an old feature dying out, the oldest speakers would use it most and each younger group less. That is not what happened. It peaked in the middle-aged group, the people who had to decide whether to build a life on a changing island. And within that group it tracked not income or education but attitude toward the island itself.

Labov's conclusion was that the centralized vowels had been reassigned a social job. They had become an unconscious signal of authentic islander identity, deployed most by those staking a claim to it against the summer influx. Nobody on the island could have told you this. Nobody was imitating anyone deliberately. The feature simply became the sound of belonging.

That result reframed the field. Before it, variation was assumed to run along a single axis from vulgar to prestigious. After it, the analyst's first question became: what does this variant mean here, to these people? Sometimes the answer is status. On Martha's Vineyard it was loyalty, and speakers were moving away from the mainland standard, not toward it.

The honest postscript: in 2007, Jennifer Pope, Miriam Meyerhoff, and D. Robert Ladd restudied the island and found the picture had shifted again in the intervening forty years, with centralization of (ay) persisting while (aw) behaved differently than a simple continuation would predict. Founding studies are starting points, not scripture, and the field's willingness to go back and check is a strength.

Key idea: On Martha's Vineyard a receding local feature was revived as an unconscious marker of island identity, showing that variants can carry local social meaning unrelated to standard prestige.

The department store study: stratification in a single word

The New York study confronted a different problem. Labov wanted to know whether (r), the pronunciation or absence of r after a vowel, stratified by social class across an entire city. Interviewing a representative sample would take a year. He wanted an answer in a week.

His solution was to let the city do his sampling. He chose three department stores that served sharply different clienteles and therefore employed sharply different staff: Saks Fifth Avenue at the top, Macy's in the middle, and S. Klein, a discount store, at the bottom. The stores' advertising, prices, and locations established the ranking independently of anything linguistic.

Then the elicitation. Labov looked up in advance a department located on the fourth floor. He approached employees and asked where to find it. The natural answer is fourth floor, which contains two tokens of (r), one before a consonant and one word-finally. Having heard it once, he leaned in and said, Excuse me? The employee repeated the phrase, this time carefully and emphatically. That gave him a second, more careful style from the same speaker, seconds later. He then stepped aside and wrote everything down, including the store, the employee's approximate age, sex, and department.

Look at what this design accomplishes. It is a rapid and anonymous survey, so the observer's paradox is nearly absent: the employees experienced a customer, not a researcher. It yields social stratification without asking anyone a single personal question. And it yields a within-speaker style contrast from a two-word intervention. In a few days Labov collected hundreds of speakers.

The findings were clean. The percentage of employees using any r was highest at Saks, intermediate at Macy's, and lowest at Klein, exactly ordering the stores by their social rank. The emphatic repetition raised r use in all three stores, confirming that (r) is a marker: it patterns socially and stylistically at once.

The most interesting result was in the middle. The increase from casual to emphatic style was sharpest at Macy's. Labov connected this to a broader finding from his full New York survey: lower middle class speakers show the steepest style shifting and, in the most formal styles, actually overshoot the group above them, producing more of the prestige variant than the highest status speakers do. He called this hypercorrection, and he linked it to linguistic insecurity, the gap between how speakers say they ought to talk and how they actually talk. The people most anxious about their speech are not those at the bottom but those close enough to the top to feel the climb.

Key idea: The department store study extracted class stratification and a style shift from strangers in seconds, and revealed the hypercorrection of the lower middle class, whose formal speech overshoots the group above it.

Apparent time: reading change from a single moment

Both studies compared age groups, and that raises a deep question. Language change takes decades, but a fieldworker has months. How can anyone study change in progress?

The apparent-time construct is the answer. It rests on a specific empirical assumption: a speaker's vernacular is largely set during adolescence and remains relatively stable thereafter. If that holds, then a seventy-year-old's speech is roughly a recording of the community norm of fifty-odd years ago. Lining up age groups at one moment therefore approximates a sequence of moments. A feature that rises steadily as speaker age falls is a change in progress.

The construct is powerful and it is not free. Its main confound is age grading: a pattern in which speakers of every generation use a feature at a certain life stage and abandon it later. If teenagers in every era use more nonstandard forms and then reduce them upon entering the workforce, an apparent-time snapshot will show a false picture of change. Slang is the obvious case, but the pattern also affects variables like (ing).

Change in progressAge grading
What the age slope meansThe community norm is genuinely shiftingIndividuals shift as they age; the community is stable
Restudy of the community 20 years laterOverall rate has movedOverall rate is unchanged; the same age slope reappears
Restudy of the same individualsIndividuals mostly stay where they wereIndividuals have moved with age

Distinguishing them requires real time: either a trend study, resurveying the same community later with a fresh sample, or a panel study, returning to the same individuals. Both exist. Labov restudied New York's Lower East Side decades after his original work. In Montreal, Gillian Sankoff and Helene Blondeau tracked the same French speakers across thirty-two years and found that most adults kept their youthful variant of r while a minority did change in adulthood, which qualifies the stability assumption without overturning it. That combination, a construct that mostly works plus documented exceptions, is roughly where the evidence stands.

Key idea: Apparent time treats age differences at one moment as a proxy for change over time, an inference that holds when vernaculars are stable after adolescence and fails when a pattern is age graded.

Change from below and change from above

One more distinction that the two studies illustrate, and that recurs throughout this course. Labov used the terms change from below and change from above, where below and above refer to the level of conscious social awareness, not to social class.

Change from below happens beneath awareness. Speakers do not notice it, cannot report it, and do not style-shift much on it. Martha's Vineyard centralization was change from below: islanders were not choosing to sound local, they simply did. Most sound change works this way and is discovered by linguists rather than by speakers.

Change from above happens with awareness, typically spreading from a prestige model and appearing first in careful styles. The rise of r pronunciation in New York after the Second World War was change from above: it was borrowed consciously enough that speakers produced far more of it when paying attention, which is exactly why the emphatic repetition trick worked so well.

Key idea: Change from below proceeds beneath conscious awareness with little style shifting, while change from above spreads from a recognized prestige model and shows up first in careful speech.

Common misconceptions

  • The Vineyarders were imitating older fishermen on purpose. The pattern was unconscious. Speakers could not report it, which is what makes it change from below.
  • Prestige always means the standard. On Martha's Vineyard the socially valued variant was the local, nonstandard one. Value is assigned locally.
  • The department store study measured customers. It measured employees, whose social position was inferred from the store that hired them.
  • Hypercorrection means using a fancy word incorrectly. In the technical sense it means a group exceeding the prestige group's rate of a prestige variant in formal styles, which is what lower middle class New Yorkers did with r.
  • Apparent time proves change. It supports an inference that must be checked against age grading, ideally with a real-time trend or panel study.

Recap

  • On Martha's Vineyard, centralized (ay) and (aw) peaked among middle-aged, up-island, island-loyal speakers, marking local identity rather than status.
  • The Vineyard study redirected the field from a single prestige axis toward locally constructed social meaning.
  • The department store study used three ranked stores and the fourth floor prompt to obtain class stratification and a style shift from strangers.
  • The r results ordered Saks above Macy's above Klein, and emphatic repetition raised r everywhere, marking (r) as a marker.
  • Lower middle class speakers style-shift most steeply and hypercorrect past the group above them, reflecting linguistic insecurity.
  • Apparent time infers change from age differences, assuming vernacular stability after adolescence; age grading is the main confound.
  • Change from below runs beneath awareness; change from above spreads from a prestige model and surfaces first in careful styles.

Sources

  1. Labov, W. (1963). The social motivation of a sound change. Word, 19(3), 273-309. doi.org/10.1080/00437956.1963.11659799
  2. Pope, J., Meyerhoff, M., & Ladd, D. R. (2007). Forty years of language change on Martha's Vineyard. Language, 83(3), 615-627. doi.org/10.1353/lan.2007.0117
  3. Sankoff, G., & Blondeau, H. (2007). Language change across the lifespan: /r/ in Montreal French. Language, 83(3), 560-588. doi.org/10.1353/lan.2007.0106
  4. Wikipedia contributors. (n.d.). The Social Stratification of English in New York City. en.wikipedia.org
  5. MacNeil, R., & Cran, W. (2005). Do you speak American? PBS. pbs.org
Key terms
Centralization index
Labov's quantitative measure of how far the onset of a diphthong is raised toward a central vowel, used on Martha's Vineyard.
Rapid and anonymous survey
A design that obtains targeted tokens from many strangers in public settings, as in the department store elicitation of fourth floor.
Hypercorrection
The pattern in which a group, typically the lower middle class, exceeds the higher status group's use of a prestige variant in formal styles.
Linguistic insecurity
The gap between how speakers believe they should speak and how they actually speak, associated with steep style shifting.
Apparent time
The inference of change in progress from differences among age groups sampled at a single point in time.
Age grading
A stable pattern in which speakers of every generation use a feature at a particular life stage and abandon it later, mimicking change.
Change from below
Change proceeding beneath conscious social awareness, with little style shifting and no speaker commentary.
Change from above
Change spreading from a consciously recognized prestige model, appearing first and most strongly in careful speech.

Social Class, Networks, and Ethnicity as Correlates of Variation

  • Read a class-by-style table of sociolinguistic data and distinguish sharp from gradient stratification.
  • Explain how social network density and multiplexity account for vernacular maintenance and innovation better than class alone in some communities.
  • Describe how ethnicity correlates with variation without appealing to descent, using the concepts of ethnolect and ethnolinguistic repertoire.

The big picture

The previous lesson showed variation stratifying by class in three Manhattan stores. This lesson asks the follow-up questions. What exactly is class, and how do you measure something that societies define differently? Does class do the explanatory work, or is it standing in for something more immediate? And how does ethnicity enter the picture without the analysis sliding into claims about descent that the evidence does not support?

The answers turn out to be layered. Class works as a broad predictor and produces beautifully regular tables. But when researchers looked closely at why it works, they found a more local mechanism underneath: who you actually talk to, how often, and in how many capacities. That discovery moved the field from social categories toward social relationships, and it is the through-line of this lesson.

Key idea: Class predicts linguistic variation reliably, but the mechanism behind the correlation is the structure of a speaker's everyday social relationships, which is what network and community of practice approaches make explicit.

Measuring class, and the hazard in it

Sociolinguists have measured class in several ways: occupation alone, education alone, income, or a composite index adding these together, sometimes with housing type or neighborhood. Labov's New York work used a composite. Trudgill's Norwich study used a six-factor index. None of these is the true measure, because class is not a natural kind; it is a construct that varies by society. An index built for industrial Britain in 1970 does not transfer intact to contemporary Los Angeles, Lagos, or Mumbai, where the salient divisions may run along caste, religion, migration history, or urban and rural origin instead.

The practical consequence is that class results are only as good as the index behind them, and a careful reader checks how the index was built before accepting what it shows. It is also why the alternatives later in this lesson matter: they were developed partly because externally imposed class categories sometimes fail to predict what locally meaningful groupings predict easily.

Key idea: Class is an analyst's construct, measured through composite indices that differ across studies and societies, so its explanatory value depends on how well the index fits the community.

Reading a stratification table

Peter Trudgill's 1974 study of Norwich, England, produced what may be the most reproduced table in sociolinguistics. He measured the (ng) variable, the same walking and walkin alternation from Lesson 2, across five class groups and four styles. The figures are percentages of the nonstandard alveolar variant.

Class groupWord listReading passageFormal speechCasual speech
Middle middle class00328
Lower middle class0101542
Upper working class5157487
Middle working class23448895
Lower working class296698100

Spend a moment reading it properly, because the shape of this table is the shape of the field's central finding. Read down any column: the nonstandard variant increases steadily as you move down the class scale. Read across any row: every group uses less of the nonstandard variant as the style becomes more formal. The two effects are independent and additive.

Three further observations. First, no cell is a categorical zero except in the most formal styles for the highest group, which means class groups differ in rate rather than in kind: everyone has both variants. Second, the lines never cross. A middle middle class speaker in casual conversation uses the nonstandard variant less than a lower working class speaker does in a word list. Third, the largest style shifts occur in the middle of the scale, the same lower middle class steepness the department store study showed.

This is gradient stratification: each class step produces a proportional change. Not all variables behave this way. Walt Wolfram's Detroit research found that multiple negation showed sharp stratification instead, with a large gap between working class and middle class groups and comparatively little difference within each block. As a rule of thumb, phonological variables tend to stratify gradiently while grammatical ones tend to stratify sharply and carry heavier stigma. That distinction matters for education, because the features teachers correct most aggressively are usually the sharply stratified grammatical ones.

Key idea: Gradient stratification changes proportionally at each class step, typical of phonological variables, while sharp stratification shows a large break between broad class blocks, typical of stigmatized grammatical variables.

Social networks: from category to relationship

Class tables raise an obvious question: what is class actually doing? Nobody's vowels are influenced by an income bracket. Something has to transmit the pattern, and Lesley Milroy's fieldwork in Belfast in the 1970s identified it.

Milroy studied three working class neighborhoods and, rather than assigning speakers to class cells, measured each person's social network on two dimensions. Density asks whether the people you know also know each other. In a dense network, your neighbor, your cousin, and your workmate are the same three people, and they all know one another. Multiplexity asks how many capacities you share with someone: if a person is simultaneously your relative, your coworker, and your drinking companion, that tie is multiplex rather than uniplex.

She built a network strength scale from indicators like working with people from the neighborhood, having kin in the area, and socializing with workmates. The result was striking: network strength predicted vernacular use directly. Speakers embedded in dense, multiplex networks used markedly more local vernacular features than speakers with looser ties, and this held within a single class group.

The mechanism is intuitive once stated. A dense, multiplex network is a norm-enforcement machine. Everyone knows everyone, deviation is noticed, and the local way of speaking is continually reinforced as a badge of membership. This is why the vernacular persists for generations despite constant pressure from schools and media.

The reverse follows too. Weak ties, the acquaintances who connect otherwise separate groups, are the routes along which innovations travel. Mark Granovetter's argument about the strength of weak ties, developed for job information, transfers directly: a change cannot spread within a closed network, because everyone there already shares the norm. It spreads through the loosely connected people who move between worlds.

One Belfast finding is worth flagging because it disturbs a common assumption. In the Ballymacarrett neighborhood, men had stable local employment and therefore denser networks than women, who worked outside the area. The men accordingly used more vernacular forms. Where the network structure reversed, so did the linguistic pattern. Network structure, not gender as such, was doing the work, which is a useful warning for the next lesson.

Key idea: Dense multiplex networks enforce local vernacular norms while weak ties carry innovation between groups, which explains the mechanism that class categories only summarize.

Communities of practice

A third approach pushes further toward the local. A community of practice, a concept Penelope Eckert borrowed from Jean Lave and Etienne Wenger, is a group defined by joint engagement in some shared activity, developing its own practices, including linguistic ones, in the course of that engagement.

Eckert's Detroit-area high school study, introduced earlier, is the exemplar. The students sorted themselves into Jocks, oriented toward the school and its institutional ladder, and Burnouts, oriented toward the local urban working class world beyond it. These were not class categories; both groups drew from similar family backgrounds. They were categories of practice, built from what students did together: where they ate lunch, whether they joined activities, whether they went downtown.

The vowel data followed the practice categories more closely than parental socioeconomic status did. Burnouts led in the urban vernacular changes, Jocks lagged, and the students most extreme on either dimension were the most linguistically extreme. Speech was not a passive readout of background. It was part of how these teenagers actively built and displayed the persons they were becoming.

That reframing has consequences for the whole field. It positions speakers as agents who use variation to construct identity, not as carriers of demographic categories. Module 2's final lesson develops that idea into a full account of style.

Key idea: A community of practice is defined by shared activity rather than demographic category, and Eckert showed that such locally built groupings can predict variation better than imported class measures.

Ethnicity: correlation without descent

Ethnicity correlates strongly with linguistic variation in many societies, and it is essential to be precise about why. The correlation is not biological. There is no genetic basis for any speech pattern, and the decisive evidence is everyday: children acquire the variety of the community that raises them, whatever their ancestry. A child of any background raised among speakers of a given variety acquires that variety natively. Cases of speakers raised across community lines confirm this repeatedly.

What ethnicity indexes instead is social: patterns of residence, schooling, friendship, marriage, and institutional segregation determine who talks to whom, and network structure does the rest. Where communities are socially separated, their speech diverges; where they mix, features cross. Ethnicity correlates with speech to the extent that it correlates with contact.

A variety associated with an ethnic group is often called an ethnolect. Labov's New York survey found distinguishable patterns among speakers of Jewish and Italian background, for instance in the raising of particular vowels. Latino English varieties have been documented in Los Angeles, Chicago, and New York, each shaped by local contact histories rather than by Spanish alone, and many of their speakers are English monolinguals. Module 4 treats African American English at length.

The term ethnolect has drawn criticism for implying that a group speaks one uniform way. Sarah Bunin Benor proposed the more flexible notion of an ethnolinguistic repertoire: a set of features associated with a group that individual speakers draw on selectively and variably, deploying more or fewer depending on situation, audience, and how much they wish to foreground that identity at that moment. That formulation matches the data better, since within-group variation in ethnic communities is typically as large as between-group variation.

Key idea: Ethnicity correlates with variation through social contact rather than descent, and speakers draw selectively on an ethnolinguistic repertoire rather than speaking a uniform ethnolect.

Common misconceptions

  • Working class speakers do not use standard forms. Every group in the Norwich table uses both variants; groups differ in rate, not in inventory.
  • Class causes pronunciation. Class summarizes patterns of contact and opportunity. Network density and shared practice are the proximate mechanisms.
  • Tight-knit communities are linguistically backward. Dense networks maintain norms efficiently, which is a functional property of the network, not a deficiency of its members.
  • Innovations spread from the most popular people. Innovations travel along weak ties between groups; the densely connected core tends to already share its norms.
  • An ethnolect is a single uniform way an ethnic group speaks. Within-group variation is large, and speakers deploy features from a repertoire selectively rather than uniformly.

Recap

  • Class is a constructed index that varies by study and society, so results depend on how well the index fits the community.
  • Trudgill's Norwich table shows gradient class stratification and parallel style shifting, with lines that never cross.
  • Phonological variables tend to stratify gradiently; stigmatized grammatical variables tend to stratify sharply.
  • Milroy's Belfast work showed that dense, multiplex networks maintain vernacular norms within a single class group.
  • Weak ties, not strong ones, are the channels along which linguistic innovations spread between groups.
  • Communities of practice, such as Eckert's Jocks and Burnouts, can predict variation better than imported demographic categories.
  • Ethnicity correlates with speech through contact and social separation, and speakers use an ethnolinguistic repertoire selectively.

Sources

  1. Wikipedia contributors. (n.d.). Peter Trudgill. en.wikipedia.org
  2. Granovetter, M. S. (1973). The strength of weak ties. American Journal of Sociology, 78(6), 1360-1380. doi.org/10.1086/225469
  3. Benor, S. B. (2010). Ethnolinguistic repertoire: Shifting the analytic focus in language and ethnicity. Journal of Sociolinguistics, 14(2), 159-183. doi.org/10.1111/j.1467-9841.2010.00440.x
  4. Wikipedia contributors. (n.d.). Penelope Eckert. en.wikipedia.org
  5. Britannica. (n.d.). Sociolinguistics. Encyclopaedia Britannica. britannica.com
Key terms
Gradient stratification
A pattern in which each step along a class scale produces a proportional change in variant frequency, typical of phonological variables.
Sharp stratification
A pattern with a large break between broad class blocks and little difference within them, typical of stigmatized grammatical variables.
Network density
The extent to which the people a speaker knows also know one another, high in tightly interconnected communities.
Multiplexity
The number of distinct capacities in which two people are connected, such as being simultaneously kin, coworkers, and neighbors.
Network strength scale
Milroy's index combining indicators of dense and multiplex ties, used to predict a speaker's use of local vernacular features.
Weak ties
Loose acquaintance connections between otherwise separate groups, which serve as the channels through which linguistic innovations spread.
Community of practice
A group defined by joint engagement in a shared activity, which develops its own practices including linguistic ones.
Ethnolinguistic repertoire
Benor's term for a set of features associated with an ethnic group that individuals draw on selectively rather than uniformly.

Module 3: Dialects and Standards

How linguistic variation distributes across space, how dialect boundaries are actually drawn, and how one variety among many came to be called the language itself, with the prestige and the attitudes that follow from it.

Regional Variation: Dialect Geography, Isoglosses, and Atlases

  • Explain why mutual intelligibility fails as a criterion for separating languages from dialects, and what a dialect continuum is.
  • Define an isogloss and explain what a bundle of isoglosses does and does not establish about dialect boundaries.
  • Summarize the major dialect atlas projects and what the Atlas of North American English found about whether American dialects are converging.

The big picture

Ask someone where the American South begins and you will get an answer with real conviction and no method behind it. Dialect geography is the branch of the field that supplies the method. It asks where particular linguistic features stop, how those stopping points relate to one another, and what the resulting map tells us about history, migration, and contact.

Two results from this work are worth flagging at the outset, because both contradict widespread assumptions. First, dialect boundaries are far messier than the tidy colored maps in textbooks suggest, and the messiness is the finding rather than a failure of measurement. Second, American regional dialects have not been flattened by television and national media. Measured acoustically, several of them have been pulling further apart. Both results emerge from taking the geography seriously.

Key idea: Dialect geography maps where linguistic features stop and start, and its central findings are that boundaries are gradient rather than sharp and that regional varieties are not converging.

Language or dialect?

Before mapping dialects, we need to know what one is, and the honest answer is that the boundary between a language and a dialect is not a linguistic question at all.

The obvious criterion is mutual intelligibility: if speakers understand one another, they share a language. It fails in both directions. Danish, Norwegian, and Swedish are largely mutually intelligible in their standard forms, yet nobody calls them dialects of one Scandinavian language, because Denmark, Norway, and Sweden are separate states with separate literary traditions. Meanwhile Mandarin and Cantonese are not mutually intelligible in speech at all, and are conventionally called dialects of Chinese, supported by a shared writing system and a unified state.

Intelligibility is also often asymmetric. Speakers of a stigmatized variety generally understand the prestige variety better than the reverse, simply because they encounter it constantly in schooling and media. And intelligibility improves with exposure, so the same two speakers may fail to understand each other on day one and manage fine by week three. A criterion that shifts with familiarity cannot draw a stable boundary.

Max Weinreich popularized the observation that, in his formulation, a language is a dialect with an army and navy. The point is that the distinction is political. Linguists therefore prefer the neutral term variety, and where a distinction is needed, Heinz Kloss's terms are useful: Abstand refers to distance, varieties so structurally different that they must count as separate languages, while Ausbau refers to development, varieties treated as separate languages because each has been elaborated with its own standard, orthography, and institutions.

Key idea: Mutual intelligibility fails as a criterion in both directions, so the language and dialect distinction is political rather than linguistic, and linguists use the neutral term variety.

Dialect continua

The reason boundaries resist drawing is that variation across geography is typically continuous. In a dialect continuum, each village understands its neighbors easily, the next village over slightly less well, and so on, until speakers at the two ends of the chain cannot understand each other at all, despite there being no point along the way where a break occurs.

The classic European examples are the West Romance continuum, running through Portugal, Spain, France, and Italy, and the continental West Germanic continuum, running from the Netherlands across Germany into Austria and Switzerland. Historically a traveler could walk from Dutch-speaking territory to German-speaking territory without ever crossing a line where comprehension collapsed. The Arabic-speaking world shows the same structure across North Africa and the Middle East, where Moroccan and Iraqi varieties are very hard to mutually understand despite the continuum between them.

National standard languages break continua up after the fact. Compulsory schooling in a standard, national broadcasting, and state borders sharpen what was previously gradual, which is why the modern Dutch and German border feels much more like a wall than it did in 1800.

Key idea: Geographic variation is typically continuous, so mutual intelligibility can fail across a continuum whose adjacent points are all mutually intelligible, and national standards impose sharp boundaries afterward.

How dialect atlases were built

The discipline began in the late nineteenth century with two projects whose contrasting methods still define the trade-offs.

Georg Wenker began surveying German dialects in 1876 by mailing a set of forty sentences to schoolteachers across the country and asking them to write the sentences in the local dialect. He collected responses from tens of thousands of locations, an extraordinary density of coverage. The weakness is equally clear: the teachers were untrained transcribers using ordinary spelling, so the phonetic detail is unreliable and inconsistent from one respondent to the next.

Jules Gilliéron took the opposite approach for the Atlas linguistique de la France. He hired a single fieldworker, Edmond Edmont, who traveled the country by bicycle between 1897 and 1901, administering the same questionnaire in person at 639 locations and transcribing every response himself in a consistent phonetic notation. The trade-off inverts Wenker's: far fewer points, but every data point recorded the same way by the same trained ear.

Every subsequent project sits somewhere on that line between coverage and consistency, and knowing which end a study came from tells you what its map can support.

In North America, Hans Kurath directed the Linguistic Atlas project from the 1930s, beginning with New England and expanding outward, using trained fieldworkers and lengthy interviews with carefully selected informants. Kurath's analysis produced the classic three-way division of the eastern United States into Northern, Midland, and Southern areas, replacing an earlier and largely imaginary two-way North and South picture. Separately, Frederic Cassidy launched the Dictionary of American Regional English in the 1960s, sending fieldworkers in distinctive vans to over a thousand communities with a questionnaire of more than sixteen hundred items. Its final volume appeared in 2013, and it remains the richest record of American regional vocabulary.

Key idea: Dialect atlas methods trade coverage against consistency, from Wenker's tens of thousands of untrained postal respondents to Gilliéron's single trained fieldworker at 639 sites.

Isoglosses and why they do not bundle neatly

An isogloss is a line on a map separating the area where one variant of a feature is used from the area where another is used. There is an isogloss for the pronunciation of a vowel, another for whether people say bucket or pail, another for whether a sandwich is a sub, a hoagie, or a grinder.

A single isogloss does not make a dialect. What analysts look for is a bundle: several isoglosses running roughly together, suggesting a real historical or social boundary, such as a mountain range, an old political border, or a settlement pattern. Kurath's Northern and Midland boundary is a bundle of this kind.

But isoglosses often refuse to bundle, and the most instructive case is the Rhenish fan in western Germany. The High German consonant shift affected different words in different areas, and its isoglosses, instead of running together, splay apart like the ribs of a fan as they approach the Rhine. The line for the word for I, in its northern and southern forms, sits in one place; the line for the word for make sits elsewhere; the Benrath line and the Uerdingen line run separately. There is no single place where the dialect changes. There are many places where individual features change.

Take that seriously and it reshapes how you read every dialect map you will ever see. The colored regions are summaries imposed on gradient data, useful for teaching and misleading if taken literally. Modern work increasingly uses statistical dialectometry, computing aggregate distances between locations across hundreds of features rather than drawing lines by eye.

Key idea: An isogloss marks the boundary for one feature, dialect areas are inferred from bundles of them, and cases like the Rhenish fan show that isoglosses frequently splay apart rather than bundling.

The Atlas of North American English and the divergence finding

The most consequential recent project is the Atlas of North American English, published in 2006 by William Labov, Sharon Ash, and Charles Boberg. Its method differed sharply from earlier atlases: rather than long interviews with rural elders, the Telsur project conducted telephone interviews with speakers in urbanized areas across the continent, and rather than relying on the fieldworker's ear, it measured vowel formants acoustically.

The results identified several large-scale changes in progress. The Northern Cities Shift, a rotation of short vowels running through Rochester, Buffalo, Cleveland, Detroit, and Chicago, makes a word like block sound to outsiders closer to black. The Southern Shift moves a different set of vowels in the opposite direction across the South. And the low back merger, which collapses the vowels of cot and caught, has been spreading through the West, Canada, and western Pennsylvania.

The headline finding was the one nobody expected. Since the arrival of national broadcasting, commentators had predicted that American speech would homogenize. Measured acoustically, the opposite happened. The vowel systems of the Inland North, the South, and the West have been moving apart from one another and from the national norm. Regional differentiation increased over the twentieth century.

The explanation is the network mechanism from Module 2. Sound change spreads through repeated face to face interaction, not through passive listening. You do not acquire a vowel system from a television, because a television neither responds to you nor makes you a member of anything. This is why immigrant children acquire the local variety from playmates rather than from their parents' media consumption, and it is why heavy media exposure has left the map more differentiated rather than less.

The atlas work also supports a particular model of how changes travel. In the wave model, a change spreads outward from its origin like ripples, reaching nearby places first. In the hierarchical or cascade model, it jumps from large city to large city and only later fills in the countryside between them. American vowel changes largely follow the cascade pattern, which is why a change may be established in two distant cities while the towns between them are unaffected.

Key idea: Acoustic measurement in the Atlas of North American English showed regional vowel systems diverging rather than converging, because change spreads through face to face interaction and often cascades between cities rather than rippling outward.

Where people think dialects are

One more strand rounds this out. Perceptual dialectology, developed especially by Dennis Preston, studies not where dialects are but where ordinary people believe they are. In the draw a map task, respondents mark on a blank map the regions where they think people speak differently, and label them.

The results diverge from the production data in revealing ways. Respondents draw sharp boundaries where linguists find gradients. They identify many more distinct areas near their own home than far away. And their labels are heavily evaluative, with the same region attracting descriptions like friendly from some respondents and uneducated from others. Large crowd-sourced instruments, notably Bert Vaux's Harvard Dialect Survey and the widely circulated newspaper quizzes built from it, tap a related kind of self-report data, trading precision for extraordinary geographic reach. All of this is the natural bridge to the next lesson, which is about attitudes toward varieties rather than the varieties themselves.

Key idea: Perceptual dialectology maps beliefs about dialects, which show sharper boundaries, finer detail near home, and heavily evaluative labels compared with measured production data.

Common misconceptions

  • A language is separated from a dialect by mutual intelligibility. The criterion fails in both directions, is asymmetric, and improves with exposure; the distinction is political.
  • Dialect maps show real lines. Colored regions summarize gradient data, and cases like the Rhenish fan show individual isoglosses splaying rather than bundling.
  • Television is flattening American accents. Acoustic measurement shows major regional vowel systems diverging, because change requires interactive face to face contact.
  • Dialect surveys mainly record quaint old words. Modern atlas work measures vowel formants acoustically and tracks changes in progress in urban populations.
  • Changes spread outward from a center like ripples. Many changes cascade hierarchically between large cities, leaving intervening rural areas unaffected until later.

Recap

  • Mutual intelligibility cannot separate languages from dialects, so linguists use variety and treat the distinction as political.
  • Dialect continua make geographic variation gradual, with intelligibility failing only across long stretches.
  • Wenker's postal survey maximized coverage and Gilliéron's single fieldworker maximized consistency, defining the field's central trade-off.
  • Kurath's Linguistic Atlas established the Northern, Midland, and Southern division, and DARE documented regional vocabulary in depth.
  • An isogloss bounds one feature; dialect areas require bundles, and the Rhenish fan shows how often bundles fail to form.
  • The Atlas of North American English found the Northern Cities Shift, the Southern Shift, and a spreading low back merger, with regions diverging rather than converging.
  • Perceptual dialectology shows that beliefs about dialect boundaries are sharper, more locally detailed, and more evaluative than the production data.

Sources

  1. Wikipedia contributors. (n.d.). Atlas of North American English. en.wikipedia.org
  2. Cassidy, F. G., et al. (n.d.). Dictionary of American Regional English. University of Wisconsin. dare.wisc.edu
  3. Britannica. (n.d.). Dialect. Encyclopaedia Britannica. britannica.com
  4. Wikipedia contributors. (n.d.). Isogloss. en.wikipedia.org
  5. MacNeil, R., & Cran, W. (2005). Do you speak American? PBS. pbs.org
Key terms
Variety
A neutral term for any identifiable form of language, used because the language and dialect distinction is political rather than linguistic.
Dialect continuum
A chain of varieties in which each is mutually intelligible with its neighbors but the endpoints are not intelligible to each other.
Abstand and Ausbau
Kloss's distinction between varieties counted as separate languages by structural distance and those counted separate by institutional elaboration.
Isogloss
A line on a map separating the area using one variant of a single feature from the area using another.
Isogloss bundle
Several isoglosses running roughly together, taken as evidence of a genuine dialect boundary, often reflecting historical or geographic barriers.
Rhenish fan
The splaying pattern of High German consonant shift isoglosses near the Rhine, the classic demonstration that isoglosses need not bundle.
Northern Cities Shift
A rotation of short vowels documented across the Inland North of the United States, from Rochester through Chicago.
Perceptual dialectology
The study of where non-linguists believe dialects are located and what evaluative labels they attach to them.

The Standard Language: How It Was Made, Prestige, and Attitudes

  • Trace the actual historical process by which English acquired a standard, using Haugen's stages of selection, codification, elaboration, and acceptance.
  • Distinguish overt from covert prestige and explain the self-report evidence that revealed the second.
  • Describe the matched guise technique and summarize what accent evaluation research finds about status and solidarity.

The big picture

Standard English is often treated as the language itself, with everything else a departure from it. This lesson replaces that picture with a history, because the standard has one, and it is surprisingly short and surprisingly local. Standard English descends from the speech of one region, spread through specific institutions, was fixed by specific technologies, and was codified by named people with datable opinions.

Establishing that does not make the standard unimportant. It is enormously important, which is why the second half of this lesson turns to prestige and to the experimental methods that measure how listeners actually react to accents. Those reactions are consistent, they are measurable, and they are consequential, which is the subject of the final module.

Key idea: A standard language is a variety that particular historical processes elevated, and understanding those processes is what separates the fact of its prestige from the myth of its inherent superiority.

What standardization is

Einar Haugen described standardization as four processes, and they are a useful checklist for any language.

Selection is the choice of one variety to develop, usually the speech of a politically or economically dominant group. Codification is fixing its forms in dictionaries, grammars, and spelling conventions, so that variation is reduced and there is a right answer. Elaboration is extending the variety to new functions: law, science, administration, literature. Acceptance is the population coming to regard it as the proper form, including those who do not speak it.

Notice that only codification and elaboration involve the language at all. Selection is a political and economic outcome, and acceptance is an ideological one. That is the structural reason standard varieties are not linguistically superior: nothing in the process tests them for expressive power, and no candidate variety was ever eliminated for lacking it.

Key idea: Standardization proceeds through selection, codification, elaboration, and acceptance, and only the middle two involve the language itself, while selection and acceptance are political and ideological.

How English standardized

Now the specific history. In the fourteenth century there was no standard English. Writers used their own regional forms, and spelling varied even within a single manuscript.

The variety that would become the standard was the East Midland dialect as it was spoken in London. Its advantage was not linguistic. London was the seat of government and the commercial center, and the East Midland region sat inside the triangle formed by London, Oxford, and Cambridge, so its speech had the densest contact with administration, trade, and learning. Selection, in Haugen's sense, was decided by that geography.

In the fifteenth century, the royal administrative offices known collectively as Chancery began producing documents with increasingly consistent spellings. Chancery Standard was a written norm imposed by bureaucratic convenience, and because government documents circulated everywhere, its choices spread.

Then came printing. William Caxton set up his press in Westminster in 1476, and printers had to choose spellings, since a page cannot hedge. Caxton himself described the problem vividly in a prologue, recounting merchants stranded on the Thames who asked a woman for eggs and were told she spoke no French, because in her dialect the word was eyren. His question, which form should a printer use, is exactly the standardizer's question, and printers answered it by settling on London forms.

Here is the accident that shaped English spelling permanently. Printing fixed spellings during the fifteenth and sixteenth centuries, precisely while the Great Vowel Shift was moving the pronunciation of every long vowel in the language. Spelling froze; pronunciation kept going. That is why the vowel letters in English do not match their values in other European languages, and why words like knight and through preserve consonants nobody has pronounced for centuries. English spelling is not illogical. It is a photograph of the fifteenth century.

Codification arrived properly in the eighteenth century. Samuel Johnson's Dictionary of 1755 gave English an authoritative wordlist. Robert Lowth's grammar of 1762 and Lindley Murray's of 1795 supplied rules, including several of the prescriptions examined in Lesson 1. Jonathan Swift proposed in 1712 an academy to regulate English on the French model. It was never established, which is why English, unlike French or Spanish, has no official body with the authority to rule on usage. Its standard is enforced entirely by publishers, schools, and social pressure.

Spoken prestige was codified later and separately. Received Pronunciation emerged in the nineteenth century as the accent of the English public schools, spread among a socially mobile elite. Alexander Ellis used the term in 1869, Daniel Jones described it phonetically, and the BBC adopted it from its founding in 1922, which is why it was long called BBC English. It is worth stating plainly that RP has never been the accent of more than a small minority of Britons, usually estimated at a few percent, and the BBC itself now broadcasts in a range of regional accents.

Meanwhile Noah Webster's American dictionaries deliberately diverged from British practice, giving American English its color, center, and defense spellings as a nationalist project. Standardization is not one process with one outcome; it happens separately in each polity that undertakes it.

PeriodDevelopmentEffect on the standard
1300sNo standard; regional writing throughoutSelection has not occurred
1400sChancery documents adopt consistent spellingsA written administrative norm emerges
1476 onwardCaxton and the printing pressSpelling fixed on London forms and frozen mid vowel shift
1755 to 1795Johnson's dictionary, Lowth's and Murray's grammarsCodification of words and rules
1800sReceived Pronunciation in the public schoolsA prestige accent, later spread by the BBC

Key idea: English standardized from the East Midland speech of London through Chancery, printing, and eighteenth century codification, with spelling frozen during the Great Vowel Shift and a prestige accent codified only in the nineteenth century.

Standard language ideology

James and Lesley Milroy gave a name to the belief system that grows around a standard. Standard language ideology is the conviction that a language has one correct form, that this form is the language proper, and that deviation is error rather than difference.

Its most visible expression is what they call the complaint tradition: the durable public genre in which commentators announce that the language is deteriorating. This genre has run continuously in English for centuries, always describing decline, always locating the golden age a generation or two before the writer, and never producing a language that has actually collapsed. The persistence of the complaint across five hundred years of allegedly terminal decline is itself the strongest evidence against its premise.

The ideology has a practical consequence worth naming. Because the standard is treated as the absence of dialect rather than as a dialect, its speakers are heard as having no accent while everyone else is heard as having one. That is an illusion of familiarity. Every speaker of every variety has an accent, including newsreaders.

Key idea: Standard language ideology treats one variety as the language itself, generating a centuries-old complaint tradition and the illusion that standard speakers have no accent.

Overt and covert prestige

If the standard is so advantageous, a puzzle follows: why does anyone keep using nonstandard forms? Speakers know the standard exists. They hear it constantly. Many can produce it. Yet local vernaculars persist for generations.

The answer is that there are two kinds of prestige. Overt prestige is the openly acknowledged kind that attaches to the standard: the forms people say are correct and that carry institutional reward. Covert prestige is the unacknowledged value attached to nonstandard forms within a community, where they signal toughness, authenticity, local loyalty, and solidarity.

Trudgill's Norwich data revealed this with an elegant piece of evidence. Alongside recording what speakers actually said, he asked them which variant they used. The self-reports diverged from the recordings in a systematic and revealing way: men frequently claimed to use nonstandard local forms that the recordings showed they did not use, while women frequently claimed standard forms that the recordings showed they did not use. Both groups misreported, in opposite directions.

That result is only interpretable if the two groups were aspiring to different targets. The men were claiming the covert prestige of local working class speech. The women were claiming the overt prestige of the standard. Nobody was confused about the facts; both were reporting an identity rather than a measurement. It also supplies part of the answer to the gender paradox from the previous lesson.

Key idea: Overt prestige attaches to the standard and covert prestige to local vernaculars, and Norwich self-report data showed men over-claiming nonstandard forms while women over-claimed standard ones.

Measuring attitudes: the matched guise technique

Attitudes toward accents are hard to survey directly, because people know which answers are socially acceptable. Wallace Lambert and colleagues solved this in 1960 with a design that has been used ever since.

In the matched guise technique, perfectly bilingual or bidialectal speakers record the same passage twice, once in each variety. Listeners hear the recordings scattered among fillers, believe they are hearing different people, and rate each voice on traits like intelligence, ambition, kindness, and height. Because the same person produced both recordings, everything except the variety is held constant: voice quality, pitch, speaking rate, and content. Any difference in ratings must be a reaction to the variety itself.

Lambert's Montreal results were sobering. Listeners rated the English guises more favorably than the French guises on traits like intelligence and ambition. The finding that mattered most was that French-Canadian listeners did so too, rating the English guises above the French ones and, in effect, above voices from their own community. Prestige hierarchies are not merely imposed from outside; they are absorbed by the groups they disadvantage.

Decades of accent evaluation research, much of it associated with Howard Giles in Britain, have refined the picture into a consistent two-dimensional pattern. Standard accents like RP score highest on status and competence: intelligent, ambitious, educated, successful. Regional and urban accents score lower on those traits but as high or higher on solidarity: friendly, warm, trustworthy, sincere. Certain urban accents score low on both, which is where the sharpest discrimination clusters.

Accent typeStatus and competence ratingsSolidarity and social attractiveness ratings
Standard prestige accentHighestLower
Rural regional accentLowerHigh
Stigmatized urban accentLowestLow to moderate

Two cautions before leaving this. A matched guise study measures reactions in a controlled setting, which is not the same as behavior in a rental office or a hiring decision. And ratings differ by who is listening, since in-group listeners often rate their own variety higher on solidarity. Module 6 takes up the studies that measure the behavior directly.

Key idea: The matched guise technique holds the speaker constant while varying the variety, revealing that standard accents are rated higher on status and lower on solidarity, and that disadvantaged groups often share the hierarchy.

Common misconceptions

  • Standard English is the original English that dialects departed from. It descends from one regional variety among many and has changed continuously since.
  • English spelling is arbitrary or badly designed. It largely records fifteenth century pronunciation, frozen by printing while the Great Vowel Shift continued.
  • An academy governs English. Swift's 1712 proposal failed, and English has no regulatory body; its norms are enforced by publishers, schools, and social pressure.
  • Received Pronunciation is how most British people speak. It has always been the accent of a small minority, and the BBC now broadcasts in many regional accents.
  • People use nonstandard forms because they do not know better. Covert prestige gives local forms real value as markers of solidarity and authenticity.

Recap

  • Haugen's four stages are selection, codification, elaboration, and acceptance, and only the middle two concern the language itself.
  • English standardized from London East Midland speech through Chancery documents and the printing press.
  • Spelling was fixed during the Great Vowel Shift, which explains most English spelling irregularity.
  • Johnson, Lowth, and Murray codified English in the eighteenth century, and no academy was ever established.
  • Received Pronunciation was codified in the nineteenth century and has never been the accent of more than a small minority.
  • Standard language ideology sustains a complaint tradition and the illusion that standard speakers have no accent.
  • Overt prestige attaches to the standard and covert prestige to vernaculars, as the Norwich self-report data revealed.
  • Matched guise studies find standard accents rated higher on status and lower on solidarity, with hierarchies internalized by disadvantaged groups.

Sources

  1. Lambert, W. E., Hodgson, R. C., Gardner, R. C., & Fillenbaum, S. (1960). Evaluational reactions to spoken languages. Journal of Abnormal and Social Psychology, 60(1), 44-51. doi.org/10.1037/h0044430
  2. Britannica. (n.d.). English language. Encyclopaedia Britannica. britannica.com
  3. Wikipedia contributors. (n.d.). Standard English. en.wikipedia.org
  4. Wikipedia contributors. (n.d.). Received Pronunciation. en.wikipedia.org
  5. MacNeil, R., & Cran, W. (2005). Do you speak American? PBS. pbs.org
Key terms
Standardization
Haugen's four-part process of selecting a variety, codifying its forms, elaborating its functions, and gaining population acceptance.
Chancery Standard
The consistent written norm developed in fifteenth century English royal administrative offices, which spread through official documents.
Great Vowel Shift
The long-vowel sound change that continued after printing fixed English spelling, producing most modern spelling irregularity.
Received Pronunciation
The prestige British accent codified in the nineteenth century public schools, adopted by the BBC and spoken by a small minority.
Standard language ideology
The Milroys' term for the belief that a language has one correct form and that other varieties are errors rather than differences.
Complaint tradition
The centuries-old public genre asserting that the language is deteriorating, always locating its golden age a generation earlier.
Overt prestige
The openly acknowledged social value attached to standard forms and to the institutional rewards they carry.
Covert prestige
The unacknowledged value attached to nonstandard local forms as markers of solidarity, authenticity, and group loyalty.
Matched guise technique
Lambert's method in which one bilingual or bidialectal speaker records the same passage in two varieties, so listener ratings isolate the variety.

Module 4: Varieties in Depth

A close structural look at African American English, the debates about its origins and its treatment in schools, and then a survey of other varieties that face similar misreadings, from Chicano English and Appalachian English to Indian English and the global Englishes.

African American English as a Rule-Governed System

  • Describe the principal phonological features of African American English and the constraints that govern them.
  • Explain the AAE aspect system, including habitual be, stressed BIN, and completive done, and the meanings each marker contributes.
  • State the copula absence rule and explain why its distribution is decisive evidence that the variety is rule-governed.

The big picture

African American English is the most thoroughly researched variety of American English and the most widely misunderstood. That combination is not an accident. Sixty years of linguistic work have described its structure in exhaustive detail, and almost none of that work has reached the classrooms, courtrooms, and offices where its speakers are evaluated.

This lesson does the structural work. It takes the features most often cited as errors and shows what rule each one follows. The argument is not that the variety deserves respect as a matter of courtesy. The argument is that it has a grammar, that the grammar can be stated precisely, that speakers obey it without exception, and that some of its distinctions are finer than anything Standard English can express with a single verb form. Those are empirical claims, and they are what the evidence shows.

Some terminology first. The variety has been called Black English, African American Vernacular English, Ebonics, and African American Language. Current usage in the field tends toward African American English, abbreviated AAE, and this course uses that. Three clarifications matter. It is not spoken by all African Americans, and many speak only standard varieties. It is spoken by some people who are not African American, typically those raised in the relevant communities. And it is not uniform: AAE has regional variation of its own, so a speaker in Philadelphia and a speaker in Houston differ. AAE is defined by a set of features and a community of practice, not by the ethnicity of any individual speaker.

Key idea: African American English is a systematically described variety defined by linguistic features and community, not by any speaker's ethnicity, and its grammar can be stated with the precision of any other grammar.

Sound patterns and their constraints

Start with pronunciation, because the constraints are easy to demonstrate and immediately undercut the idea of carelessness.

Consonant cluster reduction is the most cited feature: test can be pronounced tes, hand as han, desk as des. Notice first that this is not unique to AAE; every English speaker reduces clusters sometimes, as in the ordinary pronunciation of last night. What differs is the rate. What is decisive, though, is that the rate is not free. Reduction is much more likely when the next word begins with a consonant than when it begins with a vowel, so tes case is far more common than tes area. And reduction is disfavored when the final consonant is itself a separate meaningful unit: the final d of a past tense form like guessed resists reduction more than the final d of a single-morpheme word like guest. A speaker who was simply being lazy would have no reason to treat two identical-sounding endings differently depending on whether one of them carries grammatical meaning. AAE speakers do, consistently.

Other well documented patterns include the treatment of the th sounds, which vary by position: word-initially they are often produced as stops, so this becomes dis and thing becomes ting or thing, while word-finally they are often produced as labiodentals, so bath becomes baf and mother becomes muvver. Again the pattern is positional and regular, not random substitution. AAE is also non-rhotic in the manner of many Southern and coastal varieties, and it vocalizes l in some positions. Some words carry stress on a different syllable than in other varieties, as in POlice and UMbrella.

Key idea: AAE phonological features are governed by precise constraints, such as cluster reduction being sensitive to both the following sound and whether the final consonant carries grammatical meaning.

The copula rule, and why it settles the argument

Now the feature that produced the field's most elegant demonstration. In AAE, forms of the verb be can be absent: He tired. She the one who did it. They going home.

Called sloppy, this looks like a missing word. Labov's 1969 analysis showed it is nothing of the kind, and the argument is worth following step by step because it is a model of how the descriptive case is made.

Standard English contracts the copula: he is tired becomes he's tired. But contraction is not free in Standard English either; there are places where every speaker of every variety refuses it. You can say he's tired, but at the end of a sentence you cannot say the starred form: asked whether he is tired, nobody answers with a contracted yes he's. You must say yes he is. Similarly, contraction is blocked in other specific environments.

Labov's finding was that AAE copula absence occurs exactly and only where Standard English permits contraction. Where Standard English blocks contraction, AAE blocks absence. Speakers say yes he is, never the starred form yes he. First person am can contract to I'm but is not absent. Past tense was is never absent. The environments line up precisely.

Follow what that means. The AAE pattern is not the absence of a rule; it is an additional step applied to the output of a rule that Standard English already has. Where Standard English goes from is to 's, AAE can go one step further to nothing, and it can do so only where the first step was licensed. A speaker being careless could not produce this pattern, because carelessness has no way of knowing which environments Standard English contraction permits. Only a grammar knows that.

EnvironmentStandard English contractionAAE absence
Before an adjective: he is tiredAllowed: he's tiredAllowed: he tired
Before a noun phrase: she is the oneAllowed: she's the oneAllowed: she the one
Sentence-final: yes he isBlockedBlocked
First person: I am hereAllowed: I'm hereBlocked
Past tense: he was hereBlockedBlocked

Key idea: AAE copula absence occurs precisely where Standard English allows contraction and is blocked precisely where contraction is blocked, which no account based on carelessness can explain.

The aspect system

The heart of AAE grammar is its treatment of aspect, meaning the internal shape of an event: whether it is ongoing, habitual, completed, or extended over time. Standard English marks tense heavily and aspect lightly. AAE does the reverse, with a set of preverbal markers that Standard English simply lacks.

Habitual be is the best known. Uninflected be marks an event as recurring or characteristic. Crucially, it contrasts with the absent copula:

  • He working means he is working right now.
  • He be working means he works regularly, habitually, as a matter of routine.

These are different sentences with different truth conditions, and speakers do not interchange them. The starred sentence he be working right now is ill-formed, because the adverb contradicts the habitual marker. That an added adverb can render a sentence ungrammatical is exactly the kind of fact that only exists inside a grammar.

Experimental evidence confirms that this is real knowledge rather than analyst's invention. In comprehension studies using picture and cartoon tasks, children who speak AAE reliably interpret a sentence with be as describing a habitual situation, choosing the picture showing a repeated activity, while children who speak mainstream varieties choose the picture showing the activity happening once, right now. The two groups are not hearing the same sentence.

Stressed BIN, conventionally written in capitals because the stress is what carries the meaning, marks the remote past with continuing relevance. She BIN married does not mean she was married once. It means she has been married for a long time and still is. John Rickford's comprehension experiments found that AAE speakers and non-speakers systematically diverge on exactly this: presented with She BIN married, AAE speakers overwhelmingly conclude that she is still married, while non-speakers frequently conclude that she is not. This is a case where a listener can be certain they understood a sentence and be wrong about its meaning, which has obvious implications for classrooms and courtrooms.

Completive done marks an action as finished, often with emphasis: He done finished his homework. Combined markers push further. Be done marks a resultant or future perfect state: I be done finished by the time you get here. The system stacks.

MarkerExampleMeaningStandard English equivalent
Zero copulaHe workingHappening right nowHe is working
Habitual beHe be workingRecurring, characteristicHe usually works (no single verb form)
Stressed BINShe BIN marriedSince long ago, and stillShe has been married for a long time
Completive doneHe done finishedCompleted, emphaticHe has already finished
Be doneI be done finishedWill have completed by thenI will have finished

Read that table as an argument. Every row in the middle column expresses something Standard English needs an adverb or a multi-word construction to convey, and the habitual and remote past distinctions have no exact Standard English equivalent at all. On the specific dimension of aspect, AAE is the more finely differentiated system.

Key idea: AAE marks aspect with a set of preverbal markers, including habitual be and stressed BIN, that draw distinctions Standard English cannot express with any single verb form.

Other grammatical features

Several further features round out the system, each regular.

Third person singular present tense verbs need not carry -s: he walk, she go. This is often described as dropping an ending, but note that the ending carries no information the subject has not already supplied. Many of the world's languages, including standard varieties of Mandarin and Swedish, have no such agreement marker at all, and English itself lost most of its person and number endings centuries ago. AAE has simply carried a long-running English trend one step further.

Possession can be marked by word order rather than by -s: John book, my sister car. Because English possessive order is fixed, the relationship is unambiguous without the ending.

Negation shows several patterns. Negative concord permits multiple negatives, as covered in Lesson 1. Ain't serves as a general negator, including for did not in some contexts. Negative inversion moves the negated auxiliary to the front: Don't nobody like him, meaning nobody likes him. That last construction repays attention, because it is a syntactic rule with a specific trigger and a specific output, and no one taught it to anyone.

Existential sentences use it or dey where other varieties use there: It's a lot of people here, meaning there are a lot of people here. And embedded questions may keep inverted order: I asked him did he go, where Standard English requires whether he went.

Key idea: Features such as agreement without -s, possessive by word order, negative inversion, and existential it are each governed by specific rules, several of which continue long-running historical trends in English.

Why this matters before we get to the debates

The next lesson takes up where AAE came from and what happened when a school district tried to act on this research. Before that, hold onto the structural result, because every later argument depends on it.

The features surveyed here are not approximations of Standard English. They are the output of a grammar with its own constraints, some of which encode meanings Standard English cannot encode compactly. A listener unfamiliar with that grammar will misunderstand specific sentences, and will typically not know they have misunderstood, since the words are all familiar. That is the mechanism behind a great deal of what looks like deficiency: it is a comprehension failure on the listener's side being recorded as a production failure on the speaker's side.

Key idea: Misunderstandings of AAE typically arise from a listener's unfamiliarity with its grammar, and are then misrecorded as deficiencies in the speaker.

Common misconceptions

  • AAE is Standard English with mistakes in it. Its features follow constraints that carelessness could not produce, most clearly the copula rule's exact match with Standard English contraction environments.
  • Habitual be just means is. He working and he be working differ in meaning, and adding right now makes the be sentence ungrammatical.
  • All African Americans speak AAE, and only they do. The variety is defined by features and community, and use varies widely among individuals of any background.
  • Dropping the -s makes the sentence ambiguous. The subject already supplies person and number, and many languages, including standard Mandarin and Swedish, lack the marker entirely.
  • AAE is simpler than Standard English. On aspect it is more finely differentiated, marking habitual and remote past distinctions Standard English cannot express in one verb form.

Recap

  • AAE is defined by linguistic features and community membership, is regionally varied, and is not coextensive with any ethnic group.
  • Consonant cluster reduction is constrained by the following sound and by whether the final consonant carries grammatical meaning.
  • Copula absence occurs exactly where Standard English permits contraction and is blocked exactly where contraction is blocked.
  • Habitual be marks recurring events and contrasts systematically with the zero copula for current events.
  • Stressed BIN marks a state beginning long ago and continuing, and non-speakers systematically misinterpret it.
  • Completive done and combined markers such as be done extend the aspect system further.
  • Agreement without -s, possessive word order, negative inversion, and existential it are each rule-governed.
  • Apparent deficiency is often a listener's comprehension failure recorded as a speaker's production failure.

Sources

  1. Labov, W. (1969). Contraction, deletion, and inherent variability of the English copula. Language, 45(4), 715-762. doi.org/10.2307/412333
  2. Wikipedia contributors. (n.d.). African-American Vernacular English. en.wikipedia.org
  3. Rickford, J. R. (n.d.). Publications on African American Vernacular English. Stanford University. johnrickford.com
  4. Linguistic Society of America. (n.d.). What is sociolinguistics? linguisticsociety.org
  5. MacNeil, R., & Cran, W. (2005). Do you speak American? PBS. pbs.org
Key terms
African American English (AAE)
A variety of English defined by a set of systematic phonological and grammatical features and associated with African American communities, with its own regional variation.
Consonant cluster reduction
The simplification of word-final consonant clusters, constrained by the following sound and by whether the final consonant is a separate morpheme.
Copula absence
The omission of forms of be, occurring exactly in the environments where Standard English permits contraction.
Aspect
The grammatical marking of an event's internal shape, such as ongoing, habitual, or completed, as distinct from its location in time.
Habitual be
Uninflected be marking an event as recurring or characteristic, contrasting with the zero copula used for current events.
Stressed BIN
A stressed remote past marker indicating that a state began long ago and continues, systematically misinterpreted by non-speakers.
Completive done
A preverbal marker indicating that an action has been completed, often with emphatic force.
Negative inversion
A construction placing the negated auxiliary at the front of the clause, as in Don't nobody like him.

Origins, the Oakland Controversy, and What Actually Helps Students

  • Compare the anglicist, creolist, and substrate accounts of AAE origins and state where the evidence currently sits.
  • Give a factual account of the 1996 Oakland resolution, what it actually proposed, how it was misreported, and how linguists responded.
  • Summarize the research on educational approaches to dialect difference and identify which interventions have empirical support.

The big picture

The previous lesson established that African American English has a grammar. Two questions follow immediately. Where did that grammar come from? And what should schools do with children who bring it to class?

The first question has been debated for sixty years and remains genuinely open, which makes it a good case study in how historical linguistics argues from incomplete evidence. The second question is less open than the public argument suggests. There is a real research literature on teaching standard varieties to speakers of stigmatized ones, it points fairly consistently in one direction, and that direction is close to the opposite of what most classrooms actually do.

Between the two sits the most publicly explosive episode in the history of American linguistics, in which a school district tried to act on the research and was buried in ridicule within a week. Telling that story accurately matters, because almost everything most people know about it is wrong.

Key idea: The origins of AAE remain a live scholarly debate, while the question of what helps its speakers in school is far better settled than public argument suggests.

Where did AAE come from?

The evidence problem is severe. Nobody recorded the speech of enslaved Africans and their descendants in the seventeenth and eighteenth centuries. The written record consists of literary dialect written mostly by white authors, which is unreliable, plus court transcripts and letters. Every position in this debate is an inference from fragmentary material, and honest researchers say so.

The anglicist position, associated with dialectologists such as Hans Kurath and Raven McDavid, holds that AAE derives principally from the regional English dialects of the British settlers among whom enslaved people lived and worked. Its evidence is the substantial overlap between AAE and older Southern and Appalachian white varieties: negative concord, ain't, a-prefixing in some historical records, and various pronunciations are shared. On this account AAE is an English dialect that diverged like other English dialects.

The creolist position, associated with William Stewart, J. L. Dillard, and John Rickford, holds that AAE descends from a plantation creole similar to those of the Caribbean, which later decreolized under sustained contact with English. Its evidence is structural: the absence of copula, the preverbal aspect markers, the absence of inflectional endings, and existential constructions all have close parallels in Caribbean English creoles. The strongest single piece of support is Gullah, an English-based creole still spoken in the Sea Islands of South Carolina and Georgia, which demonstrates that creolization did occur on the North American mainland.

The neo-anglicist position emerged from a body of evidence assembled in the 1980s and 1990s, much of it by Shana Poplack and colleagues. They examined varieties spoken by descendants of nineteenth century African American emigrants who had been isolated from later developments: the Samana English of the Dominican Republic, African Nova Scotian English, and the recorded interviews with formerly enslaved people made in the 1930s. These diaspora varieties turned out to look more like older regional English than like creoles, which suggests that earlier AAE was less creole-like than the creolist account predicts, and that some of its distinctive modern features developed later.

Most researchers now hold a combined position. English dialect input is clearly the main source of the lexicon and much of the grammar. Substrate influence from West African languages plausibly shaped parts of the sound system and aspect marking. Contact conditions varied enormously between, say, a Sea Island rice plantation with a large African-born majority and a small farm in Virginia, so different degrees of creolization in different places is the expected outcome rather than an awkward compromise.

A related dispute ran through the 1980s. The divergence hypothesis, advanced by Labov and Wendell Harris and by Guy Bailey and Natalie Maynor, argued that urban AAE was moving further away from local white vernaculars rather than converging with them, driven by residential segregation, and pointed to features such as habitual be appearing more frequently among younger speakers. Others contested the interpretation. The unresolved status of that argument is itself informative: it shows that AAE is a living variety with changes in progress, not a residue of the past.

PositionCore claimPrincipal evidence
AnglicistDescended from British regional dialectsShared features with older Southern and Appalachian white varieties
CreolistDescended from a plantation creole, later decreolizedStructural parallels with Caribbean creoles; the existence of Gullah
Neo-anglicistEarlier AAE resembled older English more than creolesDiaspora varieties in Samana and Nova Scotia; ex-slave recordings
Combined viewEnglish input plus substrate influence, varying by local contact conditionsThe wide variation in plantation demographics across regions

Key idea: Anglicist, creolist, and neo-anglicist accounts each rest on partial evidence, and most researchers now accept a combined view in which the degree of creole influence varied with local contact conditions.

Oakland, 1996: what actually happened

On December 18, 1996, the school board of the Oakland Unified School District in California passed a resolution concerning the language of its African American students. Within days it was national news, and within a week the story being reported bore little resemblance to the document.

Start with the situation. African American students made up over half the district's enrollment and were substantially overrepresented in special education referrals and disciplinary actions while being underrepresented in gifted programs. Their average grade point average was well below the district mean. A district task force had spent months on the problem and concluded that one contributing factor was that teachers, largely unfamiliar with the students' home variety, were treating systematic dialect features as errors and mistaking dialect difference for deficiency.

The resolution's actual proposals were these: recognize that many students come to school speaking a systematic variety; train teachers to recognize its features so they can distinguish dialect from error; and use that knowledge to teach Standard English more effectively. The goal stated in the document was proficiency in Standard English.

Two features of the text produced the explosion. The first was the word Ebonics, coined in 1973 by the psychologist Robert Williams, which was unfamiliar to the public and sounded to many like an invented euphemism. The second was far more damaging: the original resolution described the students' language as genetically based. In linguistics, genetic is standard terminology for descent relationships among languages, as in the Indo-European genetic family. Read by a general audience, it appeared to claim a biological basis for how Black children speak. The district amended the wording in January 1997, but by then the story had set.

The reporting compounded it. The dominant national story was that Oakland would teach Ebonics to students instead of English, or would teach classes in Ebonics. The resolution proposed neither. Condemnation was swift and came from across the political spectrum, including from prominent African American public figures, and much of it was mocking. Comedians ran with it for months.

The Linguistic Society of America responded at its January 1997 annual meeting with a formal resolution. It stated that the systematic and rule-governed nature of AAE is well established among linguists; that characterizations of it as slang, mutant, or broken English are incorrect and demeaning; that the distinction between a language and a dialect is a social and political matter rather than a linguistic one; and that the Oakland decision to recognize the students' vernacular in teaching them the standard was linguistically and pedagogically sound. The statement received a fraction of the coverage the controversy did.

The aftermath is the part worth dwelling on. The program was never implemented as designed. Federal funding was not sought in the end, the task force wound down, and other districts drew the obvious lesson about what happens to administrators who raise the topic. A pedagogically defensible proposal was destroyed by a terminology choice and a misreported summary, and the chilling effect on school districts long outlasted the news cycle.

Key idea: The Oakland resolution proposed training teachers to recognize AAE in order to teach Standard English more effectively, was widely misreported as proposing to teach Ebonics instead of English, and was endorsed as sound by the Linguistic Society of America.

The deficit theory and its refutation

To see why any of this mattered, look at what the resolution was reacting against. In the 1960s a widely adopted account held that children from poor and minority households suffered verbal deprivation: that they heard little language at home, possessed impoverished vocabularies, and could not use language for logical reasoning. Carl Bereiter and Siegfried Engelmann built intervention programs on that premise, in some cases treating children as effectively without language.

Labov's 1969 paper on the logic of nonstandard English dismantled this. He showed that the evidence for verbal deprivation came from testing situations that guaranteed silence: a young Black child alone in a room with a large unfamiliar white adult holding a test protocol produces monosyllables, and this was recorded as an absence of language. Change the setting, bring in a familiar interviewer, sit on the floor, allow a friend to be present, and the same child produces fluent, complex, argumentatively sophisticated speech. Labov also demonstrated that AAE speakers construct logical arguments with full formal machinery, contrary to the claim that the variety could not support reasoning.

The methodological lesson generalizes well beyond this case: a testing situation that suppresses the behavior being measured will reliably produce evidence of its absence.

Key idea: Verbal deprivation theory rested on testing conditions that suppressed speech, and Labov showed that the same children produce fluent, logically complex language once the setting is changed.

What the evidence says helps

So what does work? Several converging lines of research point the same way.

First, dialect awareness and contrastive analysis. Rather than marking home-variety features as errors, this approach teaches students to compare the two systems explicitly: here is how your variety marks possession, here is how the school variety does, here is when each is expected. Julie Sweetland's classroom research in Cincinnati and Howard Fogel and Linnea Ehri's experimental work both found that explicit contrastive instruction improved students' Standard English writing more than exposure to the standard alone did. Making the systems visible is what allows students to move between them deliberately.

Second, dialect readers. In the 1970s Gary and Charlesetta Simpkins developed the Bridge series, reading materials that began in AAE and transitioned gradually toward Standard English. In evaluation, students using Bridge made substantially larger reading gains than comparison students on standard instruction. The materials were withdrawn after public objection. The reason for discontinuation was the reaction, not the results.

Third, the assessment problem. Research by Anne Charity, Hollis Scarborough, and Darion Griffin found that young African American children's familiarity with school English predicted early reading achievement. That correlation is easy to misread as evidence of deficit. The mechanism it points to is mismatch: a child who reads the printed word asked aloud as ask, following their own phonology, is producing a correct decoding of the word, but a scoring protocol that does not recognize dialect will mark it wrong. Holly Craig and Julie Washington's work similarly found that children who could shift toward school English in school contexts showed stronger reading outcomes, supporting explicit instruction in shifting rather than in replacement.

Fourth, and closely related, over-identification in special education. Because standardized language assessments were normed on mainstream speakers, dialect features register as errors, and AAE-speaking children have been disproportionately referred for speech and language services. Dialect-sensitive assessment instruments were developed specifically to separate genuine language disorder from dialect difference, and their existence is an acknowledgment that the earlier instruments were measuring the wrong thing.

What does not help is equally clear. Correcting a child's pronunciation while they are reading interrupts comprehension and teaches them that reading is a performance to be policed. Treating dialect features in oral reading as decoding errors produces a false record of reading difficulty. And telling children that their home speech is wrong asks them to accept that their family and community speak badly, which is both false and a poor foundation for learning anything.

The consensus position among linguists working in education is additive rather than substitutive: students should gain full command of the standard variety because its social value is real, and they should gain it as an addition to their home variety rather than as a replacement for it. That is roughly what Oakland proposed.

Key idea: Contrastive analysis, dialect readers, and dialect-sensitive assessment have empirical support, and the consensus approach is additive, teaching the standard as an addition to the home variety rather than as a correction of it.

Common misconceptions

  • Oakland proposed teaching students in Ebonics instead of English. The resolution's stated goal was proficiency in Standard English; the proposal was to train teachers to recognize the home variety.
  • The genetically based wording claimed a biological basis for Black speech. Genetic is standard linguistic terminology for descent among languages; the wording was ambiguous and was amended.
  • Linguists were divided about Oakland. The Linguistic Society of America formally endorsed the decision as linguistically and pedagogically sound.
  • The Bridge dialect readers failed. Students using them outgained comparison groups; the program ended because of public objection.
  • The origins question has been settled. Anglicist, creolist, and neo-anglicist evidence all carry weight, and the combined view accommodates varying local contact conditions.

Recap

  • The anglicist account derives AAE from British regional dialects, the creolist account from a decreolized plantation creole.
  • Neo-anglicist evidence from Samana, Nova Scotia, and ex-slave recordings suggests earlier AAE resembled older English more than creoles.
  • Most researchers accept a combined account in which creole influence varied with local plantation demographics.
  • The 1996 Oakland resolution aimed at Standard English proficiency by training teachers to recognize the students' variety.
  • The term Ebonics and the phrase genetically based drove a misreporting cascade that the January 1997 amendment could not undo.
  • The Linguistic Society of America endorsed the Oakland approach as linguistically and pedagogically sound.
  • Verbal deprivation theory was refuted by showing that the testing setting itself suppressed the children's speech.
  • Contrastive analysis, dialect readers, and dialect-sensitive assessment have support; correction during reading and replacement approaches do not.

Sources

  1. Linguistic Society of America. (1997). LSA resolution on the Oakland Ebonics issue. linguisticsociety.org
  2. Wikipedia contributors. (n.d.). Oakland Ebonics controversy. en.wikipedia.org
  3. Charity, A. H., Scarborough, H. S., & Griffin, D. M. (2004). Familiarity with school English in African American children and its relation to early reading achievement. Child Development, 75(5), 1340-1356. doi.org/10.1111/j.1467-8624.2004.00744.x
  4. Fogel, H., & Ehri, L. C. (2000). Teaching elementary students who speak Black English Vernacular to write in Standard English. Contemporary Educational Psychology, 25(2), 212-235. doi.org/10.1006/ceps.1999.1002
  5. Wikipedia contributors. (n.d.). African-American Vernacular English. en.wikipedia.org
Key terms
Anglicist hypothesis
The account deriving AAE principally from the regional British English dialects spoken by settlers in the American South.
Creolist hypothesis
The account deriving AAE from a plantation creole that later decreolized under sustained contact with English.
Gullah
An English-based creole spoken in the Sea Islands of South Carolina and Georgia, demonstrating that creolization occurred on the mainland.
Neo-anglicist evidence
Data from isolated diaspora varieties and ex-slave recordings suggesting earlier AAE resembled older English varieties more than creoles.
Divergence hypothesis
The 1980s claim that urban AAE was moving further from local white vernaculars under conditions of residential segregation.
Verbal deprivation theory
The discredited 1960s claim that poor and minority children heard little language and could not reason with it, refuted by attention to testing conditions.
Contrastive analysis
An instructional approach that explicitly compares the home variety with the school variety so students can shift deliberately between them.
Additive approach
Teaching the standard variety as an addition to a student's home variety rather than as a replacement or correction of it.

Chicano English, Appalachian English, Indian English, and the Global Englishes

  • Explain why Chicano English is a native variety rather than second-language interference, and identify its characteristic features.
  • State the grammatical constraints on Appalachian a-prefixing and evaluate the claim that the region preserves Elizabethan English.
  • Apply Kachru's three circles model to world Englishes, and describe what English as a lingua franca research has found.

The big picture

The AAE case is the most studied, but it is not unique. The same pattern recurs wherever a variety departs from a standard: the departures are systematic, the speakers are judged as deficient, and the judgment is made by people who have not looked at the system. This lesson runs the pattern past three more varieties and then scales the question up to English worldwide.

Scaling it up changes the arithmetic in a way worth stating at the outset. Speakers of English as an additional language now substantially outnumber those who learned it first. A great many, perhaps most, English conversations happening right now involve no one for whom English is a first language. Once that is true, the question of whose English counts as correct is no longer a domestic dispute inside a few countries.

Key idea: Systematic varieties are misread as deficient wherever they occur, and because English now has far more additional-language than first-language speakers, the question of whose norms govern it has become a global one.

Chicano English

Chicano English is a variety of English spoken principally in Mexican American communities in the American Southwest and beyond. The single most important fact about it is the one most often missed: it is a native variety of English, and many of its speakers are English monolinguals who do not speak Spanish at all.

This matters because Chicano English is routinely misidentified as the accented English of a Spanish speaker still learning the language. Carmen Fought's fieldwork in Los Angeles documented the confusion and its cost. Students born in the United States, speaking English as their only language, have been placed in English as a second language classes on the basis of their accent. The variety did originate historically in contact with Spanish, in the same way that many varieties originate in contact, but it is now transmitted from one generation of English speakers to the next independently of any Spanish knowledge.

Its features are consistent. Rhythmically it is more syllable-timed than most other American varieties, giving syllables more equal duration rather than compressing unstressed ones. Related to this, unstressed vowels reduce less, so a word like together retains a fuller first vowel rather than reducing it to a schwa. Intonation contours differ, with characteristic rise-fall patterns that speakers of other varieties sometimes misread as emphasis or attitude. There are consonantal patterns too, including devoicing of some final consonants. In vocabulary and grammar, barely can mean just recently, and tell appears in reported questions where other varieties use ask.

Fought's most theoretically interesting finding was that Chicano English participates in local sound changes like any other variety. The fronting of the vowel in words like dude, a change spreading across California, appeared in her Chicano English data and patterned by social group within the community, with gang-affiliated and non-affiliated speakers differing. A variety that participates in regional sound change and stratifies internally by social group is behaving exactly as a full, autonomous variety behaves.

Key idea: Chicano English is a native English variety with syllable-timed rhythm, reduced vowel reduction, and distinctive intonation, transmitted independently of Spanish and participating in local sound changes.

Appalachian English and the a-prefix

Appalachian English supplies the single most elegant classroom demonstration of rule-governedness in a stigmatized variety, and it is worth working through carefully.

The feature is a-prefixing: a prefix pronounced like a schwa attached to a verb, as in he was a-huntin or she came a-runnin. Outsiders hear this as decoration or archaism, added at random for flavor. Walt Wolfram and Donna Christian showed it is nothing of the kind. It obeys at least three constraints, and speakers apply all of them without ever having been taught any.

First, it attaches only to verbs in progressive form, not to gerunds or nouns that happen to end the same way. You can say he was a-huntin, but the starred form a-huntin is dangerous, using the same word as the subject noun of a sentence, is impossible.

Second, it cannot follow a preposition. He makes money by building houses cannot become the starred form he makes money by a-buildin houses, even though the verb is in the right form.

Third, and most striking, it attaches only to verbs stressed on the first syllable. A-follerin is fine because follering is stressed on its first syllable. The starred a-discoverin is not, because discovering is stressed on its second. A phonological condition on a syntactic process is not the kind of thing anyone invents casually, and it is exactly the kind of thing a grammar contains.

ExampleGrammatical?Constraint at work
He was a-huntinYesProgressive verb, first-syllable stress, no preposition
A-huntin is dangerousNoThe form is a subject noun, not a progressive verb
He makes money by a-buildin housesNoFollows a preposition
She was a-discoverin the truthNoStress falls on the second syllable of the verb
She was a-follerin meYesAll three conditions met

Appalachian English has other well documented features: was-leveling, so we was rather than we were; double modals such as might could and used to could, which express degrees of possibility that a single modal cannot; completive done; the second person plural forms you-uns and y'all; right as an intensifier, as in right good; and retentions such as holler for a small valley and poke for a bag.

That last category feeds a persistent myth worth confronting directly. It is often claimed that Appalachian speech is preserved Elizabethan or Shakespearean English, a pocket of the sixteenth century surviving in the mountains. It is not. The variety does retain some forms that other varieties lost, as every variety retains some things and loses others. But it has also innovated continuously for four centuries, and double modals and a-prefixing in their current form are not Elizabethan. No living variety is a museum. The romantic version of the myth is well intentioned, often deployed to defend the region, and it still gets the linguistics wrong, because it defends the variety on the grounds that it is really an old prestige variety rather than on the grounds that it is a functioning modern system.

Key idea: Appalachian a-prefixing obeys syntactic, positional, and stress constraints that speakers apply without instruction, and the claim that the region preserves Elizabethan English is a myth, since the variety has innovated continuously.

Indian English

India has one of the largest English-speaking populations of any country. Indian English is not a stage on the way to British English; it is a nativized variety with its own conventions, its own literature, and its own internal variation across a country with many first languages.

Phonologically, the stop consonants in words like tea and dare are often produced with the tongue curled back in a retroflex articulation, matching the consonant inventories of many Indian languages. The rhythm, as in Chicano English, is more syllable-timed than British or American English.

Grammatically, several patterns are stable norms rather than errors. The invariant tag isn't it can attach to any clause, as in you are coming tomorrow, isn't it, in the way that French uses n'est-ce pas and German nicht wahr. Progressive forms occur with stative verbs, as in I am knowing the answer. Article usage differs systematically from British norms. And the lexicon includes both borrowings and native coinages, among them lakh and crore for units of a hundred thousand and ten million, cousin brother and cousin sister to specify a relationship English otherwise leaves vague, and prepone, a neatly formed antonym of postpone that English arguably should have had all along.

The analytic point is the same as with the other varieties. A feature used consistently by a large speech community, transmitted to children, and used in that community's newspapers and novels is a norm of that variety. Calling it an error requires a prior decision that another community's norms govern, and that decision is political rather than linguistic.

Key idea: Indian English is a nativized variety whose retroflex consonants, invariant tag, stative progressives, and distinctive lexicon are conventions of the variety rather than deviations from British English.

Kachru's three circles

Braj Kachru proposed a model that organizes the global picture, and it remains the standard starting point.

The Inner Circle comprises countries where English is the historical first language of most of the population: the United Kingdom, the United States, Canada, Australia, New Zealand, and Ireland. Kachru described these as norm-providing.

The Outer Circle comprises countries where English was institutionalized through colonial history and now functions in government, education, law, and media alongside other languages: India, Nigeria, Kenya, Singapore, the Philippines, Malaysia, and many others. Kachru called these norm-developing, since local standards were emerging.

The Expanding Circle comprises countries where English is widely learned as a foreign language without official internal functions: China, Japan, Brazil, Russia, much of Europe. Kachru called these norm-dependent.

CircleExamplesRole of EnglishRelation to norms
InnerUK, US, Australia, IrelandHistorical first language of most speakersNorm-providing
OuterIndia, Nigeria, Singapore, PhilippinesInstitutionalized second language from colonial historyNorm-developing
ExpandingChina, Japan, Brazil, GermanyForeign language without internal official functionsNorm-dependent

The model has been productively criticized. The boundaries blur badly: many Outer Circle speakers acquire English from birth and are native speakers by any reasonable definition, while Expanding Circle countries increasingly use English internally in business and higher education. The model also builds in the assumption it was partly meant to challenge, since calling Inner Circle countries norm-providing concedes that the norms belong to them. Kachru himself argued forcefully that Outer Circle varieties should be judged by their own standards, and the terminology sits awkwardly with that argument.

Key idea: Kachru's Inner, Outer, and Expanding circles organize the global spread of English by historical role, though the boundaries blur and the labels themselves concede norm authority to the Inner Circle.

English as a lingua franca

A newer research strand takes the demographic fact seriously. If most English interactions occur between people who each learned it as an additional language, then the relevant question is not how closely they approximate a British or American model but whether they understand one another.

Jennifer Jenkins and Barbara Seidlhofer developed this into the study of English as a lingua franca. Jenkins's work on pronunciation asked which features actually carry intelligibility in such interactions and which do not, proposing what she called a lingua franca core. Some findings were counterintuitive. The th sounds, drilled relentlessly in classrooms worldwide, turn out to matter little for mutual intelligibility among non-native speakers, while consonant cluster accuracy and vowel length distinctions matter a great deal. On this evidence, a great deal of pronunciation teaching optimizes for the wrong target.

Singapore illustrates the social side of the same question. Singapore Colloquial English, widely known as Singlish, is a contact variety with substrate influence from Hokkien, Malay, and Tamil, featuring discourse particles like lah and lor, topic-prominent structure, and copula patterns. Singaporeans generally command both Singlish and a standard variety and switch between them by situation, a functional split that anticipates the concept of diglossia in the next module. The government's Speak Good English Movement, launched in 2000, has campaigned against Singlish on the grounds that it harms international intelligibility, while many Singaporeans treat it as a marker of national identity. The argument is a live one and shows that language policy debates are rarely only about communication.

Nigeria offers another configuration, with Nigerian English as an established standard variety used in education and media, alongside Nigerian Pidgin, a separate contact language with tens of millions of speakers that Module 5 takes up in its own right.

Key idea: English as a lingua franca research asks what actually supports intelligibility among additional-language speakers, finding that some heavily taught features matter little while others matter greatly.

Common misconceptions

  • Chicano English is the accent of someone still learning English. Many of its speakers are English monolinguals, and the variety is transmitted natively regardless of Spanish knowledge.
  • A-prefixing is random decoration. It requires a progressive verb, forbids a preceding preposition, and demands first-syllable stress.
  • Appalachian English preserves Elizabethan English. It retains some old forms and has innovated for four centuries; no living variety is a museum.
  • Indian English features are errors relative to British English. They are stable community norms, used in that community's education, journalism, and literature.
  • Pronunciation teaching should target native-speaker accuracy on every sound. Lingua franca research finds some heavily drilled features contribute little to actual intelligibility.

Recap

  • Chicano English is a native variety with syllable-timed rhythm and distinctive intonation, misidentified as learner English at real cost to students.
  • It participates in regional sound changes and stratifies internally, behaving as a full autonomous variety.
  • Appalachian a-prefixing obeys progressive-form, preposition, and first-syllable stress constraints applied without instruction.
  • The Elizabethan English myth is false, since Appalachian English has both retained and innovated continuously.
  • Indian English has stable norms including retroflex stops, invariant isn't it, stative progressives, and its own lexicon.
  • Kachru's three circles organize English worldwide as norm-providing, norm-developing, and norm-dependent, with boundaries that blur.
  • English as a lingua franca research identifies which features actually support intelligibility among additional-language speakers.
  • Singapore's Singlish debate shows that language policy arguments involve identity as much as communication.

Sources

  1. Wikipedia contributors. (n.d.). Chicano English. en.wikipedia.org
  2. Wolfram, W., et al. (n.d.). Language and Life Project. North Carolina State University. languageandlife.org
  3. Wikipedia contributors. (n.d.). Appalachian English. en.wikipedia.org
  4. Wikipedia contributors. (n.d.). World Englishes. en.wikipedia.org
  5. Britannica. (n.d.). English language. Encyclopaedia Britannica. britannica.com
Key terms
Chicano English
A native variety of English spoken in Mexican American communities, transmitted independently of Spanish and often misidentified as learner English.
Syllable-timed rhythm
A rhythmic pattern giving syllables roughly equal duration, contrasting with the stress-timed compression of unstressed syllables.
A-prefixing
The Appalachian construction attaching a schwa prefix to progressive verbs, constrained by verb form, preposition, and first-syllable stress.
Double modal
A construction combining two modal verbs, as in might could, expressing degrees of possibility a single modal cannot.
Nativized variety
A variety of a language that has developed stable local norms within a community, transmitted to children and used in local institutions.
Three circles model
Kachru's classification of world English into Inner, Outer, and Expanding circles by historical role and relation to norms.
English as a lingua franca
English used between speakers who each acquired it as an additional language, studied for what actually supports mutual intelligibility.
Singlish
Singapore Colloquial English, a contact variety with Hokkien, Malay, and Tamil influence, used alongside a standard variety by situation.

Module 5: Multilingual Societies

What happens when languages share a speaker, a household, or a country: bilingualism and the grammar of code-switching, the functional division of diglossia, the birth of pidgins and creoles, and the shift, loss, revival, and policy that decide which languages survive.

Bilingualism, Code-Switching, and Diglossia

  • Describe individual and societal bilingualism and explain why the balanced bilingual is the wrong benchmark.
  • State the structural constraints on code-switching and explain why switching requires greater rather than lesser competence.
  • Define diglossia, distinguish Ferguson's original formulation from Fishman's extension, and apply both to real cases.

The big picture

Begin by correcting a default assumption. If you grew up in a largely monolingual country, it is natural to treat one person with one language as the normal case and bilingualism as an interesting complication. Globally, the reverse is true. More than half the world's population uses two or more languages routinely, and in much of Africa, South Asia, and Southeast Asia, three or four is ordinary. Monolingualism is the special case that needs explaining.

This lesson covers what happens inside bilingual speakers and inside bilingual societies. Two findings drive it. The first is that code-switching, the alternation between languages within a conversation, has its own grammar, so tight that it can only be produced by someone with a strong command of both systems. The second is that societies frequently assign different varieties to different domains in a stable arrangement that can persist for centuries.

Key idea: Bilingualism is the global norm rather than the exception, and both the individual practice of code-switching and the societal arrangement of diglossia turn out to be highly structured.

What a bilingual is

Ask most people to define a bilingual and you will get something like a person who speaks two languages perfectly, equally well, like two monolinguals in one head. Almost nobody meets that definition, and Francois Grosjean argued that it was never the right one.

His complementarity principle states the alternative. Bilinguals acquire and use their languages for different purposes, with different people, in different domains. A woman might handle her family life, cooking, and childhood memories in one language and her engineering work, tax paperwork, and university education in another. She will lack technical vocabulary in the first and kinship and cooking vocabulary in the second. This is not partial competence in either language. It is the expected outcome of using languages for different things, and measuring each of her languages against a monolingual who does everything in that language is measuring against a benchmark that her life never required.

Several distinctions organize the field. Simultaneous bilinguals acquire both languages from infancy; sequential bilinguals add the second later. Additive bilingualism describes a situation where the second language is added while the first is maintained and valued, which is typical when the first language has social prestige. Subtractive bilingualism, Wallace Lambert's term, describes a situation where the second language displaces the first, typical when the first language is stigmatized and school is conducted only in the second. The educational outcomes of the two situations differ substantially, which makes the distinction more than a taxonomy.

Several persistent myths deserve direct refutation. Raising a child with two languages does not cause confusion or language delay; bilingual children hit milestones on a normal schedule, and when their vocabularies in both languages are counted together, the totals are comparable to monolingual peers. A child who mixes languages is not confused; children as young as two adjust which language they use according to who they are speaking with. And there is no critical shortage of room in the mind. The bilingual advantage in executive function, once widely claimed, has been substantially disputed in recent replication work, so the honest summary is that bilingualism carries clear social and communicative benefits while claims about general cognitive enhancement remain contested.

Key idea: Bilinguals use their languages for complementary purposes rather than duplicating a monolingual's competence twice, and bilingual acquisition causes neither confusion nor delay.

Code-switching has a grammar

Now the central structural result. Code-switching is the alternation between two languages within a single conversation, often within a single sentence. It is widely described by outsiders, and sometimes by speakers themselves, as sloppiness, laziness, or a symptom of not knowing either language properly. The data say something close to the opposite.

First distinguish it from neighbors. Borrowing is the incorporation of a word from one language into another, integrated into its phonology and grammar, as English did with tortilla and karaoke; a monolingual can use a borrowing. Interference is the intrusion of one system into another in ways the speaker cannot control. Code-switching is neither: it is fluent, intentional alternation between two systems both of which the speaker commands.

Shana Poplack's 1980 study of Puerto Rican speakers in New York established the typology and the constraints. Switches come in three types. Tag-switching inserts a tag or interjection from the other language. Intersentential switching occurs at a sentence boundary. Intrasentential switching occurs within a clause and is the type that requires the most competence in both grammars simultaneously.

Poplack proposed two constraints. The equivalence constraint says switches occur at points where the surface structures of the two languages correspond, so the switch does not violate the word order rules of either. The free morpheme constraint says a switch cannot occur between a bound morpheme and a lexical item unless that item is phonologically integrated into the first language: an English stem with a Spanish inflectional ending attached is not a possible switch, while a fully assimilated borrowing can take the ending normally.

Carol Myers-Scotton's Matrix Language Frame model offers a different and influential account. In it, one language serves as the matrix, supplying the grammatical frame and the system morphemes such as inflections and function words, while the other, the embedded language, contributes content items into that frame. The two models make different predictions and the debate between them is ongoing, but they agree on the point that matters here: switching is constrained, not free.

The evidence that this is real knowledge rather than analyst's tidiness comes from who can do it. Poplack's central observation was that the most fluent bilinguals produced the most intrasentential switching, while less proficient speakers switched mainly at tags and sentence boundaries, the structurally safest points. If switching were a symptom of deficiency, the relationship would run the other way. Bilinguals also reject constructed switch sentences that violate the constraints, giving the same kind of grammaticality judgments speakers give within a single language.

PhenomenonWhat it isWho does it
BorrowingA word integrated into the receiving language's phonology and grammarMonolinguals and bilinguals alike
InterferenceUncontrolled intrusion of one system into anotherTypically less proficient speakers
Tag-switchingAn interjection or tag from the other languageBilinguals at all proficiency levels
Intersentential switchingA switch at a sentence boundaryBilinguals at all proficiency levels
Intrasentential switchingA switch inside a clause, obeying both grammarsPredominantly the most fluent bilinguals

Key idea: Code-switching obeys structural constraints and the intrasentential type is produced predominantly by the most fluent bilinguals, so switching signals greater command of both languages rather than less.

Why speakers switch

Structure is only half the story. John Gumperz distinguished situational switching, driven by a change in setting, participants, or topic, from metaphorical switching, in which a speaker switches within an unchanged situation in order to invoke the associations of the other language: shifting to the home language to signal intimacy, or to the official language to signal authority.

Myers-Scotton's markedness model develops this. In any given interaction there is an unmarked, expected choice of language given the participants and setting. A speaker who uses it is confirming the ordinary relationship. A marked choice, using the unexpected language, does social work: it can create distance, claim authority, express solidarity, exclude a bystander, or signal that the relationship is being renegotiated. Switching back and forth can itself be the unmarked choice in communities where mixed speech is the everyday norm.

This is worth connecting to the stigma. Mixed varieties often attract dismissive labels, and speakers themselves sometimes apologize for speaking, say, Spanglish. Poplack's constraints are the answer to that: a practice that requires simultaneous command of two grammars, that fluent speakers do more than novices, and that carries fine-grained social meaning is not a failure to speak either language.

Key idea: Speakers switch situationally with changes in setting and metaphorically to invoke a language's associations, with marked choices doing identifiable social work.

Diglossia

Move now from the individual to the society. Charles Ferguson described in 1959 a stable arrangement he called diglossia, in which two varieties of the same language coexist with a strict functional division.

The High variety, conventionally H, is used for formal writing, sermons, university lectures, news broadcasts, and political speeches. It is learned through formal education, is highly codified, usually carries a prestigious literary heritage, and, crucially, has no native speakers: nobody acquires H at home. The Low variety, L, is used for conversation, family life, folk literature, and informal broadcasting. It is acquired natively and is often regarded by its own speakers as having no grammar at all.

Ferguson's four defining cases remain the standard illustrations. In the Arabic-speaking world, Modern Standard Arabic serves as H across many countries while regional colloquial varieties serve as L. In German-speaking Switzerland, Standard German is H and Swiss German dialects are L. In Greece, katharevousa, an archaizing written form, served as H against dimotiki, the spoken language, until the arrangement was ended by policy in 1976. In Haiti, French served as H against Haitian Creole as L, though Haitian Creole has since gained official status.

DomainVariety used
Sermon, political speech, university lectureHigh
Newspaper editorial, formal correspondenceHigh
Instructions to servants, waiters, workersLow
Conversation with family and friendsLow
Folk literature, popular songLow

Joshua Fishman extended the concept in 1967 to cover functional divisions between entirely distinct languages, not just varieties of one. In Paraguay, Spanish and Guarani divide domains in this way, with Guarani carrying intimacy and national identity and Spanish carrying official and educational functions. Fishman also crossed diglossia with bilingualism to produce a four-way grid: societies can have both, either one alone, or neither, and each combination describes a real configuration.

Criticism of the model has been productive. Real situations leak: H forms appear in casual talk, L forms appear in writing, and the neat domain table describes an idealization. Diglossia is also less stable than Ferguson suggested, as the Greek and Haitian cases show. And the framing can obscure inequality, since a population fluent only in L is effectively excluded from law, government, and higher education, which is a political fact the tidy functional description can make invisible.

Key idea: Diglossia is a stable functional division between a High variety learned in school with no native speakers and a Low variety acquired at home, extended by Fishman to distinct languages and criticized for idealizing leaky, unequal realities.

Common misconceptions

  • A real bilingual speaks both languages perfectly and equally. Bilinguals use their languages for complementary purposes, so uneven domain coverage is the normal outcome, not a deficiency.
  • Bilingualism confuses children and delays speech. Milestones are reached on a normal schedule, and combined vocabulary totals are comparable to those of monolinguals.
  • Code-switching means not knowing either language properly. Intrasentential switching is produced predominantly by the most fluent bilinguals and obeys the grammars of both languages.
  • Switching happens at random points. Poplack's equivalence and free morpheme constraints and the Matrix Language Frame model all predict where switches can and cannot occur.
  • In diglossia, the High variety is the everyday language of the elite. H has no native speakers at all; everyone acquires L at home and learns H through schooling.

Recap

  • Bilingualism is the global norm, and the balanced bilingual is a benchmark almost no one meets or needs to.
  • Grosjean's complementarity principle explains uneven domain coverage as the expected result of using languages for different purposes.
  • Additive bilingualism maintains the first language while subtractive bilingualism displaces it, with different educational outcomes.
  • Code-switching differs from borrowing and interference and is constrained by equivalence and free morpheme conditions.
  • The Matrix Language Frame model assigns the grammatical frame and system morphemes to one language and content items to the other.
  • The most fluent bilinguals produce the most intrasentential switching, refuting the deficiency account.
  • Gumperz distinguished situational from metaphorical switching, and marked choices carry identifiable social meaning.
  • Ferguson's diglossia divides domains between a schooled High variety with no native speakers and a natively acquired Low variety.

Sources

  1. Ferguson, C. A. (1959). Diglossia. Word, 15(2), 325-340. doi.org/10.1080/00437956.1959.11659702
  2. Poplack, S. (1980). Sometimes I'll start a sentence in Spanish y termino en espanol: Toward a typology of code-switching. Linguistics, 18(7-8), 581-618. doi.org/10.1515/ling.1980.18.7-8.581
  3. Fishman, J. A. (1967). Bilingualism with and without diglossia; diglossia with and without bilingualism. Journal of Social Issues, 23(2), 29-38. doi.org/10.1111/j.1540-4560.1967.tb00573.x
  4. Wikipedia contributors. (n.d.). Code-switching. en.wikipedia.org
  5. Wikipedia contributors. (n.d.). Diglossia. en.wikipedia.org
Key terms
Complementarity principle
Grosjean's observation that bilinguals acquire and use their languages for different purposes and domains, so coverage is naturally uneven.
Additive bilingualism
A situation in which a second language is added while the first is maintained and valued.
Subtractive bilingualism
A situation in which a second language displaces the first, typical where the first language is stigmatized and schooling ignores it.
Code-switching
Fluent alternation between two languages within a conversation or sentence by a speaker who commands both systems.
Equivalence constraint
Poplack's condition that switches occur where the surface structures of both languages correspond, violating neither one's word order.
Free morpheme constraint
Poplack's condition that a switch cannot occur between a bound morpheme and a lexical item unless that item is phonologically integrated.
Matrix Language Frame
Myers-Scotton's model in which one language supplies the grammatical frame and system morphemes while the other contributes content items.
Diglossia
A stable functional division between a High variety learned through schooling with no native speakers and a natively acquired Low variety.

Pidgins, Creoles, and Language Contact

  • Place borrowing, convergence, pidginization, and creolization on a spectrum of contact outcomes.
  • Distinguish a pidgin from a creole structurally and socially, and describe what happens during creolization.
  • Compare bioprogram, substrate, and gradualist accounts of creole origins, and explain the post-creole continuum.

The big picture

Languages in contact do things to each other, and the outcomes range from trivial to spectacular. At the mild end, a language borrows a word. At the extreme end, a genuinely new language comes into existence, sometimes within a single generation, complete with a grammar nobody designed.

That extreme case is the subject of this lesson, and it is one of the few places in linguistics where you can watch language creation rather than infer it. Creoles are also, not coincidentally, among the most stigmatized languages on earth, routinely dismissed as broken versions of the European languages that supplied their vocabulary. As with every other case in this course, the dismissal does not survive looking at the grammar.

Key idea: Contact outcomes range from single borrowings to the creation of entirely new languages, and creoles offer a rare opportunity to observe grammar coming into existence.

The contact spectrum

Start at the mild end. Lexical borrowing is universal; no language has ever been found that refuses it. English is a spectacular case, having taken the majority of its vocabulary from French, Latin, and Greek while remaining a Germanic language in its core grammar and its most frequent words.

Borrowing follows a rough hierarchy of ease. Nouns are borrowed most readily, then verbs and adjectives, then function words such as prepositions and conjunctions, and finally, rarely and only under intense contact, bound morphemes like inflectional endings. A language that has borrowed inflections has been in very deep contact indeed. Calques, or loan translations, borrow structure rather than form, as English did in taking skyscraper into other languages piece by piece.

This is where the myth of the pure language dies. Every well documented language shows contact effects, and the languages usually held up as pure are simply those whose borrowing happened long enough ago to be invisible to speakers.

Deeper contact produces convergence, in which neighboring languages come to resemble one another structurally while keeping separate vocabularies. The Balkan area is the classic case: Albanian, Greek, Romanian, Bulgarian, and Macedonian, which belong to different branches of Indo-European, have converged on features including a postposed definite article, the loss of the infinitive, and a future formed with a verb meaning want. Even more striking is the village of Kupwar in India, studied by John Gumperz and Robert Wilson, where speakers of Marathi, Urdu, and Kannada had lived together for centuries. Their languages had converged so far in syntax that sentences could be translated word for word across all three while the vocabularies remained entirely distinct.

Key idea: Contact effects run from lexical borrowing, which every language does, through structural convergence in which neighboring languages align their grammars while keeping separate vocabularies.

Pidgins

A pidgin arises when groups with no common language must communicate regularly, typically in trade, plantation labor, or maritime contexts. Its defining social property is that it has no native speakers: everyone who uses a pidgin has another first language.

Structurally, a pidgin is reduced. Its lexicon is small and drawn mostly from the socially dominant language present, called the superstrate or lexifier. Its morphology is minimal, with little or no inflection for tense, number, or case. Its syntax is simple, with limited embedding, and it leans heavily on context. Common strategies include reduplication for emphasis or plurality and multi-word paraphrase in place of vocabulary the pidgin lacks.

Historical examples show the range. Chinook Jargon served trade across the Pacific Northwest of North America. Russenorsk, used between Russian and Norwegian fishing communities in the Arctic for over a century, drew roughly equally on both languages, which is unusual. Trade pidgins have arisen repeatedly wherever commerce outran mutual intelligibility.

Pidgins are often described as developing through stages: a jargon, unstable and highly variable; a stabilized pidgin, with agreed norms; and an expanded pidgin, which has grown enough vocabulary and grammar to handle a full range of functions even while still having few or no native speakers.

Key idea: A pidgin is a reduced contact language with no native speakers, drawing its lexicon mainly from a dominant lexifier and minimizing morphology and embedding.

Creoles

A creole is what a pidgin becomes when it acquires native speakers. Children born into a community where the pidgin is the main shared language acquire it as a first language, and in doing so they expand it into a complete one.

What gets added is exactly what a full language needs. Vocabulary multiplies. Grammatical morphology develops, typically by grammaticalization, in which content words become grammatical markers. Word order fixes. Embedding and subordination appear, allowing complex sentences. The result is a language that can do everything any language does, including poetry, legislation, and abstract argument.

Tok Pisin, an official language of Papua New Guinea with millions of speakers, shows the process clearly.

Tok PisinLiteral sourceMeaning
Mi no saveme no (Portuguese saber, to know)I do not know
Haus bilong mihouse belong memy house
Bai mi goby and by me goI will go
Gras bilong fesgrass belong facebeard
Wanpela manone fellow manone man

Look at what these show. The future marker bai descends from the English phrase by and by, a content expression that has been ground down into a grammatical particle: that is grammaticalization, the same process that gave English its own future with will. The possessive bilong is a systematic grammatical device, not a quaint phrase. And gras bilong fes is not a failure to know the word beard; it is productive compounding, the same strategy that gives German and Chinese much of their vocabulary.

Haitian Creole makes the grammatical point even more sharply. It has roughly twelve million speakers, has been an official language of Haiti since 1987, and marks tense, mood, and aspect with a set of preverbal particles in a fixed order: mwen manje for I eat, mwen te manje for I ate, mwen ap manje for I am eating, and combinations that layer these meanings. A language with an ordered system of preverbal tense, mood, and aspect markers is not French with the endings knocked off; it is a different grammar. Other major creoles include Jamaican Patois, Sranan Tongo in Suriname, Papiamentu in the Caribbean Netherlands, Cape Verdean Creole, Mauritian Creole, and Gullah in the American Sea Islands.

Key idea: A creole is a pidgin with native speakers, expanded into a full language through vocabulary growth, grammaticalization of content words into markers, fixed word order, and the development of subordination.

Where does creole structure come from?

Creoles arising in widely separated parts of the world share structural features to a striking degree, notably preverbal tense, mood, and aspect markers in a consistent order. Explaining that similarity is the central theoretical dispute.

Derek Bickerton's language bioprogram hypothesis proposed the boldest answer. On his account, children presented with a variable, impoverished pidgin cannot acquire a grammar from it because the input does not contain one, so they supply structure from an innate language capacity. The similarities across unrelated creoles reflect that shared human endowment. His key evidence came from Hawaii, where he argued that the first locally born generation produced a creole with grammatical features absent from the pidgin their parents spoke.

Substrate accounts argue instead that the similarities come from the languages the original speakers already had. Many Atlantic creoles arose in communities dominated by speakers of West African languages, which have their own preverbal aspect systems, so the resemblance may reflect shared substrate rather than innate structure.

Gradualist accounts, associated with Salikoko Mufwene and Robert Chaudenson, question the abrupt scenario entirely. They argue that creoles developed over generations from the colonial varieties of the lexifier, which were themselves nonstandard and variable, and that features were selected from a feature pool contributed by all the languages present. On this view there was no single moment of creation.

Michel DeGraff added a critical dimension by attacking what he called creole exceptionalism: the assumption that creoles form a structurally distinct class of languages, definable by their simplicity or their birth. He argues that no reliable structural criterion separates creoles from other languages, that the class was defined originally on social and racial grounds rather than linguistic ones, and that treating creoles as a special simple type continues to license their stigmatization. That is not a fringe complaint; it has substantially reshaped how the field frames these questions.

Key idea: Bioprogram, substrate, and gradualist accounts each explain creole similarities differently, and creole exceptionalism has been challenged on the grounds that no structural criterion reliably sets creoles apart.

The post-creole continuum

Where a creole remains in contact with its lexifier, especially when that lexifier is the language of schooling and government, speakers develop a range of varieties between the two. This is the post-creole continuum, and Jamaica is the standard illustration.

At one end sits the basilect, the variety furthest from the standard. At the other sits the acrolect, closest to standard English. In between lies a graded series of mesolects. A Jamaican speaker may command a wide stretch of this range and move along it according to situation and interlocutor, which makes the continuum a style resource as well as a structural fact.

Two points make this more than a local curiosity. First, decreolization, the drift of a creole toward its lexifier, has been proposed as part of the history of African American English, connecting Module 4's origins debate to this one. Second, the continuum makes counting difficult: asking how many people speak Jamaican Creole assumes a boundary that the data do not contain.

Meanwhile the social standing of creoles is changing in places. Haitian Creole gained official status in 1987 and is used in education and government. Nigerian Pidgin, with tens of millions of speakers, gained a major international news service in 2017 when the BBC launched a Pidgin-language operation, a decision that treated it as a language people conduct serious business in rather than a joke.

Key idea: Continued contact with a lexifier produces a post-creole continuum from basilect through mesolects to acrolect, along which speakers shift by situation, complicating any attempt to count speakers.

Common misconceptions

  • Creoles are broken or simplified versions of European languages. They have their own grammars, including ordered preverbal tense, mood, and aspect systems that their lexifiers lack.
  • Pidgins and creoles are the same thing. A pidgin has no native speakers and is structurally reduced; a creole has native speakers and is a full language.
  • A creole's vocabulary shows it is a dialect of the lexifier. Vocabulary source and grammatical system are independent; a language can take most of its words from one source and its grammar from elsewhere.
  • Compounds like the Tok Pisin word for beard show a lack of vocabulary. Productive compounding is a normal word formation strategy used heavily by German and Chinese.
  • Creole origins are settled. Bioprogram, substrate, and gradualist accounts remain in contention, and creole exceptionalism itself has been challenged.

Recap

  • Contact outcomes run from lexical borrowing through structural convergence to the creation of new languages.
  • Borrowing follows a hierarchy from nouns to bound morphemes, and no language is free of contact effects.
  • The Balkan area and the village of Kupwar show grammars converging while vocabularies stay separate.
  • A pidgin has no native speakers, a reduced lexicon from a lexifier, minimal morphology, and limited embedding.
  • A creole is a pidgin with native speakers, expanded by grammaticalization, fixed word order, and subordination.
  • Tok Pisin and Haitian Creole illustrate grammaticalized markers such as bai and the ordered preverbal particles te and ap.
  • Bioprogram, substrate, and gradualist accounts compete to explain cross-creole structural similarities.
  • The post-creole continuum runs from basilect through mesolects to acrolect and is used as a stylistic resource.

Sources

  1. Bickerton, D. (1984). The language bioprogram hypothesis. Behavioral and Brain Sciences, 7(2), 173-188. doi.org/10.1017/S0140525X00044149
  2. DeGraff, M. (2005). Linguists' most dangerous myth: The fallacy of Creole Exceptionalism. Language in Society, 34(4), 533-591. doi.org/10.1017/S0047404505050207
  3. Britannica. (n.d.). Creole languages. Encyclopaedia Britannica. britannica.com
  4. Britannica. (n.d.). Pidgin. Encyclopaedia Britannica. britannica.com
  5. Wikipedia contributors. (n.d.). Tok Pisin. en.wikipedia.org
Key terms
Lexifier
The socially dominant language that supplies most of the vocabulary of a pidgin or creole, also called the superstrate.
Pidgin
A reduced contact language with no native speakers, arising where groups without a common language must communicate regularly.
Creole
A language that developed from a pidgin by acquiring native speakers and expanding into a full linguistic system.
Grammaticalization
The process by which content words become grammatical markers, as when Tok Pisin bai developed from the phrase by and by.
Sprachbund
A linguistic area in which unrelated or distantly related neighboring languages converge structurally while keeping separate vocabularies.
Language bioprogram hypothesis
Bickerton's proposal that children create creole grammar from innate capacity when the pidgin input lacks a consistent grammar.
Creole exceptionalism
The assumption, challenged by DeGraff, that creoles form a structurally distinct and simpler class of languages.
Post-creole continuum
The graded range from basilect through mesolects to acrolect that develops when a creole stays in contact with its lexifier.

Language Shift, Endangerment, Revitalization, and Policy

  • Explain the mechanism of language shift through domain loss and the three-generation pattern.
  • Apply vitality frameworks such as EGIDS and the UNESCO scale, and identify intergenerational transmission as the decisive factor.
  • Compare the Hawaiian, Māori, and Welsh revitalization efforts, and distinguish status, corpus, and acquisition planning.

The big picture

There are roughly seven thousand languages spoken today. A large share of them, commonly estimated at around forty percent, are considered endangered, and projections that half or more will cease to be spoken within this century are widely cited. Those figures are estimates built on contested definitions, and it is worth holding them loosely. What is not in doubt is the direction: a small number of languages are gaining speakers rapidly while a very large number are losing them.

This lesson explains the mechanism, which is more mundane and more preventable than the word death suggests, and then examines what has actually worked to reverse it. The revitalization cases are among the most encouraging material in this course, and they are encouraging in a specific and instructive way: the successes share a common design feature, and it is not the one most people would guess.

Key idea: Language endangerment is widespread and its mechanism is well understood, and the documented successes in reversing it share an identifiable common structure.

How shift actually happens

Almost nobody decides to stop speaking their language. Shift is a cumulative outcome of many small, individually reasonable choices, and it proceeds through the loss of domains.

The sequence is fairly regular. The heritage language first disappears from work and the market, where the dominant language brings economic advantage. Then it disappears from school, either by policy or by default. Then it retreats from friendship networks as young people's peer groups shift. The home is typically the last domain to go, and when the home goes, transmission stops.

That final step produces the classic three-generation pattern. The grandparents' generation speaks the heritage language, often as their only language or their strongest one. The parents' generation is bilingual, having acquired the heritage language at home and the dominant language at school. The children's generation understands the heritage language but does not speak it, and their children do not understand it. Three generations is enough for a language with thousands of speakers to arrive at a handful.

The pressures driving this are worth naming individually, because they differ in whether policy can address them. Economic opportunity concentrated in the dominant language is the most pervasive. Urbanization breaks up the dense local networks that Module 2 showed maintain vernacular norms. Media and schooling conducted only in the dominant language remove domains directly. Stigma does its own work, and it is often internalized: parents who were punished for speaking a language commonly decline to pass it on, believing they are protecting their children.

And in many cases the pressure was explicit coercion. Residential and boarding school systems in the United States, Canada, and Australia removed Indigenous children from their families and punished them for speaking their languages, over generations and as deliberate policy. In Wales, the practice remembered as the Welsh Not humiliated schoolchildren caught speaking Welsh. These were not side effects of modernization. They were programs, and their linguistic consequences were the intended ones.

A terminological note that matters to the communities involved. Language death, meaning the point at which no speakers remain, is less common than language shift, in which a community moves to another language while continuing to exist. And many communities prefer to call a language with no current speakers sleeping or dormant rather than extinct, because materials and descendants remain and reawakening is possible. Several languages once listed as extinct now have speakers again.

Key idea: Shift proceeds through domain loss ending at the home, producing a three-generation pattern, and it has often been driven by explicit coercive policy rather than by choice.

Measuring vitality

To act on endangerment you need to grade it, and two frameworks dominate.

Joshua Fishman's Graded Intergenerational Disruption Scale, published in 1991, arranged the situation of a threatened language on an eight-stage scale, with the crucial threshold being whether the language is still being transmitted to children at home. Paul Lewis and Gary Simons expanded this into EGIDS, a thirteen-level scale used by Ethnologue to classify every language it catalogs, running from international languages at one end through various levels of vitality and endangerment to dormant and extinct at the other.

UNESCO's framework assesses nine factors, including absolute speaker numbers, the proportion of speakers within the community, domains of use, response to new media, materials for education and literacy, and community attitudes. It produces a six-level classification: safe, vulnerable, definitely endangered, severely endangered, critically endangered, and extinct.

UNESCO levelSituation
SafeSpoken by all generations; transmission uninterrupted
VulnerableMost children speak it, but use may be restricted to certain domains
Definitely endangeredChildren no longer learn it as a mother tongue at home
Severely endangeredSpoken by grandparents; the parent generation may understand but not use it
Critically endangeredThe youngest speakers are grandparents who use it partially and infrequently
ExtinctNo speakers remain

Notice what every framework converges on. Raw speaker numbers matter far less than whether children are acquiring the language at home. A language with fifty thousand speakers, none of them under forty, is in worse condition than a language with two thousand speakers being raised by their parents in it. Intergenerational transmission is the variable that decides.

Key idea: Vitality frameworks such as EGIDS and the UNESCO scale converge on intergenerational transmission rather than speaker numbers as the decisive indicator of a language's condition.

Three revitalization cases

If transmission is the variable, then any effective intervention has to restore it. The best documented cases did exactly that, and by strikingly similar means.

Hawaiian. Hawaiian-medium education was banned in 1896, after the overthrow of the Hawaiian Kingdom, and the language declined steeply thereafter. By the early 1980s fewer than about two thousand speakers remained, nearly all elderly, apart from the community on the island of Niihau. In 1983 a group of educators and parents founded an organization to establish Hawaiian-medium preschools, and the first Punana Leo, meaning language nest, opened in 1984. The model is total immersion from the earliest years, staffed by fluent speakers, with families expected to participate and learn. The 1896 ban was repealed in 1986, and Hawaiian-medium schooling was extended upward through the grades, eventually to university level. The result is a generation of children who acquired Hawaiian as a first language from teachers and then, in the following generation, from parents. Hawaiian had already been named an official state language in 1978, but the constitutional recognition alone would have changed nothing without the preschools.

Māori. The New Zealand effort followed a similar design and partly inspired the Hawaiian one. Kōhanga Reo, also meaning language nests, began in 1982, staffed largely by elders who were fluent speakers and organized by communities rather than by the state. Kura Kaupapa Māori extended Māori-medium education into the school years. The Māori Language Act of 1987 made te reo Māori an official language and created a commission to promote it, and a Māori-language television service launched in 2004. The honest assessment is mixed: transmission was substantially strengthened and the language's public presence transformed, while te reo Māori remains classified as endangered and the proportion of fluent speakers has not returned to earlier levels.

Welsh. Welsh had been declining for over a century when the modern effort took shape. The Welsh Language Act of 1993 and the 2011 Measure established legal standing, requiring public bodies to treat Welsh and English equally in Wales. The television channel S4C launched in 1982. The most consequential element, again, was education: Welsh-medium schools expanded substantially, producing large numbers of speakers who did not acquire the language at home. The Welsh Government has set a target of a million speakers by 2050. Census figures have moved unevenly, and the persistent challenge is the gap between school-acquired ability and everyday use, which is a reminder that producing speakers and producing speech communities are different problems.

One case is often cited that should be handled carefully. Hebrew's revival as an everyday spoken language, after centuries of use principally in liturgy and scholarship, is the most complete case on record. But its conditions were unusual: a nationalist movement with institutional power, a state, and a population of immigrants speaking many different languages who genuinely needed a common tongue. Treating Hebrew as a general model sets an expectation that no other case has been able to meet.

Where a language has few remaining speakers and no children learning it, a different technique applies. The master-apprentice model, developed for California's Indigenous languages by Leanne Hinton and colleagues, pairs one fluent elder with one committed adult learner for many hours a week of immersion in ordinary daily activity, with no translation permitted. It produces small numbers of conversational speakers, which is the realistic goal at that stage.

Key idea: Successful revitalization has centered on immersion transmission to young children, as in the Hawaiian Punana Leo and Māori Kōhanga Reo language nests, with legal recognition supporting but never substituting for it.

What works and what does not

Assembling the record produces a few clear lessons.

Immersion beats instruction. A few hours a week of language classes produces students who know about the language, not speakers of it. Language nests work because young children spend their days in a setting where the language is the only medium.

Community control matters. Programs designed and run by the communities themselves have outperformed programs designed for them, both in sustaining commitment and in matching local realities.

Documentation is not revitalization. Recording a language, writing its grammar, and archiving its texts are all valuable and sometimes urgent, but a documented language with no speakers is still not spoken. Documentation preserves the possibility of reawakening; it does not by itself accomplish it.

Legal status is necessary but insufficient. Official recognition opens funding, schooling, and public presence. Hawaiian was official in 1978 and still declining in 1983. The recognition mattered once there were preschools for it to support.

Key idea: Immersion, community control, and restored transmission drive revitalization, while documentation and legal status enable it without substituting for it.

Language policy and planning

All of this sits inside deliberate policy, and Robert Cooper's three-way division organizes it.

Status planning decides which language is used for what: official languages, languages of instruction, court languages, broadcasting requirements. Corpus planning works on the language itself: devising or reforming an orthography, coining technical vocabulary, publishing dictionaries and grammars, standardizing among competing forms. A revitalization effort typically needs both, since a language returning to schools needs terminology for chemistry and mathematics. Acquisition planning, Cooper's addition, addresses how people will actually learn it: schools, teacher training, adult classes, materials.

National configurations vary enormously. France's constitution names French as the language of the Republic, a strongly monolingual stance. Switzerland has four national languages assigned largely by territory, so the language of public life depends on where you are. South Africa's constitution recognizes multiple official languages, including a sign language added in 2023. India names Hindi and English for central government purposes, schedules a long list of recognized languages, and has promoted a three-language formula in schooling, which has been politically contentious for decades. Canada and Belgium apply territorial principles, while some systems follow a personality principle in which a citizen's rights follow them regardless of location.

The United States long had no official language at the federal level, though many states adopted official English laws and an organized English-only movement pressed the issue for decades. An executive order issued in March 2025 designated English as the official language of the United States, a change whose practical effects will play out over time. Bilingual education has been a parallel battleground: California voters restricted it through Proposition 227 in 1998 and reversed course through Proposition 58 in 2016.

Underlying these debates is a rights framework. The United Nations Declaration on the Rights of Indigenous Peoples affirms rights to revitalize, use, develop, and transmit Indigenous languages and to establish education in them. Language policy is thus rarely only administrative. It decides who can access courts, schooling, and government in the language they actually think in, which is why these arguments generate the heat they do.

Key idea: Status, corpus, and acquisition planning are the three levers of language policy, and national configurations range from constitutional monolingualism to territorial and personality-based multilingual systems.

Common misconceptions

  • Languages die naturally, like species. Shift is driven by economic pressure, domain loss, stigma, and in many cases explicit coercive schooling policy.
  • Speaker numbers indicate how endangered a language is. Intergenerational transmission is decisive; a language with many elderly speakers and no child learners is worse off than a small language being raised in.
  • Recording a language saves it. Documentation preserves the possibility of reawakening but produces no speakers by itself.
  • Official status revives a language. Hawaiian was an official state language while still declining; the preschools changed the trajectory.
  • Hebrew shows any language can be revived the same way. Its conditions, including a state, a nationalist movement, and immigrants needing a common language, have not been replicated.

Recap

  • Shift proceeds by domain loss from work through school and friendship to the home, where transmission finally stops.
  • The three-generation pattern moves a community from monolingual heritage speakers to non-speaking grandchildren.
  • Coercive schooling, including residential school systems and the Welsh Not, was deliberate policy with intended linguistic effects.
  • EGIDS and the UNESCO six-level scale both identify intergenerational transmission as the decisive vitality factor.
  • Hawaiian and Māori revitalization centered on immersion language nests for young children, with legal recognition in support.
  • Welsh expanded through Welsh-medium schooling, leaving a gap between school-acquired ability and everyday use.
  • Documentation and official status enable revitalization but do not accomplish it; immersion and community control do the work.
  • Status, corpus, and acquisition planning are the three levers, applied very differently across national configurations.

Sources

  1. UNESCO. (n.d.). Languages and multilingualism. unesco.org
  2. Eberhard, D. M., Simons, G. F., & Fennig, C. D. (Eds.). (n.d.). Ethnologue: Languages of the world. SIL International. ethnologue.com
  3. Wikipedia contributors. (n.d.). Hawaiian language. en.wikipedia.org
  4. Wikipedia contributors. (n.d.). Maori language. en.wikipedia.org
  5. Welsh Government. (n.d.). Welsh language. gov.wales
Key terms
Language shift
The process by which a community gradually moves from using one language to another, usually through the loss of domains of use.
Domain loss
The retreat of a language from areas of life such as work, school, and friendship, ending with the home.
Three-generation pattern
The common sequence in which grandparents speak the heritage language, parents are bilingual, and children do not speak it.
Intergenerational transmission
The passing of a language to children at home, identified by every vitality framework as the decisive factor in survival.
EGIDS
The Expanded Graded Intergenerational Disruption Scale, a thirteen-level vitality classification developed from Fishman's scale and used by Ethnologue.
Language nest
An immersion early-childhood program staffed by fluent speakers, exemplified by the Hawaiian Punana Leo and Māori Kōhanga Reo.
Master-apprentice model
An intensive one-to-one immersion method pairing a fluent elder with an adult learner, used where few speakers remain.
Status planning
Policy decisions about which language is used for which official and public functions.
Corpus planning
Work on the language itself, including orthography, terminology, dictionaries, and standardization among competing forms.

Module 6: Language and Power in Daily Life

Sociolinguistics at the scale of seconds and of consequences: how conversation is organized, how politeness and address negotiate face and power, and what the audit and courtroom evidence shows about the real cost of speaking a stigmatized variety.

Talk in Interaction: Politeness, Face, Turn-Taking, and Address

  • Describe the turn-taking system, adjacency pairs, preference organization, and repair as analyzed in conversation analysis.
  • Apply Brown and Levinson's face model and its five strategies, and evaluate the critiques of its universality.
  • Explain how T and V pronouns and address terms encode power and solidarity, and how contextualization cues produce cross-cultural miscommunication.

The big picture

Everything so far has worked at the scale of communities and decades. This lesson works at the scale of seconds. It turns out that ordinary conversation, the most familiar thing humans do, is organized by machinery so precise that its timing beats human reaction time, and that nobody can describe from introspection.

Two research traditions supply the tools. Conversation analysis examines recorded talk in fine detail to find the structures participants themselves orient to. Politeness theory asks how speakers manage the constant risk that talk poses to people's sense of themselves. Together they explain a great deal of what happens when interaction goes well, and rather more about what happens when it goes badly across cultural lines.

Key idea: Ordinary conversation is governed by precise structural machinery and by continuous management of participants' public self-image, neither of which speakers can describe from introspection.

Turn-taking

Start with a fact that should be surprising. In conversation, one party talks at a time, speaker change recurs, and the gap between turns is typically around two hundred milliseconds. Overlap happens but is brief and quickly resolved.

Two hundred milliseconds is the problem. Planning even a short utterance takes considerably longer than that, and simple reaction time to an unexpected stimulus is roughly in the same range. So speakers cannot be waiting for silence and then beginning to plan. They must be projecting the end of the current turn while it is still running, planning their own turn during it, and launching at the projected completion point. The smoothness of conversation is a prediction problem, solved constantly and unconsciously.

Harvey Sacks, Emanuel Schegloff, and Gail Jefferson described the system in 1974. Turns are built from turn-constructional units, which may be a word, a phrase, a clause, or a sentence, and whose completion is projectable from grammar, intonation, and content together. At the end of each unit comes a transition-relevance place, where speaker change becomes legitimate. There the rules apply in order: if the current speaker has selected a next speaker, that party speaks; if not, any party may self-select, with the first starter getting the turn; if neither happens, the current speaker may continue.

Tanya Stivers and colleagues tested whether this is a cultural artifact by examining conversation across ten languages spanning five continents, including languages of small communities with no relation to one another. They found the same basic pattern everywhere, with modest cultural variation in the average gap length around a common design. The turn-taking system looks like a human universal with local settings.

Key idea: Turn transitions average around two hundred milliseconds, too fast for reactive planning, so speakers project turn completion in advance, and the system appears across unrelated languages worldwide.

Adjacency pairs and preference

Turns are not independent. Many come in pairs whose first part makes a particular second part expected: question and answer, greeting and greeting, invitation and acceptance or refusal, summons and response.

The key property is conditional relevance. Once a first pair part is produced, the second is not merely likely but accountably due, and its absence is itself meaningful. If you ask a question and get silence, you do not conclude that nothing happened. You conclude that something happened, and you begin to interpret it: they did not hear, they are offended, they are thinking. Absence is noticeable only because the structure made something relevant.

Second parts are not equal. Preference organization describes a structural asymmetry, and the word preference here has nothing to do with what anyone wants. Preferred responses, such as accepting an invitation or agreeing with an assessment, come quickly and plainly. Dispreferred responses, such as refusing or disagreeing, are systematically packaged: they are delayed by a pause, prefaced with a hesitation marker such as well, softened with a token appreciation, and accompanied by an account.

So a refusal rarely sounds like no. It sounds like a pause, then something along the lines of: well, I would love to, but I have got my sister staying that weekend. Every element of that is doing structural work. The consequence is that a competent participant hears the pause and knows the refusal is coming before any content arrives, which is why speakers so often withdraw an invitation during the pause itself.

Key idea: Adjacency pairs make a second part conditionally relevant so that its absence is meaningful, and dispreferred responses are systematically delayed, prefaced, softened, and accounted for.

Repair

Conversation constantly goes slightly wrong and is constantly fixed, almost always without anyone noticing. Repair is the organized set of practices for handling trouble in speaking, hearing, or understanding.

Schegloff, Jefferson, and Sacks showed that repair is ordered by a clear preference. Self-initiated self-repair, where the speaker notices and fixes their own trouble mid-turn, is the most common and the most preferred. Next comes self-initiated repair completed by the other party. Then other-initiated self-repair, where the recipient signals trouble and the original speaker fixes it. Least preferred is other-initiated other-repair, in which the recipient both flags the problem and corrects it, which is why straightforwardly correcting someone feels socially loaded.

Other-initiation has its own beautiful result. Mark Dingemanse, Francisco Torreira, and Nick Enfield examined the item English writes as huh, used to signal that a turn was not caught, across a large sample of unrelated languages. They found a strikingly similar form nearly everywhere: a short, syllabic, questioning-intoned vocalization with a similar vowel quality. Since unrelated languages do not usually agree on the form of a word, they argued this is a case of convergent evolution driven by the interactional job it does. The pressure to signal trouble instantly, with minimal articulation, in a form that clearly requests a repeat, shapes the item toward the same solution everywhere.

Key idea: Repair is ordered from most preferred self-initiated self-repair to least preferred other-initiated other-repair, and the trouble-signaling item resembling huh has converged on a similar form across unrelated languages.

Face and politeness

Now the other tradition. Erving Goffman introduced face as the positive public self-image a person claims in interaction, and face-work as the effort participants put into maintaining it, including each other's. Interaction is cooperative on this front: people generally protect one another's face, because a scene in which someone is humiliated is uncomfortable for everyone in it.

Penelope Brown and Stephen Levinson developed this into a systematic model. They split face in two. Positive face is the desire to be liked, approved of, and included. Negative face is the desire to be unimpeded, autonomous, and not imposed upon. Many ordinary acts threaten one or the other. A request threatens negative face by imposing. A criticism threatens positive face by withholding approval. Brown and Levinson call these face-threatening acts.

Faced with one, a speaker chooses among five strategies, ordered by how much redress they offer.

StrategyExample, asking to borrow a laptopWhat it does
Bald on recordGive me your laptop.No redress; used for urgency, high power, or great intimacy
Positive politenessHey, you are always so generous. Lend me your laptop?Appeals to connection and approval
Negative politenessI am sorry to bother you, but could I possibly borrow your laptop?Minimizes imposition and offers an exit
Off recordMy laptop just died and I have a deadline tonight.Hints, leaving deniability for both parties
Do not do the actSay nothing.Avoids the threat entirely

Which strategy a speaker picks is predicted by three factors combined: the social distance between the parties, the power difference between them, and how large the imposition is ranked in that culture. Increase any one and speakers move down the table toward more redress. This makes testable predictions, and it broadly holds.

The universality claim has drawn substantial criticism, and the criticism is instructive. Sachiko Ide argued that Japanese politeness is driven substantially by discernment, the obligation to use forms appropriate to one's position in a social structure, rather than by strategic calculation about face threats. Yoshiko Matsumoto argued that Japanese face is fundamentally relational, concerned with one's place in a group, so a model built on individual autonomy misdescribes it: acknowledging dependence can be face-enhancing rather than face-threatening. Similar arguments have been made about Chinese politeness. The model also carries a built-in bias in treating negative politeness as the more polite option, which reflects a particular cultural weighting of autonomy. The reasonable summary is that face and face-work are plausibly universal, while the specific two-part division and the strategic calculus are shaped by culture.

Key idea: Brown and Levinson model politeness as redress for threats to positive and negative face, scaled by distance, power, and imposition, and critics argue the model's individualist framing does not transfer to societies where face is relational.

Address terms: power and solidarity

Address terms encode social relationships directly, and Roger Brown and Albert Gilman's 1960 study of European pronouns remains the classic analysis.

Many languages have two second-person forms: French tu and vous, German du and Sie, Spanish tu and usted, Russian ty and vy. Brown and Gilman labeled the familiar form T and the formal form V, and traced a historical shift. In the older system, usage was largely asymmetric and encoded power: the superior gave T and received V, so a noble addressed a servant with T and was addressed with V. Over time the system shifted toward symmetry and encoded solidarity instead: intimates exchange mutual T, non-intimates exchange mutual V, and the asymmetric power usage receded. The change tracks broader social change, and the moment when two people switch to mutual T is a marked social event in many communities, sometimes explicitly negotiated.

English lost this distinction in an unexpected direction. Thou was the T form and you the V form, and it was thou that dropped out, leaving the polite plural to cover everything. English speakers are therefore all addressing each other with the historically formal pronoun.

English does the same work with names instead. Roger Brown and Marguerite Ford analyzed the choice between title plus last name and first name. Reciprocal first names signal intimacy or solidarity. Reciprocal title plus last name signals mutual distance and respect. Asymmetric usage, where one party gives a first name and receives a title, marks power, which is why a physician may be Doctor while the patient is addressed by first name, and why the same asymmetry appears between teachers and students, employers and employees, and adults and children. Languages with elaborate honorific systems, such as Japanese and Korean, encode this grammatically and pervasively rather than only in address.

Key idea: Address forms encode power when used asymmetrically and solidarity when used reciprocally, a system visible in the historical shift of European T and V pronouns and in English name choices.

When conventions collide

Because these conventions are unconscious and vary between communities, they are a rich source of misunderstanding that gets attributed to personality instead.

John Gumperz called the small signals that tell listeners how to interpret an utterance contextualization cues: intonation, rhythm, pausing, formulaic wording, code choice. They do not carry content; they carry instructions about how to take the content. When two speakers use different cue systems, both understand every word and neither understands the message.

Gumperz's most cited case involved cafeteria staff of South Asian background serving employees at a British airport. Offering gravy, the servers said the word with falling intonation, which in their variety marked a routine offer. British listeners, whose system marks offers with rising intonation, heard the falling contour as a flat statement and found it rude and uninterested. The staff, meanwhile, experienced the customers as inexplicably hostile. Neither side could identify the cause, because nobody involved was conscious of the rule being applied. The dispute was read on both sides as an attitude problem.

That mechanism generalizes far beyond an airport. Differences in what counts as an acceptable pause, in how directly a request may be made, in whether overlap signals engagement or interruption, and in how disagreement is packaged all produce reliable, repeated misreadings of character. The next lesson takes up what happens when such misreadings occur in situations where one party has power over the other.

Key idea: Contextualization cues instruct listeners how to interpret utterances, and when speakers use different cue systems the resulting misunderstanding is routinely attributed to personality rather than to convention.

Common misconceptions

  • Conversational turn-taking is loosely coordinated. Gaps average around two hundred milliseconds, requiring projection of turn endings rather than reaction to them.
  • A dispreferred response reflects the speaker's psychological reluctance. Preference here is structural: refusals are packaged with delay, prefaces, and accounts regardless of feelings.
  • Correcting someone is a neutral act. Other-initiated other-repair is the least preferred repair type, which is why correction carries social weight.
  • Politeness means indirectness. Bald on record is appropriate in urgency and intimacy, and treating negative politeness as most polite reflects one culture's weighting of autonomy.
  • Cross-cultural friction in conversation reflects attitude. Differing contextualization cue systems produce systematic misreadings that neither party can consciously detect.

Recap

  • Turns are built from projectable units ending at transition-relevance places, where ordered rules allocate the next turn.
  • Average gaps of about two hundred milliseconds require speakers to project completion and plan in advance.
  • Cross-linguistic study across ten languages found the same turn-taking design with modest local variation.
  • Adjacency pairs make second parts conditionally relevant, so their absence is itself interpretable.
  • Dispreferred responses are delayed, prefaced, mitigated, and accounted for, which makes them recognizable before their content.
  • Repair runs from most preferred self-initiated self-repair to least preferred other-initiated other-repair.
  • Brown and Levinson scale politeness strategies by social distance, power, and ranked imposition, with critiques targeting the individualist framing of face.
  • T and V pronouns and English name choices encode power asymmetrically and solidarity reciprocally.
  • Mismatched contextualization cues generate misunderstandings that both parties attribute to character.

Sources

  1. Sacks, H., Schegloff, E. A., & Jefferson, G. (1974). A simplest systematics for the organization of turn-taking for conversation. Language, 50(4), 696-735. doi.org/10.2307/412243
  2. Stivers, T., Enfield, N. J., Brown, P., Englert, C., Hayashi, M., Heinemann, T., et al. (2009). Universals and cultural variation in turn-taking in conversation. PNAS, 106(26), 10587-10592. doi.org/10.1073/pnas.0903616106
  3. Dingemanse, M., Torreira, F., & Enfield, N. J. (2013). Is Huh? a universal word? Conversational infrastructure and the convergent evolution of linguistic items. PLoS ONE, 8(11), e78273. doi.org/10.1371/journal.pone.0078273
  4. Wikipedia contributors. (n.d.). T-V distinction. en.wikipedia.org
  5. Wikipedia contributors. (n.d.). Politeness theory. en.wikipedia.org
Key terms
Transition-relevance place
The point at the end of a turn-constructional unit where speaker change becomes legitimate and the turn allocation rules apply.
Adjacency pair
A two-turn sequence in which a first part makes a particular second part conditionally relevant, such as question and answer.
Preference organization
The structural asymmetry by which preferred second parts come quickly and plainly while dispreferred ones are delayed, prefaced, and accounted for.
Repair
The organized practices for handling trouble in speaking, hearing, or understanding, ordered from self-initiated self-repair to other-initiated other-repair.
Face
Goffman's term for the public self-image a person claims in interaction, which participants collaboratively maintain.
Positive face
The desire to be liked, approved of, and included, threatened by acts such as criticism and disagreement.
Negative face
The desire to be unimpeded and autonomous, threatened by acts such as requests, orders, and impositions.
T-V distinction
The contrast between familiar and formal second person pronouns, historically asymmetric for power and later reciprocal for solidarity.
Contextualization cue
Gumperz's term for a signal such as intonation or rhythm that instructs listeners how to interpret an utterance rather than adding content.

Consequences: Linguistic Profiling, the Courts, Digital Language, and Careers

  • Describe the audit evidence for linguistic profiling in housing and the legal treatment of accent discrimination in employment.
  • Explain how dialect misunderstanding operates in courtrooms, using the transcription accuracy research and the Jeantel testimony analysis.
  • Evaluate the claim that digital communication degrades language, and identify applied careers that use sociolinguistic training.

The big picture

Lesson 8 established that listeners rate accents consistently in laboratory settings. This lesson asks the harder question: does that translate into behavior with material consequences? The answer is yes, it has been measured, and the measurements are among the most important results the field has produced.

This is where the descriptive stance from Lesson 1 pays off. If varieties were genuinely deficient, unequal outcomes would be unfortunate but explicable. Because they are not, unequal outcomes are discrimination on the basis of a proxy for race, national origin, or class. That reframing is what makes sociolinguistics evidence in a legal sense and not merely an academic position.

Key idea: Attitudes measured in the laboratory translate into measurable behavior in housing, hiring, and courtrooms, and because varieties are not deficient, the resulting disparities constitute discrimination rather than consequence.

Linguistic profiling in housing

John Baugh coined the term linguistic profiling for the practice of identifying a speaker's race or ethnicity from their voice and acting on that identification. He came to the topic through experience: apartment-hunting by telephone, he found that appointments arranged over the phone repeatedly evaporated when he arrived in person.

The resulting study, published in 1999 by Thomas Purnell, William Idsardi, and Baugh, is a model of design. Baugh is a native speaker of three varieties: African American English, Chicano English, and Standard American English. He telephoned landlords across several San Francisco Bay Area communities with an identical script, varying only the guise. Because it was the same person with the same voice, the same script, the same stated qualifications, and the same time of day, the design isolates the variety exactly as a matched guise study does, but the outcome measured is a real appointment rather than a rating on a scale.

Appointment rates differed sharply by guise. They also varied by neighborhood: the gap between the standard guise and the other two was largest in the areas with the smallest minority populations. So the effect was not uniform prejudice but a locally calibrated response.

A second experiment answered the obvious objection that identification from a phone call must be unreliable. Listeners heard a single word, hello, spoken in each guise, and identified the speaker's ethnicity at rates well above chance. Discrimination therefore does not require a conversation. One word is enough for a listener to categorize a caller and act on it.

The legal significance is direct. Housing discrimination on the basis of race is prohibited in the United States under the Fair Housing Act, and this line of research demonstrates a mechanism by which it can operate over the telephone, invisibly, without any explicit reference to race by either party.

Key idea: Baugh's audit study held the speaker, script, and qualifications constant while varying only the variety, and found appointment rates differing by guise and by neighborhood composition, with listeners able to categorize a speaker from a single word.

Accent and employment

Employment presents a harder legal problem, because employers may lawfully require effective communication for a job. The difficulty is that the judgment of whether communication is effective is made by the same listeners whose attitudes are under study.

Two frequently discussed United States cases illustrate the tension. In one, a job applicant of Filipino background who had scored highest on the written examination for a clerk position was denied the job on the grounds that his accent would impede communication with the public; the court accepted the employer's account. In another, a speaker of Hawaii Creole English with strong meteorological qualifications was passed over for an on-air weather broadcasting position in favor of a candidate with a mainland accent, with the reasoning resting on how he sounded. Courts have generally held that national origin discrimination is unlawful while accent-based decisions may be permissible where accent genuinely and materially interferes with job performance, and the gap between those two statements is where most of the disputed cases live.

Rosina Lippi-Green analyzed this pattern and introduced a useful concept: the communicative burden. Understanding across an accent difference is a shared task, and a listener can choose how much of it to take up. Listeners routinely accommodate accents they find prestigious and decline to accommodate accents they do not, then report the failure as the speaker's unintelligibility. Framing intelligibility as a property of the speaker alone conceals a decision the listener made.

Key idea: Accent-based employment decisions occupy a contested legal space, and the concept of communicative burden shows that judged unintelligibility often reflects a listener's willingness to accommodate rather than an objective property of the speech.

Language in the courtroom

The stakes rise again in legal proceedings, where a misunderstanding can decide a case.

John Rickford and Sharese King analyzed the testimony of Rachel Jeantel in the 2013 trial arising from the killing of Trayvon Martin. Jeantel, a speaker of African American English, testified for hours as the prosecution's key witness. Jurors afterward described her testimony as difficult to understand and reported finding it not credible, and one juror indicated it was essentially disregarded. Rickford and King documented that the official transcript contained numerous errors in which AAE grammatical features were mistranscribed or rendered as something other than what was said, and that specific constructions were systematically misread by participants unfamiliar with the variety. Their broader argument is that the courts' treatment of vernacular speakers raises a question of access to justice, not merely of style.

Taylor Jones, Jessica Kalbfeld, Ryan Hancock, and Robin Clark tested the transcription problem experimentally. They played recordings of African American English to certified court reporters in Philadelphia, professionals required to meet a ninety-five percent accuracy standard for certification. On the AAE passages, accuracy fell dramatically below that threshold, with reporters transcribing only around sixty percent of the utterances accurately and performing worse still when asked to paraphrase what they had heard. Note the implication carefully: the official record of what a witness said, the document appeals courts read, is unreliable in a patterned way that disadvantages speakers of one variety.

Other courtroom issues follow the same shape. Comprehension of rights warnings has been shown to depend on wording that many people do not parse as intended, particularly juveniles and speakers of other varieties or languages. The linguistics of consent during police encounters raises comparable questions about whether a request is heard as a request or an instruction. And the right to an interpreter, though established, is undermined in practice when no certified interpreter exists for a language or creole, when ad hoc interpreters such as family members are used, or when interpreters render summaries rather than complete testimony. In asylum proceedings, some jurisdictions have used language analysis to test claimed origins, a practice that linguists have criticized heavily where it is performed by unqualified analysts or ignores the linguistic effects of displacement and long residence elsewhere.

Key idea: Court reporters transcribing African American English fall far below their certification accuracy standard, and analysis of the Jeantel testimony shows how a vernacular speaker's evidence can be misheard, mistranscribed, and discounted.

Does digital communication degrade language?

Turn now to a claim you have encountered many times, and apply the tools this course has built. Texting and social media are said to be ruining language: spelling is collapsing, abbreviations are replacing words, young people can no longer write.

First, recognize the genre. This is the complaint tradition from Lesson 8, now running on a new target. The same structure appeared for the telephone, the telegraph, and the novel.

Second, look at the evidence, which is unusually clear. David Crystal's analysis of text message corpora found that abbreviations and non-standard spellings make up a small proportion of what people actually text, far below the impression the complaints create. Sali Tagliamonte and Derek Denis analyzed a large corpus of teenagers' instant messaging and found it to be a hybrid register, drawing on both speech and formal writing, dominated by standard forms and structured by the same variationist patterns found in speech. The abbreviations everyone cites, including the famous one written as three letters and meaning laughing out loud, were far rarer in the corpus than the public discussion implies.

Third, the developmental evidence points the opposite way from the complaint. Studies of children's texting have repeatedly found that use of abbreviations correlates positively with spelling and literacy scores, not negatively. The mechanism makes sense on reflection: to abbreviate a word you must know how it is spelled and how it sounds, and the manipulation involved is phonological play. A child who writes l8r has performed an analysis, not forgotten one.

Fourth, the register argument. Texting is not degraded formal writing; it is a distinct register, closer to speech than to essays and used for purposes that essays never served. Its features solve problems that written text has always had: emoji and repeated letters substitute for the intonation, facial expression, and volume that speech carries and that plain writing loses. Speakers who text this way generally write formal prose perfectly well when the situation calls for it, which is exactly the style shifting from Module 2.

One further finding is worth noting because it contradicts a common assumption. Regional variation has not been erased online. Analyses of geolocated social media data have found strong regional patterning in vocabulary and spelling, and have tracked innovations spreading between cities in patterns resembling the cascade diffusion from Lesson 7. Digital communication is another domain in which variation lives, not a solvent that dissolves it.

Key idea: Corpus evidence shows digital communication is dominated by standard forms, abbreviation correlates positively with literacy, and online language shows the same regional patterning and register shifting found in speech.

Where this training goes

Sociolinguistics is unusually applied for a humanities-adjacent field, and the applications are worth naming concretely.

Forensic linguistics puts sociolinguists in courts as expert witnesses on transcript accuracy, dialect comprehension, authorship, and the reliability of speaker identification. Speech-language pathology depends on distinguishing dialect difference from genuine disorder, which requires exactly this training. Education uses it for teacher preparation, dialect-informed reading instruction, and assessment design. Language documentation and revitalization work with communities on the projects described in Module 5. Policy and planning roles exist in governments and international organizations. Lexicography, translation, and interpreting all draw on it.

Technology has become a major destination, and one recent finding shows why the training matters there. Allison Koenecke and colleagues tested five major commercial automated speech recognition systems and found that word error rates for Black speakers were roughly double those for white speakers, with the gap traceable substantially to the acoustic models and to training data that underrepresented African American English. As speech interfaces spread into hiring platforms, medical documentation, captioning, and customer service, a system that transcribes one group's speech at half the accuracy of another's is a fairness problem being built into infrastructure. Recognizing that problem, diagnosing its cause, and specifying the data needed to fix it are sociolinguistic tasks.

Which brings the course back to where it started. The central empirical finding is that all varieties studied turn out to be rule-governed, and that judgments about bad grammar track social position rather than linguistic quality. The central practical finding is that those judgments have consequences anyway: in apartments, in jobs, in transcripts, in classrooms, and now in the training data of systems that will make decisions at scale. Holding both of those facts at once, without letting either one cancel the other, is what it means to think about language sociolinguistically.

Key idea: Sociolinguistic training applies directly in forensic, clinical, educational, policy, and technology settings, where speech recognition systems have been shown to transcribe some varieties at far lower accuracy than others.

Common misconceptions

  • Telephone discrimination requires the caller to state their race. Listeners identify varieties from as little as a single word at rates well above chance.
  • Unintelligibility is a property of the speaker. The communicative burden is shared, and listeners accommodate prestigious accents while declining to accommodate others.
  • Court transcripts are objective records. Certified reporters transcribing African American English fall far below their own certification standard.
  • Texting is ruining spelling. Abbreviations are a small share of actual messages and correlate positively with literacy scores in children.
  • The internet is erasing regional variation. Geolocated social media data show robust regional patterning and city-to-city diffusion of innovations.

Recap

  • Baugh's audit study varied only the variety and found appointment rates differing by guise and by neighborhood composition.
  • Listeners identified speaker ethnicity from the single word hello at above-chance rates, so profiling needs no conversation.
  • Accent-based employment decisions sit in a contested legal space between lawful communication requirements and unlawful national origin discrimination.
  • The communicative burden is shared, so judged unintelligibility often reflects a listener's choice not to accommodate.
  • Court reporters transcribed African American English at roughly sixty percent accuracy against a ninety-five percent certification standard.
  • Analysis of the Jeantel testimony documented mistranscription and systematic misunderstanding of a vernacular witness.
  • Corpus evidence shows digital messaging is dominated by standard forms and that abbreviation correlates positively with literacy.
  • Automated speech recognition has shown word error rates roughly twice as high for Black speakers, a fairness problem in deployed infrastructure.

Sources

  1. Purnell, T., Idsardi, W., & Baugh, J. (1999). Perceptual and phonetic experiments on American English dialect identification. Journal of Language and Social Psychology, 18(1), 10-30. doi.org/10.1177/0261927X99018001002
  2. Rickford, J. R., & King, S. (2016). Language and linguistics on trial: Hearing Rachel Jeantel (and other vernacular speakers) in the courtroom and beyond. Language, 92(4), 948-988. doi.org/10.1353/lan.2016.0078
  3. Jones, T., Kalbfeld, J. R., Hancock, R., & Clark, R. (2019). Testifying while Black: An experimental study of court reporter accuracy in transcription of African American English. Language, 95(2), e216-e252. doi.org/10.1353/lan.2019.0042
  4. Tagliamonte, S. A., & Denis, D. (2008). Linguistic ruin? LOL! Instant messaging and teen language. American Speech, 83(1), 3-34. doi.org/10.1215/00031283-2008-001
  5. Koenecke, A., Nam, A., Lake, E., Nudell, J., Quartey, M., Mengesha, Z., et al. (2020). Racial disparities in automated speech recognition. PNAS, 117(14), 7684-7689. doi.org/10.1073/pnas.1915768117
  6. Linguistic Society of America. (n.d.). Careers in linguistics. linguisticsociety.org
Key terms
Linguistic profiling
Baugh's term for identifying a speaker's race or ethnicity from their voice and acting on that identification, often over the telephone.
Audit study
A design that holds an applicant's qualifications and script constant while varying one characteristic, measuring a real-world outcome such as an appointment.
Communicative burden
Lippi-Green's concept that understanding across an accent difference is a shared task, so judged unintelligibility reflects listener effort as well as speaker output.
Forensic linguistics
The application of linguistic analysis to legal questions, including transcript accuracy, authorship, and speaker identification.
Language analysis for determination of origin
The use of speech analysis to test an asylum seeker's claimed nationality, criticized where analysts are unqualified or displacement effects are ignored.
Textism
An abbreviation or respelling characteristic of digital messaging, whose use correlates positively rather than negatively with literacy in children.
Hybrid register
Tagliamonte and Denis's characterization of instant messaging as drawing on features of both speech and formal writing.
Word error rate
The standard accuracy measure for speech recognition systems, found to be roughly twice as high for Black speakers in a study of five commercial systems.

Open the interactive version with quizzes and progress →