🗣️ Linguistics · Undergraduate · LING 350

Phonetics & Phonology

A complete, college-level course in phonetics and phonology, built around analysis rather than description. Phonetics asks what speech physically is: how the vocal folds vibrate, where the tongue goes, what the pressure wave looks like, and what a listener does with it. Phonology asks a different question about the same sounds: which differences a particular language treats as differences, and…

Start the interactive course (quizzes, progress, videos) →

Free forever. No sign-up, no ads. 15 lessons. The full lesson text is below so you can read it right here.

Module 1: The Sounds Themselves

The division of labour between phonetics and phonology stated on real data, the vocal tract from the lungs to the lips, and a working command of the consonant chart: place, manner, voicing, and the IPA used as a tool rather than admired as a poster.

Phonetics and Phonology: Two Questions About One Sound

  • Demonstrate English aspiration for yourself and state the rule that predicts where it appears.
  • Distinguish a phonetic question from a phonological question, and decide which one a given piece of evidence answers.
  • Explain why the same physical difference is contrastive in Korean and predictable in English, using slashes and brackets correctly.

A strip of paper and two words

Tear a strip off a sheet of paper, about two fingers wide, and hold it upright a couple of centimetres in front of your lips. Say the word pin. The strip jumps forward. Now say spin, at the same volume, with the strip in the same place. It barely twitches. Do it a few times, alternating. The effect is large and it is completely reliable, and if you are a native speaker of English you have been producing it every day of your life without once noticing.

What the paper is registering is a puff of air. In pin, the lips come apart and the vocal folds do not start vibrating for something like sixty thousandths of a second; during that gap, air rushes out unimpeded. Phoneticians call this aspiration and write it with a raised h: [pʰ]. In spin, voicing begins almost as soon as the lips open, so there is no puff. That sound is written [p], with no raised h.

So the p of pin and the p of spin are not the same. They differ by a measurable, repeatable physical property, one your paper strip just detected. Yet ask any English speaker how many different p sounds they just made and the answer is one. That gap, between what your mouth did and what your mind counted, is the subject of this whole course.

Key idea: Aspiration is a real, measurable difference between the p of pin and the p of spin, and English speakers produce it consistently while reporting that they made the same sound twice.

Where the puff shows up, and where it does not

Before deciding what to make of that, collect more data. Say each of these words with the paper strip and mark whether it moves. Do it honestly, and do not let spelling talk you into an answer.

WordTranscriptionPaper stripPosition of the stop
pin[pʰɪn]JumpsStart of a stressed syllable
spin[spɪn]StillAfter s
appear[əˈpʰɪɹ]JumpsStart of a stressed syllable
happy[ˈhæpi]Little or nothingStart of an unstressed syllable
top[tʰɑp]Jumps on the t, not on the final pStart of a stressed syllable, and word end
stop[stɑp]StillAfter s
kill[kʰɪɫ]JumpsStart of a stressed syllable
skill[skɪɫ]StillAfter s

The pattern is not random and it is not about the letter p. It covers all three English voiceless stops, [p], [t] and [k], and it is stated in terms of position: an English voiceless stop is aspirated at the beginning of a stressed syllable, and unaspirated elsewhere, most obviously after s. You did not learn this rule. Nobody taught it to you. You obey it anyway, which is the first sign that something in your head, and not just in your mouth, is doing work.

Here is the second sign, and it is the decisive one. There is no pair of English words that differ only in aspiration. None. You cannot find two words, one with [pʰ] and one with [p] in the same position, that mean different things. Aspiration in English never carries a message. It is a free rider on position.

The point: English aspiration is fully predictable from position, and no English word pair is distinguished by aspiration alone, so the puff of air conveys nothing about which word you said.

What a Korean speaker hears

Now take the same physical difference to Seoul. Korean has three series of stops at each place of articulation, and at the labial place they are written 불, 풀 and 뿔. The first, 불 [pul], means fire. The second, 풀 [pʰul], means grass. The third, 뿔 [p͈ul], the so-called tense series, means horn.

Look at what separates fire from grass: nothing but the puff of air. The lips do the same thing, the vowel is the same, the final consonant is the same. Only aspiration differs, and the words mean entirely different things. In Korean, aspiration is not a free rider. It is the whole cargo.

The consequence for listeners is the part people find hardest to believe. A Korean speaker learning English hears the p of pin and the p of spin as clearly different, in the way you hear the p of pin and the b of bin as different. An English speaker learning Korean often cannot hear 불 and 풀 apart at all at first, and produces one when aiming for the other. The ears are the same. The acoustic signal is the same. What differs is the system each listener is running the signal through.

Hindi pushes the point further, with a four-way contrast at the same place: an unaspirated voiceless stop, an aspirated voiceless stop, a voiced stop, and a breathy-voiced stop, so that [pʰal] fruit and [bal] hair are only two members of a larger set. Languages do not agree about how many distinctions to carve out of the same articulatory space, and nothing in the acoustics tells them what to do.

Worth holding on to: One physical difference, aspiration, is predictable and meaningless in English and contrastive and meaningful in Korean, which means the difference between them is not in the sounds but in the systems.

The two questions, stated precisely

You can now say exactly what the two halves of this course do, without hand waving.

Phonetics studies speech sounds as physical events, independently of any particular language. Its questions are answered with instruments and measurements. How long is the delay before voicing starts? Where does the tongue tip contact? What frequencies carry the energy in this fricative? A phonetician can study a sound from a language they do not speak, and two phoneticians measuring the same recording should get the same numbers. The field divides into articulatory phonetics, which studies how sounds are produced, acoustic phonetics, which studies the pressure wave, and auditory or perceptual phonetics, which studies what listeners do with it. Modules 1 through 3 of this course are phonetics.

Phonology studies the sound system of a particular language: which differences that language treats as differences, how its sounds pattern, which combinations it permits, and what changes when sounds meet. Its questions are answered by analysis of distribution and by what speakers of that language accept or reject. Is the difference between [pʰ] and [p] used to tell words apart here? What contexts predict which one appears? Why does no English word start with the sequence ng? Modules 4 through 6 are phonology.

The relationship is not that phonology is the theory and phonetics the data. Both are empirical. The distinction is in what the question is about: phonetics asks what the sound is, phonology asks what the language does with it. And a fact can only be phonological relative to a language, which is why the sentence Korean has a three-way stop contrast makes sense and the sentence [pʰ] is a phoneme, full stop, does not.

Slashes, brackets, and why the difference matters

The notation carries the distinction, so it is worth getting right on the first day.

  • Square brackets enclose phonetic transcription: what was actually said, as precisely as you care to record it. [spɪn] and [pʰɪn] are claims about physical events.
  • Slashes enclose phonemic transcription: the units of contrast in a particular language. English /spɪn/ and /pɪn/ record which mental categories were used, and deliberately leave out the aspiration, because in English the aspiration is not a choice.

So the English word pin is /pɪn/ phonemically and [pʰɪn] phonetically. Both are correct; they answer different questions. Writing /pʰɪn/ for English would be a mistake, because it asserts that English speakers choose aspiration, which they do not. Writing /pʰul/ for Korean grass is right, because they do.

A short exercise. Decide which notation each statement calls for. First: the second sound of the English word stop lasts about fifteen milliseconds before voicing begins. Second: English distinguishes the words pin and bin. Third: in this recording the speaker released the final t audibly. The first and third are phonetic claims about events and belong in brackets; the second is a claim about the English system and belongs in slashes.

Remember: Brackets record what was said; slashes record what was chosen. The same word gets both, and the difference between the two transcriptions is exactly the part of the pronunciation the language did not leave up to the speaker.

A second case, where English speakers get fooled

Aspiration is the standard first example because English speakers cannot hear it. Here is one where English speakers hear a difference that is not, on the surface, there.

In most North American English, say the pairs latter and ladder, writer and rider, atom and Adam, at ordinary conversational speed. The consonant in the middle is neither a clean [t] nor a clean [d]. It is a tap, written [ɾ]: the tongue tip strikes the alveolar ridge once, very fast, and the vocal folds keep vibrating throughout. In both members of each pair, the same tap appears. Phonetically, latter and ladder can be identical in that consonant.

Yet speakers report hearing a difference, and often they are right, because the difference has migrated onto the vowel. Vowels before voiced consonants in English are longer than vowels before voiceless ones, and that length difference survives the tap. Time your own: the a of ladder is measurably longer than the a of latter. In dialects with Canadian raising, writer and rider differ more visibly still, because the diphthong raises before the underlying voiceless consonant: [ɹʌɪɾɚ] against [ɹaɪɾɚ]. Module 5 uses exactly this pair to argue for an ordering of rules, so file it away.

The phonetic question here is what the tongue did and how long the vowel lasted. The phonological question is why English speakers still feel there are two different consonants when the recording shows one. Neither question answers the other, and you need both.

Common misconceptions

  • Phonetics is about sounds and phonology is about spelling. Neither is about spelling. Orthography is a separate system, often centuries out of date with pronunciation; the gh of night records a consonant English lost. Nothing in this course is a claim about letters.
  • A phoneme is a sound. A phoneme is a category of sounds that a particular language treats as one unit of contrast. English /t/ covers at least [tʰ], [t], [ɾ], [ʔ] and an unreleased [t], depending on position. The phoneme is the label on the box, not any one item in it.
  • Some languages have more sounds than others because they are more complex. Inventory size varies enormously, from about eleven distinctive sounds in Rotokas and thirteen in Hawaiian to well over a hundred in some Khoisan languages, and it correlates with nothing about the cultures or the expressive power of the languages involved.
  • If I cannot hear a distinction, it is not really there. Your inability to hear the Korean plain and aspirated stops apart is a fact about your phonological system, not about the acoustics. Instruments detect the difference easily, and Korean infants have learned it before their first birthday.
  • Careful speech is the real pronunciation and casual speech is sloppy. Both are governed by rules, and the casual forms are usually the ones that reveal the system, which is why phoneticians work hard to record unselfconscious speech.

What to carry forward

  • English voiceless stops are aspirated at the start of a stressed syllable and unaspirated after s, a rule you follow without ever having learned it.
  • No English word pair is distinguished by aspiration alone, so aspiration in English is predictable rather than contrastive.
  • The same difference is contrastive in Korean, where 불 fire and 풀 grass differ by nothing else, so contrast is a property of a language and not of a sound.
  • Phonetics asks what the sound physically is and can be studied language-independently; phonology asks what a particular language does with it.
  • Brackets record what was said, slashes record the units of contrast, and the gap between the two transcriptions is the predictable part.
  • The flapping of latter and ladder shows the reverse case: one surface consonant, two categories, with the evidence displaced onto vowel length and quality.

Sources

  1. International Phonetic Association. (n.d.). The International Phonetic Alphabet and the IPA chart. internationalphoneticassociation.org
  2. Wikipedia contributors. (n.d.). Aspirated consonant. Wikipedia. en.wikipedia.org
  3. Wikipedia contributors. (n.d.). Korean phonology. Wikipedia. en.wikipedia.org
  4. Britannica. (n.d.). Phonetics. Encyclopaedia Britannica. britannica.com
  5. Ladefoged, P., & Johnson, K. (2015). A course in phonetics (7th ed.). Cengage Learning.
Key terms
Phonetics
The study of speech sounds as physical events, covering their production, their acoustics, and their perception, independently of any one language.
Phonology
The study of the sound system of a particular language: which differences are contrastive, how sounds pattern, and what rules govern them.
Aspiration
A delay between the release of a stop and the start of vocal fold vibration, heard as a puff of air and written with a raised h.
Contrast
The relation between two sounds in a language when swapping one for the other can change which word is being said.
Phoneme
A unit of contrast in a particular language, written between slashes; it covers a set of physically different sounds that the language treats as one.
Allophone
One of the physically distinct sounds that realize a single phoneme, whose appearance is predictable from context.
Tap
A consonant made by a single very fast strike of the tongue tip against the alveolar ridge, written with the symbol for a flap.
Broad transcription
Transcription that records only the contrastive units, leaving out predictable phonetic detail.

Inside the Vocal Tract: Air, Voicing, and Articulators

  • Trace a breath of speech from the lungs to the lips and name what happens to it at each stage.
  • Describe the states of the glottis and use voice onset time to compare the stop systems of English, Spanish and Thai.
  • Identify the active and passive articulators and explain what the velum controls.

Where the buzz comes from

Put two fingers flat on the front of your throat, just above the notch of your collarbone, and hiss a long ssssss. Nothing. Now, without pausing for breath, slide into zzzzzz. A buzz starts under your fingers, and it stops the instant you slide back to ssssss. Your tongue has not moved. The airflow has not changed much. What switched on is a pair of folds of muscle and mucous membrane sitting horizontally across the top of your windpipe, roughly 17 to 25 millimetres long in an adult man and 12 to 17 in an adult woman.

During ordinary speech those folds slam together and blow apart something like 100 to 120 times a second in a typical adult male voice and around 200 times a second in a typical adult female voice. That rate is the fundamental frequency, and it is what you hear as the pitch of a voice. Everything else in this lesson is either the pump that drives those folds or the tube that shapes what comes out of them.

In short: Voicing is vocal fold vibration, it can be switched on and off independently of everything the tongue is doing, and its rate is what you hear as pitch.

The power supply: air going out of the lungs

Speech runs on exhaled air. You raise the pressure below the glottis by contracting the intercostal muscles and letting the elastic recoil of the ribcage do most of the work, and the pressure difference drives air upward. This is the pulmonic egressive airstream mechanism: pulmonic because the lungs are the piston, egressive because the air is going out. Every language on earth uses it for most of its sounds, and many languages use nothing else.

Two consequences follow immediately. First, speech happens on the outbreath, which is why a long sentence leaves you short of air and why speakers time their breathing to syntactic boundaries. Second, the subglottal pressure is fairly steady over a phrase, which means loudness and pitch have to be controlled elsewhere, mostly at the larynx.

The folds do not vibrate because a muscle is contracting rhythmically. Nothing in the body twitches 200 times a second on command. They vibrate because of an interaction between air pressure and tissue elasticity, the myoelastic-aerodynamic account: air pressure below forces the closed folds apart, the air rushing through the narrow gap drops in pressure and, together with the elastic recoil of the stretched tissue, sucks them shut again, whereupon pressure rebuilds and the cycle repeats. Pull the folds tighter and thinner and the cycle runs faster, which raises pitch. This is why singing high is muscular work at the larynx and has nothing to do with breathing harder.

The larynx as a valve, in several settings

Think of the glottis, the space between the folds, as a valve with more than two settings. Speech uses at least five, and languages differ in which ones they exploit.

SettingWhat the folds doHeard asExample
VoicelessHeld apart, no vibrationWhispery or silentThe s of sip
VoicedVibrating regularlyA buzz, a definite pitchThe z of zip
Breathy voiceVibrating while still leaking airSighing, murmuredHindi breathy-voiced stops
Creaky voiceVibrating slowly and irregularly, folds thickVocal fry, a rattlePhrase-final English
ClosedPressed fully shut, then releasedA catch in the throatThe glottal stop in uh-oh

The glottal stop deserves a moment. Say uh-oh slowly. Between the two syllables your vocal folds close completely and the sound stops dead. That closure is a consonant, written [ʔ], and in Arabic and Hawaiian it is a full member of the consonant inventory. In many varieties of English it turns up as a pronunciation of /t/ in words like button [ˈbʌʔn̩] or in the Cockney bottle. English speakers produce it constantly and rarely notice it, which should be a familiar situation by now.

Voice onset time: turning a language into a number

The last lesson used a paper strip. Now use a clock. Voice onset time, or VOT, is the interval between the release of a stop closure and the start of voicing. Take the release as time zero. If voicing starts afterwards, VOT is positive; if the folds were already buzzing before the release, VOT is negative, which phoneticians call prevoicing or voicing lead.

Leigh Lisker and Arthur Abramson measured VOT across eleven languages in a 1964 study that reorganized how the field thought about stop contrasts. Their result was that the continuous VOT dimension gets carved up into a small number of regions, and that languages choose two or three of them.

VOT regionApproximate rangeEnglishSpanishThai
Voicing lead (prevoiced)About minus 60 to minus 130 msOccasional, not requiredThe b, d, g seriesThe b, d series
Short lagAbout 0 to plus 30 msThe b, d, g seriesThe p, t, k seriesThe p, t, k series
Long lag (aspirated)About plus 40 to plus 100 msThe p, t, k seriesNot usedThe aspirated series

Read the table across and a well known difficulty falls out. English /b/ and Spanish /p/ occupy the same region: both are short lag. So when a Spanish speaker says the Spanish word pan with a short-lag stop, English ears often hear ban, and when an English speaker aims for Spanish pan with a long-lag aspirated stop, Spanish ears hear something oddly breathy that is not quite either category. Neither speaker is careless. They are dividing the same physical continuum at different places.

Thai uses all three regions at once, which is what a three-way laryngeal contrast looks like when you plot it. Nothing about the human vocal tract required two categories rather than three; the dimension is continuous, and the number of cuts is a fact about the language.

The upshot: VOT converts a laryngeal contrast into a measurable number, and comparing the numbers shows that languages cut one continuous dimension into different numbers of pieces at different places.

Above the larynx: a tube you can reshape

The buzz produced at the glottis is not yet speech. It is a rich, flat-sounding noise. What turns it into vowels and consonants is the tube above, about 17 centimetres long in an adult man and 14 to 15 in an adult woman, running from the vocal folds up through the pharynx and out through either the mouth or the nose.

The switch between those two exits is the velum, the soft palate at the back of the roof of your mouth. Raised, it seals the nasal cavity off and all the air goes out through the mouth: an oral sound. Lowered, the nasal passage opens: a nasal sound. Test it. Pinch your nostrils shut and try to say the word morning. You cannot; the m and the ng need an open nose, and the word comes out strangled. Now pinch your nose and say the word cackle. No difficulty at all, because every sound in it is oral.

The same test on a vowel is more surprising. Pinch your nose and say bad, then ban. The vowel of ban is affected, because English speakers lower the velum during a vowel that precedes a nasal consonant. That is anticipatory nasalization, and it will come back in Module 4 as a piece of data to be analyzed.

Who moves, and what they move towards

Inside the tube, the standard way to describe a constriction is to name an active articulator, the thing that moves, and a passive articulator, the thing it moves towards.

  • Active: the lower lip, and the tongue, which is not one articulator but several. Phoneticians divide it into the tip or apex, the blade or lamina just behind it, the front, the back or dorsum, and the root, which faces the back wall of the pharynx.
  • Passive: the upper lip, the upper teeth, the alveolar ridge (run your tongue tip back from your teeth until you feel the bony shelf), the hard palate behind it, the soft palate or velum, the uvula hanging at its rear edge, and the pharyngeal wall.

Naming both halves is what gives a sound its place label. The tongue tip against the alveolar ridge gives an alveolar; the lower lip against the upper teeth gives a labiodental; the tongue dorsum against the velum gives a velar. You can locate every one of these on yourself in under a minute, and doing so is far more useful than memorizing the words. Say aaa and then move the tongue slowly back until it touches: you have just found your velum by hitting it.

What matters here: Every consonant place label is just the name of the passive articulator, with the active one usually left implicit because it is predictable from the passive one.

When the lungs are not the pump

Pulmonic egressive air is the default, not the only option. Three other mechanisms occur in the world's languages, and all three are used contrastively somewhere.

  • Ejectives, written with an apostrophe as in [pʼ], [tʼ], [kʼ]. Close the glottis, make an oral closure, then jerk the closed larynx upward. The trapped air is compressed and bursts out with a sharp popping quality. Ejectives are contrastive in Amharic, Georgian, Quechua, Hausa and Navajo, among many others.
  • Implosives, written [ɓ], [ɗ], [ʄ]. The larynx moves down instead of up while the folds continue to vibrate, lowering the pressure in the mouth so that air is drawn slightly inward at the release. Sindhi and Hausa contrast implosives with plain voiced stops.
  • Clicks, written with symbols such as [ʘ], [ǀ], [ǃ], [ǂ] and [ǁ]. Two closures are made in the mouth, the tongue body is pulled back to rarefy the air between them, and the front closure is released with a sharp suction sound. English speakers make clicks all the time as gestures, the tsk of disapproval and the giddy-up used to a horse, but only in southern Africa are they organized into a phonological system, as in Zulu, Xhosa and the Khoisan languages.

Common misconceptions

  • Whispering is just quiet speech. Whisper turns off vocal fold vibration entirely and replaces it with turbulent noise at the glottis. This is why a whispered sip and zip are still distinguishable: the difference has to be carried by duration and by the surrounding vowel rather than by voicing.
  • Pitch is controlled by breathing harder. Louder is breathing harder. Higher is tightening and thinning the vocal folds at the larynx. The two are separable, which is why you can sing a high note softly.
  • English b, d, g are voiced throughout. In word-initial position they are usually short lag, which means voicing starts at or just after the release rather than during the closure. The label voiced is a phonological one; the phonetics is often better described as unaspirated.
  • Nasal sounds are made in the nose. They are made in the mouth like any other consonant; what makes them nasal is that the velum is lowered so the nasal cavity is coupled in as an extra resonator. The place of the oral closure still distinguishes m from n from ng.
  • Clicks are not real speech sounds. In Xhosa and Zulu clicks appear in ordinary vocabulary and distinguish words exactly as p and b do in English.

Putting it together

  • Speech normally runs on pulmonic egressive air, and the steady subglottal pressure means pitch and loudness are controlled at the larynx rather than at the lungs.
  • Vocal fold vibration is aerodynamic, not muscular twitching, which is why it can run at 100 to 250 cycles per second.
  • The glottis is a valve with several settings: voiceless, voiced, breathy, creaky and fully closed, and languages recruit different subsets of them.
  • Voice onset time makes a laryngeal contrast measurable, and English, Spanish and Thai cut the same continuum in different places, which explains a common bilingual mismatch.
  • The velum switches the nasal cavity in and out, as the nose-pinch test on morning against cackle demonstrates.
  • Place labels name the passive articulator; the tongue supplies most of the active ones and is divided into tip, blade, front, back and root.
  • Ejectives, implosives and clicks use the larynx or the tongue body as the pump instead of the lungs, and all three are contrastive somewhere.

Sources

  1. National Institute on Deafness and Other Communication Disorders. (n.d.). What is voice? What is speech? What is language? NIH. nidcd.nih.gov
  2. Wikipedia contributors. (n.d.). Vocal tract. Wikipedia. en.wikipedia.org
  3. Wikipedia contributors. (n.d.). Voice onset time. Wikipedia. en.wikipedia.org
  4. University of Glasgow. (n.d.). Seeing Speech: an articulatory web resource for the study of phonetics. seeingspeech.ac.uk
  5. Lisker, L., & Abramson, A. S. (1964). A cross-language study of voicing in initial stops: Acoustical measurements. Word, 20(3), 384-422.
Key terms
Pulmonic egressive airstream
The default power supply for speech: air pushed outward from the lungs, used for most or all sounds in every language.
Glottis
The space between the vocal folds, whose configuration determines voicing, whisper, breathy voice, creak or complete closure.
Fundamental frequency
The rate at which the vocal folds complete a cycle of opening and closing, perceived as the pitch of the voice.
Voice onset time
The interval between the release of a stop closure and the onset of voicing, measured in milliseconds and negative when voicing leads the release.
Velum
The soft palate, which seals off the nasal cavity when raised and opens it when lowered, making a sound oral or nasal.
Active articulator
The moving structure that forms a constriction, normally the lower lip or some part of the tongue.
Passive articulator
The stationary surface that an active articulator approaches, such as the alveolar ridge, hard palate or velum, and the usual source of a place label.
Ejective
A consonant made with a glottalic egressive airstream, in which a closed larynx is raised to compress trapped air, written with an apostrophe.
Click
A consonant made with a velaric ingressive airstream, in which the tongue body rarefies air between two oral closures before the front one is released.

Consonants by Place, Manner, and Voicing

  • Give the full three-term label for any consonant symbol on the IPA chart, and produce the symbol from a label.
  • Locate the eleven places of articulation on your own vocal tract and identify the manner categories by what happens to the airflow.
  • Read the shaded and empty cells of the IPA chart correctly, and transcribe the consonants of English.

A sound English has no letter for

The Welsh town of Llanelli begins with a consonant that English does not have. English speakers reaching for it produce either a plain l, or a cluster that sounds like thl, or a k, and none of these is right. What Welsh speakers do is put the tongue in position for an l, leaving the sides of the tongue clear as usual, and then push air through hard enough that it becomes turbulent, without any voicing. The result hisses down the sides of the tongue.

The Latin alphabet has no letter for that. The IPA has one symbol: [ɬ]. And the symbol comes with a label that tells you exactly how to make it, without a recording and without a native speaker in the room: voiceless alveolar lateral fricative. Voiceless tells you the folds are apart. Alveolar tells you the tongue tip is at the ridge behind your teeth. Lateral tells you the air goes around the sides. Fricative tells you the constriction is tight enough to make noise. Four pieces of information, one symbol, and any trained phonetician anywhere reconstructs the same sound.

That is what the International Phonetic Alphabet is for. It was created by the International Phonetic Association, founded in Paris in 1886 under the leadership of the French phonetician Paul Passy, with the first version of the alphabet published in 1888. The chart has been revised repeatedly since, most substantially in 2005. It is not a decorative poster. It is a coordinate system, and this lesson teaches you to use it as one.

Key idea: An IPA symbol is shorthand for a set of articulatory instructions, which is why the label and the symbol are interchangeable once you know the system.

The three questions that name a consonant

Every consonant on the chart is named by answering three questions in a fixed order.

  1. Voicing. Are the vocal folds vibrating? Voiced or voiceless. You tested this with your fingers on your throat in the last lesson.
  2. Place. Where is the constriction? Named after the passive articulator, from the lips at the front to the glottis at the back.
  3. Manner. What kind of constriction is it? Complete closure, close enough to make noise, or merely narrowed.

So [z] is a voiced alveolar fricative, [m] is a voiced bilabial nasal, [k] is a voiceless velar plosive, [ð] is a voiced dental fricative. Say each label to yourself and then make the sound from the instructions rather than from the symbol. That is the exercise that turns the chart from something you look things up in into something you can think with.

Place: eleven positions along one tube

Work from the front of the mouth backwards. You can find every one of these on yourself.

PlaceConstrictionSymbolsWhere to hear it
BilabialLower lip to upper lipp b m ɸ βEnglish pat, bat, mat
LabiodentalLower lip to upper teethf v ɱEnglish fan, van
DentalTongue tip to upper teethθ ðEnglish thin, this
AlveolarTongue tip or blade to the ridge behind the teetht d n s z l ɹ ɾEnglish tin, din, sin, nil
PostalveolarTongue blade just behind the ridgeʃ ʒEnglish ship, and the middle of measure
RetroflexTongue tip curled back towards the hard palateʈ ɖ ɳ ʂ ʐ ɻHindi and Tamil stop series
PalatalTongue front to the hard palatec ɟ ɲ ç jSpanish ano against año, German ich
VelarTongue back to the soft palatek ɡ ŋ x ɣEnglish cat, sing; German Bach
UvularTongue back to the uvulaq ɢ ɴ χ ʁ ʀFrench rouge, Arabic Qatar
PharyngealTongue root to the pharynx wallħ ʕArabic, in the letters ha and ayn
GlottalThe vocal folds themselvesʔ hEnglish uh-oh, hat

Two entries deserve comment. The palatal nasal [ɲ] is the difference between Spanish ano and año, which is a genuine minimal pair and one worth remembering as evidence that palatals are not exotic. And German splits its back fricatives by the neighbouring vowel: ich is [ɪç], with a palatal fricative, while Bach is [bax], with a velar one. English speakers hear both as an h-like noise and cannot place them, which by now should be a predictable reaction.

Manner: how much the airflow is obstructed

Place tells you where; manner tells you what kind of interference happens there. Arrange the manners by how much they block the air, from total closure to almost none.

MannerWhat the air doesExamples
Plosive (stop)Complete closure, pressure builds, then releasep b t d k ɡ ʔ
NasalComplete oral closure, but the velum is lowered so air exits the nosem n ŋ ɲ ɴ
TrillAn articulator is set vibrating by the airflow, several closures per secondr ʀ ʙ
Tap or flapOne very fast contact, too brief to build pressureɾ ɽ
FricativeConstriction narrow enough to make the airflow turbulentf v θ ð s z ʃ ʒ x ħ h
Lateral fricativeTurbulent airflow down the sides of the tongueɬ ɮ
ApproximantNarrowed, but not enough for turbulenceɹ j w ɰ
Lateral approximantCentral closure, free airflow around the sidesl ʎ

Spanish gives you the trill and the tap as a minimal pair, which is the cleanest demonstration available: perro, with a trill, means dog; pero, with a single tap, means but. Same place, same voicing, different manner, different word. English uses the tap constantly, in the middle of butter, but never contrastively.

Affricates sit slightly outside this list because they are a stop released into a fricative at the same place: [tʃ] in church, [dʒ] in judge. The IPA writes them as two symbols, sometimes joined by a tie bar, and phonologists argue about whether they are one segment or two. The argument is real, and Module 4 gives you the tools to have it.

The point: Voicing, place and manner are three independent dimensions, and the chart is simply their cross-product, with each cell holding a voiceless symbol on the left and a voiced one on the right.

Reading the gaps in the chart

The consonant chart has two kinds of empty space and they mean different things, which is one of the most useful things to know about it.

Shaded cells mark articulations judged to be impossible. There is no velar trill, because the tongue body is too heavy and too poorly supported to be set vibrating by airflow. There is no pharyngeal nasal, because a nasal requires the velum to be lowered while the oral tract is closed, and a pharyngeal closure is behind the velum, so the air has nowhere to go. There is no glottal lateral, because there are no sides to the glottis. These are claims about anatomy, and they are testable.

Blank unshaded cells mark articulations that are possible but not known to be used contrastively in any documented language. A voiced velar trill is impossible; a voiced bilabial trill is merely rare, and it does occur, in languages of Vanuatu and in Brazilian indigenous languages. The blank is a fact about the sample of documented languages, not about the human body, and blanks have been filled in as documentation has improved.

The consonants of English, placed

Most varieties of English have twenty-four contrastive consonants. Here they are, arranged the way the chart arranges them, with a keyword each.

MannerBilabialLabiodentalDentalAlveolarPostalveolarPalatalVelarGlottal
Plosivep bt dk ɡ
Affricatetʃ dʒ
Nasalmnŋ
Fricativef vθ ðs zʃ ʒh
Approximantɹjw
Laterall

Keywords: pat, bat, tin, din, cat, gap, church, judge, mat, nap, sing, fan, van, thin, this, sin, zip, ship, measure, hat, red, yes, wet, let. Note three traps that spelling sets for you. The letter j spells [dʒ] in English but stands for the palatal approximant [j] in the IPA, the sound at the start of yes. The digraph ng spells a single consonant [ŋ], not a sequence of n and g, which is why sing has three sounds and not four. And th spells two different consonants, voiceless in thigh and voiced in thy, a genuine minimal pair that most speakers have never noticed.

The sound [w] is the awkward one. It involves both a velar constriction and lip rounding, so it is properly a labial-velar approximant and appears in a separate part of the chart. Charts often list it under velar for convenience, and this one does too.

Naming practice, in both directions

Cover the right column and work down the left, then cover the left and work down the right. Say each sound aloud as you go.

LabelSymbol and a word
Voiced velar nasalŋ, as in sing
Voiceless postalveolar fricativeʃ, as in ship
Voiced dental fricativeð, as in this
Voiceless velar fricativex, as in German Bach
Voiced alveolar trillr, as in Spanish perro
Voiced palatal nasalɲ, as in Spanish año
Voiceless uvular plosiveq, as in Arabic Qatar
Voiced labiodental fricativev, as in van
Voiceless glottal plosiveʔ, as in the middle of uh-oh
Voiced alveolar lateral approximantl, as in let

Common misconceptions

  • The IPA is an alphabet for writing languages. It is a transcription system for recording speech. It is not designed for everyday writing, and no community uses it as an orthography.
  • One letter equals one sound. Not in English orthography, and not even in the IPA when you consider affricates and diphthongs, which use two symbols for what may be one phonological unit.
  • Voiced and voiceless pairs like s and z differ only in voicing. They usually differ in several correlated ways at once: voiced fricatives tend to be shorter, weaker in frication and preceded by longer vowels. Voicing is the label, not the whole story.
  • An empty cell on the chart means the sound cannot be made. Only the shaded cells claim that. An unshaded blank means nobody has documented a language that uses it contrastively, which is a much weaker claim and has been overturned before.
  • English r is a trill. The English r of red is an approximant, [ɹ], with no contact at all. The trill [r] is the Spanish and Italian sound, and using the wrong symbol for English is one of the most common transcription errors.

The short version

  • Every consonant is named by voicing, then place, then manner, and the label is a set of instructions you can execute.
  • Place names the passive articulator, running from bilabial at the lips to glottal at the vocal folds, with eleven positions in ordinary use.
  • Manner ranks the degree of obstruction, from complete closure in plosives and nasals through turbulence in fricatives to free flow in approximants.
  • Spanish perro and pero isolate manner, trill against tap, with everything else held constant.
  • Shaded cells on the chart are anatomical impossibilities; blank unshaded cells are gaps in the documented sample.
  • English has twenty-four consonants, and its spelling hides three of them behind digraphs and one behind a letter the IPA uses differently.

Sources

  1. International Phonetic Association. (n.d.). Full IPA chart. internationalphoneticassociation.org
  2. Wikipedia contributors. (n.d.). Place of articulation. Wikipedia. en.wikipedia.org
  3. Wikipedia contributors. (n.d.). Manner of articulation. Wikipedia. en.wikipedia.org
  4. UCLA Phonetics Laboratory. (n.d.). UCLA Phonetics Lab Archive. University of California, Los Angeles. archive.phonetics.ucla.edu
  5. International Phonetic Association. (1999). Handbook of the International Phonetic Association: A guide to the use of the International Phonetic Alphabet. Cambridge University Press.
Key terms
Place of articulation
Where in the vocal tract a consonant constriction is made, named after the passive articulator involved.
Manner of articulation
The degree and kind of obstruction to the airflow, ranging from complete closure in a plosive to free flow in an approximant.
Plosive
A consonant made with a complete closure in the oral tract and a raised velum, so pressure builds and is released as a burst.
Fricative
A consonant whose constriction is narrow enough to make the airflow turbulent, producing audible noise.
Approximant
A consonant whose articulators approach one another without producing turbulence, such as the English r of red.
Lateral
A consonant in which the airstream passes around one or both sides of the tongue rather than over the centre.
Affricate
A stop released into a fricative at the same place of articulation, such as the consonants of church and judge.
Three-term label
The standard description of a consonant, giving voicing, place and manner in that order, from which the sound can be reconstructed.

Module 2: Vowels, Transcription, and Prosody

The vowel space and how languages divide it, transcription from broad to narrow with the diacritics that earn their keep, and the properties that ride on top of segments: stress, tone and intonation.

Vowels and the Vowel Space

  • Place any vowel using height, backness and rounding, and read a vowel quadrilateral.
  • Transcribe the stressed vowels of your own English variety and explain the tense and lax distinction.
  • Compare how Spanish, Arabic, French, Finnish and Turkish divide the same articulatory space.

Five words, one frame

Say these five words in order, slowly, with a mirror in front of you: beat, bit, bait, bet, bat. The first consonant is the same in all five. The last consonant is the same in all five. Everything that distinguishes them lives in the middle, and in the mirror you can watch what it is: your jaw drops progressively, by roughly a centimetre in total from the first word to the last, and your tongue body falls with it.

Five words, five distinct meanings, one changing part. English uses at least eleven distinct vowels in stressed syllables, and some varieties use more. But unlike consonants, you cannot describe any of them by naming a place of contact, because there is no contact. The tongue never touches anything during a vowel. That is the definitional problem this lesson solves.

Key idea: Vowels are made without any constriction narrow enough to obstruct the airflow, which means the consonant vocabulary of place and manner cannot describe them and a different coordinate system is needed.

Three parameters instead of three

The system replaces voicing, place and manner with height, backness and rounding.

  • Height is how close the tongue body comes to the roof of the mouth. Four steps are conventionally recognized: close, close-mid, open-mid and open. The vowel of beat is close; the vowel of bat is open. Height correlates directly with jaw position, which is why the mirror test works.
  • Backness is where along the front-to-back axis the highest point of the tongue sits. Three steps: front, central and back. The vowel of beat is front; the vowel of boot is back; the vowel at the start of about is central.
  • Rounding is whether the lips are protruded and narrowed. Say beat and then boot in front of the mirror and watch the lips. In English, rounding is not independent: back vowels are rounded and front vowels are not, so it carries no information. In French it is independent, and that changes everything, as you will see below.

Try isolating one parameter at a time. Say the vowel of beat and, holding the tongue still, round your lips. What comes out is close to the French vowel of tu. Now say the vowel of beat and slide the tongue backwards without changing the height; you approach the vowel of boot minus the lip rounding, which is roughly the Turkish and Japanese high back unrounded vowel [ɯ]. These are not tricks. They are the coordinate system doing its job.

The quadrilateral, and where it came from

The familiar four-sided vowel chart is not a picture of the mouth, though it is often mistaken for one. It is a reference frame, and it was built deliberately. In the 1910s the British phonetician Daniel Jones defined a set of eight cardinal vowels, fixed reference points chosen so that no language would be needed to define them: number 1 is the closest and frontest vowel that can be made without turbulence, number 5 is the openest and backest, and the rest are spaced at roughly equal auditory intervals around the periphery. They are [i e ɛ a] down the front and [ɑ ɔ o u] up the back. A further eight secondary cardinals reverse the rounding of each.

The shape is a quadrilateral rather than a rectangle because the tongue simply cannot reach as far back at the open end as it can at the close end. The space is genuinely narrower at the bottom, so the chart's slanted left edge encodes an anatomical fact.

Cardinal vowels are not the vowels of any language. They are the graph paper on which a language's vowels get plotted, so that a phonetician can say that the Italian e sits close to cardinal 2 and the French one slightly higher, and be understood.

The vowels of American English, placed

Here is a typical General American inventory, with a keyword each. Your own variety may differ, and where it does, trust your own mouth.

SymbolKeywordHeightBacknessRoundingTense or lax
ibeatCloseFrontUnroundedTense
ɪbitNear-closeFrontUnroundedLax
baitClose-mid, glidingFrontUnroundedTense
ɛbetOpen-midFrontUnroundedLax
æbatNear-openFrontUnroundedLax
ɑbot, fatherOpenBackUnroundedTense
ɔboughtOpen-midBackRoundedTense
boatClose-mid, glidingBackRoundedTense
ʊbookNear-closeBackRoundedLax
ubootCloseBackRoundedTense
ʌbutOpen-midCentral to backUnroundedLax
əthe first vowel of aboutMidCentralUnroundedReduced

Add the three clear diphthongs, [aɪ] in bite, [aʊ] in bout and [ɔɪ] in boy, and the rhotic vowel [ɝ] in bird, and you have most of what you need to transcribe American English.

Two facts about this table are worth arguing over. First, the vowels of bait and boat are transcribed here as diphthongs, because in American English they genuinely glide, which you can hear by drawing out the word bait and noticing that it ends nearer the vowel of bit. Some transcription traditions write them as plain [e] and [o]. Both are defensible; what is not defensible is failing to say which convention you are using.

Second, for a great many North American speakers the bot and bought rows have merged, so that cot and caught, Don and dawn, stock and stalk are homophones. The merger is spreading. If your two rows sound the same, do not force a distinction into your transcription: record what you say.

What matters here: The vowel chart is a reference frame, not a list of correct pronunciations, and your job when transcribing is to plot your own speech on it rather than to match a table.

Tense, lax, and the test that separates them

The tense and lax labels look impressionistic, but English gives them a sharp distributional test: a lax vowel cannot end a stressed syllable, and a tense vowel can.

Consider the possible one-syllable words. See, saw, sue, spa, day and though are all fine, and their vowels are tense. Now try to end a word with the vowel of bit, or bet, or book, or but. There is no English word pronounced [sɪ], [sɛ], [sʊ] or [sʌ] as a complete stressed syllable. Lax vowels require a following consonant. That is a genuine phonological generalization, arrived at by looking at distribution rather than at feel, and it is a preview of the reasoning you will do throughout Module 4.

The same space, divided differently

Every language uses the same tongue in the same mouth. What varies is how many regions it carves the space into, and along which parameters.

LanguageVowel qualitiesWhat it exploits that English does not
Spanish5: i e a o uNothing extra; a small, evenly spread system with no tense and lax pairs
Modern Standard Arabic3 qualities: i a uA length contrast on each, doubling the inventory without adding qualities
Finnish8 qualitiesContrastive length, in a system where it distinguishes ordinary words
FrenchAbout 12 oral qualitiesFront rounded vowels, and a set of contrastive nasal vowels
Turkish8Independent front and back rounding, organized by vowel harmony

Finnish makes the length point unmistakably, because the contrast is carried in ordinary vocabulary: tuli means fire, tuuli means wind, tulli means customs. Only the length of the vowel or the consonant changes, and three different words result. English speakers, whose vowel length is predictable from the following consonant, tend to hear these as the same word said at different speeds. Japanese does the same thing: obasan means aunt and obaasan, with a long vowel, means grandmother.

French supplies the rounding point. The vowel of tu is front, close and rounded, and it contrasts with the front unrounded vowel of the word dit and the back rounded vowel of tout. Three words, one parameter each. English never separates rounding from backness, so English speakers learning French famously substitute the back rounded vowel and say tout for tu. And French nasal vowels are contrastive rather than allophonic: beau and bon differ only in whether the velum is lowered.

Turkish organizes its eight vowels into a harmony system. The plural suffix has two shapes, and which one appears is determined by the last vowel of the stem: ev, house, takes evler, while kitap, book, takes kitaplar. The suffix has no fixed vowel of its own; it copies a feature from the stem. Module 5 returns to harmony as a phonological process.

Common misconceptions

  • The vowel chart is a cross-section of the mouth. It is an auditory and articulatory reference frame based on Jones's cardinal vowels. Modern acoustic measurements plot vowels in a similar shape, which is reassuring, but the chart is not an anatomical drawing.
  • Long vowels are just short vowels held longer. In English the tense and lax pairs differ in quality as well as duration, and the quality difference survives when the durations are equalized. In Finnish and Japanese, by contrast, duration really is the contrast.
  • Everyone speaking the same language has the same vowels. Vowel systems vary more across dialects than consonant systems do. The cot and caught merger, the trap and bath split, and Canadian raising all mean that two speakers of English may have different numbers of vowel phonemes.
  • Schwa is a lazy version of another vowel. Schwa is a vowel in its own right, the most frequent one in English, and its distribution is governed by the stress system rather than by effort. It is what unstressed syllables reduce to, not what tired speakers produce.
  • A language with only three vowels is impoverished. Arabic, with three vowel qualities, has an enormous consonant inventory and a rich morphology built on consonantal roots. Systems trade off; they do not simply rank.

What to remember

  • Vowels have no contact point, so they are described by tongue height, tongue backness and lip rounding rather than by place and manner.
  • Daniel Jones's cardinal vowels are language-independent reference points, and the quadrilateral is graph paper rather than anatomy.
  • General American has around eleven stressed vowel qualities plus schwa, three diphthongs and a rhotic vowel, and speakers differ, notably over the cot and caught merger.
  • Tense and lax is not a feeling but a distribution: lax vowels cannot end a stressed syllable in English.
  • Languages divide the same space differently, using three qualities or twelve, and adding length, nasality or independent rounding as extra dimensions.
  • Finnish tuli, tuuli and tulli show length carrying lexical contrast, and French tu against tout shows rounding separated from backness.

Sources

  1. Wikipedia contributors. (n.d.). Vowel. Wikipedia. en.wikipedia.org
  2. Wikipedia contributors. (n.d.). Vowel diagram. Wikipedia. en.wikipedia.org
  3. International Phonetic Association. (n.d.). The IPA chart. internationalphoneticassociation.org
  4. Britannica. (n.d.). Vowel. Encyclopaedia Britannica. britannica.com
  5. UCLA Phonetics Laboratory. (n.d.). Course in phonetics: contents and sound files. University of California, Los Angeles. phonetics.ucla.edu
Key terms
Vowel height
How close the tongue body comes to the roof of the mouth, conventionally divided into close, close-mid, open-mid and open.
Vowel backness
The front-to-back position of the highest point of the tongue, divided into front, central and back.
Rounding
Whether the lips are protruded and narrowed during a vowel; independent of backness in French and Turkish but predictable from it in English.
Cardinal vowel
One of the fixed, language-independent reference points defined by Daniel Jones, used to locate the vowels of any language on a common frame.
Tense vowel
In English, a vowel that can occur at the end of a stressed syllable, such as the vowels of see, saw and sue.
Lax vowel
In English, a vowel that requires a following consonant in a stressed syllable, such as the vowels of bit, bet, book and but.
Diphthong
A vowel whose quality changes measurably during its production, such as the vowels of bite, bout and boy.
Schwa
The mid central unrounded vowel of unstressed syllables, the most frequent vowel in English.
Vowel harmony
A pattern in which vowels within a word must agree in some feature, so that suffixes change shape to match the stem, as in Turkish.

Transcribing Real Speech: Broad, Narrow, and the Diacritics That Matter

  • Produce a broad transcription of ordinary English words and sentences without being misled by spelling.
  • Add narrow detail using aspiration, nasalization, syllabicity, dentalization and release diacritics, and justify each addition.
  • Diagnose the six errors that beginning transcribers make most often.

Three words from one root

Say photograph, photography and photographic, one after another, at normal speed. Then transcribe them.

  • photograph: [ˈfoʊtəɡɹæf]
  • photography: [fəˈtɑɡɹəfi]
  • photographic: [ˌfoʊtəˈɡɹæfɪk]

The three words share a spelling, photograph, and share almost nothing phonetically. The first vowel is a full diphthong in the first and third words and a schwa in the second. The second vowel is a schwa in the first, a full open back vowel in the second, and a schwa again in the third. Nothing about this is visible in the spelling, and nothing about it is optional: move the stress and the vowels follow.

That is why transcription is a skill rather than a chore. A transcription records what a mouth did. Spelling records what a printing convention settled on, in English mostly before 1700. This lesson is about keeping the two apart under pressure.

The point: English vowel quality in unstressed syllables is determined by the stress pattern, so the same written string can be transcribed three different ways in three related words.

Broad and narrow: two settings on one instrument

A broad transcription records only the contrastive units, one symbol per phoneme, and is conventionally written between slashes. A narrow transcription adds as much predictable phonetic detail as you have chosen to record, and goes in square brackets. Neither is more correct. They answer different questions, and the right level of detail depends on what you are studying.

WordBroadNarrowWhat the narrow version added
pin/pɪn/[pʰɪ̃n]Aspiration on the stop, nasalization on the vowel
spin/spɪn/[spɪ̃n]Nasalization only; no aspiration after s
little/ˈlɪtəl/[ˈlɪɾɫ̩]Tap for t, dark and syllabic l, no schwa
button/ˈbʌtən/[ˈbʌʔn̩]Glottal stop for t, syllabic n
tenth/tɛnθ/[tʰɛn̪θ]Aspiration, and a dental n assimilated to the following th
cat/kæt/[kʰæt̚]Aspiration, and an unreleased final stop

Look at the pin and spin rows together. The broad transcriptions differ by exactly one symbol, the s, which is the only difference an English speaker chose. The narrow transcriptions differ by two things, because the aspiration comes along automatically. Every diacritic in a narrow transcription is a claim that you actually observed something, so add them deliberately, not decoratively.

The diacritics worth learning first

There are dozens. These eight will carry you through most English work.

DiacriticMeaningExample
Raised h, as in tʰAspiratedtop [tʰɑp]
Tilde above, as in æ̃Nasalizedban [bæ̃n]
Vertical stroke below, as in n̩Syllabicrhythm [ˈɹɪðm̩]
Bridge below, as in n̪Dentalwidth [wɪd̪θ]
Small circle below, as in ɹ̥Devoicedpray [pʰɹ̥eɪ]
Corner above right, as in t̚No audible releaseapt [æp̚t]
Triangular colon, as in aːLongFinnish tuuli [ˈtuːli]
Vertical line before a syllablePrimary stress, or secondary if loweredphotographic [ˌfoʊtəˈɡɹæfɪk]

Note that the dark l of feel and little is written with a separate symbol, [ɫ], rather than a diacritic, though the velarization diacritic is also available. Both are acceptable; pick one and be consistent.

Six transcriptions that are wrong, and where the reasoning failed

These are the errors that show up in the first week of every phonetics course. Each one has a diagnosis.

  1. box as [bɑx]. The letter x spells a sequence of two consonants, so the word is [bɑks]. The symbol [x] is a voiceless velar fricative, the sound in German Bach, which English does not have. Diagnosis: the transcriber copied a letter instead of listening.
  2. red as [red]. The IPA symbol [r] is an alveolar trill, the Spanish sound in perro. English red begins with an approximant, [ɹ]. Diagnosis: the transcriber assumed the IPA uses letters the way English spelling does.
  3. yes as [yɛs]. The symbol [y] is a close front rounded vowel, the vowel of French tu. The consonant at the start of yes is [j]. Diagnosis: the same assumption, in the opposite direction.
  4. sing as [sɪŋɡ]. In most varieties of English, sing ends in a single velar nasal with no following stop. But finger really is [ˈfɪŋɡɚ], with the stop. The pair singer [ˈsɪŋɚ] and finger [ˈfɪŋɡɚ] is worth memorizing, because it shows that the presence of the [ɡ] is a real fact about particular words, not about the spelling ng. Diagnosis: the transcriber generalized from the spelling to all cases.
  5. knee as [kni] and lamb as [læmb]. Both silent letters were once pronounced, and English spelling remembers them. The transcriptions are [ni] and [læm]. Diagnosis: the transcriber was recording the history of the word rather than the word.
  6. cats as [kætz]. The plural suffix is voiceless after a voiceless consonant: [kʰæts]. It is voiced after a voiced one: dogs is [dɔɡz]. This is not a spelling fact, and Module 4 derives the whole pattern from a single rule. Diagnosis: the transcriber assumed one spelling means one sound.

Bottom line: Almost every beginner error is a case of transcribing the spelling, and the cure is to say the word aloud at natural speed before writing anything down.

A worked passage

Take the sentence: the winter athlete wanted a little water. Here it is at two levels of detail, in a General American accent, spoken at conversational speed.

Broad: /ðə ˈwɪntɚ ˈæθlit ˈwɑntəd ə ˈlɪtəl ˈwɔtɚ/

Narrow: [ðə ˈwɪɾ̃ɚ ˈæθlit ˈwɑnɪd ə ˈlɪɾɫ̩ ˈwɔɾɚ]

Four things happened between the two lines, and each is a rule you can state.

  • Flapping. The t of little and water became a tap, because it sits between a stressed vowel and an unstressed one.
  • Nasal flapping and deletion. In winter the t after n is commonly lost or nasalized entirely, which is why winter and winner can be homophones for many American speakers. Note the nasalized tap in the narrow line.
  • Syllabic consonants. The final syllable of little has no vowel at all in fast speech; the dark l carries the syllable by itself.
  • Reduction. The unstressed vowel of wanted, and the article a, both reduce, and in the case of wanted the t drops out too.

Notice that the athlete kept its full form. It is stressed, and stressed syllables resist all four processes. Stress is the single best predictor of how much reduction a syllable will undergo, which is why the next lesson takes it up directly.

Transcribing yourself, not a dictionary

A dictionary transcription is a claim about a standardized reference accent. Yours is a claim about you. When the two disagree, and they will, the transcription of your own speech is not an error.

Three places where they routinely disagree. First, rhoticity: if you speak a non-rhotic variety, the r in car and card is not there, and the transcription is [kʰɑː] and [kʰɑːd]. Second, the cot and caught merger: if the two words are the same for you, use one symbol. Third, the vowel of bath: many British and Southern Hemisphere speakers have the vowel of father there, while most North American speakers have the vowel of trap.

The discipline is the same one from Lesson 1. Brackets record what happened. If you want to record the reference accent instead, say so, and use its transcription throughout rather than mixing the two, which produces a transcription of nobody.

Common misconceptions

  • A narrow transcription is a better transcription. A narrow transcription is a more detailed one. If the detail is not relevant to the question you are asking, it is clutter, and every diacritic you add is a claim you must be able to defend.
  • Transcription is a way of writing pronunciation for learners. That is one use of it. Its primary use in linguistics is as data: the transcription is the observation on which analysis is built.
  • Fast speech forms are sloppy and should be transcribed in their careful form. Fast speech is systematic, and its systematicity is precisely what phonology studies. Transcribing what a speaker would have said if they had been reading aloud discards the data.
  • The IPA has one right symbol for every English sound. Several English sounds have competing conventions, notably the vowels of bait and boat and the English r. What matters is declaring your convention and holding to it.
  • Stress marks are optional decoration. In English they are load-bearing, because vowel reduction follows stress. A transcription of a polysyllabic English word without stress marked is incomplete.

Recap

  • Photograph, photography and photographic share a spelling and differ in almost every vowel, because English vowel quality follows stress.
  • Broad transcription records contrastive units between slashes; narrow transcription adds observed detail between brackets, and each added diacritic is a claim.
  • Eight diacritics carry most English work: aspiration, nasalization, syllabicity, dentalization, devoicing, no audible release, length and stress.
  • The six standard beginner errors are all versions of one mistake, transcribing the spelling instead of the speech.
  • Singer against finger shows that the presence of a stop after the velar nasal is a fact about particular words.
  • Flapping, nasal deletion, syllabic consonants and vowel reduction all target unstressed material and leave stressed syllables alone.
  • Transcribe your own variety, declare your conventions, and do not mix a reference accent with your own.

Sources

  1. International Phonetic Association. (n.d.). IPA chart with diacritics and suprasegmentals. internationalphoneticassociation.org
  2. Wikipedia contributors. (n.d.). English phonology. Wikipedia. en.wikipedia.org
  3. Wikipedia contributors. (n.d.). Flapping. Wikipedia. en.wikipedia.org
  4. Interactive IPA Chart. (n.d.). IPA chart with audio. ipachart.com
  5. Ladefoged, P., & Johnson, K. (2015). A course in phonetics (7th ed.). Cengage Learning.
Key terms
Broad transcription
Transcription recording only contrastive units, one symbol per phoneme, conventionally enclosed in slashes.
Narrow transcription
Transcription that adds observed predictable detail using diacritics, enclosed in square brackets.
Diacritic
A mark added to an IPA symbol to record a modification such as aspiration, nasalization, dentalization or syllabicity.
Syllabic consonant
A consonant that forms the nucleus of a syllable with no vowel, marked with a vertical stroke below, as in the second syllable of rhythm.
Flapping
The realization of English t and d as a single fast tap between a stressed and an unstressed vowel, as in water and little.
Vowel reduction
The shift of an unstressed vowel towards schwa, which is why the same written syllable is pronounced differently as stress moves.
Rhoticity
Whether a variety of English pronounces r after a vowel; non-rhotic varieties do not, so car and card lack the consonant.
Dark l
The velarized lateral of syllable codas in English, as in feel and little, written with a crossed l or with a velarization diacritic.

Above the Segment: Stress, Tone, and Intonation

  • Name the four acoustic correlates of English stress and rank them by how strongly they cue it.
  • Read and produce the four Mandarin tones, and distinguish level from contour tone systems.
  • Explain how tone, stress and intonation coexist in one utterance without conflicting.

Seven meanings, one sentence

Read this sentence aloud seven times, each time leaning on a different word: I never said she stole my money.

  • I never said it, though someone did.
  • I never said it, not once.
  • I never said it, though I may have implied it.
  • I never said she stole it, I said someone else did.
  • I never said she stole it, only that she took it.
  • I never said she stole my money, it was somebody else's.
  • I never said she stole my money, she stole something else.

Seven different claims, seven different legal positions, out of an identical string of words. The segments did not change. What changed rides on top of them: pitch, loudness, duration. These are the suprasegmentals, and they carry a startling proportion of what a sentence actually communicates.

Key idea: Properties that extend over syllables and phrases rather than over single segments can change meaning by themselves, which is why a full description of speech cannot stop at the consonant and vowel.

What stress physically is

Ask a speaker what they do to stress a syllable and they will usually say they say it louder. Measurement says otherwise, or at least says loudness is the least of it. Four properties distinguish a stressed English syllable from an unstressed one.

CorrelateWhat happensHow strongly it cues stress
Fundamental frequencyA pitch movement, usually a rise or a peak, lands on the syllableStrong
DurationThe syllable, especially its vowel, is longerStrong
Vowel qualityThe vowel keeps its full quality instead of reducing to schwaStrong in English
IntensityThe syllable is somewhat louderWeakest of the four

Dennis Fry established the ranking experimentally in the 1950s by synthesizing noun and verb pairs with the correlates manipulated independently, and finding that duration and fundamental frequency shifted listeners' judgments far more reliably than intensity did. The finding matters practically: if you are trying to hear stress in an unfamiliar language, listen for length and pitch movement rather than for volume.

Vowel quality is the correlate English leans on hardest, and it is why the previous lesson kept returning to reduction. In a language like Spanish, whose unstressed vowels keep their quality, stress has to be carried by duration and pitch alone, which is one reason Spanish stress is harder for English speakers to hear.

English noun and verb pairs, and what they prove

English stress is not merely decorative. It distinguishes words.

SpellingNounVerb
permit[ˈpʰɝmɪt][pʰɚˈmɪt]
record[ˈɹɛkɚd][ɹɪˈkʰɔɹd]
contract[ˈkʰɑntɹækt][kʰənˈtɹækt]
insult[ˈɪnsʌlt][ɪnˈsʌlt]
rebel[ˈɹɛbəl][ɹɪˈbɛl]

Look at what follows the stress shift. It is never only the stress; the vowels change too, because reduction targets unstressed syllables. That is the point of the table: stress in English is the switch, and vowel quality is the visible consequence.

The same works above the word. A greenhouse, stressed on the first element, is a glass building for plants. A green house, stressed on the second, is a house that is green. A blackbird is a species; a black bird is any bird that is black. English compounds take stress on the first element and phrases take it on the second, reliably enough that the pattern is a diagnostic for compoundhood.

Why this matters: Stress in English is contrastive, distinguishing both word classes and compound from phrase, which means it belongs in a transcription rather than being left to the reader.

Tone: pitch inside the word

In English, changing the pitch of a syllable changes the attitude, not the word. In a tone language it changes the word. Mandarin is the standard demonstration, because one syllable carries four contrastive tones.

PinyinCharacterTone shapeChao numeralsMeaning
High level55mother
Rising35hemp
Low dipping214horse
Falling51to scold

The Chao numerals in the fourth column are a notation invented by the linguist Yuen Ren Chao: the speaker's pitch range is divided into five levels, 1 for the lowest and 5 for the highest, and a tone is written as the sequence of levels it passes through. So 35 is a rise from mid to high, and 214 is a fall to the bottom followed by a rise. The system is language-independent, which is why it is used for tone the way the cardinal vowels are used for vowel quality.

Tone is not rare. In the World Atlas of Language Structures sample of 527 languages, roughly 42 percent have some form of tone, with heavy concentrations in sub-Saharan Africa, East and Southeast Asia, and parts of the Americas. English speakers who find tone exotic are the ones in the minority.

Two kinds of tone system

Systems divide into level or register systems, where each tone is a pitch height, and contour systems, where a tone is a movement.

Yoruba is the classic level system: three tones, high, mid and low, written with an acute accent, no accent and a grave accent respectively. Igbo works with two. In such languages a word's tone pattern is as much a part of it as its consonants, and dictionaries mark it.

Mandarin, Cantonese, Thai and Vietnamese are contour systems. Cantonese has six tones, Thai has five and Vietnamese has six. Thai gives a compact demonstration on one syllable: มา, with mid tone, means come; หมา, with rising tone, means dog; ม้า, with high tone, means horse. Three animals and a verb separated by nothing but pitch.

The division is not absolute. Many systems mix level and contour tones, and contour tones often turn out, on analysis, to be sequences of level tones associated to one syllable, which is the insight behind the autosegmental treatment you will meet in Module 6.

Intonation: pitch over the phrase

Intonation is pitch used at the level of the phrase or utterance rather than the word. English uses it for at least four jobs.

  • Sentence type. The string you are coming falls at the end as a statement and rises at the end as a question. Nothing else changes, and the difference is not optional.
  • Focus. The seven readings that opened this lesson are all done with intonation: a pitch accent lands on the focused word.
  • Grouping. Intonation phrases mark syntactic boundaries. Compare my brother who lives in Leeds called, with no break, which implies other brothers, against my brother, who lives in Leeds, called, with breaks, which does not.
  • Continuation. In a list, non-final items typically rise and the final item falls, which is how a listener knows the list has ended before the speaker says so.

Languages differ in the details, and the differences are a common source of cross-cultural misreading. A pattern that signals politeness in one language can read as impatience or sarcasm in another, and unlike a mispronounced consonant, the error is usually attributed to the speaker's character rather than to their accent.

Three systems, one pitch contour

An obvious question: if Mandarin uses pitch for tone, how can it also use pitch for intonation? The answer is that they operate at different scales and are superimposed. Tone controls the shape within the syllable; intonation shifts the overall register, compresses or expands the range, and adjusts the height of the whole contour. A Mandarin question raises and expands the pitch range across the utterance while each syllable keeps its lexical tone shape inside that adjusted frame. Speakers of tone languages have no difficulty being sarcastic.

Rhythm is worth one honest note. Languages have traditionally been sorted into stress-timed (English, German, Dutch), syllable-timed (Spanish, French, Italian) and mora-timed (Japanese), with the idea that the relevant units recur at equal intervals. Attempts to find that isochrony in measurements have largely failed: the intervals are not equal in any of the three types. What survives is a real perceptual difference driven by whether the language reduces unstressed vowels and how much its syllable structures vary. The categories are a useful shorthand for something real, but not the clockwork the names suggest.

Common misconceptions

  • Stress is loudness. Duration, pitch movement and vowel quality all cue stress more reliably than intensity does, which is why whispered speech still has audible stress.
  • Tone languages sound like singing. The pitch range used for lexical tone is ordinary speech range. Tone languages sound like speech, and their speakers are not doing anything effortful.
  • Tone is a primitive or exotic feature. Around 42 percent of a large representative sample of languages use tone, including languages with tens of millions of speakers and long literary traditions.
  • A tone language cannot have intonation. It has both, layered: tone determines the shape within a syllable, intonation adjusts the register and range of the whole phrase.
  • Intonation is emotional rather than linguistic. Some of it is, but sentence type, focus and phrase grouping are grammatical functions with systematic form, and getting them wrong changes what a sentence means, not merely how it feels.

Where this leaves us

  • Suprasegmental properties extend over syllables and phrases and can change meaning on their own, as the seven readings of one sentence show.
  • English stress is cued by pitch movement, duration and full vowel quality much more than by loudness.
  • Stress is contrastive in English, separating noun and verb pairs like permit and permit, and compounds like greenhouse from phrases like green house.
  • In a tone language pitch is part of the word: Mandarin distinguishes mother, hemp, horse and to scold on one syllable.
  • Level systems such as Yoruba assign pitch heights; contour systems such as Mandarin and Thai assign pitch movements, and Chao numerals notate either.
  • Intonation handles sentence type, focus, grouping and continuation, and coexists with lexical tone by operating on the whole phrase.
  • The stress-timed and syllable-timed labels describe something real about vowel reduction and syllable structure, but not literal equal timing.

Sources

  1. Wikipedia contributors. (n.d.). Tone (linguistics). Wikipedia. en.wikipedia.org
  2. Wikipedia contributors. (n.d.). Intonation (linguistics). Wikipedia. en.wikipedia.org
  3. Wikipedia contributors. (n.d.). Standard Chinese phonology. Wikipedia. en.wikipedia.org
  4. Wikipedia contributors. (n.d.). Prosody (linguistics). Wikipedia. en.wikipedia.org
  5. Fry, D. B. (1955). Duration and intensity as physical correlates of linguistic stress. Journal of the Acoustical Society of America, 27(4), 765-768.
Key terms
Suprasegmental
A property such as stress, tone, length or intonation that extends over more than a single segment.
Stress
The relative prominence of a syllable, cued in English by pitch movement, duration and unreduced vowel quality more than by loudness.
Tone
The use of pitch to distinguish words or grammatical forms, present in roughly two fifths of the world's languages.
Level tone system
A tone system whose tones are pitch heights, such as the high, mid and low of Yoruba.
Contour tone system
A tone system whose tones are pitch movements, such as the rising and falling tones of Mandarin and Thai.
Chao tone numerals
A notation dividing a speaker's pitch range into five levels and writing a tone as the sequence of levels it traverses.
Intonation
The use of pitch over a phrase or utterance to mark sentence type, focus, grouping and continuation.
Pitch accent
A pitch movement associated with a prominent syllable, which in English marks the focused word in an utterance.

Module 3: Acoustics and Perception

What speech looks like when you plot it: waveforms, spectrograms and the cues that identify segments; the source-filter model and the formants that make a vowel; and what listeners actually do with the signal, from categorical perception to the McGurk effect.

Reading Waveforms and Spectrograms

  • Identify voicing, silence and frication in a waveform, and explain what each looks like and why.
  • Read a wideband spectrogram: locate bursts, aspiration, formants, nasal murmur and fricative noise.
  • Choose between wideband and narrowband analysis for a given question and say what each one costs.

What the word cats looks like

Open a spectrogram of a single English word, cats, and read it left to right. You see a blank rectangle lasting perhaps 70 milliseconds: nothing is happening, because the tongue back is sealed against the velum. Then a thin vertical spike, the burst of the released stop. Then about 50 milliseconds of grey hash spread up the frequency scale, which is the aspiration. Then the signal turns dark and organized: three or four heavy horizontal bands, ruled through by closely spaced vertical lines, running for roughly 150 milliseconds. That is the vowel. Then another blank rectangle, shorter this time, as the tongue tip seals the alveolar ridge, another spike, and finally 120 milliseconds of dense noise piled up above 4000 Hz, which is the s.

Everything in that paragraph is a physical measurement, and every claim in it is checkable by anyone with the same recording. That is the point of acoustic phonetics: it replaces what you think you heard with something on a screen that two people can argue about.

The tool that made this possible was built at Bell Telephone Laboratories during the Second World War, and the results were published in 1947 by Ralph Potter, George Kopp and Harriet Green under the title Visible Speech. Their hope was that deaf readers might learn to read spectrograms as text. That did not work, for reasons this lesson ends with, but the instrument transformed the field.

Key idea: A spectrogram converts the invisible into a picture with axes, so that claims about speech become measurements rather than impressions.

The waveform: pressure against time

Start simpler. A waveform plots air pressure against time: how far the air at the microphone was pushed above or below its resting value, moment by moment. Three patterns cover most of what you need.

What you seeWhat it isWhy
A repeating shape, the same bump recurring at even intervalsVoicingEach repetition is one cycle of the vocal folds; the interval between them is the period, and its reciprocal is the fundamental frequency
Dense, irregular, small-amplitude scribbleFrication or aspirationTurbulent airflow is aperiodic by nature, so no shape repeats
A flat line, or near itA stop closure, or a pauseThe tract is sealed, so no pressure variation reaches the microphone

You can measure fundamental frequency directly off a waveform without any further processing. Find two adjacent identical points, measure the interval, and invert it. If successive cycles are 8 milliseconds apart, the period is 0.008 seconds and the frequency is 125 Hz. This is worth doing by hand once, because it makes the abstraction concrete: pitch is a count of repetitions per second, nothing more.

What the waveform is bad at is telling you which vowel you are looking at. Two different vowels at the same pitch produce repeating shapes at the same rate, differing only in the fine detail inside each cycle, and the eye is poor at that. For quality you need the frequency axis.

From waveform to spectrogram

A spectrogram is built by taking a short slice of the waveform, working out how much energy it contains at each frequency, drawing that as a vertical strip of light and dark, sliding the slice along and repeating. The result has three dimensions on a flat page: time runs left to right, frequency runs bottom to top, and amplitude is darkness. Speech work usually shows 0 to 5000 Hz, because that band contains almost everything that distinguishes one English segment from another.

The dark horizontal bands in a vowel are formants, the resonances of the vocal tract, numbered upward from F1. They are the subject of the next lesson. For now, treat them as the visible signature of vowel quality: they move when the tongue moves, and their pattern is what lets you name a vowel from a picture.

Wideband and narrowband: a real trade-off

You cannot have precise timing and precise frequency at once. Analyze a longer slice of signal and you resolve frequency finely but smear events in time; analyze a shorter slice and you locate events precisely but cannot separate nearby frequencies. Every spectrogram is a choice on that scale, and the two standard settings have names.

WidebandNarrowband
Analysis bandwidthAbout 300 HzAbout 45 Hz
Window lengthAround 3 to 5 msAround 20 to 30 ms
Vertical striationsYes, one per glottal pulseNo
Horizontal harmonicsNoYes, evenly spaced, one per multiple of the fundamental
Best forFormants, bursts, segment boundaries, timingReading fundamental frequency and studying pitch contours

The vertical striations on a wideband spectrogram are individual glottal pulses. If you can count them, you can count the fundamental frequency: striations 8 milliseconds apart mean 125 pulses per second. On a narrowband spectrogram you get the same information the other way round, from the spacing of the horizontal harmonic lines, since they sit at every multiple of the fundamental. Two routes to one number, and each is a good check on the other.

The upshot: Wideband is the default for segmental work because timing and formants matter most there; narrowband is for questions about pitch.

A field guide to the marks

Here is what to look for, in the order you will meet these on a typical utterance.

Segment typeSignature on a wideband spectrogram
Voiceless stopA blank gap, then a narrow vertical spike at the release
Voiced stopThe same gap, but with a dark band at the very bottom, below about 250 Hz: the voice bar
AspirationGrey noise after the burst, shaped by the formants of the following vowel
VowelStrong dark formant bands with vertical striations; F1 and F2 clearly separable
The fricative sIntense noise concentrated above roughly 4000 Hz
The fricative shIntense noise starting lower, from roughly 2500 Hz, and often louder overall
The fricatives f and thWeak, diffuse noise with no clear concentration; the hardest segments to spot
NasalA sudden drop in energy, a low murmur band near 250 to 300 Hz, and faint or missing higher formants
ApproximantFormants that glide smoothly, with no noise and no break in voicing

Place of articulation for stops is read not from the burst but from the formant transitions into the following vowel. As the tongue moves from the closure into the vowel, the formants sweep from a starting point characteristic of the place. Labials show a low second formant at release, rising into the vowel. Alveolars point towards a locus somewhere near 1700 to 1800 Hz. Velars show a distinctive convergence of the second and third formants at the release, close enough together that phoneticians call it the velar pinch, and once you have seen it once you never miss it.

Reading an utterance end to end

Take the phrase see the man. Working left to right, expect: intense high-frequency noise for the s, lasting well over 100 milliseconds; an abrupt switch to a voiced region with a very low first formant and a very high second formant, which is the vowel of see; a short weak fricative with a voice bar, the th; a mid central vowel with F1 and F2 close together, which is the reduced vowel of the; a low murmur and a sudden loss of higher energy, the m; a vowel with a high first formant and a mid second formant, and visibly nasalized at both ends, the vowel of man; and finally another murmur for the n.

Now notice what you cannot see anywhere in that description: a gap between the words. There is no acoustic marker for a word boundary in ordinary connected speech. The phrase see the man and the invented phrase seethe uh man have essentially the same signal. Word boundaries are supplied by the listener, not by the talker, and that fact drives most of the next two lessons.

What a spectrogram will not tell you

Potter, Kopp and Green hoped spectrograms could be read like text. They cannot, and the reasons are instructive.

  • There is no invariance. The same phoneme has different acoustic shapes in different contexts. The burst spectrum for a k before the vowel of key is nothing like the burst for a k before the vowel of caw, because the tongue is in a different place.
  • There is no segmentation. Speech is not a row of beads. Because articulators move continuously and overlap, information about a consonant is smeared across the neighbouring vowel, and there is often no point at which one segment can be said to end.
  • Speakers differ. A child's formants sit far higher than an adult man's for the same vowel, and a spectrogram does not normalize for that. Listeners do it effortlessly and the machine does not.

Skilled spectrogram readers do exist, and they get impressively far, but they work the way a codebreaker works, using context, statistics and knowledge of the language, rather than reading symbols off a page.

Common misconceptions

  • Darker means higher pitch. Darkness is amplitude at that frequency. Pitch is read from the striation spacing on a wideband display or the harmonic spacing on a narrowband one.
  • Formants are harmonics. Harmonics come from the source, at multiples of the fundamental, and move up and down with pitch. Formants come from the filter, depend on the shape of the tract, and stay put when you change pitch on the same vowel.
  • Silence on a spectrogram means silence in speech. Silence is usually a stop closure, an active articulation, not the absence of speech. Deleting it makes the utterance unintelligible.
  • You can read a spectrogram off like text. The lack of invariance and the lack of segmentation defeat symbol-by-symbol reading, which is why the visible speech project did not achieve its original goal.
  • Higher sampling or better software would make word boundaries visible. The boundaries are genuinely absent from the signal in connected speech; better instruments cannot recover information that was never encoded.

Pulling it together

  • A waveform shows periodicity for voicing, aperiodic noise for frication, and near-silence for a stop closure, and fundamental frequency can be measured off it directly.
  • A spectrogram puts time on one axis and frequency on the other, with darkness as amplitude, usually over 0 to 5000 Hz for speech.
  • Wideband and narrowband trade time resolution against frequency resolution; the first shows glottal pulses and formants, the second shows harmonics and pitch.
  • Segments have characteristic signatures: bursts and gaps for stops, high noise for s, low murmur for nasals, gliding formants for approximants.
  • Stop place is read from formant transitions rather than from the burst, with the converging second and third formants of the velar pinch as the clearest case.
  • Connected speech contains no acoustic word boundaries, and the lack of invariance and of segmentation is why spectrograms cannot be read as text.

Sources

  1. Wikipedia contributors. (n.d.). Spectrogram. Wikipedia. en.wikipedia.org
  2. Wikipedia contributors. (n.d.). Acoustic phonetics. Wikipedia. en.wikipedia.org
  3. Boersma, P., & Weenink, D. (n.d.). Praat: doing phonetics by computer. University of Amsterdam. fon.hum.uva.nl
  4. UCLA Phonetics Laboratory. (n.d.). UCLA Phonetics Lab Archive. University of California, Los Angeles. archive.phonetics.ucla.edu
  5. Potter, R. K., Kopp, G. A., & Green, H. C. (1947). Visible speech. D. Van Nostrand Company.
Key terms
Waveform
A plot of air pressure against time, in which repeating cycles indicate voicing and irregular scribble indicates turbulent noise.
Spectrogram
A display with time on the horizontal axis, frequency on the vertical axis and amplitude shown as darkness.
Formant
A resonance of the vocal tract, visible as a dark horizontal band, numbered upward from the lowest as F1, F2 and so on.
Wideband spectrogram
An analysis with a broad filter and short window, giving good time resolution, visible glottal striations and clear formants.
Narrowband spectrogram
An analysis with a narrow filter and long window, resolving individual harmonics and making the fundamental frequency easy to read.
Voice bar
A dark band of energy at the very bottom of a spectrogram during a closure, indicating that the vocal folds are still vibrating.
Velar pinch
The convergence of the second and third formants at the release of a velar stop, a reliable visual cue to that place of articulation.
Lack of invariance
The fact that one phoneme has no single constant acoustic shape, since its realization changes with the surrounding sounds.
Segmentation problem
The absence of clean acoustic boundaries between segments and between words in connected speech.

Formants and the Source-Filter Model

  • Explain the source-filter model and give three pieces of evidence that the source and the filter are independent.
  • Calculate the resonances of a uniform tube and relate them to the formants of a neutral vowel.
  • Read the Peterson and Barney formant table, plot vowels in F1 and F2 space, and explain why the categories overlap.

Seventy-six speakers and ten words

In 1952 Gordon Peterson and Harold Barney published measurements of a data set that is still being used. They recorded 76 speakers, 33 men, 28 women and 15 children, each saying ten English words twice: heed, hid, head, had, hod, hawed, hood, who would, hud and heard. Every word had the same frame, an h before and a d after, so that whatever differed between them was the vowel and nothing else. Then they measured the first three formants of each vowel, and played all 1520 utterances to a panel of listeners.

Two results came out, and they pull in opposite directions. The first: plot each vowel by its first formant against its second and the ten vowels fall into ten distinguishable regions, arranged in a shape that looks strikingly like the articulatory vowel quadrilateral drawn upside down. The second: the regions overlap. One speaker's hid can sit exactly where another speaker's head sits. Yet listeners identified the intended vowel in the great majority of trials.

That combination, a clear acoustic pattern that nonetheless fails to separate the categories cleanly, is the central fact of vowel acoustics, and this lesson builds the model that explains it.

Key idea: Vowel quality is carried by the first two formants, but the mapping from formants to vowel category is not one to one across speakers, which forces the listener to do work.

Two things happening at once

The model that organizes all of this is the source-filter model, set out most fully by Gunnar Fant in 1960. Its claim is that speech production has two stages that can be described, and manipulated, independently.

  • The source is what the larynx produces: for voiced sounds, a train of pulses as the vocal folds slam shut, repeating at the fundamental frequency. Such a pulse train is rich in harmonics, containing energy at every whole multiple of the fundamental, with the higher harmonics weaker, falling off at roughly 12 decibels per octave. If you could hear it directly it would be a flat, buzzy rasp with no vowel quality at all.
  • The filter is the vocal tract above the larynx, a tube of air with resonances of its own. Frequencies near a resonance pass through strongly; frequencies between resonances are damped. Those resonances are the formants.

Three observations show that the two really are separable, and they are more convincing than any argument.

  1. Whisper. Whispering removes the periodic source entirely and replaces it with noise at the glottis. The filter is unchanged, so the formants are still there, and whispered vowels remain perfectly identifiable. Vowel quality therefore does not depend on the source.
  2. Singing a scale on one vowel. Sing a long ah up an octave. The fundamental doubles; the vowel does not change, because you did not move your tongue. Source varies, filter constant.
  3. Helium. Breathing helium raises the speed of sound in the vocal tract, which pushes all the resonances upward while leaving vocal fold vibration largely as it was. The result is the familiar cartoon voice. Notice what this means: the effect is not that the speaker's pitch went up. Their filter changed.

A tube 17.5 centimetres long

Now the useful part, which is that you can calculate the formants of a simple vocal tract with arithmetic you already have.

Model the tract as a uniform tube of air, closed at the glottis and open at the lips, about 17.5 centimetres long in an adult man. Such a tube is a quarter-wave resonator: its lowest resonance has a wavelength four times its length, and further resonances occur at odd multiples. Take the speed of sound as 35,000 centimetres per second.

First resonance: 35,000 divided by (4 times 17.5) equals 35,000 divided by 70, which is 500 Hz. The next two are three times and five times that: 1500 Hz and 2500 Hz.

Those three numbers are close to the measured formants of a schwa, the vowel you make when the tongue does nothing in particular. That is the model earning its keep: an idealized tube of the right length predicts the neutral vowel of a real speaker to within a few tens of hertz.

Run it again for a vocal tract of 14.5 centimetres, typical for an adult woman. 35,000 divided by 58 is about 603 Hz, with further resonances near 1810 and 3020 Hz. So the same neutral vowel has formants roughly 20 percent higher, purely because the tube is shorter. No difference in effort, intention or language is involved, and this is the root of the normalization problem below.

What matters here: Formant frequencies scale inversely with vocal tract length, so a woman's and a child's formants for the same vowel are systematically higher than a man's.

How the tongue moves the resonances

Constricting the tube at different points shifts the resonances in predictable directions, and two rules of thumb cover most of what you need.

  • F1 varies inversely with tongue height. High vowels have low first formants; low vowels have high ones. The vowel of heed has F1 near 270 Hz; the vowel of had has F1 near 660 Hz.
  • F2 varies with tongue backness. Front vowels have high second formants; back vowels have low ones. The vowel of heed has F2 near 2290 Hz; the vowel of who would has F2 near 870 Hz.
  • Lip rounding lowers all formants, because protruding the lips lengthens the tube.

Because F1 goes up as the vowel goes down, and F2 goes up as the vowel goes forward, a plot with F1 increasing downward and F2 increasing leftward reproduces the articulatory quadrilateral. That is the standard vowel plot, and it is why acoustic and articulatory descriptions agree without either having been fitted to the other.

The numbers, and what to do with them

Here are the Peterson and Barney mean values for their adult male speakers, in hertz.

WordVowelF1F2F3
heedi27022903010
hidɪ39019902550
headɛ53018402480
hadæ66017202410
hodɑ73010902440
hawedɔ5708402410
hoodʊ44010202240
who wouldu3008702240
hudʌ64011902390
heardɝ49013501690

Work the table rather than reading it. Go down the first four rows: F1 climbs from 270 to 660 while F2 falls gently. Those are the front vowels, descending the front edge of the quadrilateral. Now compare heed and who would: nearly the same F1, 270 against 300, because both are high vowels, but F2 differs by more than 1400 Hz, because one is front and the other back and rounded. And look at the last row: heard is the only vowel with a low third formant, near 1690 Hz, which is the acoustic signature of American r and the reason F3 is worth measuring at all.

Overlap, and the problem it creates

Now the awkward result. Take the mean values above and compare them to a child's. A child's tract may be 11 centimetres long, so the same vowel comes out with formants perhaps 40 percent higher. A child's hood can land, in absolute hertz, exactly where an adult man's hud does. If listeners simply matched incoming formants to stored numbers, they would be wrong constantly. They are not.

What listeners appear to do is normalize, judging a vowel relative to the speaker rather than absolutely. Several cues feed this: the fundamental frequency, which correlates with body size; the third formant, which is relatively stable within a speaker; and the range of vowels heard earlier from the same voice. A classic demonstration is that the identity of an ambiguous vowel can be flipped by changing the formants of the carrier sentence around it, so the same acoustic token is heard as one word after one voice and another word after another.

This is the first appearance of a theme that dominates the next lesson: the listener is not a passive measuring device. The signal underdetermines the message, and perception fills the gap.

Common misconceptions

  • Formants are what the vocal folds produce. The folds produce the source. Formants are properties of the tube above them, which is why whispering, with no fold vibration at all, still yields identifiable vowels.
  • A higher voice has higher formants. Pitch and formants are independent. A bass singing high and a soprano singing low can have their formant patterns the other way round from their fundamentals.
  • Helium makes your voice higher. It raises the speed of sound and therefore the resonances of the filter. The vocal folds keep vibrating at close to their usual rate.
  • Vowel categories are separated cleanly in the acoustics. Peterson and Barney's own data show the regions overlapping, which is precisely why normalization is required and why formant values alone cannot label a vowel.
  • You need F3 and above to identify a vowel. F1 and F2 carry nearly all the vowel quality information in English. F3 matters chiefly for r-coloured vowels, where it drops sharply.

The takeaway

  • Peterson and Barney measured ten vowels from 76 speakers in a fixed frame, and found distinguishable but overlapping regions in formant space.
  • The source-filter model separates the glottal pulse train, with its harmonics and its fundamental, from the resonances of the tract above.
  • Whisper, singing a scale on one vowel, and the helium effect are three independent demonstrations that source and filter can vary separately.
  • A uniform tube of 17.5 centimetres has resonances at 500, 1500 and 2500 Hz, which is close to the measured schwa of an adult male speaker.
  • F1 varies inversely with tongue height, F2 with backness, and rounding lowers everything, which is why an F1 by F2 plot reproduces the vowel quadrilateral.
  • Shorter tracts give proportionally higher formants, so vowel categories overlap across speakers and listeners must normalize rather than match absolute values.

Sources

  1. Wikipedia contributors. (n.d.). Formant. Wikipedia. en.wikipedia.org
  2. Wikipedia contributors. (n.d.). Source-filter model. Wikipedia. en.wikipedia.org
  3. Wikipedia contributors. (n.d.). Fundamental frequency. Wikipedia. en.wikipedia.org
  4. Boersma, P., & Weenink, D. (n.d.). Praat: doing phonetics by computer. University of Amsterdam. fon.hum.uva.nl
  5. Peterson, G. E., & Barney, H. L. (1952). Control methods used in a study of the vowels. Journal of the Acoustical Society of America, 24(2), 175-184.
  6. Fant, G. (1960). Acoustic theory of speech production. Mouton.
Key terms
Source-filter model
The account of speech production that separates the glottal sound source from the resonant filtering of the vocal tract above it.
Glottal source
The pulse train produced by vocal fold vibration, containing energy at every multiple of the fundamental frequency and falling off with height.
Quarter-wave resonator
A tube closed at one end and open at the other, whose lowest resonance has a wavelength four times its length.
F1
The lowest vocal tract resonance, which varies inversely with tongue height: low for close vowels, high for open ones.
F2
The second resonance, which is high for front vowels and low for back vowels, and is lowered further by lip rounding.
F3
The third resonance, largely uninformative for most vowels but dropping sharply for r-coloured vowels such as the one in heard.
Normalization
The listener's adjustment of vowel judgments to the speaker's vocal tract, using cues such as fundamental frequency and previously heard vowels.
Vowel plot
A graph of F1 against F2, conventionally with both axes reversed, which reproduces the shape of the articulatory vowel quadrilateral.

How Listeners Hear: Categorical Perception and the McGurk Effect

  • State the two criteria for categorical perception and say what the classic identification and discrimination curves look like.
  • Explain the McGurk effect, infant discrimination and perceptual narrowing, and what each shows about the listener.
  • Lay out the case for and against the motor theory of speech perception and name the evidence that bears on it.

Fourteen synthetic syllables

In 1957 Alvin Liberman and colleagues at Haskins Laboratories built a set of fourteen artificial syllables. Each was identical except for the starting frequency of the second formant, which stepped in equal increments from one end of the set to the other. Physically, step 3 differs from step 4 by exactly as much as step 8 differs from step 9. Played to listeners, the fourteen sounded like a run of syllables beginning b, then d, then g.

Two things happened when the researchers tested them. Asked to label each stimulus, listeners did not report a gradual shift. They said b, b, b, b, b, then abruptly d, d, d, and then abruptly g. The identification curve had steep walls. Then, asked whether two stimuli were the same or different, listeners were excellent at telling apart a pair that straddled one of those walls and close to hopeless at telling apart an equally different pair that sat inside a single category.

Equal physical steps, wildly unequal perceptual consequences. The pattern was named categorical perception, and it launched fifty years of argument about what a listener is doing.

Key idea: Discrimination of speech sounds is predicted not by the size of the acoustic difference but by whether the two sounds fall on opposite sides of a category boundary.

The criteria, stated so they can be tested

Categorical perception is a claim with two testable parts, and it is worth stating them precisely because the term gets used loosely.

  1. The identification function is sharp. Plot the percentage of trials labelled b against stimulus number and you get a curve that stays near 100 percent, drops steeply over one or two steps, and stays near zero, rather than sloping evenly.
  2. Discrimination is predictable from identification. Listeners can distinguish two stimuli only about as well as they can label them differently. Within a category, discrimination falls towards chance; across the boundary, it peaks.

The second criterion is the surprising one. For most perceptual dimensions it is false. You can easily tell two shades of blue apart even when you call them both blue, and two pitches apart even when you call them both high. Speech, at least for consonants, behaves differently.

The original strong claim has since been softened, and honestly. Using more sensitive tasks, including speeded same-different judgments, later researchers found that within-category discrimination is reliably better than chance: listeners retain some access to the acoustic detail they were said to have discarded. Categorical perception is real but it is a matter of degree, and it is much stronger for stop consonants than for vowels, which are perceived far more continuously.

Moving the boundary

The same design works on voice onset time, and this is where it becomes a fact about languages rather than about ears. Synthesize a continuum from a prevoiced stop through short lag to long lag, in ten millisecond steps, and ask listeners where b becomes p.

An English listener puts the boundary somewhere around plus 25 to 30 milliseconds. A Spanish listener puts it close to zero. Both are listening to the same stimuli. Both show sharp identification functions and discrimination peaks; what differs is where the wall stands, and it stands where their language put it. Module 1 predicted exactly this from the production data, and the perception data confirms it independently.

The point: The boundary is learned, not given, which is why the same continuum is heard differently by speakers of English and Spanish while both show equally categorical behaviour.

Hearing with your eyes

In 1976 Harry McGurk and John MacDonald published a two-page paper in Nature with a title that says the whole thing: Hearing lips and seeing voices. They dubbed an audio recording of the syllable ba onto a video of a face articulating ga. Adult viewers, watching the face while hearing the sound, overwhelmingly reported hearing neither: they reported da.

The result is not a trick of attention or an error of report. Close your eyes and you hear ba, correctly. Open them and you hear da, and you cannot stop yourself. Knowing exactly how the illusion is constructed does not weaken it, which sets it apart from most visual illusions and makes it uncomfortable to sit with. The perceptual system integrates visual and auditory information about the same articulatory event before anything reaches awareness, and it does not offer you the raw audio.

The fusion is not arbitrary. The lips visibly close for ba, so a listener who sees no closure has evidence against a labial. The auditory signal is inconsistent with a velar. What comes out is the alveolar compromise, which fits both sources tolerably. Perception is behaving like an inference from imperfect evidence, not like a transcription of the acoustic input.

What infants can do, and what they stop doing

If category boundaries are learned, when? Peter Eimas and colleagues answered part of this in 1971 with a method that works on babies. Infants suck harder on a pacifier when they hear something new; the rate drops as they habituate. Present a syllable repeatedly until sucking declines, then switch to a stimulus 20 milliseconds away on the VOT continuum. If sucking rises, the infant noticed.

One-month-old and four-month-old infants noticed a 20 millisecond change that crossed the adult English boundary, and did not notice an equally large change that stayed inside a category. Infants far too young to have learned English words were already treating the continuum categorically.

Then comes the striking part. Janet Werker and Richard Tees showed in 1984 that infants in English-speaking homes at six to eight months could discriminate contrasts English does not use, including a Hindi dental against retroflex distinction and a velar against uvular distinction in Nthlakapmx, a Salish language of British Columbia. By ten to twelve months, the same infants could no longer do it, while infants learning Hindi or the Salish language still could.

So the first year of life is not a period of acquiring distinctions. It is largely a period of losing the ones the ambient language does not use, tuning a general capacity into a specific one. Patricia Kuhl later described a related finding, the perceptual magnet effect, in which a good example of a native vowel category behaves like a magnet, making nearby sounds harder to tell apart than the same acoustic distance would be near a category edge.

The lexicon leans on the ear

Perception is also pushed from above. William Ganong showed in 1980 that an ambiguous sound halfway between d and t is identified as whichever produces a real English word: the same acoustic token is heard as d before ash, giving dash, and as t before ask, giving task. Word knowledge is reaching down and biasing a phonetic judgment.

Richard Warren's phonemic restoration work makes the same point more dramatically. Excise a single consonant from a recorded sentence and replace it with a cough. Listeners report hearing the sentence intact, complete with the missing consonant, and cannot say which sound the cough covered. The brain is not merely interpreting the signal. It is supplying material that is not there.

The dispute: what is the listener recovering?

These findings, taken together, prompted a genuine and long-running disagreement.

The motor theory, developed by Liberman and Ignatius Mattingly, holds that the objects of speech perception are not sounds but intended articulatory gestures. Its case: the acoustic signal has no invariant properties corresponding to phonemes, but the gestures do; the McGurk effect shows perception integrating any evidence about the gesture, whatever the sense; and categorical perception looks like the output of a system tuned to discrete articulatory targets rather than continuous acoustics.

General auditory accounts hold that speech perception uses ordinary auditory mechanisms plus learning, with no special gesture-recovery module. Their case rests on three findings that are hard for the motor theory. First, chinchillas trained by Patricia Kuhl and James Miller in 1975 placed their voice onset time boundary close to where English-speaking humans place theirs, and chinchillas do not have vocal tracts that make human stops. Similar results have come from quail and other species. Second, non-speech sounds with the right temporal structure can show boundary effects of the same kind. Third, people with severe congenital impairments of speech production nonetheless perceive speech normally.

What would settle it? A clean demonstration that some perceptual effect requires representing an articulatory gesture and cannot be produced by any auditory system with the right learning history. Nothing yet meets that bar in both directions, which is why the field has largely moved to positions in the middle: perception is multimodal and heavily shaped by experience, gestures are one useful description of what the categories track, and no separate module has been demonstrated to be necessary.

Common misconceptions

  • Categorical perception means listeners cannot hear within-category differences at all. The strong version has been tested and is false; sensitive tasks recover within-category sensitivity. The effect is a strong bias, not a deletion.
  • Vowels are perceived as categorically as consonants. They are not. Vowel perception is markedly more continuous, which is one of the reasons vowel boundaries shift so easily across dialects.
  • The McGurk effect works because people are paying attention to the face instead of the sound. It survives instruction, practice and full knowledge of the setup, which is the signature of integration below the level of decision.
  • Infants learn to hear the contrasts of their language during the first year. The main change is the reverse: they start able to discriminate contrasts from many languages and narrow to the ones their input uses.
  • If a sound was not in the signal, the listener cannot report hearing it. Phonemic restoration shows exactly that happening, reliably, with listeners unable to locate the interruption.

What you now know

  • Categorical perception has two criteria: a sharp identification function, and discrimination that tracks identification rather than acoustic distance.
  • The 1957 Haskins continuum showed equal physical steps producing very unequal discrimination, with peaks at the category boundaries.
  • Voice onset time boundaries sit at different places for English and Spanish listeners, so the boundary is learned rather than given.
  • The McGurk effect shows visual and auditory evidence being fused before awareness, and it survives knowing about it.
  • Infants discriminate categorically within months of birth, and over the first year they lose sensitivity to contrasts their language does not use.
  • Word knowledge biases phonetic decisions, and phonemic restoration shows listeners supplying sounds that were physically absent.
  • The motor theory and general auditory accounts both have real evidence; the chinchilla and non-speech results are the strongest pressure on the motor theory.

Sources

  1. Liberman, A. M., Harris, K. S., Hoffman, H. S., & Griffith, B. C. (1957). The discrimination of speech sounds within and across phoneme boundaries. Journal of Experimental Psychology, 54(5), 358-368. doi.org
  2. McGurk, H., & MacDonald, J. (1976). Hearing lips and seeing voices. Nature, 264, 746-748. doi.org
  3. Werker, J. F., & Tees, R. C. (1984). Cross-language speech perception: Evidence for perceptual reorganization during the first year of life. Infant Behavior and Development, 7(1), 49-63. doi.org
  4. Wikipedia contributors. (n.d.). Categorical perception. Wikipedia. en.wikipedia.org
  5. Eimas, P. D., Siqueland, E. R., Jusczyk, P., & Vigorito, J. (1971). Speech perception in infants. Science, 171(3968), 303-306.
Key terms
Categorical perception
A pattern in which identification of a stimulus continuum is sharp and discrimination is good only across a category boundary.
Identification function
A plot of how often listeners assign each stimulus on a continuum to one category, steep-walled when perception is categorical.
Discrimination peak
The rise in same-different accuracy for stimulus pairs that straddle a category boundary rather than sitting inside a category.
McGurk effect
The fusion of conflicting auditory and visual speech information into a third percept, as when audio ba and video ga are heard as da.
High-amplitude sucking paradigm
A method for testing infant discrimination in which a rise in sucking rate after a stimulus change indicates that the change was noticed.
Perceptual narrowing
The loss over the first year of life of sensitivity to contrasts not used by the ambient language, from an initially broader capacity.
Perceptual magnet effect
The reduced discriminability of sounds near a good exemplar of a native category, as though the prototype pulled neighbours towards it.
Phonemic restoration
The reported perception of a speech sound that has been physically replaced by noise, with listeners unable to locate the interruption.
Motor theory of speech perception
The claim that listeners recover intended articulatory gestures rather than acoustic patterns, proposed by Liberman and Mattingly.

Module 4: The Phoneme

The analytic core of the course: how to take a data set in a language you do not speak, decide whether two sounds contrast or are variants of one unit, state the environment that predicts them, and then generalize from single sounds to natural classes defined by features.

Phonemes, Allophones, and Complementary Distribution

  • Run the five-step phonemic analysis procedure on an unfamiliar data set and state the resulting rule.
  • Distinguish contrast, complementary distribution and free variation, and give the evidence that separates them.
  • Choose an underlying form and defend the choice, and recognize the cases where the procedure needs care.

Eight Spanish words

Here is a data set. These are ordinary Spanish words in careful phonetic transcription, as a Spanish speaker actually pronounces them.

SpellingTranscriptionMeaning
dedo[ˈdeðo]finger
donde[ˈdonde]where
nada[ˈnaða]nothing
falda[ˈfalda]skirt
cada[ˈkaða]each
andar[anˈdaɾ]to walk
lado[ˈlaðo]side
verdad[beɾˈðað]truth

Two sounds are in play: the stop [d] and the fricative [ð], the sound at the start of English this. A Spanish speaker will tell you these are the same letter and the same sound. An English speaker will tell you they are obviously different, since English uses them to distinguish den from then. Both intuitions are about a system, not about the acoustics, and this lesson gives you a procedure that settles the matter without asking anyone's opinion.

Key idea: Whether two sounds are one unit or two is a question about their distribution in a particular language, and distribution can be checked.

The three relations two sounds can be in

For any two sounds in any language, exactly one of these holds.

  • Contrast. They appear in the same environment and swapping them changes the word. English [d] and [ð] contrast: den and then, dare and there, breeding and breathing. They are separate phonemes.
  • Complementary distribution. They never appear in the same environment. Wherever one is possible, the other is impossible. They are allophones of a single phoneme, and something in the environment predicts which one turns up.
  • Free variation. They can appear in the same environment, but swapping them does not change the word. The final stop of English cat may be released or unreleased with no consequence at all.

The words minimal pair name the strongest evidence for the first case: two words differing in exactly one sound in the same position, with different meanings. One good minimal pair proves contrast and ends the discussion. The absence of minimal pairs proves nothing on its own, because it might be an accident of your data, which is why the rest of the procedure exists.

The procedure, in five steps

  1. Look for minimal pairs. If you find one, the two sounds are separate phonemes. Stop; you are done.
  2. If there are none, list the environments. For every token of each sound, write down what precedes it and what follows it, and whether it is initial, medial or final. Use a two-column table so the pattern can be seen rather than remembered.
  3. Look for a generalization. Ask whether the environments of one sound share something the environments of the other never has. Nasals, voiced sounds, word beginnings, stressed syllables and vowel neighbours are the usual suspects.
  4. Decide which allophone is basic. Normally the one with the wider or less specifiable distribution, since it is what appears everywhere else.
  5. State the rule. Write it as a change from the basic form to the derived one, with the triggering environment named. Then test it against every line of your data, including the lines you did not use to form it.

Working the Spanish data

Step one. Scan for a minimal pair with [d] and [ð] in the same position. There is none in this data set, and there is none in Spanish generally. So the sounds might be allophones.

Step two. Tabulate the environments.

SoundPreceded byFrom which words
[d]A pause, that is, word-initial after silencededo
[d][n]donde, andar
[d][l]falda
[ð][e]dedo
[ð][a]nada, cada, lado, verdad
[ð][ɾ]verdad

Step three. What do the [d] environments share? A pause, an [n] and an [l]. All three involve either silence or a consonant made with the tongue tip at the alveolar ridge, and in each case there is a complete closure or a pause immediately before. What do the [ð] environments share? They follow a vowel or a tap, both of which are continuants: the airflow was never fully stopped.

Step four. Which is basic? Count contexts: [ð] appears after any vowel, and vowels are everywhere. [d] appears in three specifiable places. The narrower distribution is the derived one only if the wider one can be described as an elsewhere case, and here it can: after a pause, [n] or [l] we get the stop; everywhere else we get the fricative. But most analyses of Spanish take [d] as basic, because the environment for the stop is a short, statable list while the environment for the fricative would have to be everything not on that list. Both formulations describe the same facts; the standard one writes the rule as a weakening of the stop.

Step five. State it: the phoneme is /d/, and it is realized as [ð] except after a pause, a nasal or a lateral, where it stays [d]. Now test it. Every line of the table fits, and so does verdad, which contains one of each in a single word, exactly where the rule predicts. That last check is the important one; a rule that only fits the lines it was built from has not been tested.

Spanish does the same thing with /b/ and /ɡ/, giving [β] and [ɣ] in the same environments, which is a strong argument that the rule targets a class of sounds rather than one segment. Module 4's next lesson is about how to say that properly.

The upshot: Contrast is proved by a minimal pair; complementary distribution is proved by tabulating environments and finding that the two lists never overlap and can be told apart by a general property.

A second data set: English vowels before nasals

Now one from your own language, where you can check the intuitions directly. Say these aloud and pay attention to your velum.

WordTranscriptionWordTranscription
bee[bi]bean[bĩn]
bead[bid]beam[bĩm]
seat[sit]seen[sĩn]
deed[did]dean[dĩn]
bud[bʌd]bun[bʌ̃n]

Step one: is there a minimal pair distinguished by nasalization alone? No. There is no English word pair separated by whether a vowel is nasalized, because whenever a vowel is nasalized a nasal consonant follows it.

Step two and three: the nasalized vowels occur before nasal consonants, and only there; the plain vowels occur everywhere else. That is complementary distribution, and the generalization is immediate.

Step four and five: the plain vowel is basic, because its distribution is the residue. The rule is that a vowel becomes nasalized before a nasal consonant in the same syllable. Physically the cause is transparent: lowering the velum takes time, so a speaker starts lowering it during the vowel in order to have it down when the nasal arrives. That is coarticulation, and a great many allophonic rules are coarticulation that a language has made obligatory.

Compare French, where nasalization is contrastive: beau and bon differ by nothing else. The physical mechanism is the same in both languages. What differs is whether the language put the difference to work.

Where the procedure needs care

Complementary distribution is not sufficient. In English, [h] occurs only at the beginning of a syllable and [ŋ] only at the end. They never appear in the same environment, so they are in perfect complementary distribution. Nobody analyzes them as allophones of one phoneme, and rightly, because they share no phonetic similarity and no rule relating them would make any physical sense. The requirement of phonetic similarity is an extra condition, and this pair is the standard reason it exists.

Free variation is a real third option. If two sounds occur in the same environment but the word does not change, you have neither contrast nor complementary distribution. Word-final stops in English may be released or not; many speakers vary between a glottal stop and a full [t] in words like button. Do not force such cases into a rule.

Neutralization can hide a contrast. German has a genuine /d/ and /t/ contrast, but in word-final position both surface as [t]. So Rad, wheel, and Rat, advice, are homophones in the singular, while the plural forms keep them apart. A snapshot of word-final position alone would show no contrast at all and would mislead you. Always look at more than one position, and at related forms of the same word.

Near-minimal pairs are usable evidence. If the perfect pair does not exist in your data, two words differing in the target sounds plus something irrelevant elsewhere still argue for contrast, provided the irrelevant difference could not plausibly be conditioning the sounds you care about.

Common misconceptions

  • A phoneme is the sound in the speaker's head and the allophone is what comes out. That is a picture, not a definition. A phoneme is defined by its contrastive behaviour in a language; whether speakers store one form or many is a separate empirical question, and the evidence is mixed.
  • If you cannot find a minimal pair, the sounds are allophones. Absence of minimal pairs is weak evidence: your data may be small, or the contrast may be rare. You still have to show complementary distribution positively.
  • Allophonic differences are too small to hear. Some are large. Spanish [d] and [ð] differ as much as English den and then, which English speakers hear instantly. The difference is in what the language does with them, not in their size.
  • Every phonetic difference belongs to some rule. Free variation is real, and so is ordinary random variability. Not every observation is a pattern.
  • Rules are learned as rules. Speakers have no conscious access to them. You produce the Spanish or English patterns above without being able to state them, which is why the analysis has to be done from data rather than by asking.

Summing up

  • Two sounds are in contrast, in complementary distribution, or in free variation, and the procedure decides which.
  • One minimal pair settles contrast; the absence of minimal pairs settles nothing by itself.
  • Spanish [d] and [ð] are allophones of one phoneme, with the stop appearing after a pause, a nasal or a lateral and the fricative everywhere else.
  • English vowels are nasalized before nasal consonants and nowhere else, an obligatory version of ordinary coarticulation, while French uses the same difference contrastively.
  • The basic allophone is normally the one whose environment is the residue after the specifiable cases are listed.
  • Complementary distribution must be paired with phonetic similarity, which is why English [h] and [ŋ] are not analyzed as one phoneme.
  • Neutralization, as in German final devoicing, can hide a contrast in one position, so always check related forms.

Sources

  1. Wikipedia contributors. (n.d.). Phoneme. Wikipedia. en.wikipedia.org
  2. Wikipedia contributors. (n.d.). Allophone. Wikipedia. en.wikipedia.org
  3. Wikipedia contributors. (n.d.). Complementary distribution. Wikipedia. en.wikipedia.org
  4. Wikipedia contributors. (n.d.). Spanish phonology. Wikipedia. en.wikipedia.org
  5. Britannica. (n.d.). Phonology. Encyclopaedia Britannica. britannica.com
Key terms
Minimal pair
Two words differing in exactly one sound in the same position and differing in meaning, which proves that the two sounds contrast.
Complementary distribution
The relation between two sounds whose environments never overlap, so that each appears exactly where the other cannot.
Free variation
The relation between two sounds that may occur in the same environment without changing the word, such as released and unreleased final stops in English.
Environment
The context in which a sound occurs, normally stated in terms of the neighbouring sounds and its position in the word or syllable.
Basic allophone
The realization taken as underlying, usually the one whose distribution is the residue after the specifiable contexts have been listed.
Phonetic similarity
The additional requirement that two sounds be articulatorily or acoustically alike before complementary distribution can justify grouping them.
Coarticulation
The overlap of articulatory gestures for neighbouring sounds, the physical origin of many allophonic rules.
Neutralization
The loss of a contrast in a particular position, as when German final obstruents all surface voiceless.

Distinctive Features and Natural Classes

  • Decompose segments into distinctive features and identify the major class, laryngeal, manner and place groups.
  • Define a natural class by counting features, and show that a given set of segments is not one.
  • Derive the three English plural allomorphs from a single underlying form using feature-based rules.

One suffix, three pronunciations

Say these plurals aloud and listen to the ending.

Ends in [s]Ends in [z]Ends in [ɪz]
cats, books, cliffs, months, ropesdogs, cars, pens, boys, lambs, wavesbuses, buzzes, churches, judges, dishes

Three different endings, one grammatical job, and no English speaker ever gets the choice wrong, including on words they have never met. Hand someone the invented nouns wug, gorch and flass and they will produce wugs with a [z], gorches with an [ɪz] and flasses with an [ɪz], instantly and without deliberation.

So there is a rule. The question is what it is a rule about, and that turns out to be the interesting part. The [s] column is not a list of arbitrary words. Its members all end in a voiceless sound. The [z] column all end in a voiced one. The [ɪz] column all end in one of six particular consonants. Any rule worth writing has to refer to those groups, and a theory of phonology has to say what makes a group like that available to a rule in the first place.

Key idea: Phonological rules do not target individual segments; they target groups of segments that share properties, so the theory needs a vocabulary of properties.

Why the segment is the wrong unit

Suppose you insisted that rules operate on whole segments. Then the plural rule would have to be written as a list: after p, t, k, f or the voiceless th, use [s]; after b, d, g, v, the voiced th, m, n, ng, l, r, w, j or any vowel, use [z]. That is a long, unstructured list, and it misses the point in a way you can demonstrate.

Consider the rule you would get by shuffling the lists. After p, m and l use [s]; after t, b and n use [z]. Nothing in a list-based theory says this is a worse rule. It has the same length and the same form. Yet no language does anything like it, while rules conditioned by voicing are found everywhere. A theory in which possible rules and impossible rules look identical is not explaining anything.

Distinctive features fix this. A segment is not a unit but a bundle of properties, mostly binary, and rules refer to features. Instantly, the rule that mentions voicing is one feature long and the shuffled rule cannot be stated at all without listing every segment. The theory now predicts which rules are natural.

The framework comes from Roman Jakobson, Gunnar Fant and Morris Halle in 1952, and was recast in the form still mostly used by Noam Chomsky and Morris Halle in The Sound Pattern of English in 1968.

The features you need

Different textbooks differ in the details. This set will handle everything in this course.

GroupFeaturePositive value picks out
Major classconsonantalSounds with a real constriction: all consonants except the glides and h
Major classsonorantVowels, glides, liquids and nasals: sounds made with free enough airflow for spontaneous voicing
Major classsyllabicSegments that form a syllable nucleus
LaryngealvoiceSounds made with vocal fold vibration
Laryngealspread glottisAspirated sounds and h
MannercontinuantSounds with continuous airflow through the mouth: fricatives, approximants, vowels
MannernasalSounds made with the velum lowered
MannerlateralSounds with airflow around the sides of the tongue
Mannerdelayed releaseAffricates, as opposed to plain stops
MannerstridentNoisy fricatives and affricates: s, z, sh, zh, ch, j, f, v
PlacelabialSounds involving the lips
PlacecoronalSounds made with the tongue tip or blade: dental through postalveolar and retroflex
PlacedorsalSounds made with the tongue body: palatals, velars, uvulars, and all vowels

Two notes. Nasals are [+sonorant] even though they involve a complete oral closure, because the open nasal passage keeps the pressure across the glottis low enough for voicing to continue without effort. And vowels are specified for place under dorsal, using [high], [low] and [back], which is how the vowel chart of Module 2 is translated into features.

Natural classes, defined by counting

A natural class is the set of all and only the segments in a language that share a given specification. The operational test is arithmetic: it takes fewer features to pick out the whole class than to pick out any single member of it.

Work through some examples in English.

SetSpecificationFeatures neededNatural class?
p t kminus voice, minus sonorant, minus continuant3Yes, and specifying p alone needs more
m n ngplus nasal1Yes
s z sh zh ch jplus strident, plus coronal2Yes
p b f v m wlabial1Yes
l r w jplus sonorant, minus syllabic, minus nasal3Yes, the class of glides and liquids
p and ngno shared specification excludes everything elseRequires listing bothNo
b d g m n ngplus voice, minus continuant2Yes

The failed case is the instructive one. To pick out exactly p and ng you would have to say something like: labial voiceless stop, or dorsal nasal. The word or is doing the work, and a disjunction is not a specification. That is the formal reason no phonological rule anywhere targets that pair, and the prediction is testable: search the world's languages and you will not find one.

What matters here: A natural class is cheaper to describe than its members; a set that costs more to describe than its members is a list, and lists are not what rules refer to.

Solving the plural

Now do the work. Start from a single underlying suffix, /z/. Three facts have to fall out.

  1. After a [+strident, +coronal] segment, that is after s, z, sh, zh, ch or j, a vowel is inserted between stem and suffix. This is epenthesis, and its motivation is that English does not tolerate two adjacent sibilants at the end of a word.
  2. After a [-voice] segment, the suffix loses its voicing. This is assimilation: the suffix takes a feature value from its neighbour.
  3. Otherwise nothing happens and the suffix surfaces as [z].

Now derive three words and watch the order matter.

Stagecatsdogsbuses
Underlyingkaet + zdog + zbus + z
After epenthesisno changeno changebus + ɪz
After devoicingkaetsno changeno change
Surface[kʰæts][dɔɡz][ˈbʌsɪz]

Look hard at the buses column. Devoicing does not apply, because by the time it looks, the segment before the suffix is the epenthetic vowel, which is voiced. Now reverse the order. Devoicing applies first, since the stem ends in a voiceless s, giving bus plus s. Then epenthesis inserts a vowel, giving [ˈbʌsɪs]. That is not an English word. The wrong order predicts the wrong output, which means the order is not a matter of taste. Epenthesis must precede devoicing, and the next lesson is entirely about arguments of this shape.

The same analysis, feature for feature, handles the past tense: walked with [t], hugged with [d], wanted with [ɪd], from an underlying /d/, with epenthesis after coronal stops and devoicing after voiceless segments.

Assimilation as feature spreading

Features also explain what assimilation is. Say these three words at conversational speed and pay attention to the nasal in the prefix: impossible, intolerable, incongruous. The first has [m], the second [n], and the third, for most speakers in casual speech, has the velar nasal. Three different nasals, one prefix.

Written with segments, that is three separate rules. Written with features, it is one: the nasal takes the place features of the following consonant. Labial before a labial, coronal before a coronal, dorsal before a dorsal. A single statement, and it also predicts the fourth case correctly, since a following labiodental gives a labiodental nasal in careful speech.

That is what a feature theory buys. It does not merely restate the data more compactly. It says which generalizations are available to a language and which are not, and it makes claims that a survey of the world's languages can falsify.

Common misconceptions

  • Features are just a compact notation. They are a hypothesis about what the phonological system can refer to. The claim that no rule can target a disjunction is empirical and could be overturned by finding one.
  • Every set of sounds is a natural class if you look hard enough. The counting test rules sets out. If a specification requires or, the set is a list.
  • Nasals cannot be sonorants because the mouth is closed. Sonority here is about whether voicing continues without effort, and the open nasal passage vents the pressure that would otherwise stop the folds vibrating.
  • The plural has three suffixes. One suffix and two rules accounts for all three surface forms, and it explains why speakers apply the pattern correctly to words they have never heard.
  • Feature theories are settled. The inventory and the geometry are actively debated: whether features are binary or privative, whether place features are organized in a hierarchy, and whether features are innate or emerge from learning are all open questions.

Looking back

  • The English plural has three surface forms conditioned by the last segment of the stem, and speakers extend the pattern to invented words without hesitation.
  • Rules stated over lists of segments cannot distinguish the patterns languages actually use from the ones they never use.
  • Segments decompose into features grouped as major class, laryngeal, manner and place.
  • A natural class is a set that costs fewer features to specify than its individual members; a set requiring a disjunction is a list, not a class.
  • One underlying suffix plus epenthesis after sibilants plus devoicing after voiceless segments derives all three plural forms.
  • Epenthesis must apply before devoicing, or buses comes out wrong, which is a genuine argument for rule ordering.
  • Nasal place assimilation in impossible, intolerable and incongruous is one rule when stated over features and three when stated over segments.

Sources

  1. Wikipedia contributors. (n.d.). Distinctive feature. Wikipedia. en.wikipedia.org
  2. Wikipedia contributors. (n.d.). Natural class. Wikipedia. en.wikipedia.org
  3. Wikipedia contributors. (n.d.). Assimilation (phonology). Wikipedia. en.wikipedia.org
  4. Wikipedia contributors. (n.d.). Epenthesis. Wikipedia. en.wikipedia.org
  5. Chomsky, N., & Halle, M. (1968). The sound pattern of English. Harper and Row.
  6. Jakobson, R., Fant, G., & Halle, M. (1952). Preliminaries to speech analysis: The distinctive features and their correlates. MIT Press.
Key terms
Distinctive feature
A minimal property of a segment, usually binary, over which phonological rules and contrasts are stated.
Natural class
The set of all and only the segments sharing a feature specification, identifiable because it costs fewer features to name than its members do.
Sonorant
A segment produced with airflow free enough that voicing continues without special effort: vowels, glides, liquids and nasals.
Continuant
A segment with continuous airflow through the oral cavity, which excludes stops, affricates and nasals.
Strident
A feature picking out the noisier fricatives and affricates, used in English to define the class that triggers plural epenthesis.
Coronal
A place feature covering segments articulated with the tongue tip or blade, from dental through postalveolar and retroflex.
Epenthesis
The insertion of a segment, such as the vowel that separates an English stem from a following sibilant suffix.
Assimilation
A change in which a segment takes on a feature value from a neighbour, as when a nasal adopts the place of the following consonant.
Allomorph
One of the surface forms of a single morpheme, such as the three pronunciations of the English plural.

Module 5: Rules and Syllables

How phonological processes are written down, why their order can be argued from surface data, and how syllable structure and phonotactics decide what a language will let a word look like.

Phonological Rules and Rule Ordering

  • Write a phonological rule with a target, a change and an environment, and classify it by process type.
  • Build a derivation from underlying form to surface form and use it to argue for an order between two rules.
  • Identify feeding, bleeding, counterfeeding and counterbleeding relations, and explain what opacity is.

Two words, one consonant, two vowels

For a large number of speakers in Canada and the northern United States, the words writer and rider are pronounced [ˈɹʌɪɾɚ] and [ˈɹaɪɾɚ]. Read those transcriptions carefully. The consonant in the middle is the same in both: a voiced tap. What differs is the diphthong. In writer it starts higher, closer to the vowel of but; in rider it starts lower, at the vowel of father.

This is a genuine puzzle, and it is one of the tidiest arguments in phonology. The raised diphthong occurs before voiceless consonants: write, rice, ripe, light all have it, while ride, rise, bribe, lied do not. But in writer the consonant is not voiceless. It is a tap, and taps are voiced. So the rule that raises the diphthong is looking at something that is not there on the surface.

The only way to make that work is for the raising to happen at a stage when the consonant was still a voiceless /t/, and for the flapping to happen afterwards. Which is to say: the two processes are ordered, and the surface data can prove it. This lesson is about how that kind of argument is built.

Key idea: When a process applies in an environment that is invisible on the surface, the environment must have been present at an earlier stage, which forces an ordering between processes.

How a rule is written

A phonological rule has three parts, and every rule you will meet can be read off in this shape.

  • A target: what the rule applies to, usually stated as a natural class rather than a single segment.
  • A change: what happens to it, usually a change in one or two feature values, or deletion, or insertion.
  • An environment: the context that triggers it, with a blank marking where the target sits.

The standard notation writes these as target, then an arrow, then the change, then a slash, then the environment with an underline for the target's position. Written out in words, the English nasalization rule from the last module reads: a vowel becomes nasalized in the environment before a nasal consonant in the same syllable. The plural devoicing rule reads: the suffix becomes voiceless after a voiceless segment. Each names a class, a change and a context, and nothing else.

Two conventions save trouble. A rule with no environment applies everywhere. And a rule that fails to name a feature is silent about it, which means the feature is unchanged, not that it is unspecified.

The processes languages actually use

Rules are not arbitrary. A small set of process types accounts for the overwhelming majority.

ProcessWhat happensA real case
AssimilationA segment takes a feature from a neighbourEnglish impossible, intolerable, incongruous
DissimilationA segment becomes less like a nearby oneThe Latin suffix that gives navalis and mortalis appears as solaris and militaris after a stem containing l
DeletionA segment is removedEnglish schwa loss in chocolate, camera and family in ordinary speech
InsertionA segment is addedThe vowel in buses; the p many speakers add in warmth and something
MetathesisTwo segments swap orderLatin miraculum and periculum become Spanish milagro and peligro
LenitionA segment weakens towards a more open articulationSpanish /d/ surfacing as a fricative between vowels; English flapping
FortitionA segment strengthensWord-initial strengthening of glides in some languages

Assimilation is by far the most common, which is unsurprising once you remember Module 1: articulators move continuously and overlap, so making a segment more like its neighbour is what happens if nobody stops it. Many rules are coarticulation that a language has made obligatory and then extended beyond what the physics requires.

Deriving writer and rider

Now build the argument properly. Two rules, in a variety with Canadian raising.

  • Raising: the diphthong of ride becomes raised before a voiceless consonant.
  • Flapping: /t/ and /d/ become a tap between a stressed vowel and an unstressed one.

Take the underlying forms to be the obvious ones, with a /t/ in writer and a /d/ in rider, since the related words write and ride have those consonants plainly. Now run the two orders.

Stagewriter, order Arider, order Awriter, order Brider, order B
Underlyingɹaɪt + ɚɹaɪd + ɚɹaɪt + ɚɹaɪd + ɚ
First ruleRaising: ɹʌɪtɚRaising: does not applyFlapping: ɹaɪɾɚFlapping: ɹaɪɾɚ
Second ruleFlapping: ɹʌɪɾɚFlapping: ɹaɪɾɚRaising: does not applyRaising: does not apply
Predicted surface[ˈɹʌɪɾɚ][ˈɹaɪɾɚ][ˈɹaɪɾɚ][ˈɹaɪɾɚ]

Order A, raising before flapping, predicts a difference. Order B, flapping before raising, predicts that the two words are homophones, because by the time raising looks for a voiceless consonant there is none left anywhere. Speakers of this variety distinguish the words. So order A is right and order B is wrong, and the argument rests entirely on data anyone can collect.

Notice what you have just done. You inferred something about the internal organization of a grammar from pronunciations, without any access to the speaker's mind and without asking them anything. That is the standard shape of a phonological argument.

The upshot: Rule ordering is not a stipulation but a conclusion, forced when one order predicts the observed forms and the other does not.

Four relations between two rules

Once rules can be ordered, exactly four relations become possible, and they have names worth knowing because they organize the literature.

RelationDefinitionEffect
FeedingThe first rule creates new inputs for the secondThe second applies more than it otherwise would
BleedingThe first rule destroys inputs for the secondThe second applies less than it otherwise would
CounterfeedingThe order is reversed, so a rule that could have created inputs applies too lateA rule fails to apply where its trigger is visible
CounterbleedingThe order is reversed, so a rule that could have removed inputs applies too lateA rule has applied where its trigger is no longer visible

The English plural from the last lesson is a bleeding relation. Epenthesis inserts a voiced vowel between the stem and the suffix, and by doing so it destroys the environment devoicing needed. Apply them the other way and buses comes out wrong.

Writer and rider is a counterbleeding relation. Flapping would have bled raising if it went first, by removing the voiceless consonant. Because raising goes first, flapping applies too late to prevent it, and the raised vowel survives with its cause erased.

Feeding is easiest to see in a constructed example, so here is one, labelled as constructed. Imagine a language with two rules: first, a word-final mid vowel raises to a high vowel; second, a velar stop becomes a fricative before a high vowel. Take an underlying form ending in a velar plus a mid vowel. Rule one raises the vowel, which creates the environment rule two needs, and rule two then applies to a form that did not originally meet its description. That is feeding: the first rule fed the second.

Opacity, and why it is evidence rather than a nuisance

A rule is opaque when its effect cannot be read off the surface. Counterbleeding produces one kind of opacity: raising has applied in writer, but the voiceless consonant that motivated it is gone, so a learner hearing only [ˈɹʌɪɾɚ] has no surface reason for the raised vowel. Counterfeeding produces the other kind: a rule visibly fails to apply even though its trigger is sitting right there.

Opacity matters because it is the strongest evidence that phonology involves intermediate stages rather than a direct statement about surface forms. If the grammar only ever compared underlying forms to surface forms, there would be no place for the raised vowel in writer to come from. It has also been the central difficulty for theories that dispense with ordered rules, including early Optimality Theory, which compares candidate surface forms and therefore has to work hard to reproduce effects whose causes are no longer visible.

French supplies a historical version of the same thing. Latin words ending in a nasal consonant developed nasalized vowels before the nasal, and the nasal consonant was later weakened or lost. What remains is a nasal vowel with nothing to have caused it, which is exactly why modern French has contrastive nasal vowels rather than allophonic ones. Module 6 returns to this as a mechanism of sound change.

Common misconceptions

  • Rule order is decided by the analyst for convenience. Where the two orders make different predictions, the data decides. Where they make the same predictions, the order is genuinely unsettled and should be reported as such.
  • Underlying forms are how the word used to be pronounced historically. They are a synchronic hypothesis about what speakers store, justified by the alternations in the current language. Sometimes they resemble an older form and sometimes they do not.
  • A rule applies whenever its environment appears anywhere in the word. The environment is a precise local statement, including whether it is bounded by a syllable or a word edge, and getting the boundary wrong is the most common error in writing rules.
  • Opacity is a sign that the analysis is wrong. Opacity is a well-documented property of real languages, and reproducing it is a requirement on theories rather than an embarrassment.
  • Every alternation needs a rule. Irregular alternations such as English foot and feet are stored, not derived. Rules are for patterns that generalize to new words, which is why the wug test matters.

What to carry forward

  • A rule names a target class, a change and an environment, and nothing more.
  • The recurring process types are assimilation, dissimilation, deletion, insertion, metathesis, lenition and fortition, with assimilation by far the most common.
  • Raising must precede flapping, because the reverse order predicts that writer and rider are homophones, and for many speakers they are not.
  • A derivation lays out underlying form, each rule application and the surface form, so two candidate orders can be compared directly.
  • Two ordered rules can stand in a feeding, bleeding, counterfeeding or counterbleeding relation.
  • Opacity, where a rule has applied without a visible cause or has failed to apply with one, is real and is the strongest argument for intermediate stages.

Sources

  1. Wikipedia contributors. (n.d.). Phonological rule. Wikipedia. en.wikipedia.org
  2. Wikipedia contributors. (n.d.). Canadian raising. Wikipedia. en.wikipedia.org
  3. Wikipedia contributors. (n.d.). Feeding order. Wikipedia. en.wikipedia.org
  4. Wikipedia contributors. (n.d.). Optimality Theory. Wikipedia. en.wikipedia.org
  5. Chomsky, N., & Halle, M. (1968). The sound pattern of English. Harper and Row.
Key terms
Phonological rule
A statement pairing a target class with a change and the environment in which the change occurs.
Underlying representation
The stored form of a morpheme, posited to account for the alternations it shows across contexts.
Derivation
The step by step application of rules from an underlying form to a surface form, used to compare competing analyses.
Feeding
An order in which an earlier rule creates new inputs for a later one, so the later rule applies more widely.
Bleeding
An order in which an earlier rule destroys inputs for a later one, so the later rule applies less widely.
Counterbleeding
An order in which a rule that could have removed inputs applies too late, leaving an effect whose cause is no longer visible.
Opacity
The property of a rule whose application, or failure to apply, cannot be justified from the surface form alone.
Lenition
A weakening process moving a segment towards a more open articulation, as with intervocalic Spanish stops or English flapping.
Metathesis
A process in which two segments exchange positions, as in the Latin to Spanish development of miraculum into milagro.

Syllable Structure and Phonotactics

  • Distinguish accidental from systematic gaps in a lexicon and explain what each shows.
  • Divide a syllable into onset, nucleus and coda, state the English onset constraints, and apply the sonority sequencing principle.
  • Predict how a language will repair an illegal sequence in a loanword, using Japanese and Hawaiian as worked cases.

Three words that do not exist

Here are three strings: blick, bnick, lbick. None of them is an English word. Ask any English speaker which of them could be one and the answers are immediate and unanimous. Blick could be. Bnick could not. Lbick is not even worth considering.

That reaction is the whole subject of this lesson, because it cannot come from experience with the words themselves. All three are equally absent from the language. The difference must come from knowledge about what English words are allowed to look like, and that knowledge is detailed, unconscious and shared. Blick is an accidental gap, a possible word that happens not to exist. Bnick and lbick are systematic gaps, excluded by the rules.

The study of those rules is phonotactics, and the unit they are stated over is the syllable.

Key idea: Speakers can sort nonexistent strings into possible and impossible words, which shows that a language contains constraints on sound sequences independent of its actual vocabulary.

Inside the syllable

A syllable has a nucleus, almost always a vowel, which is obligatory. Consonants before it form the onset; consonants after it form the coda. Nucleus and coda together form the rhyme.

WordOnsetNucleusCoda
atnoneæt
teatinone
strengthss t ɹɛŋ k θ s
plantp læn t

Why group the nucleus and coda together rather than the onset and nucleus? Three independent reasons. Verse rhyme matches exactly that constituent, which is where the name comes from: cat and hat rhyme because their rhymes are identical while their onsets differ. Syllable weight, which drives the stress systems of the next lesson, is computed from the rhyme and ignores the onset entirely. And rules that refer to part of a syllable overwhelmingly refer to the rhyme or the coda rather than to a group consisting of the onset and the nucleus.

What English allows in an onset

English permits onsets of zero to three consonants, and each size is constrained.

  • One consonant: almost any, with one clean exception. No English word begins with the velar nasal, though plenty end with it. That is a systematic gap you can verify against any dictionary.
  • Two consonants: mostly an obstruent followed by a liquid or glide, as in play, brew, twin, cure, or /s/ followed by a stop or nasal, as in spy, stay, ski, small, snow.
  • Three consonants: only /s/, then a voiceless stop, then a liquid or glide: split, strong, scream, squid, spew.

Now look at the gaps inside those patterns, because they are more informative than the patterns. English has /pl/ and /bl/, and /kl/ and /ɡl/, but not /tl/ or /dl/. That is why tlick is as impossible as bnick, even though /t/ and /l/ are both perfectly ordinary English consonants and the sequence occurs happily across a syllable boundary in atlas. The constraint is about the onset, not about the sounds.

Sonority, and the principle it supports

The patterns above are not a random list. Rank segments by sonority, roughly their loudness and openness at a fixed length, pitch and stress.

RankClassExamples
Most sonorousVowelsa, i, u
Glidesw, j
Liquidsl, r
Nasalsm, n, ng
Fricativesf, s, sh, v, z
Least sonorousStopsp, t, k, b, d, g

The sonority sequencing principle says that sonority rises from the edge of the onset to the nucleus and falls from the nucleus to the edge of the coda. A syllable is a sonority hill with the vowel at the top.

Test it. In plant, the onset runs stop then liquid, which rises; the coda runs nasal then stop, which falls. Perfect hill. In blick, the onset runs stop then liquid: rising, fine. In bnick, the onset runs stop then nasal, also rising, so the principle allows it, and yet English rejects it. That is worth being honest about: sonority sequencing is a necessary condition, not a sufficient one. Languages add their own further restrictions, and English happens to bar stop plus nasal onsets while Greek and Russian permit them, as English speakers discover when they meet the name Dmitri or the word pneumonia, whose spelling records a cluster English deleted.

In lbick the onset runs liquid then stop, which falls before the nucleus. That violates sonority sequencing outright, and it is why the string feels not merely unusual but unpronounceable.

What matters here: Sonority sequencing rules out lbick everywhere in the world, while the ban on bnick is a fact about English in particular, and keeping those two kinds of impossibility apart is most of the skill.

Where s breaks the rules

Now look at spy, stay, ski, and street. In each, a fricative precedes a stop in the onset, so sonority falls and then rises again. That directly violates the principle.

The violation is not confined to English. Sequences of /s/ plus a stop misbehave in language after language, and they misbehave in the same direction. The standard response is to treat that /s/ as sitting outside the syllable proper, attached to a higher level of structure as an appendix. Supporting evidence: English three-consonant onsets always have /s/ first and never anything else; the /p/ of spy is unaspirated exactly as a non-initial stop would be; and in many languages that otherwise forbid clusters, an /s/ plus stop sequence is the one exception tolerated.

Whether that is an explanation or a relabelling is a fair question, and it is argued about. What is not in doubt is that the sequences are exceptional and that the exception is systematic across languages.

Dividing a word into syllables

Where does the boundary in the word extra fall? Speakers give conflicting answers, and phonology gives a principle: the maximal onset principle says to put as many consonants as the language permits into the onset of the following syllable. English allows /str/ as an onset, so extra divides as e plus stra rather than ex plus tra.

Two complications are real. First, English disallows a lax vowel at the end of a stressed syllable, which pulls in the opposite direction: in happy, maximal onset says the /p/ belongs to the second syllable, but the lax vowel in the first syllable wants a coda. The usual conclusion is that the consonant belongs to both, which is called ambisyllabicity. Second, morphology interferes: speakers often syllabify at a morpheme boundary even where the maximal onset principle would not, dividing mis plus lead rather than mi plus slead.

Other languages, other limits

English is generous by world standards. Compare two extremes and one repair strategy.

  • Hawaiian allows only open syllables with at most one consonant in the onset. There are no consonant clusters anywhere and no codas at all. Every word ends in a vowel. When English words are borrowed, they are rebuilt to fit: Merry Christmas becomes Mele Kalikimaka, with vowels inserted to break every cluster and to supply an ending.
  • Japanese allows a simple onset, a vowel, and a coda restricted to a moraic nasal or the first half of a doubled consonant. Loanwords are repaired by inserting a vowel, usually [u]: strike becomes sutoraiku, Christmas becomes kurisumasu, and McDonald's becomes makudonarudo. Note that the repair is insertion rather than deletion, and that the inserted vowel is consistent, which shows the process is grammatical rather than improvised.
  • Georgian goes the other way, permitting onsets of four, five and six consonants that English speakers can barely attempt. And Nuxalk, a Salishan language of British Columbia, has words containing no vowels at all, which forces a rethink of what a syllable nucleus has to be.

Loanword adaptation is a useful window precisely because the speaker has no stored form to fall back on. Whatever they produce is generated by the grammar on the spot, which makes it evidence about the grammar rather than about the lexicon.

Knowledge that is gradient, not binary

One refinement, because the blick test as usually presented is too clean. Ask people to rate a range of invented words rather than to sort them into two boxes, and the ratings do not split into possible and impossible. They spread out, and they line up with how frequently the relevant sequences occur in the existing vocabulary. Strings built from common clusters are rated better than strings built from rare but legal ones.

That is a real result and it constrains theories. Phonotactic knowledge cannot be only a filter that passes or blocks; it is at least partly statistical, tracking the frequencies of the patterns a learner has heard. The categorical judgments about blick and lbick are the extreme ends of a graded scale rather than the whole of it.

Common misconceptions

  • A syllable is defined by a beat or a pulse of the chest. The chest pulse theory was tested in the 1960s and did not survive. The syllable is a unit of phonological organization, defined by structure and by the rules that refer to it.
  • Every word English lacks is impossible in English. Most absent strings are accidental gaps. Blick was available the whole time and could be adopted tomorrow.
  • Speakers of languages with simple syllables cannot hear or produce clusters. They can be taught to, and their repairs in loanwords are systematic rather than failed attempts. What is at issue is the grammar, not the anatomy.
  • Sonority sequencing explains all cluster restrictions. It rules out lbick, but it permits bnick, which English rejects anyway. Language-specific constraints do the rest.
  • Syllable boundaries are always determinate. Ambisyllabic consonants and conflicts between maximal onset and morphology mean some boundaries genuinely have no single right answer.

The short version

  • Blick is an accidental gap and bnick and lbick are systematic gaps, and speakers sort them without ever having heard any of them.
  • The syllable divides into onset and rhyme, with the rhyme containing nucleus and coda, and rhyme is the constituent that verse and syllable weight refer to.
  • English onsets run to three consonants, and only in the shape s plus voiceless stop plus liquid or glide, and it bars /tl/ and /dl/ and initial velar nasals.
  • Sonority sequencing makes a syllable a hill with the vowel at the peak, and it rules out onsets of falling sonority.
  • Sequences of s plus stop violate that principle across many languages and are usually treated as an appendix outside the syllable proper.
  • The maximal onset principle divides words, with ambisyllabicity and morphology as genuine complications.
  • Hawaiian and Japanese repair illegal sequences by inserting vowels, and the consistency of the repair is what shows it is grammatical.
  • Ratings of invented words are graded rather than binary and track the frequency of the sequences in the existing vocabulary.

Sources

  1. Wikipedia contributors. (n.d.). Syllable. Wikipedia. en.wikipedia.org
  2. Wikipedia contributors. (n.d.). Phonotactics. Wikipedia. en.wikipedia.org
  3. Wikipedia contributors. (n.d.). Sonority Sequencing Principle. Wikipedia. en.wikipedia.org
  4. Wikipedia contributors. (n.d.). Japanese phonology. Wikipedia. en.wikipedia.org
  5. Wikipedia contributors. (n.d.). Hawaiian language. Wikipedia. en.wikipedia.org
Key terms
Phonotactics
The constraints a language places on which sequences of sounds may occur, and where in the syllable or word they may occur.
Accidental gap
A string that the phonotactics of a language permit but that happens not to be a word, such as blick in English.
Systematic gap
A string excluded by the constraints of a language, such as bnick or lbick in English.
Onset
The consonants preceding the nucleus of a syllable, limited in English to three and to a small set of shapes.
Rhyme
The nucleus and coda taken together, the constituent that verse rhyme matches and that syllable weight is computed from.
Sonority
A ranking of segments by openness and carrying power, from stops at the bottom through fricatives, nasals and liquids to vowels at the top.
Sonority sequencing principle
The generalization that sonority rises towards the nucleus and falls away from it, making the syllable a sonority peak.
Maximal onset principle
The rule for syllabification that assigns as many consonants as the language permits to the following onset.
Ambisyllabicity
The analysis in which one consonant belongs simultaneously to the coda of one syllable and the onset of the next.

Module 6: Prosodic Systems, Acquisition, and Change

How languages decide where stress falls and how tone is assigned, and then how a child builds a sound system in the first years of life and how such systems change across generations.

Stress Systems and Tone Languages

  • Apply the Latin stress rule to unfamiliar words using syllable weight, and contrast it with fixed-stress systems.
  • Explain the mora, metrical feet and quantity sensitivity, and say why English stress resists a simple rule.
  • Analyze tone on a separate tier, and work Mandarin third tone sandhi and Tokyo Japanese pitch accent.

A rule the Romans followed without writing it down

Latin stress can be stated in two lines, and the two lines predict the stress of every Latin word of three syllables or more.

  1. If the second-to-last syllable, the penult, is heavy, stress it.
  2. Otherwise, stress the syllable before it, the antepenult.

A syllable is heavy if it contains a long vowel, a diphthong, or a vowel followed by a consonant in the same syllable. Otherwise it is light. Run four words.

WordSyllablesPenultWeightStress falls on
amīcus, frienda mī cusmī, long vowelHeavyThe penult: a MI cus
refectus, restoredre fec tusfec, closed by a consonantHeavyThe penult: re FEC tus
dominus, masterdo mi nusmi, short and openLightThe antepenult: DO mi nus
facilis, easyfa ci lisci, short and openLightThe antepenult: FA ci lis

Notice what the rule needs in order to work. It needs to count syllables from the end of the word, and it needs to know whether a syllable is heavy. Neither is about which particular sounds are present. This lesson is about the machinery those two requirements imply, and about the parallel machinery that assigns tone.

Key idea: Stress assignment refers to position within the word and to syllable weight, not to the identity of the segments, which is why it needs a level of structure above the segment.

Fixed stress, and what it buys

Latin is quantity sensitive. Many languages are not: they simply put stress in a fixed position.

LanguagePosition
Finnish, Hungarian, Czech, IcelandicFirst syllable
PolishPenultimate syllable
TurkishFinal syllable, for most native vocabulary
FrenchFinal syllable of the phrase rather than of the word

Fixed stress has a side effect worth noticing: it turns stress into a boundary signal. A Finnish listener knows a new word has begun every time they hear a stressed syllable, and a Polish listener knows a word is one syllable from ending. Languages with unpredictable stress, like English and Russian, cannot use it that way, and their listeners rely on other cues to segment the stream.

The French entry deserves its qualification. French stress is a property of the phrase, not the word, which is why the same word carries prominence in isolation and loses it inside a longer group. This is one reason French is often described as having no word stress at all.

Weight, and the mora

The unit that makes weight precise is the mora. A light syllable has one mora; a heavy syllable has two. A short vowel contributes one mora, a long vowel or diphthong contributes two, and in many languages a coda consonant contributes one. Onsets never count, which is the same result the rhyme constituent gave you in the last lesson, arrived at independently.

Languages differ on whether a closed syllable is heavy. In Latin it is. In some languages only vowel length counts, so CVC patterns with CV. That parameter, set one way or the other, accounts for a large part of the variation in stress systems across the world.

Japanese makes the mora audible. The name Tokyo is two syllables but four moras, because both vowels are long: to, o, kyo, o. Japanese verse counts moras, not syllables, which is why a haiku line of five is five moras. Speakers of mora-counting languages find the unit as obvious as English speakers find the syllable.

Feet: how the beat gets built

Stress systems are usually described by grouping syllables into feet, small units of one strong and one weak syllable, and then designating one foot as the strongest in the word.

  • A trochee is strong then weak. English content words overwhelmingly favour it, which is why HAP-py, TA-ble and WIN-dow feel natural and the reverse feels foreign.
  • An iamb is weak then strong, common in languages of the Americas and in English verse, though not in the English lexicon.

A stress system is then largely a set of choices: which foot shape, whether feet are built from the left edge or the right, whether the syllable at the edge is skipped, and whether feet are sensitive to weight. Latin, in these terms, builds a weight-sensitive trochee at the right edge with the final syllable skipped. Once you have those parameters, the two-line Latin rule turns out to be a special case of a general theory rather than a curiosity of one language.

English, the hard case

English stress does not reduce to two lines, and it is worth being honest about why. Three sources of complication interact.

  • Weight matters, but not simply. English does show a penultimate weight effect in many nouns, but exceptions are numerous and often lexical.
  • Word class matters. The noun and verb pairs of Module 2, such as permit, record and contract, differ in stress alone.
  • Morphology matters, and suffixes divide into two kinds. Some are stress-neutral and leave the stem's stress alone: happy and happiness, sharp and sharply. Others force a shift: atom becomes atomic, photograph becomes photography and then photographic, electric becomes electricity. The stress moves to a position determined by the suffix, and the vowels reduce accordingly.

That last pattern is where the previous modules come together. The stress shift is one process; the vowel reduction it triggers is another; and the alternation between the full and reduced vowels of the same morpheme is exactly the kind of evidence that argues for a single underlying form. Photograph and photography share a stored root, and the differences between them are computed.

The upshot: A stress system is a set of parameters, not a list, and English looks irregular because several of its parameters interact with morphology and word class rather than because it lacks a system.

Tone on its own tier

Now the other half. In Module 2 you met Mandarin, Thai and Yoruba tone as data. The organizing insight came in the 1970s, principally from John Goldsmith, and it is called autosegmental phonology: tones are represented on a separate tier from the segments, and linked to tone-bearing units by association lines.

Four kinds of evidence push in that direction, and they are the reason the proposal stuck.

  1. Stability. In many languages, when a vowel is deleted its tone does not disappear. It reattaches to a neighbouring syllable. If the tone were simply a feature of the vowel, deleting the vowel would take the tone with it.
  2. Contours as sequences. A falling tone behaves in many languages exactly like a high tone followed by a low tone that happen to sit on one syllable, and it splits into two when the syllable does.
  3. One to many. A single tone can spread across several syllables, and several tones can crowd onto one, which no theory of tone as a segment feature handles gracefully.
  4. Floating tones. Some morphemes consist of a tone and nothing else, marking a grammatical distinction with no segmental material at all. A separate tier gives such a morpheme somewhere to live.

Tone sandhi, worked

Sandhi is the general name for changes that happen when forms are put together. Mandarin has the best known tone sandhi rule, and it is easy to check.

The rule: when two third tones are adjacent, the first becomes a second tone. So the greeting 你好, which is nǐ hǎo with two third tones in isolation, is pronounced ní hǎo, with the first syllable rising. The same applies to 很好, hěn hǎo, pronounced hén hǎo.

Two details make this more than a curiosity. First, speakers are generally unaware of it; ask them and they will tell you they said two third tones. Second, in strings of three or more third tones the outcome depends on how the words group syntactically, so the rule has to look at structure and not merely at adjacency. Sandhi is grammar.

Mandarin has a second, lexically specific rule: the negator 不, bù, which carries a fourth tone, becomes bú before another fourth tone, as in 不是, bú shì. Rules restricted to a single morpheme are unusual and worth flagging as such.

Pitch accent, the third option

Between a stress language and a tone language sits a third arrangement. In a pitch accent system, at most one syllable or mora per word is marked, and the pitch of the whole word follows from where the mark is.

Tokyo Japanese is the standard example. The syllable sequence hashi has three lexical entries, distinguished by accent placement: 箸, chopsticks, is accented on the first mora; 橋, bridge, is accented on the second; and 端, edge, is unaccented. In isolation the last two can sound alike, but add a following particle and they separate cleanly, because the accent on bridge causes a fall immediately after it while the unaccented word lets the pitch stay high onto the particle. That is a real minimal set, and it shows why the accent has to be a property of the word rather than of any single syllable's pronunciation.

Swedish and Norwegian do something related with two contrasting word accents: the written form anden is the duck with one accent and the spirit with the other. Speakers of these languages hear the difference immediately, and learners find it among the hardest things to acquire.

Rather than three sealed categories, it is more accurate to see stress, pitch accent and tone as a space of options: how many units per word may bear a prosodic mark, and whether the mark specifies pitch, prominence or both.

Common misconceptions

  • Fixed stress means the language is simpler. It means one parameter is set to a constant. Finnish has an intricate vowel harmony system and a rich morphology; nothing is saved overall.
  • English stress is unpredictable. Large parts are predictable from weight, word class and suffix type. What is true is that no single short rule covers it, and that some placements are lexical.
  • Onsets affect syllable weight. In the overwhelming majority of systems they do not, which is one of the strongest arguments for the rhyme as a constituent.
  • Tone sandhi is sloppy speech. The Mandarin third tone rule is obligatory, structure-sensitive and applied by careful speakers reading aloud.
  • Pitch accent languages are tone languages with fewer tones. The defining property is different: at most one mark per word, with the rest of the contour predictable, rather than a tone specified on every syllable.

Putting it together

  • Latin stresses a heavy penult and otherwise the antepenult, which requires counting from the word edge and computing syllable weight.
  • Fixed-stress languages put stress in a constant position, which turns it into a word boundary cue for listeners.
  • Weight is measured in moras, with onsets never counting, and languages differ over whether a closed syllable counts as heavy.
  • Stress systems are described by foot shape, direction of parsing, edge effects and weight sensitivity, which makes Latin a parameter setting rather than an oddity.
  • English stress interacts with weight, word class and two kinds of suffix, and the shifts it produces drive the vowel reduction of photograph and photography.
  • Autosegmental phonology puts tone on its own tier, supported by tone stability, contours splitting into level tones, spreading, and floating tone morphemes.
  • Mandarin third tone sandhi is obligatory and sensitive to syntactic grouping, and speakers are largely unaware of it.
  • Pitch accent systems mark at most one unit per word, as in the Tokyo Japanese triple for chopsticks, bridge and edge.

Sources

  1. Wikipedia contributors. (n.d.). Stress (linguistics). Wikipedia. en.wikipedia.org
  2. Wikipedia contributors. (n.d.). Syllable weight. Wikipedia. en.wikipedia.org
  3. Wikipedia contributors. (n.d.). Autosegmental phonology. Wikipedia. en.wikipedia.org
  4. Wikipedia contributors. (n.d.). Tone sandhi. Wikipedia. en.wikipedia.org
  5. Wikipedia contributors. (n.d.). Metrical phonology. Wikipedia. en.wikipedia.org
  6. Hayes, B. (1995). Metrical stress theory: Principles and case studies. University of Chicago Press.
Key terms
Syllable weight
The classification of a syllable as light or heavy, computed from the rhyme and used by quantity-sensitive stress systems.
Mora
A unit of syllable weight; a short vowel contributes one, a long vowel or diphthong two, and in many languages a coda consonant one.
Quantity sensitivity
The property of a stress system that consults syllable weight, as Latin does, rather than only counting positions.
Foot
A grouping of a strong syllable with a weak one, the unit from which stress patterns are built.
Trochee
A foot with the strong syllable first, the dominant shape in English content words.
Autosegmental phonology
The framework in which tones occupy a separate tier and are linked to tone-bearing units by association lines.
Tone stability
The survival of a tone when its vowel is deleted, the classic argument for representing tone on its own tier.
Tone sandhi
A change in tone caused by neighbouring tones, such as the Mandarin rule turning the first of two third tones into a rising tone.
Pitch accent
A system in which at most one unit per word carries a prosodic mark, from which the rest of the pitch contour follows.

Learning a Sound System, and How Sound Systems Change

  • Describe the course of phonological acquisition, including babbling, the common child processes and the perception-production gap.
  • Explain the listener-based account of sound change and apply it to nasalization and palatalization.
  • State Grimm's Law and the Great Vowel Shift as regular correspondences, and explain phonologization with real cases.

The child who would not accept his own word

In 1960 Jean Berko and Roger Brown reported a small exchange that has been quoted ever since. A young child referred to his plastic fish as his fis. An adult, trying to meet him halfway, said: this is your fis? The child said no. The adult tried again: your fis? No again. Finally the adult said: your fish? Yes, the child replied. My fis.

Read that carefully, because it settles something. The child could hear the difference between fis and fish perfectly well, and rejected the adult's imitation of his own pronunciation. What he could not do was produce it. Perception was ahead of production, and the gap between them was invisible from the outside until someone thought to test it.

This lesson follows a sound system twice: once as a child builds one in a few years, and once as a community changes one across centuries. The two processes are connected, and the last section says how.

Key idea: A child's mispronunciations are not perceptual failures; the phonological representation is more accurate than the output, which means production has its own separate limitations.

The first year: losing options

Module 3 gave you the perception evidence. Newborn and one-month-old infants already discriminate speech contrasts categorically. Between six and twelve months, infants stop discriminating contrasts their ambient language does not use, while infants exposed to those contrasts keep them. The first year is a narrowing.

Production runs in parallel. Around six to ten months infants begin canonical babbling, producing well-formed consonant and vowel syllables, usually reduplicated: ba ba ba, da da da. The inventory of babbling is remarkably similar across language communities at first and gradually comes to resemble the ambient language in its consonants and its rhythm.

One finding shows that babbling is linguistic rather than merely oral-motor. Deaf infants exposed to a signed language from birth babble with their hands, producing repetitive, meaningless sequences built from the phonological units of the sign language, on a comparable developmental schedule. Whatever babbling is, it is not about the mouth.

The order production follows, and why

Children do not acquire the sounds of their language in a random order, and the pattern is broadly consistent. Stops, nasals and glides come early. Fricatives come later. The English consonants that give the most trouble longest are the interdental fricatives, the r, and the sibilant s, some of which are not mastered until well into the school years.

The errors themselves are highly systematic, which is the point. Each of these has a name because each recurs across children and languages.

ProcessWhat happensExample
Cluster reductionA consonant cluster loses a memberspoon said as poon
StoppingA fricative becomes the corresponding stopthat said as dat
FrontingA velar is produced at the alveolar ridgekey said as tea
GlidingA liquid becomes a gliderabbit said as wabbit
Final consonant deletionThe coda is droppeddog said as daw
Consonant harmonyOne consonant assimilates to another at a distancedog said as gog

Compare that list with Module 5. Every one of these is a process that occurs as a rule in some adult language: cluster reduction in Hawaiian, stopping in varieties of English and in many creoles, final consonant deletion in French liaison contexts and in Mandarin's history. Children are not doing something alien to phonology. They are applying the same repair operations the grammars of the world use, before the target grammar is in place to stop them.

The engine of change: a listener who does not undo enough

Now shift from the individual to the community. Why do sound systems change at all, given that every generation is trying to talk like the one before?

John Ohala's answer, which has organized much of the field's thinking, locates the mechanism in perception rather than in laziness. Every utterance carries coarticulatory distortion: a vowel before a nasal is nasalized, a velar before a front vowel is fronted, a consonant next to a rounded vowel is rounded. Listeners normally correct for this automatically, attributing the distortion to the neighbour and recovering what the speaker intended. But sometimes they do not, and then they store the distorted version as the target and produce it deliberately.

Two consequences follow, and both are testable. First, the changes that happen most often should be exactly the ones that coarticulation produces, and they are: nasalization before nasals, palatalization before front vowels, vowel rounding next to labials. Second, changes should sometimes run in the other direction, when a listener over-corrects and removes a property the speaker actually intended. That kind of change, called hypercorrection in this technical sense, is also attested.

The account explains something that a laziness account cannot: why change is systematic rather than sporadic, and why it targets particular environments. Nobody decides to nasalize. The physics does it, and occasionally the listener takes it literally.

What matters here: Sound change starts as ordinary coarticulation that a listener reinterprets as intended, which is why the recurring changes across unrelated languages are the ones the vocal tract produces for free.

Grimm's Law: what regularity looks like

The nineteenth century discovered that such changes are strikingly regular. Jacob Grimm, building on an observation by Rasmus Rask, set out in 1822 the consonant correspondences between the Germanic languages and their relatives.

Proto-Indo-EuropeanBecomes in GermanicLatin or GreekEnglish
pfLatin pater, piscisfather, fish
tthLatin tresthree
khLatin cornu, centumhorn, hundred
dtLatin duo, decemtwo, ten
gkLatin genuknee

Look down the last two columns. Father and pater are not similar by chance and not borrowed from one another; they are the same inherited word run through a systematic set of changes. The pattern holds across the vocabulary, which is what makes it a law rather than a list of resemblances.

Its exceptions produced the next advance. A set of words did not behave as Grimm predicted, and in 1875 Karl Verner showed that the exceptions were themselves regular once the position of the original accent was taken into account: where the accent did not immediately precede, the expected voiceless fricative came out voiced instead. An apparent counterexample turned into a second law, which is a good working model of how a science of this kind makes progress.

The Great Vowel Shift, and why English spelling looks the way it does

Between roughly 1400 and 1700 the long vowels of English moved. High vowels became diphthongs, and everything below them moved up a step.

Middle English long vowelModern EnglishExample
Long iThe diphthong of bitebite, mine, time
Long uThe diphthong of househouse, mouse, out
Long eThe vowel of meetmeet, feet, see
Long oThe vowel of bootboot, moon, food
Long aThe vowel of namename, take, made

This is why English spelling assigns values to the vowel letters that no other European language recognizes. English orthography was largely fixed, by scribal practice and then by printing, before and during the shift. The letters kept their older continental values on the page while the vowels moved underneath them. The word name is spelled for a pronunciation that ended around the sixteenth century.

Whether the shift was one connected chain, with each vowel pushing or pulling the next, or a set of separate changes that happened to overlap in time, is still argued. What is not in doubt is the correspondence set, which any Middle English text will confirm.

Phonologization: how an allophone becomes a phoneme

The mechanism that links this lesson back to Module 4 is phonologization. An allophonic difference is predictable, so it carries no information. But if the environment that predicted it disappears, the difference is stranded, and speakers have no choice but to treat it as contrastive.

Three cases, all real.

  • French nasal vowels. Vowels nasalized before nasal consonants, exactly as English vowels do now. Then the nasal consonants weakened or were lost in many contexts. The nasal vowel remained with nothing to explain it, and beau and bon became a minimal pair.
  • English f and v. In Old English these were allophones of one phoneme, with the voiced one appearing between voiced sounds. Later changes, including the loss of final vowels and a flood of French loanwords, put them into the same environments, and modern English has feel against veal and fine against vine.
  • English foot and feet. A vowel in the plural suffix caused the stem vowel to front, an assimilation. The suffix vowel was then lost. What survives is a vowel alternation with no visible cause, which is why English has a handful of irregular plurals that behave nothing like the productive one.

In each case the sequence is the same: a rule creates a predictable difference, the trigger for the rule disappears, and the difference is reinterpreted as part of the stored form. That is exactly the opacity of Module 5 seen over historical time, and it is the strongest reason those two lessons belong in one course.

Common misconceptions

  • Children mispronounce because they hear badly. The fis phenomenon shows the opposite: children reject adult imitations of their own errors, which means their stored representation is more accurate than their output.
  • Sound change is decay, and languages are getting sloppier. Change is directional but not degenerative. Every stage of every language is a fully functioning system, and changes that simplify one part of a grammar routinely complicate another.
  • Sound change happens because people are lazy. Effort plays some part, but the listener-based account explains far more, including why the same changes recur in unrelated languages and why they target specific environments.
  • Grimm's Law had exceptions, so it was not a law. Its exceptions were shown by Verner to be regular under a further condition, which strengthened the whole approach rather than weakening it.
  • English spelling is simply irrational. Much of it is a faithful record of a pronunciation that has since changed. The gh of night, the k of knee and the vowel of name were all once pronounced as written.

Where this leaves us

  • The fis phenomenon shows perception running ahead of production, so child errors are output limitations rather than perceptual ones.
  • The first year narrows perception to the ambient language, while canonical babbling begins around six to ten months and is linguistic rather than merely oral, as manual babbling in signing infants shows.
  • Child production errors, from cluster reduction to consonant harmony, are the same repair operations adult grammars use as rules.
  • Ohala's account locates sound change in listeners who fail to undo ordinary coarticulation, which predicts which changes recur across unrelated languages.
  • Grimm's Law states regular consonant correspondences between Germanic and its relatives, and Verner's Law made its exceptions regular too.
  • The Great Vowel Shift moved the English long vowels between about 1400 and 1700, after the spelling had largely settled.
  • Phonologization turns an allophonic difference into a contrast when its conditioning environment is lost, as in French nasal vowels, English f and v, and foot against feet.

Sources

  1. Wikipedia contributors. (n.d.). Sound change. Wikipedia. en.wikipedia.org
  2. Wikipedia contributors. (n.d.). Grimm's law. Wikipedia. en.wikipedia.org
  3. Wikipedia contributors. (n.d.). Great Vowel Shift. Wikipedia. en.wikipedia.org
  4. Wikipedia contributors. (n.d.). Babbling. Wikipedia. en.wikipedia.org
  5. National Institute on Deafness and Other Communication Disorders. (n.d.). Speech and language developmental milestones. NIH. nidcd.nih.gov
  6. Berko, J., & Brown, R. (1960). Psycholinguistic research methods. In P. H. Mussen (Ed.), Handbook of research methods in child development. Wiley.
Key terms
Fis phenomenon
The observation that a child who mispronounces a word rejects an adult's imitation of that mispronunciation, showing perception ahead of production.
Canonical babbling
The production of well-formed consonant and vowel syllables, often reduplicated, beginning around six to ten months.
Cluster reduction
A child process removing a member of a consonant cluster, as when spoon is produced without its initial consonant.
Stopping
A child process replacing a fricative with the corresponding stop, as when that is produced with an initial d.
Listener-based sound change
Ohala's account in which change begins when a listener fails to attribute coarticulatory distortion to its cause and stores it as intended.
Grimm's Law
The regular set of consonant correspondences between Proto-Indo-European and Germanic, published by Jacob Grimm in 1822.
Verner's Law
The 1875 explanation of Grimm's Law exceptions in terms of the position of the original accent.
Great Vowel Shift
The movement of the English long vowels between roughly 1400 and 1700, after English spelling had largely stabilized.
Phonologization
The process by which a predictable allophonic difference becomes contrastive after the environment that conditioned it disappears.

Open the interactive version with quizzes and progress →