🗣️ Linguistics · Graduate · LING 410

Syntax

A graduate course in syntactic theory for a reader who has met syntax once in a survey and now wants to argue it. The evidence is on the page: numbered example sentences, an asterisk on every string that native speakers reject, and labelled bracketings set in monospace, because a text course cannot draw trees. You begin with grammaticality as data, the competence and performance distinction, and…

Start the interactive course (quizzes, progress, videos) →

Free forever. No sign-up, no ads. 17 lessons. The full lesson text is below so you can read it right here.

Module 1: What a Sentence Is, and How You Can Tell

Grammaticality judgments as evidence, the five tests that find a constituent, and the distributional diagnostics that assign a word to a category without appealing to what it means.

Colourless Green Ideas: Grammaticality as Evidence

  • State what an asterisk on a string claims, and distinguish it from a claim about frequency, meaning or correctness.
  • Separate a competence question from a performance one using centre embedding, garden paths and comparative illusions.
  • Explain why a descriptive syntactician has no use for prescriptive rules, and what the two enterprises are each good for.

Two strings, five words each

In 1957 a publisher in The Hague brought out a monograph of barely more than a hundred pages by a twenty-eight-year-old at MIT. In section 2.3 of Syntactic Structures, Noam Chomsky printed two strings built from the same five words:

(1) Colorless green ideas sleep furiously.
(2) *Furiously sleep ideas green colorless.

Neither string had ever been spoken by anybody. Neither describes a possible state of affairs. On any statistical model of English available in 1957, and on most of the ones people were still building forty years later, both have an estimated probability indistinguishable from zero. That was Chomsky's point: whatever separates them, it is not frequency, because both are equally unattested, and it is not meaning, because neither has one.

And yet you do not treat them alike. Read (1) aloud and it takes ordinary declarative intonation, falling at the end. Read (2) aloud and you have to chop it into pieces. More usefully, (1) submits to every operation English performs on a sentence and (2) submits to none:

(3) Do colorless green ideas sleep furiously?
(4) Colorless green ideas do not sleep furiously.
(5) I doubt that colorless green ideas sleep furiously.
(6) *Do furiously sleep ideas green colorless?
(7) *I doubt that furiously sleep ideas green colorless.

That asymmetry is the first piece of evidence in this course, and it is the kind of evidence the whole course runs on. You have a system in your head that assigns a structure to (1) and refuses to assign one to (2), and it does this for strings you have never encountered. Describing that system is what syntax is.

Key idea: A syntactic theory is not a description of what people say. It is a description of what a speaker's internalised system will and will not generate, tested against strings nobody has ever said.

What the asterisk claims

The prefixed asterisk is the field's oldest piece of notation and the most misread. It records that native speakers, asked whether a string is a possible sentence of their language, reject it. It does not record that the string is rare, that it is illogical, that it is bad style, or that no one has ever produced it. Those are four different properties and they come apart constantly.

Practice has grown finer than a single mark. Most contemporary papers use a small scale:

MarkWhat it recordsExample
(none)Fully acceptableWho did you say that Mary saw?
?Mildly degraded, most speakers still accept?Who do you wonder whether Mary saw?
??Seriously degraded, speakers hesitate??What did you meet the man who bought?
*Rejected as not a sentence of the language*Who did you say that saw Mary?
#Well formed but semantically or pragmatically anomalous#Colorless green ideas sleep furiously.

That last row matters more than it looks. Chomsky's famous sentence is not ungrammatical at all. It is impeccably built and it is nonsense, and the two facts are independent. Many authors would mark it with a hash and not a star. The reverse case is just as instructive: a string can be perfectly interpretable and still be rejected.

(8) The child seems asleep.
(9) The child seems to be sleeping.
(10) *The child seems sleeping.

You know exactly what (10) is trying to say. Every content word is in place, the word order is the ordinary one, and a learner who produced it would be understood at once. It is still not English. The verb seem accepts an adjective phrase and it accepts an infinitive, and it refuses a bare present participle, and no fact about meaning explains that refusal. Facts of this kind, where interpretation is unaffected and acceptability collapses anyway, are the reason syntax is studied as a system in its own right rather than read off semantics.

Competence, performance, and three sentences that confuse them

Chomsky's second methodological move, made in 1965, was to separate competence, the knowledge of the language, from performance, what a speaker does with it under real conditions of memory, attention and time. The distinction earns its keep because judgments do not sort cleanly into good and bad. Three cases show why.

First, centre embedding. Build a relative clause inside a relative clause inside a relative clause:

(11) The malt lay in the house.
(12) The malt that the rat ate lay in the house.
(13) The malt that the rat that the cat killed ate lay in the house.

Sentence (13) is close to unreadable, and yet it is generated by exactly the rules that generate (12), applied one more time. Nothing in the grammar distinguishes them. What distinguishes them is a working memory that cannot hold three unfinished subjects at once. Mark (13) with a star and your grammar has to contain a rule counting embeddings, a rule that would be unlike any other rule in the language and that no child could learn. Treat it as a performance failure and the grammar stays simple.

Second, the garden path:

(14) The horse raced past the barn fell.

Most readers hit a wall at fell and conclude the string is broken. It is not. Compare The horse ridden past the barn fell, which nobody stumbles over: raced past the barn is a reduced relative clause, and the difficulty is that raced can also be a simple past tense, so the parser commits early and commits wrongly. The grammar is fine; the processor was ambushed.

Third, and running the other way, the comparative illusion:

(15) More people have been to Russia than I have.

Almost everyone accepts (15) on first hearing. Try to say what it means and the sentence dissolves: more people than what? More people have been to Russia than the number of people I have been to Russia? There is no coherent reading. Here the processor delivers a verdict of acceptable for a string the grammar arguably does not generate at all. Acceptability, which is a behavioural response, and grammaticality, which is a theoretical claim about the system, are not the same variable, and the gap between them runs in both directions.

The upshot: Judgments are data about performance from which grammaticality is inferred. Any argument that treats a raw judgment as the theoretical claim itself will eventually be embarrassed by a centre embedding or by (15).

Why the rules you were taught are not the rules

A descriptive grammar states what speakers do; a prescriptive rule states what an authority wants them to do. Three of the best-known prescriptions have no basis in the syntax of English at all.

The ban on the split infinitive (to boldly go) was an invention of nineteenth-century handbooks, defended by analogy with Latin, where the infinitive is a single word and cannot be split. The ban on preposition stranding (the man who I spoke to) is usually traced to a remark by John Dryden in 1672. And the condemnation of negative concord (He don't know nothing) rests on an argument from logic that no language has ever obeyed: standard French, Italian, Spanish, Russian and Polish all use two negative elements for one negation, and so did Chaucer.

The decisive observation is not that prescriptions are snobbish. It is that they are irrelevant to the phenomena that need explaining. No teacher ever told you that (2) or (10) was wrong. Nobody had to. Meanwhile the rules teachers do state are exactly the ones that vary between speakers, need repeating every year, and are broken by the best writers in the language. That is the signature of a social convention, not of a mental system.

None of this makes prescription pointless. Standard varieties are real, gatekeepers use them, and knowing the conventions of formal written English is a practical skill. The point is division of labour: an editor's job is one thing, and stating the rules that let you reject (10) without ever having been taught anything about seem is another.

Are judgments good enough to build on?

A reasonable objection: this field rests on introspection by people who have a theory to defend. Jon Sprouse and Diogo Almeida tested the worry directly. They extracted 469 data points from Adger's textbook Core Syntax and put them to hundreds of naive participants in formal experiments with proper controls and statistics. The informally reported judgments replicated at about 98 per cent. A follow-up with Carson Schütze sampled sentence types at random from ten years of Linguistic Inquiry and reported replication around 95 per cent. The wholesale objection fails: textbook syntax is not built on artefacts.

The retail objection survives, and it is the one you should carry. The contested cases, the ones where a theory turns on a contrast between two mildly degraded strings, are exactly the cases where informal judgment is least reliable, and those are the places where formal acceptability experiments now do real work. When you meet a starred example in this course that you are not sure about, that reaction is data too. Write it down.

Common misconceptions

  • Grammatical means correct, the way school meant it. It means generated by the speaker's system. He ain't got none is grammatical in the variety that produces it and stigmatised in another; those are separate facts.
  • A grammatical sentence is one that makes sense. Sentence (1) is well formed and meaningless. Sentence (10) is interpretable and ill formed. The two dimensions cross.
  • If a sentence is hard to understand, it is ungrammatical. Sentence (13) is generated by the grammar and defeats your memory. Sentence (15) is easy and may not be generated at all.
  • An asterisk means nobody says it. It means speakers reject it when asked. Rare and rejected are different measurements, which is why corpus frequency alone cannot settle a syntactic question.
  • Syntax is just word order. Sentences (8) to (10) have identical word order in the relevant respect and differ in acceptability, and later lessons will show two readings sharing one order.

What to carry forward

  • Chomsky's pair separates grammaticality from both frequency and meaning: two equally unattested, equally meaningless strings, one of which the grammar generates.
  • The asterisk records rejection by native speakers, and the finer marks (?, ??, #) record degradation and semantic anomaly rather than ungrammaticality.
  • Acceptability is behaviour, grammaticality is theory, and they diverge in both directions: centre embeddings are grammatical and unusable, comparative illusions are usable and arguably ungrammatical.
  • Prescriptive rules are social conventions about a standard variety; they never explain the judgments a speaker makes without instruction, which is the explanandum here.
  • Large-scale replication has shown informal judgments to be reliable in bulk, at roughly 95 to 98 per cent, and least reliable exactly where theories are most contested.

Sources

  1. Wikipedia. (n.d.). Colorless green ideas sleep furiously. Wikimedia Foundation. en.wikipedia.org
  2. Scholz, B. C., Pelletier, F. J., Pullum, G. K., and Nefdt, R. (2024). Philosophy of linguistics. Stanford Encyclopedia of Philosophy. plato.stanford.edu
  3. Wikipedia. (n.d.). Grammaticality. Wikimedia Foundation. en.wikipedia.org
  4. Chomsky, N. (1957). Syntactic structures. Mouton.
  5. Sprouse, J., and Almeida, D. (2012). Assessing the reliability of textbook data in syntax: Adger's Core Syntax. Journal of Linguistics, 48(3), 609-652.
Key terms
Grammaticality
A theoretical property: the string is generated by the speaker's internalised grammar. Not the same as being frequent, meaningful or socially approved.
Acceptability
A behavioural response: what a speaker reports when asked whether a string is a possible sentence. The observable from which grammaticality is inferred.
Competence
A speaker's knowledge of the language, abstracted away from the memory, attention and time limits that shape actual use.
Performance
Language use under real conditions, where a perfectly grammatical sentence can still be unusable, as in triple centre embedding.
Asterisk
The prefix marking a string that native speakers reject as not a sentence of the language; a hash instead marks a well formed but anomalous one.
Descriptive rule
A statement of the regularity that speakers actually observe, answerable to judgments and corpora.
Prescriptive rule
A statement of what an authority holds speakers ought to do, such as the bans on split infinitives and stranded prepositions.
Comparative illusion
A string such as More people have been to Russia than I have, accepted on first hearing but with no coherent interpretation.

Constituency and the Five Tests, Run on Real Sentences

  • Apply substitution, movement, coordination, ellipsis and clefting to decide whether a string of words is a constituent.
  • Use the tests to separate the two structures behind a single ambiguous sentence, and write each as a labelled bracketing.
  • State why a failed test is weaker evidence than a passed one, naming two constructions that make the coordination test overgenerate.

An elephant, and two places it could be

In the 1930 film Animal Crackers, Groucho Marx reports having shot an elephant in his pajamas, and then wonders aloud how the elephant got into them. The joke works because eleven words with no ambiguous word in them support two different sentences. Nobody misreads shot, elephant or pajamas. What is ambiguous is not any word but how the words are grouped.

(1) I shot an elephant in my pajamas.

On the ordinary reading, the phrase in my pajamas tells you about the shooting: it is attached to the verb phrase. On Groucho's reading it tells you which elephant: it is attached to the noun phrase. Write the two groupings as labelled bracketings, which is how this course will show structure from here on, since a text page cannot draw a tree:

(1a) [TP I [VP [VP shot [NP an elephant]] [PP in my pajamas]]]
(1b) [TP I [VP shot [NP [NP an elephant] [PP in my pajamas]]]]

Two claims are being made there, and only one of them is obvious. The obvious claim is about attachment. The unobvious one is that an elephant in my pajamas forms a unit in (1b) and does not in (1a). A unit of that kind is a constituent, and this lesson is the procedure for finding out whether a given stretch of words is one.

What matters here: A constituent is not a stretch of words that feels like it belongs together. It is a stretch that behaves as a unit under operations the language performs, and behaviour is testable.

The procedure: five tests

The tests descend from immediate constituent analysis, developed by Leonard Bloomfield in the 1930s and made procedural by Zellig Harris in the 1940s and 1950s. Each has the same logic: certain operations in English take a constituent as their input and refuse anything else. If your string can undergo the operation, it is a constituent; that is the reasoning, and the whole of syntactic argument works like this.

TestOperationWorks onExample
SubstitutionReplace the string with a single pro-formNP (it, they), PP (there, then), VP (do so, do it)I shot it in my pajamas.
MovementTopicalise, front, or shift the stringMost phrasal categoriesIn my pajamas, I shot an elephant.
CoordinationJoin the string to a like string with and or orAnything, in principleI shot an elephant and a giraffe.
EllipsisDelete the string under identityVP mainly (VP ellipsis)I shot an elephant and Harpo did too.
CleftingPut the string in the focus of it-was-X or what-I-did-was-XNP, PP, VP, APIt was an elephant that I shot.

A sixth, informal but very useful, is the fragment test: a string that can stand alone as the answer to a question is a constituent. Ask what I shot, and an elephant is a possible complete answer. Ask what I shot and get shot an elephant and you have an answer to a different question.

Running the tests on the elephant

Take the candidate string an elephant in my pajamas and ask whether it is a constituent. Under the tests it behaves as one, but only on the reading where the elephant is wearing the pajamas:

(2) It was an elephant in my pajamas that I shot. (clefting: passes, Groucho reading only)
(3) An elephant in my pajamas, I shot. (topicalisation: marginal but passes on the same reading)
(4) I shot an elephant in my pajamas and a hippo in my dressing gown. (coordination: passes)
(5) I shot one. (substitution by one: forced onto the ordinary reading, since one here stands for elephant, not for the whole phrase)

Now take the rival candidate, the string shot an elephant, and ask whether it is a constituent on the ordinary reading:

(6) I shot an elephant in my pajamas and Chico did so in his dressing gown. (do so replaces shot an elephant: passes)
(7) What I did in my pajamas was shoot an elephant. (pseudo-cleft: passes)
(8) In my pajamas I shot an elephant. (the PP moves independently: passes)

Sentence (6) is the crucial one. The pro-form do so replaces a verbal constituent, and here it has replaced shot an elephant while leaving in my pajamas outside, which is only possible if the smaller VP is a unit. That is (1a). Sentence (2) is crucial for the other structure, because clefting has picked up the noun phrase and the prepositional phrase together, which is only possible if they are one unit. That is (1b). One string, two constituent structures, and the tests separate them.

The same frame, one verb changed

Now change a single word and watch the results invert. Compare:

(9) The student read the book in the library.
(10) The student put the book in the library.

Sentence (9) is ambiguous in the way (1) is: the reading may have happened in the library, or the book may be the library's. Sentence (10) is not ambiguous at all, and, more interestingly, it fails tests that (9) passes:

(11) The student read the book in the library and the teacher did so in the staff room. (fine)
(12) *The student put the book in the library and the teacher did so in the staff room. (bad)
(13) The student read the book, and in the library too. (fine)
(14) *The student put the book, and in the library too. (bad)

The failure of (12) says that put the book is not a constituent, so the structure of (10) has no smaller VP inside it:

(9a) [VP [VP read [NP the book]] [PP in the library]]
(10a) [VP put [NP the book] [PP in the library]]

The prepositional phrase in (10) is demanded by the verb; the one in (9) is optional and attaches outside. English has a name for the two relations, complement and adjunct, and Lesson 5 builds a whole theory of phrase structure on the difference. For now, notice that you discovered it with a pro-form and a coordination, not by consulting intuitions about meaning.

The point: Change one lexical item and the constituent structure of the rest of the sentence changes with it. Structure is not a property of the word string; it is a property of the analysis the grammar assigns to that string.

Where the tests mislead

A test that never failed would be suspicious, and these fail in both directions. Three cases are worth memorising, because a reader who has not met them will draw wrong conclusions with great confidence.

First, coordination overgenerates. Look at:

(15) John gave Mary a book and Sue a record.
(16) I lent my bicycle to Kim on Monday and my car to Lee on Friday.

The strings Sue a record and my car to Lee on Friday are coordinated with something, yet no standard analysis makes them constituents of a simple clause. These are gapping and its relatives, where an elided verb licenses coordination of what is left. If you take coordination as a sufficient test, you will posit constituents that no other test supports.

Second, right node raising does the same to the left edge:

(17) Kim wrote, and Lee illustrated, the entire book.

Here Kim wrote and Lee illustrated look coordinated, and a subject plus a verb is not a constituent by any other diagnostic. A shared element at the right edge is licensing the pattern.

Third, and most important as a habit of mind: a failed test is weak evidence. Movement is blocked by many things beyond constituency, so a string that resists fronting may be a perfectly good constituent that has been stopped by something else. Consider that a subject NP cannot be clefted in every position, or that some VPs resist topicalisation for reasons of information structure. Passing a test shows the string can act as a unit, which is a positive fact. Failing one shows only that the string did not act as a unit in that frame, which has several possible causes.

The working rule: converging positive results from two or three independent tests are a real argument. A single failure is a question, not a verdict.

Two ambiguities to finish on

Not every structural ambiguity involves a prepositional phrase. Try coordination:

(18) old men and women

(18a) [NP [AP old] [N-bar [N-bar men] and [N-bar women]]] (everyone is old)
(18b) [NP [NP [AP old] men] and [NP women]] (only the men are old)

The test that separates them is substitution with a plural pronoun plus a modifier, or simply adding a second adjective: old men and young women forces (18b), while very old men and women most naturally keeps (18a). And try a participle:

(19) Visiting relatives can be tiresome.

Here visiting is a verb in one structure, with relatives as its object and the whole thing a gerund phrase, and an adjective in the other, modifying relatives. The number of the verb settles it, which is itself a constituency argument: Visiting relatives is tiresome forces the gerund reading, because a gerund phrase is singular, while Visiting relatives are tiresome forces the modifier reading. Agreement is reaching into the structure and reporting what it finds there.

Common misconceptions

  • Words that go together in meaning form a constituent. In (10), put and the book go together in meaning as tightly as any verb and object, and they are not a constituent, because the verb also requires the location.
  • If a string can be coordinated it must be a constituent. Gapping and right node raising both coordinate non-constituents, which is why the tests have to converge before they convince.
  • A failed test proves the string is not a constituent. Movement and clefting are blocked by information structure, weight and locality as well as by constituency, so a negative result underdetermines the structure.
  • Ambiguity is a matter of ambiguous words. No word in the elephant sentence has two senses in play. The ambiguity is entirely in the grouping.

The short version

  • A constituent is a string that behaves as a unit under substitution, movement, coordination, ellipsis and clefting; the behaviour, not the intuition, is the evidence.
  • An ambiguous sentence with no ambiguous words has two constituent structures, and the tests pick them apart one reading at a time.
  • Do so substitution and coordination distinguish an adjunct prepositional phrase, as with read, from a complement one, as with put, which is the foundation of the X-bar theory two lessons ahead.
  • Coordination overgenerates because of gapping and right node raising, so no single test is decisive; convergence is.
  • Passing a test is positive evidence that a string can act as a unit; failing one has too many possible causes to be conclusive.

Sources

  1. Wikipedia. (n.d.). Constituent (linguistics). Wikimedia Foundation. en.wikipedia.org
  2. Wikipedia. (n.d.). Immediate constituent analysis. Wikimedia Foundation. en.wikipedia.org
  3. Wikipedia. (n.d.). Syntactic ambiguity. Wikimedia Foundation. en.wikipedia.org
  4. Carnie, A. (2021). Syntax: A generative introduction (4th ed.), chapter 3. Wiley Blackwell.
  5. Sportiche, D., Koopman, H., and Stabler, E. (2014). An introduction to syntactic analysis and theory, chapters 2 and 3. Wiley Blackwell.
Key terms
Constituent
A string of words that behaves as a single unit under the syntactic operations of the language, such as substitution, movement and coordination.
Pro-form
A single word standing in for a whole phrase: it and they for noun phrases, there and then for prepositional phrases, do so for verbal ones.
Do so substitution
Replacement of a verbal constituent by do so, which diagnoses the smaller verb phrase inside a larger one and separates complements from adjuncts.
Clefting
The it-was-X-that frame, and the pseudo-cleft what-I-did-was-X, both of which place a single constituent in focus.
Gapping
Coordination with an elided verb, as in John gave Mary a book and Sue a record, which makes the coordination test appear to accept a non-constituent.
Right node raising
Coordination sharing a rightmost element, as in Kim wrote, and Lee illustrated, the entire book, another case of apparent non-constituent coordination.
Structural ambiguity
Two readings of one word string produced by two groupings rather than by any ambiguous word.
Attachment
Which node a modifier hangs from, as with a prepositional phrase joined either to the verb phrase or to the noun phrase before it.

Lexical Categories, Diagnosed by Distribution

  • Assign a nonsense word to a lexical category using distributional frames rather than meaning.
  • Separate the lexical categories from the functional ones, and state why the split matters for clause structure.
  • Apply the NICE diagnostics to distinguish auxiliary verbs from lexical verbs, and use the gerund contrast to show that category is assigned to a use, not to a form.

What a reader knows about a tove

Lewis Carroll opened Through the Looking-Glass in 1871 with a stanza in which almost every content word is invented. Its first line reports that it was brillig and that the slithy toves did gyre and gimble in the wabe. You have no idea what a tove is. You know, immediately and without effort, that toves is a plural noun, that slithy is an adjective, that gyre and gimble are verbs, and that wabe is a noun. Ask yourself how.

Not from meaning: there is none. You know it from position and from ending. Toves sits after the and an adjective, and carries a plural -s. Gyre follows the auxiliary did. Wabe follows the and a preposition. Category membership in English is read off the environments a word can occur in, and this lesson makes that procedure explicit, because everything later in the course depends on it: phrase structure rules refer to categories, X-bar theory projects them, and selection is stated over them.

Why the school definitions fail

The definitions most readers were given are semantic, and each one breaks on ordinary vocabulary.

School definitionCounterexampleWhat goes wrong
A noun is a person, place or thingdestruction, arrival, whiteness, absence, tomorrow, threeThese name events, properties, times and numbers, and are nouns by every syntactic test
A verb is an action wordknow, resemble, cost, own, deserveNo action occurs; meanwhile assassination names an action and is a noun
An adjective describes a nounbrick in brick wall, the relative clause in the man who leftNouns and clauses describe nouns too, so the property does not pick out a class
An adverb ends in -lyfast, well, very, soon, quite; and friendly, lovely, cowardlyThe ending is neither necessary nor sufficient: the last three are adjectives

The counterexamples are not exotic. Destruction and know are among the commonest words in academic English. A definition that misclassifies the ordinary cases is not a definition with a few exceptions; it is the wrong kind of definition. Parts of speech are syntactic classes, and syntactic classes are defined by syntactic behaviour.

The core of it: A word's category is a summary of where it can go, not of what it means. Two words that mean nearly the same thing, such as destroy and destruction, belong to different categories because they occur in different places.

The frames

A distributional frame is a sentence with a hole in it. Words that fit the hole share a category. Learn these and you can categorise a word you have never seen.

CategoryFrames it fitsMorphologyTest words
Noun (N)The ___ is here. I saw two ___s. Kim's ___.Plural -s, possessive -stove, sincerity, arrival
Verb (V)They will ___. They ___ed. They are ___ing.-ed, -ing, third singular -sgyre, resemble, elapse
Adjective (A)a very ___ thing. It seems ___. ___er than that.Comparative -er, superlative -estslithy, tall, absent
Adverb (Adv)She ran ___. She ___ ran. Very ___.Often -ly, but not alwaysquickly, fast, soon
Preposition (P)right ___ the house. straight ___ the wall.Noneunder, through, during
Determiner (D)___ book is here. Cannot stack: *the my book.Nonethe, this, every, my
Complementiser (C)I think ___ she left. I wonder ___ she left.Nonethat, whether, if, for
Auxiliary (T)___ she leave? She ___n't leave.Modals take no -s, no -ingcan, will, must, have, be

Two warnings about using frames. First, a word may fit several frames because English freely converts one category into another: google, impact, friend and chair are all nouns and all verbs. The frame tells you the category of a word in a use, which is what syntax needs. Second, fitting one frame is weak; fitting a whole cluster is the argument. Very fits the adjective slot in the phrase the very model and is not an adjective, since it fails the comparative and the predicative frames: *very-er, *It seems very.

The NICE properties, or how to catch an auxiliary

English auxiliaries look like verbs and are not distributed like them. Four diagnostics, known since Rodney Huddleston's work as the NICE properties, separate them cleanly.

PropertyAuxiliaryLexical verb
Negation with n-apostrophe-tShe cannot leave. She hasn't left.*She likesn't fish.
Inversion in questionsCan she leave? Has she left?*Likes she fish?
Code, that is, stranding under ellipsisShe can, and he can too.*She likes, and he likes too.
Emphatic affirmationShe HAS left, whatever you heard.*She LIKES fish, requiring do support

A lexical verb needs do to perform any of the four, which is the phenomenon called do-support and the subject of Lesson 6. Within the auxiliaries, the modals form a subclass with morphology of their own: no third singular -s (*he cans), no infinitive (*to can), no participle (*having could), and no stacking (*he will can swim). That set of gaps is the reason modals are standardly analysed as sitting in a functional head rather than as ordinary verbs, which is where the clause spine of Lesson 6 comes from.

One form, two categories, in the same sentence

The strongest demonstration that category is a property of a use, not of a word shape, comes from English gerunds. Compare:

(1) John's constant reading of the book annoyed us.
(2) *John's constant reading the book annoyed us.
(3) John's constantly reading the book annoyed us.
(4) *John's constantly reading of the book annoyed us.

The word reading is identical in all four. In (1) it is a noun: it is modified by an adjective and takes an of phrase, exactly like destruction. In (3) it is a verb: it is modified by an adverb and takes a bare direct object, exactly like read. The starred lines show that you cannot mix the two profiles, which is what proves there are two structures rather than one flexible word. Write them out:

(1a) [NP John's [N-bar [AP constant] [N-bar [N reading] [PP of the book]]]]
(3a) [NP John's [VP [AdvP constantly] [VP [V reading] [NP the book]]]]

Why this matters: When you ask what category a word belongs to, you are asking about a token in a structure. The lexicon may list reading once or twice; the syntax cares only about which set of frames the token is satisfying.

Open classes, closed classes, and why syntax cares

Nouns, verbs, adjectives and adverbs are open classes: English acquired selfie, doomscroll and rewild within living memory. Determiners, complementisers, auxiliaries and prepositions are closed: the last genuinely new English preposition is hard to name, and the whole modal system has been shrinking rather than growing. Closed classes are also short, unstressed, and resistant to being coined.

The split is not a curiosity of vocabulary. Modern syntax puts functional items at the top of every phrase: a determiner takes a noun phrase as its complement, an auxiliary takes a verb phrase, a complementiser takes a clause. That is why Lesson 6 will call an ordinary clause a TP and an embedded one a CP rather than calling both S. The functional item is the head, and the lexical material is what it selects.

One honest caveat before you generalise. The four open classes are not universal in the same shape. Many languages have a closed adjective class of a dozen or so items, with property concepts otherwise expressed by verbs or nouns, a finding associated with R. M. W. Dixon's cross-linguistic survey. Whether Salish languages such as Straits Salish distinguish nouns from verbs at all has been argued for decades. The distributional method still applies; what it returns differs from language to language, which is precisely why you run the frames rather than assuming the English answer.

Common misconceptions

  • A noun names a person, place or thing. Then destruction, absence and tomorrow would not be nouns, and they pass every noun frame there is.
  • Adverbs are the words ending in -ly. Friendly, lovely and cowardly are adjectives; fast, soon and well are adverbs with no ending at all.
  • A word has one category, listed in the dictionary. English converts freely, and the gerund pair shows that one form can be a noun in one structure and a verb in another.
  • Auxiliaries are just verbs with vague meanings. They have a distinct distribution captured by the four NICE properties, and the modals have a morphology no lexical verb shows.
  • Every language has the same categories English has. Adjective classes of a dozen items are common, and the noun and verb distinction has been seriously contested for some languages.

Putting it together

  • Category is diagnosed by distribution: the set of frames a word can appear in, plus the inflection it accepts. Meaning-based definitions misclassify very common words.
  • Use clusters of frames, never one, because a single environment can be satisfied by a word of another category, as very shows.
  • The NICE properties (negation, inversion, code, emphasis) separate auxiliaries from lexical verbs, and the modals form a defective subclass within them.
  • The gerund quartet proves categories attach to tokens in structures: adjective plus of phrase gives a noun, adverb plus bare object gives a verb, and mixing the profiles is ungrammatical.
  • Open classes carry lexical content; closed functional classes head the phrases that lexical material sits inside, which is the architecture the next module builds.

Sources

  1. Wikipedia. (n.d.). Part of speech. Wikimedia Foundation. en.wikipedia.org
  2. Wikipedia. (n.d.). Auxiliary verb. Wikimedia Foundation. en.wikipedia.org
  3. Wikipedia. (n.d.). Conversion (word formation). Wikimedia Foundation. en.wikipedia.org
  4. Huddleston, R., and Pullum, G. K. (2002). The Cambridge grammar of the English language, chapters 1 and 3. Cambridge University Press.
  5. Baker, M. C. (2003). Lexical categories: Verbs, nouns, and adjectives. Cambridge University Press.
Key terms
Distributional frame
A sentence with a gap, used to test category membership: words that fit the gap share a syntactic class.
Open class
A category that readily admits new members, such as noun, verb, adjective and adverb.
Closed class
A small fixed category such as determiner, complementiser or auxiliary, which does not admit coinages.
NICE properties
Negation, inversion, code and emphasis: the four environments in which English auxiliaries behave differently from lexical verbs.
Modal
A defective auxiliary such as can, will or must, lacking third singular -s, an infinitive and a participle, and unable to stack.
Conversion
The use of a word of one category in another without any change of form, as with google, chair and impact.
Gerund
A word in -ing that heads either a nominal structure, taking adjectives and an of phrase, or a verbal one, taking adverbs and a direct object.
Functional head
A closed-class item such as D, T or C that takes a lexical phrase as its complement and heads the resulting phrase.

Module 2: Building Structure

From rewrite rules to X-bar theory to the clause spine: how a finite grammar generates unbounded structure, why every phrase turns out to have the same shape, and what sits above the verb phrase.

Phrase Structure Rules, Derivations and Labelled Bracketings

  • Write a small phrase structure grammar and derive a sentence from it step by step, ending in a labelled bracketing.
  • Show how recursion in a finite rule set yields unbounded sentences, and how one grammar assigns two structures to one ambiguous string.
  • State four defects of a flat phrase structure grammar that motivate the X-bar theory of the next lesson.

A paper for radio engineers

In 1956 Chomsky published an argument in IRE Transactions on Information Theory, a journal read by communications engineers, that no finite-state machine could generate English. His evidence was a pattern of nested dependencies. English pairs if with then, either with or, and either pair can be wrapped around a sentence containing the other:

(1) If either the tide turns or the wind drops, then we sail.

To get (1) right, a device has to remember, while processing the material in the middle, that an if is still open. Nest the pairs three deep and it has to remember three. A machine with a fixed number of states eventually runs out. What is needed is a device that can put a structure inside a structure of the same kind, and the simplest such device is a set of rewrite rules. This lesson is the procedure for writing them, running them, and finding out where they fail.

Writing a grammar

A phrase structure rule has the form X goes to Y Z, read as: a node labelled X may have as its immediate daughters a Y and a Z, in that order. Parentheses mark an optional element. Here is a grammar of six rules, which we will call G1:

S -> NP VP
NP -> (D) (AP) N (PP)
VP -> V (NP) (PP)
PP -> P NP
AP -> (Adv) A
CP -> C S

Notice what the rules are made of: category labels, all of them earned in the previous lesson by distributional testing. A rule that mentioned things and actions would be unusable, since you would have no test for whether a given word satisfied it.

Now derive a sentence. A derivation starts at S and rewrites one symbol at a time until only words remain.

StepRule appliedResult
1StartS
2S -> NP VPNP VP
3NP -> D ND N VP
4VP -> V NP PPD N V NP PP
5NP -> D ND N V D N PP
6PP -> P NPD N V D N P NP
7NP -> D ND N V D N P D N
8Insert wordsThe student read the book in the library

The derivation records more than the word order. It records which symbol each word came from and which symbols were rewritten together, and that record is the structure. Written as a labelled bracketing, it is:

[S [NP [D The] [N student]] [VP [V read] [NP [D the] [N book]] [PP [P in] [NP [D the] [N library]]]]]

A bracketing and a tree carry exactly the same information. Every opening bracket with a label is a node; matching brackets enclose everything that node dominates; two constituents whose brackets open at the same level under one parent are sisters. Reading a bracketing is a skill worth ten minutes of practice, because it is the only way this course can show you structure.

Bottom line: A derivation is not a model of how a speaker builds a sentence in real time. It is a proof that the grammar generates the string, together with the structural description it assigns.

Two derivations, one string

G1 will also derive the same word string a second way. Instead of step 4 attaching the prepositional phrase inside the verb phrase, take the option in the noun phrase rule:

[S [NP [D The] [N student]] [VP [V read] [NP [D the] [N book] [PP [P in] [NP [D the] [N library]]]]]]

The words are in the same order. The structures differ, and so does the meaning, exactly as Lesson 2's tests predicted. A grammar that assigns two structural descriptions to one string is telling you the string is structurally ambiguous, and it does so as a consequence of the rules rather than by any special provision. That is what makes phrase structure grammar an explanation rather than a notation.

Recursion, and why the grammar is finite while the language is not

Look at the third and fourth rules. A PP contains an NP, and an NP may contain a PP. Nothing stops the cycle:

(2) the book on the table in the kitchen of the house at the end of the lane

Six rules, applied over and over, generate a phrase of any length. The same holds for clauses through the last rule: a CP contains an S, and a VP can contain a CP.

(3) Kim said that Lee believed that Sam suspected that the tide had turned.

This is recursion, and it is the formal property that answers the question Lesson 1 raised. You can recognise sentences you have never heard because a finite set of rules generates an unbounded set of structures. No list, however long, could do it; nor could the finite-state machine Chomsky ruled out in 1956. Formally, G1 is a context-free grammar, and whether natural language stays inside that class is a real question with a known answer of no, established for Swiss German and for Bambara in the 1980s on the strength of crossing dependencies.

Where G1 goes wrong: the verb decides

Run G1 forwards and it produces strings no speaker accepts:

(4) *The student elapsed the book.
(5) *The student put the book.
(6) *Kim devoured.

Every one of these is generated by VP -> V (NP) (PP), because the parentheses make the object and the locative optional for every verb alike. The grammar has no way to say that elapse refuses an object, that put demands both an object and a locative, and that devour demands an object. That information belongs to the individual word, and it is stated in the lexicon as a subcategorisation frame.

VerbFrameGoodBad
elapseV, no complementThree hours elapsed.*Three hours elapsed the meeting.
devourV, NPKim devoured the pie.*Kim devoured.
putV, NP PPKim put the pie on the sill.*Kim put the pie.
thinkV, CPKim thinks that Lee left.*Kim thinks Lee the book.
tellV, NP CPKim told Lee that it rained.*Kim told that it rained.
wonderV, interrogative CPKim wonders whether it rained.*Kim wonders that it rained.

The last row matters: think and wonder both take a clause, and they select different kinds of clause. Selection is finer than category. This is the first appearance of an idea the whole course leans on, that heads impose requirements on their complements, and it will return as theta roles in Lesson 7 and as case in Lesson 8.

Four defects, and what fixes them

G1 works, and it is nonetheless the wrong shape, for reasons that took the field from 1957 to about 1970 to articulate.

  • Redundancy. The rules for NP, VP, AP and PP are written separately, yet all four say the same thing: an optional element, a head of the phrase's own category, and optional material after the head. Stating four rules where one generalisation exists is a failure of the theory, not a fact about English.
  • No intermediate node. The flat rule VP -> V NP PP gives no node covering just read the book, and Lesson 2 showed with do so that such a node has to exist.
  • Exocentricity. Every rule but one builds a phrase named after a category inside it. The rule S -> NP VP does not: the sentence is named after nothing it contains. That exception was eventually removed, and Lesson 6 shows what replaced it.
  • Optionality explosion. The rule NP -> (D) (AP) N (PP) abbreviates sixteen rules. Worse, it treats the determiner and the prepositional phrase as the same kind of optional, when Lesson 2 already showed that some post-head material is required and some is not.

So what?: Each of these defects is a place where the notation lets you write a grammar that no language would ever have. A better theory is one that cannot express the impossible rules, and that is the argument for X-bar theory.

Common misconceptions

  • A derivation models how a speaker builds a sentence. It is a static proof that the string is in the language, together with the structure assigned. Nothing in it claims speakers rewrite S first.
  • Trees show meaning. They show constituency and category. Meaning is computed over the structure, and the ambiguity in (2) shows structure constraining meaning, not encoding it.
  • A flat structure is simpler and so preferable. Fewer nodes is not the relevant kind of simplicity. The flat verb phrase cannot state the do so facts, so it is simpler and wrong.
  • Optional in a rule means optional in a sentence. Parentheses in G1 make an object optional for every verb, which is why the grammar generates *Kim devoured. Optionality is a property of a verb, stated in its lexical entry.

What you now know

  • Rewrite rules give a finite grammar that generates unboundedly many structures, which is why recursion answers the question of how you understand new sentences.
  • A derivation delivers both a word string and a structural description, and a string with two derivations is structurally ambiguous by consequence rather than by stipulation.
  • Labelled bracketings and trees are notational variants; sisterhood, dominance and precedence can be read off either.
  • Category is not enough: verbs subcategorise for particular complements, and selection distinguishes think from wonder though both take clauses.
  • Flat phrase structure rules are redundant, headless in the case of S, lack an intermediate node the data require, and can express rules no language uses.

Sources

  1. Wikipedia. (n.d.). Phrase structure rules. Wikimedia Foundation. en.wikipedia.org
  2. Wikipedia. (n.d.). Context-free grammar. Wikimedia Foundation. en.wikipedia.org
  3. Wikipedia. (n.d.). Subcategorization. Wikimedia Foundation. en.wikipedia.org
  4. Chomsky, N. (1956). Three models for the description of language. IRE Transactions on Information Theory, 2(3), 113-124.
  5. Carnie, A. (2021). Syntax: A generative introduction (4th ed.), chapters 4 and 5. Wiley Blackwell.
Key terms
Phrase structure rule
A rewrite rule stating that a node of one category may have a given sequence of daughters, as in VP goes to V NP PP.
Derivation
A sequence of rule applications from the start symbol to a string of words, which simultaneously fixes the structure assigned.
Labelled bracketing
A linear notation for a tree in which each constituent is enclosed in brackets tagged with its category.
Recursion
The property by which a category can contain a category of the same type, so that a finite grammar generates unboundedly many structures.
Context-free grammar
A grammar whose rules rewrite a single symbol without reference to its surroundings; natural languages are known to exceed this class.
Subcategorisation frame
A lexical specification of the complements a particular head requires or permits, such as put taking both an object and a locative.
Selection
A head's requirement on the type of its complement, finer than category, as with think taking a declarative and wonder an interrogative clause.
Exocentric
Of a phrase named after no category it contains, as with the old rule S goes to NP VP.

X-bar Theory: Heads, Complements, Specifiers, and Why Branching Is Binary

  • Derive the X-bar schema from the one-replacement and do-so facts, and write any phrase as an X-bar bracketing.
  • Apply six diagnostics that separate a complement from an adjunct, and predict the ordering facts that follow.
  • State the empirical argument for binary branching from double object anaphora, and say what happened to X-bar theory afterwards.

Two books, and a problem with the word one

You are in a bookshop holding a heavy anthology and you say to the assistant that you want the big book of poems with the red cover, not the one with the blue cover. The assistant understands perfectly: you want a big book of poems, in blue. Now try to say the other thing. You want a book of poems, not a book of recipes, and you reach for the same trick:

(1) I want the big book of poems with the red cover, not the one with the blue cover.
(2) *I want this book of poems, not that one of recipes.

Sentence (2) fails. The word one can stand for big book of poems, leaving the cover phrase behind, and it cannot stand for book alone, leaving of poems behind. Since one is a pro-form, and Lesson 2 established that pro-forms replace constituents, the noun phrase must contain a node covering book of poems and a node covering big book of poems, and no node covering book to the exclusion of of poems. The flat rule NP -> (D) (AP) N (PP) from the last lesson provides none of these. Something between the noun and the noun phrase is missing.

(These judgments are textbook judgments, and (2) is milder for some speakers than the star suggests. Run it past two people before you build on it, which is the habit Lesson 1 recommended.)

The missing level, and the shape it turns out to have

Call the missing node N-bar. The structure that satisfies the one facts is:

[NP [D the] [N-bar [AP big] [N-bar [N-bar [N book] [PP of poems]] [PP with the red cover]]]]

Three N-bar nodes, nested. The innermost contains the noun and its of phrase; each modifier wraps another N-bar around it. Now one has three possible antecedents, and the facts fall out: it can be book of poems, or big book of poems, or book of poems with the red cover, and it can never be book on its own, because no node has that extension.

The do so facts from Lesson 2 have exactly the same shape one category over. Do so could replace read the book and leave in the library, and it could not replace put and leave the book on the sill. So verb phrases have V-bar nodes with the same nesting. Adjectives behave alike: very proud of his son allows so to replace proud of his son. So do prepositional phrases. Four categories, one pattern. X-bar theory, developed by Chomsky in 1970 and given its general form by Ray Jackendoff in 1977, states the pattern once, with X a variable over categories:

XP -> (specifier) X-bar
X-bar -> X-bar (adjunct) or (adjunct) X-bar
X-bar -> X (complement)

Three schemata replace the whole rule list. Every phrase now has a head of its own category, which kills the exocentricity defect; the intermediate node exists, which kills the second defect; and the four rules that repeated themselves are now one, which kills the first. The optionality explosion goes too, because the positions are defined by the schema rather than listed.

Key idea: X-bar theory is not a new set of rules for English. It is a claim that human phrases have only one shape, and that what varies between categories and between languages is which items fill the slots.

Complement or adjunct: six ways to tell

The schema puts complements as sisters of the head and adjuncts as sisters of X-bar. Six consequences follow, and each is a test you can run.

DiagnosticComplementAdjunct
Distance from the headAdjacent: student of physics with long hairOutside: *student with long hair of physics
How manyAt most one per headFreely iterated: a book on the shelf in the corner of the room
Pro-form replacementIncluded in what one or do so stands forCan be left outside it
ObligatorinessOften required: *Kim put the bookAlways optional
SelectionThe head fixes its category and typeCombines with any head of the right kind
ExtractionEasy: Which subject did you meet a student of?Hard: *Which colour did you meet a student with hair of?

The ordering fact in the first row is not a stipulation about English. It follows from the geometry: if the complement is the head's sister and the adjunct attaches higher, the complement must be nearer the head on the side where the head takes its complement. That is the kind of result a good theory produces, a fact you did not put in coming out of the structure you did.

The specifier is the third slot, sister to X-bar and daughter of XP: the determiner in a noun phrase on the older analysis, the possessor in John's book, the degree word in so proud, and, crucially for Lesson 6, the subject in a clause. Unlike adjuncts, there is at most one, and unlike complements, it does not have to be selected.

Why two daughters and never three

The schema as written allows a head to take one complement, which forces every branch to be binary. That was a deliberate choice, argued for by Richard Kayne in 1984 on formal grounds, and the strongest empirical support comes from a construction that looks flat: the double object.

(3) I showed Mary herself in the mirror.
(4) *I showed herself Mary in the mirror.
(5) I showed the boys each other in the photograph.
(6) *I showed each other the boys in the photograph.
(7) I gave no one anything.
(8) *I gave anyone nothing.

A reflexive needs an antecedent that c-commands it, a relation Lesson 13 defines precisely; for now, take it that a node c-commands its sister and everything inside it. If ditransitive verbs had a flat structure, with the verb and both objects as sisters, then each object would c-command the other, and (3) and (4) would be equally good, as would (5) and (6). They are not. The first object behaves as though it is structurally higher than the second, in three independent phenomena: reflexive binding, reciprocal binding, and negative polarity licensing. Barss and Lasnik established this asymmetry in a three-page note in Linguistic Inquiry in 1986, and Richard Larson's 1988 response built the layered verb phrase, the so-called VP shell, that is still the standard analysis.

flat: [VP [V showed] [NP Mary] [NP herself]] both objects sisters, wrongly symmetric
shell: [VP [NP Mary] [V-bar [V showed] [NP herself]]] first object higher, asymmetry predicted

Why this matters: Binary branching is not tidiness. It is what makes structural relations asymmetric, and asymmetric relations are what binding, polarity and scope turn out to need.

Writing one out

Take the phrase the recent students of syntax from Utrecht. Work from the head outward. The head is students. Its complement is of syntax, selected by the noun. From Utrecht and recent are adjuncts, one after and one before. The is the specifier.

[NP [D the] [N-bar [AP recent] [N-bar [N-bar [N students] [PP of syntax]] [PP from Utrecht]]]]

Check it against the tests. Iterate an adjunct and the structure accommodates it: the recent students of syntax from Utrecht with grants. Try to iterate the complement and it fails: *the students of syntax of semantics. Replace with one and it works at each N-bar: the recent ones from Utrecht, where ones is students of syntax. The analysis and the diagnostics agree, which is the standard by which a structure is defended.

What happened next

X-bar theory dominated the 1980s and then dissolved into something simpler. In the minimalist program of the mid-1990s, Chomsky replaced the schema with a single operation, Merge, that takes two objects and forms a set whose label is that of one of them. Binary branching is then not a stipulation but the definition of the operation, and X-bar levels become relational: a head is what has not projected yet, a phrase is what projects no further. This is bare phrase structure. The X-bar labels survive in practice because they are legible, and this course keeps using them for exactly that reason.

Common misconceptions

  • The bar level is a kind of category. N-bar is not a new part of speech. It is a projection of N, and in bare phrase structure it is not even that: it is a position in a derivation.
  • Complements are obligatory and adjuncts are optional, and that is the difference. Optionality is only one of six diagnostics, and it fails on eat, whose object is a complement and optional.
  • Binary branching is imposed for elegance. The double object asymmetries in binding, reciprocals and polarity are evidence, and a flat structure makes the wrong prediction three times over.
  • Any modifier that comes first is a specifier. Adjectives in English precede the noun and are adjuncts, since they iterate freely and can be left outside one.

Where this leaves us

  • One-replacement and do so replacement force an intermediate projection between the head and the phrase, and the same node is needed in four categories.
  • The X-bar schema states the shape once: a specifier under XP, adjuncts adjoined to X-bar, a complement sister to the head.
  • Six diagnostics separate complements from adjuncts, and the ordering restriction follows from the geometry rather than being stipulated.
  • The double object asymmetries in reflexives, reciprocals and negative polarity argue that even apparently flat structures are layered, and that branching is binary.
  • Minimalism replaced the schema with Merge, keeping binary branching as a consequence; the labels remain as a convenient notation.

Sources

  1. Wikipedia. (n.d.). X-bar theory. Wikimedia Foundation. en.wikipedia.org
  2. Wikipedia. (n.d.). Bare phrase structure. Wikimedia Foundation. en.wikipedia.org
  3. Barss, A., and Lasnik, H. (1986). A note on anaphora and double objects. Linguistic Inquiry, 17(2), 347-354.
  4. Larson, R. K. (1988). On the double object construction. Linguistic Inquiry, 19(3), 335-391.
  5. Jackendoff, R. (1977). X-bar syntax: A study of phrase structure. MIT Press.
Key terms
X-bar schema
The category-neutral template stating that a phrase consists of a specifier and an X-bar, that adjuncts adjoin to X-bar, and that the head takes a complement.
Head
The lexical item that determines the category of the phrase built on it and selects its complement.
Complement
The phrase merged as sister to the head, selected by it, at most one per head, and included in pro-form replacement.
Adjunct
A modifier adjoined to X-bar, freely iterated, never selected, and excludable from what a pro-form stands for.
Specifier
The single position sister to X-bar and daughter of XP, filled by a possessor, a degree word, or, in a clause, the subject.
One-replacement
Substitution of one or ones for a nominal constituent, which diagnoses the N-bar level and shows that complements are inside it.
Binary branching
The requirement that every node have at most two daughters, supported by the c-command asymmetries of the double object construction.
VP shell
Larson's layered verb phrase, in which the first object of a ditransitive is structurally higher than the second.
Merge
The minimalist operation combining two syntactic objects into one, from which binary branching follows by definition.

The Clause: TP, CP, the Subject Position, and Head Movement

  • Replace the exocentric rule for S with a TP headed by tense, and place the subject in its specifier.
  • Explain do-support and subject-auxiliary inversion as head movement, and say why English lexical verbs do not raise.
  • State the Extended Projection Principle and give three kinds of evidence for a subject position that no argument fills.

Where the tense goes

Chomsky's 1957 grammar contained a rule that looks eccentric now and turned out to be the seed of everything in this lesson. It said that the auxiliary of an English clause consists of a tense marker, then optionally a modal, then optionally have followed by the ending -en, then optionally be followed by the ending -ing. The oddity is that the endings are listed one step to the left of the words they end up attached to. Written out for a fully loaded clause:

Tense + have + -en + be + -ing + watch gives has been watching

Each affix hops rightwards onto the next word. The past tense hops onto have, the -en hops onto be, the -ing hops onto watch. The rule is called affix hopping, and it explains, with one statement, why the order of auxiliaries in English is rigid and why each one determines the shape of the next word rather than of itself.

Now watch what happens when something gets in the way.

(1) John left.
(2) *John not left.
(3) John did not leave.
(4) *Left John?
(5) Did John leave?

In (2) the negation sits between the tense marker and the verb, and the affix cannot hop across it. English then does something no other Germanic language does: it inserts a meaningless verb, do, purely to carry the stranded tense. That is do-support, and its existence is an argument that tense is a syntactic object in its own right, separate from the verb it usually appears on. You cannot strand something that is not there.

The point: Do-support is the visible residue of an invisible head. English inserts a dummy verb because a tense feature needs a host, which tells you that tense occupies a structural position of its own.

From S to TP

Give that position a label. The head is T, for tense; it takes the verb phrase as its complement; and the phrase it projects is TP. The subject sits in the specifier of TP. The last exocentric rule in the grammar disappears, and the clause takes the same shape as every other phrase:

[TP [NP John] [T-bar [T did] [VP [V leave]]]]
[TP [NP John] [T-bar [T -ed] [VP [V leave]]]]

The second line is the ordinary clause, with the tense affix in T and no auxiliary to carry it, so it lowers onto the verb. Modals are simply lexical items generated in T, which is why *He will can swim is bad: there is one T per clause, so there is room for one modal. That single fact, which needed a special stipulation in the old grammar, now falls out of the structure. The label IP, for inflection phrase, is the same object under an older name, and you will meet both in the literature.

1989, and a French adverb

Jean-Yves Pollock published a paper in Linguistic Inquiry in 1989 whose central data are four sentences and an adverb. In French the finite verb precedes the adverb; in English it follows it, and swapping them fails in both languages.

FrenchGlossEnglish
Jean embrasse souvent Marie.Jean kisses often Marie*John kisses often Mary.
*Jean souvent embrasse Marie.Jean often kisses MarieJohn often kisses Mary.
Jean ne mange pas de chocolat.Jean not eats not of chocolate*John eats not chocolate.
Marie a souvent mange.Marie has often eatenMary has often eaten.

Assume the adverb sits in the same place in both languages, adjoined at the left edge of the verb phrase. Then the word order difference is not a difference in where the adverb goes; it is a difference in where the verb goes. In French the finite verb raises out of the verb phrase into T, crossing the adverb. In English it stays put, and the tense affix comes down to it instead. The fourth row is the control: a French participle does not raise, and lands on the English side of the adverb.

This is head movement, and it is a second kind of displacement alongside the phrasal movement of Lesson 10. A head leaves its position and adjoins to the next head up:

French: [TP Jean [T-bar [T embrasse] [VP souvent [VP [V --] Marie]]]]
English: [TP John [T-bar [T -s] [VP often [VP [V kiss] Mary]]]]

English is not wholly without verb raising. Auxiliary be and, for many speakers, have do raise, which is why they behave like French verbs:

(6) John is not happy. (*John does not be happy.)
(7) John is often late. (*John often is late is possible but means something else.)
(8) Is John happy?

So the generalisation is not that English lacks head movement. It is that in English only auxiliaries raise, while in French all finite verbs do. Pollock connected the difference to the richness of verbal agreement morphology, a link that has been argued about ever since and that Lesson 15 revisits when it deals with parameters.

Above T: the complementiser phrase

Now (5). If did is in T, where does it go when it precedes the subject? It moves to the next head up, and that head is C, the position occupied by that, whether and if in embedded clauses. So an English yes-no question is T-to-C movement, otherwise known as subject-auxiliary inversion:

[CP [C did] [TP [NP John] [T-bar [T --] [VP leave]]]]

Three predictions follow immediately, and all three are correct. First, only auxiliaries invert, because only what is in T can move to C: *Left John? Second, an embedded clause with an overt complementiser cannot also invert, because C is already filled: *I wonder whether did he leave. Third, exactly one element inverts, because there is one T: *Has been John working?

The clause is now two layers: a CP that carries clause type and hosts complementisers and fronted material, and a TP that carries tense and hosts the subject. Every embedded clause you met in Lesson 4 fits:

[CP [C that] [TP [NP the letter] [T-bar [T had] [VP arrived]]]]

The subject that is not an argument

One question remains about the specifier of TP: does anything have to be in it? English says yes, insistently, even when there is no participant to put there.

(9) It rains. (*Rains.)
(10) There is a man in the garden. (*Is a man in the garden, as a statement.)
(11) It seems that John left. (*Seems that John left.)
(12) It is likely that Kim will win.

The it of (9) refers to nothing; the there of (10) is not a place. These are expletives, pronounced elements with no semantic content at all, inserted to satisfy a purely structural requirement. That requirement is the Extended Projection Principle, or EPP: every clause must have a subject.

Three further pieces of evidence show the EPP is structural rather than semantic. Infinitival clauses obey it too, with a silent subject: Kim wants [to leave], where something has to be understood as the leaver. Weather verbs take the expletive even in languages that mark it differently, and languages that permit a null subject, such as Italian and Spanish, are analysed as satisfying the EPP with an unpronounced pronoun rather than as lacking the requirement, which is the pro-drop parameter of Lesson 15. And, decisively for the structural reading, the expletive appears exactly where a real subject would and takes the same case and the same agreement: There are three men in the garden, with plural agreement reaching past the expletive to the associate.

Worth holding on to: A subject is a position, not a role. Some subjects bear a semantic role, some bear none, and the grammar treats the position as obligatory either way.

Common misconceptions

  • Do in do-support means something. It contributes nothing. Compare emphatic do, which does contribute, and note that the dummy appears only where a stranded tense needs a host.
  • English word order shows that English has no head movement. Auxiliaries raise to T and on to C; the restriction is to auxiliaries, not to the language.
  • The subject is the doer of the action. The subject of (9) to (12) does nothing and refers to nothing, which is why the position and the role must be separated.
  • TP is just S renamed. The relabelling brings consequences: one modal per clause, only auxiliaries inverting, one inverted element per question, and a specifier that must be filled.

Looking back

  • Affix hopping and do-support together show that tense occupies a structural position independent of the verb that usually carries it.
  • Calling that position T makes the clause endocentric: TP has a head, a complement verb phrase, and a specifier holding the subject.
  • Pollock's adverb data show French finite verbs raising to T and English ones staying in the verb phrase, with English auxiliaries as the exception that proves the mechanism.
  • Subject-auxiliary inversion is T-to-C movement, and it predicts correctly that only auxiliaries invert, that only one does, and that inversion is blocked when C is filled.
  • The EPP requires a subject position to be filled whether or not anything means anything there, which is what expletive it and there exist to do.

Sources

  1. Wikipedia. (n.d.). Do-support. Wikimedia Foundation. en.wikipedia.org
  2. Wikipedia. (n.d.). Extended projection principle. Wikimedia Foundation. en.wikipedia.org
  3. Wikipedia. (n.d.). Subject-auxiliary inversion. Wikimedia Foundation. en.wikipedia.org
  4. Pollock, J.-Y. (1989). Verb movement, Universal Grammar, and the structure of IP. Linguistic Inquiry, 20(3), 365-424.
  5. Emonds, J. (1978). The verbal complex V-prime to V in French. Linguistic Inquiry, 9(2), 151-175.
Key terms
T (tense)
The functional head carrying tense and finiteness, taking the verb phrase as complement and the subject in its specifier.
TP
The projection of T, which replaces the old exocentric S and makes the clause endocentric like every other phrase.
CP
The projection of C above TP, hosting complementisers, fronted wh-phrases and inverted auxiliaries, and encoding clause type.
Affix hopping
The lowering of an inflectional affix onto the adjacent verb, blocked when negation intervenes.
Do-support
Insertion of a semantically empty do to host a tense feature stranded by negation, inversion, ellipsis or emphasis.
Head movement
Displacement of a head to the next head up, as in French verb raising to T and English T-to-C inversion.
Expletive
A pronounced element with no semantic content, such as it in it rains or there in there is a man, inserted to fill the subject position.
EPP
The Extended Projection Principle: the requirement that the specifier of TP be filled in every clause, whether or not an argument is available.

Module 3: Arguments, Roles and Case

What a verb demands of its surroundings: the roles it assigns, the two kinds of intransitive it turns out to hide, the case its arguments receive, and the two constructions that look identical and are not.

Theta Roles, Argument Structure, and the Two Kinds of Intransitive

  • Assign theta roles to the arguments of a verb and apply the theta criterion to rule out extra and missing arguments.
  • Run five English diagnostics that separate unaccusative from unergative verbs, and interpret the result structurally.
  • Read auxiliary selection and ne-cliticisation data from Italian, Dutch and German as evidence for the same split.

Don't giggle me

In diary records of her own children's speech, Melissa Bowerman collected utterances that no adult produces and that no adult ever said to the child. Among them: a request not to be giggled, and a report of intending to fall something on somebody. Set them beside the adult sentences the child had certainly heard:

(1) The ice melted. / Kim melted the ice.
(2) The gate opened. / Kim opened the gate.
(3) The child giggled. / *Kim giggled the child.
(4) The vase fell. / *Kim fell the vase.

The child has spotted a real pattern in (1) and (2), where the same verb appears with one argument and with two, and has extended it to (3) and (4), where English refuses. So the pattern is real and it is restricted, and no obvious feature of meaning explains the restriction: falling and melting are both things that happen to an object without its cooperation, and only one of them alternates. Something about the verbs is different, and this lesson is about what.

What matters here: The number of noun phrases a verb appears with is a fact about the verb, and the facts group verbs into classes with several properties each. Argument structure is where the lexicon touches the syntax.

Roles, and the criterion that governs them

Start with the vocabulary. A verb assigns theta roles, semantic relations that its arguments bear to the event it describes.

RoleDefinitionExample, role-bearer in italics
AgentDeliberate initiator of the eventKim opened the gate.
Theme or patientEntity moved, affected or predicated ofKim opened the gate.
ExperiencerEntity in a mental stateLee feared the dog.
GoalEndpoint of a transfer or motionKim sent the parcel to Lee.
SourceStarting pointKim took the book from the shelf.
InstrumentMeans by which the event happensKim opened the gate with a crowbar.
BeneficiaryParty for whom the event is doneKim baked a cake for Lee.

The roles are not the point; the constraint over them is. The theta criterion says that every argument bears exactly one theta role and every theta role of a predicate is borne by exactly one argument. That single statement rules out both kinds of failure at once:

(5) *Kim devoured. (a role is unassigned)
(6) *Kim arrived the parcel the shelf. (arguments with no role available)
(7) *It rained Kim. (rain assigns no role, so nothing can be there)

Adjuncts are outside all of this, which is a useful test in itself. With a crowbar can be omitted from any sentence with no loss of grammaticality and can be stacked; the object of devour can do neither. That is the same complement-and-adjunct distinction of Lesson 5, now stated in semantic terms.

One intransitive class, or two?

Return to (3) and (4). The subject of giggle is an agent: the child does the giggling. The subject of fall is a theme: the vase undergoes the falling and initiates nothing. David Perlmutter proposed in 1978 that this semantic split corresponds to a structural one, and the proposal, known as the Unaccusative Hypothesis, is that the two verbs have different underlying structures despite identical surface strings:

unergative: [TP [NP the child] [T-bar [T -ed] [VP [V giggle]]]]
unaccusative: [TP [NP the vase] [T-bar [T -ed] [VP [V fall] [NP --]]]]

In an unergative clause the single argument starts as the subject. In an unaccusative clause the single argument starts as the object and moves up to the subject position, which the EPP of Lesson 6 requires it to fill. The dashes mark where it came from. That is a strong claim about an invisible difference, so it needs diagnostics, and English supplies five.

Five English tests

TestUnaccusative (fall, arrive, melt, die)Unergative (giggle, dance, work, run)
Existential thereThere arrived three visitors.*There giggled three children.
Prenominal participlethe fallen leaves, a recently arrived guest*the giggled child, *a danced student
Bare resultativeThe river froze solid.*Dora shouted hoarse (needs herself)
Causative alternationKim melted the ice. (for a subclass)*Kim giggled the child.
Locative inversionOut of the house ran the dog.Restricted, and worse for many speakers

The second row deserves a moment, because it is the cleanest. A prenominal past participle in English modifies a noun that would be the object of the corresponding verb: the stolen jewels, the written letter, the eaten cake. That it also gives the fallen leaves and a recently arrived guest, and never *the giggled child, is exactly what you expect if the subject of fall is an underlying object and the subject of giggle is not. The test is measuring the invisible structure directly.

The third row explains the child's error better than any appeal to meaning does. Only a subclass of unaccusatives alternates: melt, open, break, sink do, while arrive, fall, die, appear do not. Verbs that alternate describe changes of state that can plausibly be caused externally; verbs that do not describe changes that happen internally to the entity. The child in Bowerman's records had the syntax right and the lexical semantics not yet sorted, which is exactly the kind of error you would predict on this analysis.

The same split in three other languages

If the distinction were an artefact of English, it would not turn up in unrelated constructions elsewhere. It does, and the cross-linguistic evidence is what made the hypothesis stick.

LanguageDiagnosticUnaccusativeUnergative
ItalianPerfect auxiliaryMaria e arrivata (be)Maria ha telefonato (have)
Italianne-cliticisationNe sono arrivati tre. (three of them arrived)*Ne hanno telefonato tre.
GermanPerfect auxiliaryEr ist gefallen (he is fallen)Er hat gearbeitet (he has worked)
DutchImpersonal passive*Er wordt gevallenEr wordt hier veel gedanst (there is much dancing here)

The Italian clitic ne is the sharpest of these. It can be extracted from an object and not from a subject: Ne ho letti tre, I have read three of them, is fine, and the corresponding extraction from a subject is not. That ne can be extracted from the surface subject of arrivare and not from the surface subject of telefonare says the first is an underlying object. The Dutch impersonal passive runs the other way: it requires an agent, so it accepts unergatives and refuses unaccusatives. Four diagnostics in three languages, sorting the same verbs into the same two piles.

The upshot: A single surface pattern, noun phrase plus intransitive verb, hides two structures, and half a dozen unrelated constructions can tell them apart. That convergence is the standard of evidence syntax aims at.

Linking, and how far it goes

If roles predict positions, the mapping should be stateable. Mark Baker's 1988 Uniformity of Theta Assignment Hypothesis says that identical thematic relationships are represented by identical structural relations at the level where structure is built, which is what licenses the inference from role to underlying position that this lesson has been making. A thematic hierarchy, with agent above experiencer above goal above theme, then predicts which argument becomes the subject.

It is not a complete theory, and the honest place to stop is with the residue. Psychological predicates split awkwardly: Kim fears the dog puts the experiencer in subject position, and The dog frightens Kim puts it in object position, with the same two participants and the same roles. Beth Levin and Malka Rappaport Hovav argued in a book-length study in 1995 that unaccusativity is semantically determined and syntactically encoded, and also that no single semantic feature draws the line: telicity, agentivity and internal causation each account for part of the split. Expect a class of verbs to have borderline members. Run is unergative in Kim ran for an hour and behaves unaccusatively in Kim ran to the store, which is a fact about the construction as much as about the verb.

Common misconceptions

  • The subject is the agent. The subject of fall, arrive and seem is not an agent, and Lesson 6 already showed subjects that are not arguments at all.
  • Unaccusative means the verb takes no object. Both classes are intransitive on the surface. The claim is about where the single argument starts, not about how many there are.
  • Every intransitive verb belongs cleanly to one class. Verbs of manner of motion switch class with a goal phrase, and the two classes have contested borders.
  • The child who says do not giggle me has made a random error. The child has generalised a rule English restricts, and the restriction is lexical-semantic rather than syntactic.

Pulling it together

  • Theta roles name the relations arguments bear to an event; the theta criterion, one role per argument and one argument per role, rules out both surplus and missing arguments.
  • Intransitive verbs fall into two classes: unergatives, whose argument begins as the subject, and unaccusatives, whose argument begins as the object and raises.
  • Five English diagnostics converge on the split, and the prenominal participle test is the most direct, since that participle otherwise modifies objects.
  • Italian auxiliary selection and ne-cliticisation, German auxiliary selection and the Dutch impersonal passive sort the same verbs the same way.
  • Linking generalisations map roles onto positions, but psychological predicates and manner-of-motion verbs show the mapping is not fully determinate.

Sources

  1. Wikipedia. (n.d.). Unaccusative verb. Wikimedia Foundation. en.wikipedia.org
  2. Wikipedia. (n.d.). Theta role. Wikimedia Foundation. en.wikipedia.org
  3. Wikipedia. (n.d.). Causative alternation. Wikimedia Foundation. en.wikipedia.org
  4. Perlmutter, D. M. (1978). Impersonal passives and the Unaccusative Hypothesis. Proceedings of the Fourth Annual Meeting of the Berkeley Linguistics Society, 157-189.
  5. Levin, B., and Rappaport Hovav, M. (1995). Unaccusativity: At the syntax-lexical semantics interface. MIT Press.
Key terms
Theta role
A semantic relation a verb assigns to one of its arguments, such as agent, theme, experiencer or goal.
Theta criterion
The requirement that each argument bear exactly one theta role and each theta role be borne by exactly one argument.
Argument
A phrase required or licensed by a predicate and assigned a role by it, as opposed to an adjunct, which is neither.
Unergative verb
An intransitive verb whose single argument is an agent generated in subject position, such as giggle, dance or work.
Unaccusative verb
An intransitive verb whose single argument is a theme generated in object position and raised to subject, such as fall, arrive or melt.
Causative alternation
The pattern in which one verb appears both intransitively and transitively, as with melt and open, available only to a subclass of unaccusatives.
Auxiliary selection
The choice of a be or have perfect auxiliary, which tracks the unaccusative and unergative split in Italian, German and Dutch.
UTAH
Baker's hypothesis that identical thematic relations are represented by identical structural relations where structure is built.

Case: Structural, Inherent, Exceptional, and Burzio's Generalisation

  • State the Case Filter and use it to explain why an infinitival subject needs for and why a noun needs of.
  • Distinguish structural from inherent case using German passives and Icelandic quirky subjects.
  • Derive the passive and the unaccusative from Burzio's generalisation, and analyse an exceptional case marking clause.

Six words

Modern English marks case on six words. I and me, we and us, he and him, she and her, they and them, who and whom. You and it have given up the distinction, every noun in the language gave it up centuries ago, and the possessive -s is now a clitic that attaches to phrases (the king of Spain's daughter) rather than a case ending. Old English had five cases on every noun; the residue is six pronouns and a stigmatised whom.

You would expect a theory of case to be a small chapter in the morphology of a few pronouns. It is not. Case turns out to control where a noun phrase is allowed to appear, in English as strictly as in Latin, and the argument for that is the pattern in these four:

(1) I believe that he is honest.
(2) I believe him to be honest.
(3) *I believe he to be honest.
(4) *I believe him is honest.

Whatever governs the choice, it is not meaning: (1) and (2) mean the same thing. It is not the verb believe, which occurs in both. It is the clause the pronoun sits in, or more exactly the head above it. A finite clause gives its subject nominative; a non-finite one cannot, and the pronoun has to get its case from somewhere else.

Key idea: Case is a licensing condition on positions, not a decoration on words. English shows it on six pronouns and obeys it everywhere.

The Case Filter, and two constructions it explains

The generalisation, stated in Chomsky's 1981 lectures, is the Case Filter: every noun phrase with phonological content must be assigned case. A noun phrase in a position where no head assigns case is ungrammatical no matter how well it fits semantically.

Two English constructions exist for no other reason. First, infinitival subjects:

(5) *John to leave early would be a surprise.
(6) For John to leave early would be a surprise.
(7) *I prefer Kim to stay. (fine for most speakers, but compare the next line)
(8) I prefer for Kim to stay.

Non-finite T assigns no case, so the subject of an infinitive is stranded. English rescues it by inserting the complementiser for, a preposition-like element that assigns accusative. Second, nominals:

(9) The army destroyed the city.
(10) *the destruction the city
(11) the destruction of the city
(12) *proud his son / proud of his son

Nouns and adjectives assign theta roles and do not assign case. The theta criterion is satisfied in (10) and the Case Filter is not, so English inserts a semantically empty of to license the noun phrase. Two otherwise unrelated repairs, both explained by one filter, is the sort of result that makes a theory worth having.

Two kinds of case, compared

Not all case behaves alike, and the cleanest evidence is what happens under passivisation.

Structural caseInherent case
Assigned byFinite T gives nominative; V and P give accusativeA particular head, listed in its lexical entry
Depends on theta roleNo: the expletive there gets nominative and no roleYes: assigned along with a role by the same head
Survives passivisationNo: the object becomes nominativeYes: the argument keeps its case
Typical examplehim in I saw himGerman dative with helfen; Icelandic quirky subjects

German makes the contrast visible in four short sentences. Sehen, to see, assigns structural accusative; helfen, to help, assigns dative.

ActiveGlossPassiveWhat happened to the case
Der Mann sieht das Kind.the man sees the-ACC childDas Kind wird gesehen.Accusative becomes nominative
Der Mann hilft dem Kind.the man helps the-DAT childDem Kind wird geholfen.Dative is kept, subject position stays empty

The second passive is impersonal: nothing becomes the nominative subject, because the dative argument keeps the case helfen gave it and has no need to move. Case that is tied to a particular verb travels with the argument; case that comes from a position is lost when the position changes.

Icelandic pushes this further. Its quirky subjects bear dative or accusative and nonetheless behave as subjects on every diagnostic the language has: they raise, they control the understood subject of an infinitive, they invert in questions, and they bind reflexives. Zaenen, Maling and Thrainsson showed in 1985 that grammatical function and morphological case are separable dimensions, which is why a theory needs both a Case Filter and an EPP rather than one requirement doing both jobs.

Case across a clause boundary

Now return to (2). The pronoun him is the subject of the embedded infinitive and carries accusative, which no part of the embedded clause can supply. The standard analysis is exceptional case marking: the matrix verb assigns accusative across the clause boundary into the specifier of the embedded TP.

[TP I [T-bar [T -- ] [VP [V believe] [TP [NP him] [T-bar [T to] [VP be honest]]]]]]

Is him really the embedded subject rather than the object of believe? Three tests say yes, and each will return in the next lesson.

(13) I believe there to be a problem. (believe does not take an expletive object: *I believe there.)
(14) I believe the cat to be out of the bag. (the idiom survives, so its parts must sit together in the embedded clause)
(15) I believe John to be lying. does not entail I believe John.

Expletives and idiom chunks cannot receive theta roles, so their appearance in that position shows the matrix verb is not assigning one there. What it assigns is case and nothing else. Passivising the matrix verb confirms it: There is believed to be a problem, with the expletive now nominative.

Burzio's generalisation, and two things it explains at once

Luigi Burzio noticed in his 1986 study of Italian that two properties of verbs travel together: a verb assigns accusative case if and only if it assigns an external theta role. No verb has one property without the other.

Apply it to the passive. A passive participle suppresses the external role: nobody is named as the reader in The book was read. By the generalisation it therefore cannot assign accusative. Its object is then a noun phrase with no case, which the Case Filter forbids, so the object must move somewhere that has case, and the only such position is the subject, which the EPP wants filled anyway.

[TP [NP the book] [T-bar [T was] [VP [V read] [NP --]]]]

Now apply it to the unaccusatives of Lesson 7. They too assign no external role, so they too assign no accusative, so their single argument must raise. Passive and unaccusative are the same derivation arrived at two different ways, which is precisely the unification Burzio's paper claimed, and it explains why the two constructions share so many properties across languages.

The core of it: One biconditional links a semantic property, the external role, to a syntactic one, accusative case, and from it the passive and the unaccusative both follow without any further stipulation.

Where the theory has moved since

Two honest qualifications. First, Burzio's generalisation is a generalisation, not a principle: it describes a robust correlation nobody has fully derived, and every few years a paper argues that some class of verbs violates it. Second, the whole apparatus was rebuilt in minimalism, where case is checked through the Agree relation rather than assigned under government, and the government that held it together in 1981 no longer exists. A more radical line, associated with Alec Marantz's 1991 paper, argues that case is not what licenses noun phrases at all, and that morphological case is computed after the syntax by comparing the noun phrases in a domain. That approach, dependent case theory, is very much alive. The empirical generalisations in this lesson survive all of it; the mechanism attached to them does not.

Common misconceptions

  • English has no case, so case theory cannot matter for English. English has six case-marked pronouns and the same distributional restrictions as a richly inflected language, which is the whole argument of this lesson.
  • In I believe him to be honest, him is the object of believe. Expletives and idiom chunks appear there, and the sentence does not entail believing him, so the noun phrase is the embedded subject.
  • Case marks the semantic role. Nominative goes to expletives with no role; Icelandic quirky subjects keep dative while functioning as subjects. Case and role are separate dimensions.
  • Of in the destruction of the city is a meaningful preposition. It is inserted to satisfy the Case Filter, since nouns assign roles and not case.

The takeaway

  • The Case Filter requires every pronounced noun phrase to receive case, and it explains both for before infinitival subjects and of inside nominals.
  • Structural case comes from a position and is lost when the position changes; inherent case comes from a particular head and survives passivisation, as German dative and Icelandic quirky subjects show.
  • Exceptional case marking has a matrix verb assigning case into an embedded subject position, which the expletive and idiom tests confirm is not an object position.
  • Burzio's generalisation ties accusative case to the external theta role, and from it the passive and the unaccusative derivations both follow.
  • The mechanism has been rebuilt twice since 1981, and dependent case theory now denies that case licenses noun phrases at all; the descriptive generalisations have outlasted every version.

Sources

  1. Wikipedia. (n.d.). Exceptional case-marking. Wikimedia Foundation. en.wikipedia.org
  2. Wikipedia. (n.d.). Burzio's generalization. Wikimedia Foundation. en.wikipedia.org
  3. Wikipedia. (n.d.). Quirky subject. Wikimedia Foundation. en.wikipedia.org
  4. Burzio, L. (1986). Italian syntax: A government-binding approach. Reidel.
  5. Zaenen, A., Maling, J., and Thrainsson, H. (1985). Case and grammatical functions: The Icelandic passive. Natural Language and Linguistic Theory, 3(4), 441-483.
Key terms
Case Filter
The requirement that every noun phrase with phonological content be assigned case, which rules out otherwise well formed structures.
Structural case
Case assigned by a position, such as nominative from finite T or accusative from V, independent of theta role and lost when the position changes.
Inherent case
Case assigned by a particular head together with a theta role, listed lexically and preserved under passivisation.
Quirky subject
A subject bearing dative or accusative rather than nominative, as in Icelandic, which passes all subject diagnostics nonetheless.
Exceptional case marking
Assignment of accusative by a matrix verb into the subject position of an embedded infinitive, as in I believe him to be honest.
Burzio's generalisation
The claim that a verb assigns accusative case if and only if it assigns an external theta role.
Of-insertion
The appearance of a semantically empty of before the complement of a noun or adjective, which cannot assign case.
Dependent case
Marantz's alternative view on which morphological case is computed after the syntax by comparing noun phrases in a domain, not by licensing them.

Raising Against Control: Debugging an Analysis That Looks Right

  • Apply the expletive, idiom-chunk, selection and passive-synonymy tests to tell a raising predicate from a control predicate.
  • Write the two structures, with a trace in one and PRO in the other, and say what each position is doing.
  • Extend the diagnosis to object control, exceptional case marking and tough-movement, and identify the verb that breaks the usual rule.

The analysis a careful reader would write

Two sentences, seven words each, identical in every part of speech and in every position:

(1) John seems to have left.
(2) John hopes to have left.

Here is the analysis a careful reader writes, and it is wrong. Both sentences have John as the subject of the matrix verb. Both have an infinitival complement whose understood subject is also John. Therefore both have the same structure, differing only in which verb fills the matrix position. The reasoning is clean, the surface evidence supports it entirely, and by the end of this lesson you will be able to say exactly which step fails and why.

Start with the step that looks safest: that John is the subject of the matrix verb in both. Test it by asking what role the matrix verb assigns to its subject. Hope requires a hoper, an entity capable of hoping. Does seem require a seemer?

(3) The rock seems to have fallen.
(4) *The rock hopes to have fallen.
(5) It seems to be raining.
(6) *It hopes to be raining.

Sentences (3) and (5) are fine, and their subjects are not doing any seeming. That is the crack in the analysis, and the rest of the lesson widens it.

Four tests, each fatal on its own

Expletives. A theta role cannot be assigned to an element that refers to nothing. So a predicate that allows an expletive subject is assigning no role to that position.

(7) There seems to be a problem. / *There hopes to be a problem.
(8) It seems to be obvious that Kim lied. / *It hopes to be obvious that Kim lied.

Idiom chunks. Parts of an idiom have no independent meaning and so can bear no independent role. An idiom keeps its idiomatic reading only if its parts start out together.

(9) The cat seems to be out of the bag. (idiomatic reading available)
(10) The cat hopes to be out of the bag. (only the literal reading, about an actual cat)
(11) The shit seems to have hit the fan. / #The shit hopes to have hit the fan.

Selectional restrictions. Whatever the embedded verb demands of its subject is what the matrix subject must satisfy, and the matrix verb adds nothing. The rock seems to have fallen is fine because fall is happy with a rock. It is the embedded verb doing all the selecting.

Passive synonymy. The sharpest test of the four. Passivise the embedded clause and see whether the meaning survives.

(12) The doctor seems to have examined Bill. = Bill seems to have been examined by the doctor.
(13) The doctor hopes to examine Bill. is not equal to Bill hopes to be examined by the doctor.

Line (12) records the same fact twice; line (13) describes two different people's hopes. If the matrix subject were assigned a role by the matrix verb, changing which noun phrase occupies that position would change who has that role, which is exactly what happens in (13) and does not happen in (12).

The point: Four independent diagnostics, drawing on reference, idioms, selection and truth conditions, all say the same thing: seem assigns no role to its subject, and hope does.

The two structures

If seem assigns no external role, then by the theta criterion its subject position starts empty and something must move into it, because the EPP requires it filled. The mover is the subject of the embedded clause, which cannot stay where it is because non-finite T assigns no case. Raising is therefore forced twice over, once by the EPP and once by the Case Filter of Lesson 8.

raising: [TP [NP John] [T-bar [T -s] [VP [V seem] [TP [NP --] [T-bar [T to] [VP have left]]]]]]

Control is different at every point. Hope assigns a role to its subject, which therefore starts where it stands. The embedded clause has a subject of its own, an unpronounced pronoun called PRO, which receives the embedded verb's role and is interpreted as identical to the matrix subject.

control: [TP [NP John] [T-bar [T -s] [VP [V hope] [CP [TP [NP PRO] [T-bar [T to] [VP have left]]]]]]]

Count the theta roles. In the raising sentence, John has exactly one, from leave. In the control sentence, John has one from hope, and PRO has one from leave: two roles, two arguments, and the theta criterion is satisfied without either argument doing double duty. That count is the whole difference, and every test above is a way of measuring it.

PropertyRaising (seem, appear, happen, likely, tend)Control (hope, try, persuade, eager, promise)
Roles borne by matrix subjectOne, assigned by the embedded verbOne, assigned by the matrix verb
Expletive subjectPermittedBlocked
Idiom reading survivesYesNo
Embedded passive synonymousYesNo
Finite paraphrase with expletiveIt seems that John left.*It hopes that John left.
Embedded subject positionTrace of movementPRO

The same debugging on adjectives, and on objects

Adjectives split the same way. Likely raises; eager controls:

(14) John is likely to win. / It is likely that John will win. / There is likely to be a riot.
(15) John is eager to win. / *It is eager that John will win. / *There is eager to be a riot.

Objects split three ways, which is where most readers go wrong. Compare:

(16) Kim persuaded Lee to leave. (object control: Lee is persuaded, and Lee leaves)
(17) Kim believed Lee to have left. (exceptional case marking: Lee is not believed, the proposition is)
(18) Kim promised Lee to leave. (subject control: Kim leaves, not Lee)

Run the tests. Kim believed there to be a problem is fine and *Kim persuaded there to be a problem is not, which shows that persuade assigns a role to its object and believe does not; that is the exceptional case marking of Lesson 8, sometimes called raising to object after Paul Postal's 1974 monograph. Sentence (18) is the odd one out in a different way. The general rule, Rosenbaum's Minimal Distance Principle, is that the controller is the nearest noun phrase, which gives the object for persuade. Promise violates it, and the violation is visible in acquisition: in a study published in 1969, Carol Chomsky found that children up to about eight years old interpret John promised Bill to leave as though Bill leaves, applying the general rule to the verb that breaks it.

One more construction, and a warning

A third pattern looks like both and is neither. Chomsky's 1964 pair:

(19) John is eager to please.
(20) John is easy to please.

In (19) John does the pleasing; in (20) John is pleased. Tough-movement relates the subject to the object position of the embedded verb, which is why It is easy to please John is a paraphrase and *It is eager to please John is not, and why the embedded verb has no object of its own in (20). The warning: two sentences can share a word order, a part-of-speech sequence and an adjective slot and still involve three different dependencies. Surface identity is not evidence of structural identity, which is the lesson the whole of this course keeps re-teaching.

Bottom line: The failed step in the opening analysis was the assumption that occupying the subject position means bearing the matrix verb's role. Position and role were separated in Lesson 6 by expletives and in Lesson 8 by quirky case; raising is the same separation seen a third time.

Common misconceptions

  • Raising and control differ in meaning, so the difference is semantic. The difference is structural, and the semantic effects follow from the number of theta roles rather than the other way round.
  • PRO is just an omitted pronoun. It occupies a case-less subject position no pronounced pronoun can occupy, which is why *John hopes him to leave fails while John believes him to have left succeeds.
  • The nearest noun phrase always controls. That is the default, and promise breaks it, which is why children get it wrong for years.
  • John is easy to please and John is eager to please are the same construction. In the first, John is the understood object of please; the two differ in which position the subject is related to.

Summing up

  • A raising predicate assigns no external theta role, so its subject position is filled by movement from the embedded clause; a control predicate assigns one, so its subject is generated in place.
  • Four independent tests detect the difference: expletive subjects, idiom chunk preservation, selectional restrictions, and synonymy under embedded passivisation.
  • The structures differ in what occupies the embedded subject position: a trace under raising, the null pronoun PRO under control.
  • Objects divide three ways: object control with persuade, exceptional case marking with believe, and subject control with promise, which breaks the Minimal Distance Principle.
  • Tough-movement relates the matrix subject to an embedded object position, a third dependency wearing the same surface shape as the other two.

Sources

  1. Wikipedia. (n.d.). Raising (linguistics). Wikimedia Foundation. en.wikipedia.org
  2. Wikipedia. (n.d.). Control (linguistics). Wikimedia Foundation. en.wikipedia.org
  3. Wikipedia. (n.d.). Tough movement. Wikimedia Foundation. en.wikipedia.org
  4. Chomsky, C. (1969). The acquisition of syntax in children from 5 to 10. MIT Press.
  5. Davies, W. D., and Dubinsky, S. (2004). The grammar of raising and control: A course in syntactic argumentation. Blackwell.
  6. Landau, I. (2013). Control in generative grammar: A research companion. Cambridge University Press.
Key terms
Raising predicate
A predicate such as seem, appear or likely that assigns no theta role to its subject, whose subject position is filled by movement.
Control predicate
A predicate such as hope, try, persuade or eager that assigns a theta role to its subject or object, which then controls the reference of PRO.
PRO
The unpronounced subject of a control infinitive, occupying a position no case-marked pronoun can fill and interpreted through its controller.
Expletive test
Substituting there or non-referential it for the matrix subject; only a predicate assigning no role to that position accepts it.
Idiom chunk test
Checking whether an idiomatic reading survives; it does under raising, because the parts of the idiom start out together.
Passive synonymy test
Passivising the embedded clause; the meaning is preserved under raising and changes under control.
Minimal Distance Principle
The default that the nearest noun phrase controls PRO, which gives object control for persuade and is violated by promise.
Tough-movement
The dependency in John is easy to please, relating the matrix subject to the object position of the embedded verb.

Module 4: Displacement, Islands and Silence

Words pronounced far from where they belong, the boundaries that stop them travelling, and the constructions in which the structure is there but nothing is pronounced at all.

Wh-Movement: Finding the Gap and Proving Something Was There

  • Identify the gap in a wh-question or relative clause and show that the fronted phrase satisfies the requirements of the gap position.
  • Give five independent arguments that the dependency is movement rather than base-generation, including wanna-contraction and parasitic gaps.
  • Explain successive cyclicity and cite two languages whose morphology makes the intermediate landing site visible.

A verb with a missing object

Lesson 4 established that put demands both an object and a locative, and that leaving either out is fatal:

(1) Kim put the book on the table.
(2) *Kim put on the table.

Now consider a question that speakers of English produce a hundred times a week:

(3) What did Kim put on the table?

Nothing follows put but the locative. By the reasoning of (2), sentence (3) should be as bad as (2) is, and it is perfect. Either the subcategorisation frame for put is wrong, which would be surprising given how well it worked, or there is something in the object position that is not pronounced. The second option is wh-movement, and this lesson is an argument that it is real.

[CP [NP What] [C did] [TP [NP Kim] [T-bar [T --] [VP [V put] [NP --] [PP on the table]]]]]

The wh-phrase begins in the object position, where put can satisfy its frame and assign it a theta role, and moves to the specifier of the CP built above the clause. The auxiliary moves to C, which is the T-to-C movement of Lesson 6, and the two operations together give the word order. The dash in the object position is a trace, or in more recent terms an unpronounced copy.

Why this matters: The gap is not an absence. It is a position that is filled for every purpose the grammar cares about, and unpronounced for the one purpose the ear can check.

Five arguments that something moved

A sceptic could say the wh-phrase is simply generated at the front and understood as connected to a genuinely empty position, with no movement at all. Five kinds of evidence make that hard to maintain.

One: the fronted phrase satisfies requirements of the gap site. Selection, case and preposition stranding all reach across the distance.

(4) Who did you talk to? (the stranded preposition needs an object)
(5) Whom did you see? (accusative from the verb, in the conservative variety that still marks it)
(6) *What did Kim elapse? (the frame of the gap-site verb still rules)

Two: idiom chunks travel. An idiom's parts must originate together, as Lesson 9 established. They do:

(7) What headway did they make on the problem?
(8) How much heed did anyone pay to the warning?

Headway and heed occur essentially nowhere except with make and pay. That they can be fronted while keeping the idiom shows they started in the idiom's object slot.

Three: reconstruction. A reflexive inside the fronted phrase is interpreted as though it were still in the gap:

(9) Which picture of himself does John like best?

Himself is bound by John, which follows it in the string. Lesson 13 will show that a reflexive must be c-commanded by its antecedent, and John does not c-command the front of the sentence. It does c-command the object position, which is where the phrase must therefore be evaluated.

Four: parasitic gaps. An extra gap can be licensed by a real one:

(10) Which book did you file without reading?
(11) *You filed the book without reading.

The gap after reading is impossible on its own and becomes possible when a wh-dependency runs past it. A parasitic gap is licensed by the presence of another gap, which is hard to state at all without traces.

Five: wanna-contraction. The most quoted piece of evidence in the literature, and the most fun to test on yourself. The string Who do you want to succeed? has two readings: you want someone to succeed, or you want to succeed someone. Now contract:

(12) Who do you wanna succeed? (only the second reading survives)

On the first reading the gap sits between want and to, and something unpronounced is blocking the contraction, exactly as an overt word would: *I wanna Kim succeed. On the second reading the gap is after succeed and nothing intervenes. An invisible element that blocks a phonological process is about as close to a physical demonstration as syntax gets. Read the caveat too: some speakers accept both readings contracted, and the fact has been argued about for fifty years.

Where it lands, and where it stops on the way

The landing site is the specifier of CP. Two facts confirm it. First, the fronted phrase precedes an auxiliary that has itself moved to C, so it must be higher than C. Second, in embedded questions the wh-phrase and the complementiser compete for the same region: standard English allows I wonder what Kim said and not *I wonder what that Kim said, an effect known as the doubly filled COMP filter, which several dialects and older stages of English do not obey.

Movement can also cross more than one clause:

(13) What do you think that Kim said that Lee put on the table?

The gap is three clauses down. Does the phrase travel in one leap, or in short steps through each intermediate specifier of CP? Successive cyclic movement is the standard answer, and two languages make the intermediate steps audible.

LanguageWhat is visibleWhat it shows
IrishThe complementiser has one form in a clause a wh-dependency passes through and another where it does notThe path of the movement is marked on every clause it crosses
Colloquial GermanThe wh-word is repeated in the intermediate clause, as in the pattern who do you think who she lovesA copy is pronounced at the intermediate landing site
Child EnglishChildren produce the same repetition, as in what do you think what the monster eatsThe intermediate copy is available before the pronunciation rule is learned
West Ulster EnglishA quantifier can be stranded at the intermediate clause edgeThe moving phrase was there long enough to leave something behind

Irish is the cleanest case, worked out by James McCloskey: the language has two complementisers, and the choice between them tracks exactly which clauses a filler-gap dependency has crossed. A theory in which the phrase jumps directly to the front has nothing to say about why an embedded complementiser three clauses down should change shape.

Not every language does it, and not every phrase does

English fronts one wh-phrase and leaves the rest in place:

(14) Who bought what?
(15) *Who what bought?

Other languages make other choices, and the three-way pattern is one of the standard illustrations of parametric variation, which Lesson 14 takes up properly.

StrategyLanguageExample pattern
Front exactly oneEnglish, German, IrishWho bought what?
Front all of themBulgarian, Romanian, Serbo-CroatianWho whom saw?
Front noneMandarin, Japanese, KoreanYou bought what? with ordinary question intonation or a question particle

Wh-in-situ languages are not simpler. Mandarin obeys many of the same locality restrictions English does, which suggests the dependency is there and only the pronunciation differs.

Finally, wh-questions are one member of a family. Relative clauses, clefts and topicalisation all form the same kind of long-distance dependency, and all of them will be stopped by the same boundaries in the next lesson. Collectively this is A-bar movement, so named because the landing site is not an argument position, and it contrasts with the A-movement of raising and passive, where a noun phrase moves into a position that could host an argument.

A-movementA-bar movement
ExamplesPassive, raising, unaccusative subjectWh-questions, relatives, clefts, topicalisation
Landing siteSpecifier of TP, an argument positionSpecifier of CP, a non-argument position
DistanceLocal, one clause at a timeUnbounded, though constrained as Lesson 11 shows
MotivationCase and the EPPClause typing and information structure

Common misconceptions

  • The gap is simply nothing. It satisfies subcategorisation, receives a theta role, hosts case, licenses a parasitic gap, and blocks a contraction. It is as busy as any pronounced position.
  • Wh-movement is a fact about question words. Relatives, clefts and topicalisation move the same way and obey the same constraints; the wh-word is incidental.
  • Languages without wh-fronting lack the dependency. Mandarin and Japanese show locality effects in the same places, which is evidence for a dependency that is simply not pronounced at the front.
  • Movement goes straight from the gap to the front. Irish complementiser choice and German wh-copying make the intermediate stops visible.

What to remember

  • A wh-question contains a gap that behaves in every respect like an ordinary occupied position, which is why the verb's subcategorisation frame is still satisfied.
  • Five independent arguments support movement: satisfaction of gap-site requirements, idiom preservation, reconstruction of binding, parasitic gap licensing, and blocked wanna-contraction.
  • The landing site is the specifier of CP, above the C position that hosts inverted auxiliaries and complementisers.
  • Long movement is successive cyclic, and Irish complementisers and colloquial German wh-copying make the intermediate landing sites audible.
  • Languages front one wh-phrase, all of them, or none, and the in-situ languages still show the locality signature of a dependency.

Sources

  1. Wikipedia. (n.d.). Wh-movement. Wikimedia Foundation. en.wikipedia.org
  2. Wikipedia. (n.d.). Parasitic gap. Wikimedia Foundation. en.wikipedia.org
  3. Wikipedia. (n.d.). Wh-in-situ. Wikimedia Foundation. en.wikipedia.org
  4. McCloskey, J. (2002). Resumption, successive cyclicity, and the locality of operations. In S. D. Epstein and T. D. Seely (Eds.), Derivation and explanation in the minimalist program (pp. 184-226). Blackwell.
  5. Adger, D. (2003). Core syntax: A minimalist approach, chapters 9 and 10. Oxford University Press.
Key terms
Wh-movement
Displacement of an interrogative or relative phrase to the specifier of CP, leaving a gap in the position where it was interpreted.
Gap
The unpronounced position from which a phrase has moved, which still satisfies subcategorisation, bears a theta role and can carry case.
Trace
The formal representation of that position, treated in more recent work as an unpronounced copy of the moved phrase.
Reconstruction
Interpretation of a fronted phrase as though it were still in its gap position, visible in the binding of a reflexive inside it.
Parasitic gap
A second gap, impossible on its own, licensed by the presence of a genuine wh-dependency running past it.
Wanna-contraction
The reduction of want to, blocked when a gap intervenes, which makes an unpronounced element phonologically detectable.
Successive cyclicity
Movement in short steps through each intermediate specifier of CP rather than in one leap.
A-bar movement
Movement to a non-argument position such as the specifier of CP, as in questions, relatives, clefts and topicalisation.

Islands: Six Hundred Pages of What Cannot Be Done

  • State the complex noun phrase, coordinate structure, sentential subject, adjunct and wh-island constraints, and show each one failing on a minimal pair.
  • Explain how subjacency compresses Ross's list into one locality condition, and what Rizzi's Italian data did to the bounding nodes.
  • Describe the that-trace effect, the French que to qui alternation, and the adverb effect that rescues it.

Six hundred pages of what cannot be done

In 1967 John Robert Ross handed MIT a dissertation called Constraints on Variables in Syntax, published nineteen years later as a book titled Infinite Syntax!, exclamation mark and all. It was almost entirely a catalogue of failures: strings that every rule then on the books said should be well formed, and that no speaker of English accepts. Ross had noticed what the previous lesson left dangling. Wh-movement, we said, is unbounded. If it were, extraction from anywhere should be fine. It is not.

(1) What do you think that Kim bought?
(2) *What do you believe the claim that Kim bought?

The two sentences differ by four words. In (1) the gap is the object of bought in a plain complement clause; in (2) it is the object of bought in a clause hanging off the noun claim. Both are object positions, both are the same distance away in words, and you can work out in half a second what (2) is trying to ask. It is still dead. Ross's word for the offending structure was island: a region a dependency cannot escape from.

Remember: An island is not a semantic problem or a length problem. It is a structural region, and the same words in a different structure let the same dependency through.

Ross's islands, one minimal pair at a time

In each pair below, the good sentence and the bad one ask the same question of the world.

The complex noun phrase constraint. Nothing may be extracted from a clause that sits under a noun, whether that clause is a complement or a relative.

(3) Who did Kim report that the committee had hired?
(4) *Who did Kim report the rumour that the committee had hired?
(5) *Which candidate did Kim meet the professor who had hired?

The coordinate structure constraint. Nothing may be extracted out of one conjunct of a coordination, and no conjunct may be extracted whole.

(6) *What did Kim buy a hammer and?
(7) *What did Kim buy and a hammer?
(8) *Which shelf did Kim paint the wall and put the atlas on?

The sentential subject constraint. A clause in subject position is sealed; the same clause as a complement is not.

(9) Which candidate did it bother the committee that nobody had interviewed?
(10) *Which candidate did that nobody had interviewed bother the committee?

The adjunct island. Adverbial clauses seal too, which is why the parasitic gaps of the last lesson were surprising.

(11) What did Kim leave the party because she wanted? (fine as an echo, bad as a real question)
(12) *Which book did Kim fall asleep while reading?

The wh-island. A clause whose specifier of CP is already occupied by a wh-phrase resists a second dependency crossing it.

(13) What do you think Kim bought?
(14) ??What do you wonder whether Kim bought?
(15) ??Which atlas do you wonder who bought?

The double question mark is deliberate. Wh-islands are the weakest of the set: many speakers rate (14) clumsy rather than impossible. Ross's list also included the left branch condition, which blocks stranding a noun by extracting its possessor or determiner.

(16) Whose atlas did you borrow?
(17) *Whose did you borrow atlas?

IslandThe sealed regionStrength in English
Complex noun phraseAny clause under a noun, complement or relativeVery strong, no gradience
Coordinate structureEither conjunct of a coordination, and the conjunct itselfVery strong, with one systematic exception
Sentential subjectA clause occupying the subject positionStrong
AdjunctAdverbial and purpose clausesStrong, weakened by parasitic gap contexts
Wh-islandAn embedded question whose specifier is filledWeak to moderate, argument gaps survive better
Left branchThe specifier of a noun phrase, in EnglishStrong in English, absent in Russian and Polish

The exception that shows the constraint is about structure

Sentence (6) is impossible, and yet this is fine:

(18) What did Kim buy and Lee sell?

The difference is that (18) has a gap in every conjunct. Extraction across the board is permitted; extraction from one conjunct alone is not. If the constraint were about processing load, (18) should be worse than (6), because it carries two dependencies rather than one. It is instead perfect.

[CP [NP What] [C did] [TP [NP Kim] [VP [VP [V buy] [NP --]] and [VP [NP Lee] [V sell] [NP --]]]]]

Why this matters: Across-the-board extraction is the cleanest evidence in the whole island literature that the constraints are structural. A memory-based account has to explain why doubling the dependencies repairs the sentence.

Chomsky's compression, and what it cost

Six constraints, each stipulated, is a description rather than an explanation. Chomsky's 1973 paper Conditions on Transformations proposed one condition to derive most of them. Call two categories bounding nodes, in English NP and the clause node we now write TP. Subjacency says a single movement step may not cross more than one of them.

Successive cyclic movement is what makes this work. In (1) the wh-phrase hops to the embedded specifier of CP, crossing one TP, then to the matrix specifier of CP, crossing one more: each step is legal. In (4) there is no usable hatch, because leaving the CP under rumour means crossing that NP and the matrix TP in a single step. In (14) the hatch is physically occupied by whether.

One statement thus buys the complex noun phrase constraint, the wh-island and the sentential subject constraint together. It does not buy the coordinate structure constraint, which stayed a stipulation, and it says nothing about why subject extraction differs from object extraction. That second gap is what the empty category principle was built to fill: a trace must be properly governed, and an object trace gets that from the verb while a subject trace does not.

Italian does not have the same islands

If subjacency were a fixed fact about the human parser, every language should show the same island profile. In 1982 Luigi Rizzi published Issues in Italian Syntax and reported that Italian speakers accept sentences whose English translations are the marginal (14) and (15).

ItalianGlossEnglish counterpart
Tuo fratello, a cui mi domando che storie abbiano raccontato, era molto preoccupato.your brother, to whom I wonder what stories they told, was very worriedMarginal at best
Il solo incarico che non sapevi a chi avrebbero affidato...the only assignment that you did not know to whom they would entrustMarginal at best

Rizzi's proposal was that the bounding nodes are not universal. English takes NP and TP; Italian takes NP and CP. Extraction from an embedded question in Italian crosses two TPs but only one CP, so it survives, while the complex noun phrase constraint holds in Italian as in English because both settings include NP. One parameter, set two ways. Lesson 14 takes the idea seriously.

The upshot: Islands are not a single universal shape. The complex noun phrase constraint looks close to universal; the wh-island demonstrably varies, and it varies in a way that a single binary choice describes.

Two words that cannot stand next to a gap

English has a smaller effect that has resisted every attempt to make it go away. Extract an object across an overt complementiser and nothing happens. Extract a subject across one and the sentence dies.

(19) Who do you think that Kim saw?
(20) Who do you think Kim saw?
(21) Who do you think left?
(22) *Who do you think that left?

This is the that-trace effect, and the pattern is exact: optional in (19) and (20), obligatory to omit in (21) and (22). The empty category principle was the standard account, since a subject trace has no governing verb and an adjacent complementiser blocks licensing from above. Two facts constrain any replacement. French first: que surfaces as qui precisely when the subject of its clause has been extracted, which suggests the complementiser is doing something for the gap rather than merely sitting near it.

FrenchGlossStatus
Qui crois-tu que Marie a vu?who do you think that Marie has seenObject extraction, que
*Qui crois-tu que a vu Marie?who do you think that has seen MarieSubject extraction, que, fails
Qui crois-tu qui a vu Marie?who do you think qui has seen MarieSubject extraction, qui, fine

Second, the adverb effect, reported by Peter Culicover in 1993: put an adverbial between the complementiser and the gap and the English sentence recovers.

(23) Who do you think that under no circumstances would leave the party early?

That sentence is not elegant, but almost every speaker prefers it to (22), which has fewer words in the same configuration. A pure adjacency ban would predict it, a pure government account struggles with it, and thirty years on it is still one of the sharper tests a proposal has to pass.

Is it the grammar or is it memory?

A long-running alternative says the grammar contains no island constraints at all, and that island sentences are simply hard. Philip Hofmeister and Ivan Sag argued this in a 2010 paper in Language, showing that manipulations which lighten the processing burden, such as making the filler more informative, measurably improve island sentences.

The strongest reply is experimental. Jon Sprouse, Matthew Wagers and Colin Phillips reported in Language in 2012 that island effects are superadditive: the penalty for the two factors combined exceeds the sum of the penalty for a long dependency alone and the penalty for an embedded question alone. They then measured each participant's working memory capacity on two independent tests. If islands were memory limitations, the readers with more memory should have tolerated them better. Across hundreds of participants the correlation was absent. Neither camp has won outright, and Lesson 15 will put a harder version of this same argument under the microscope.

Common misconceptions

  • Island violations are bad because they are hard to understand. You can usually recover the intended meaning at once, which is what makes the judgment so sharp. Comprehensibility and grammaticality came apart in Lesson 1 and they come apart here.
  • Distance causes islands. Sentence (1) has a gap three clauses down and is fine; (4) has a gap two clauses down and is dead. What matters is the kind of node crossed.
  • All islands are equally strong. Complex noun phrases behave categorically, while wh-islands are gradient, sensitive to argument status, and variable across speakers and languages.
  • The that-trace effect is a rule of style. It has a morphological reflex in French and an escape hatch in the adverb effect, neither of which a style rule would have.

Recap

  • Ross's 1967 dissertation catalogued the structures a dependency cannot escape: complex noun phrases, coordinate structures, sentential subjects, adjuncts, embedded questions and left branches.
  • Across-the-board extraction is legal, which shows the coordinate structure constraint tracks structure rather than processing load.
  • Subjacency derives several of Ross's constraints from one condition, given successive cyclic movement and a designated pair of bounding nodes.
  • Rizzi's Italian data forced the bounding nodes to be parameterised: NP and TP for English, NP and CP for Italian.
  • The that-trace effect blocks subject extraction across an overt complementiser, appears as the que to qui alternation in French, and is rescued by an intervening adverbial.

Sources

  1. Wikipedia. (n.d.). Syntactic island. Wikimedia Foundation. en.wikipedia.org
  2. Wikipedia. (n.d.). Subjacency. Wikimedia Foundation. en.wikipedia.org
  3. Ross, J. R. (1967). Constraints on variables in syntax [Doctoral dissertation, Massachusetts Institute of Technology]. Published in 1986 as Infinite syntax! Ablex.
  4. Rizzi, L. (1982). Issues in Italian syntax, chapter 2. Foris Publications.
  5. Sprouse, J., Wagers, M., and Phillips, C. (2012). A test of the relation between working memory capacity and syntactic island effects. Language, 88(1), 82-123.
Key terms
Island
A structural region from which a movement dependency cannot escape, even when the intended meaning is perfectly recoverable.
Complex noun phrase constraint
Ross's ban on extracting anything out of a clause that sits under a noun, whether that clause is a complement or a relative.
Across-the-board extraction
Movement that leaves a gap in every conjunct of a coordination, which is permitted where extraction from one conjunct alone is not.
Subjacency
The condition that one movement step may not cross more than one bounding node, which derives several of Ross's constraints at once.
Bounding node
A category that counts for subjacency; English is analysed with NP and TP, Italian with NP and CP.
That-trace effect
The ban on extracting a subject across an overt complementiser, visible in the contrast between who do you think left and the version with that.
Empty category principle
The requirement that a trace be properly governed, introduced largely to explain why subject and object extraction differ.
Superadditivity
The finding that the penalty for an island violation exceeds the sum of the penalties for its two component factors measured separately.

Ellipsis: Reading the Structure Nobody Pronounced

  • Run a four-step procedure for recovering an ellipsis site, and use it to derive the strict and sloppy readings of one sentence.
  • Give three arguments that an ellipsis site contains syntactic structure: missing antecedents, case matching under sluicing, and the preposition stranding generalisation.
  • Distinguish verb phrase ellipsis, sluicing, gapping and pseudogapping, and say what each one deletes.

A pronoun with nothing to point at

In 1971 John Grinder and Paul Postal published a paper in Linguistic Inquiry built around a sentence of this shape:

(1) Kim has never ridden a camel, but Lee has, and it bit him.

Read the last clause again. What does it refer to? Not the camel Kim never rode, because there is no such camel. It refers to the camel Lee rode, and the phrase a camel in that reading is nowhere on the page. A pronoun needs an antecedent, and here the antecedent is inside the gap after has. Compare a version where the same content is expressed by a pro-form rather than by a gap:

(2) *Kim has never ridden a camel, but Lee has done it, and it bit him.

Sentence (2) is much worse on the intended reading, and the difference is instructive. Something is present in (1) that is absent in (2), and it is present without being pronounced. This lesson is a procedure for reading it.

The point: Ellipsis is the one place in syntax where you can interrogate structure that has no phonological form at all, which is why so much of the argument for abstract structure runs through it.

The procedure, run once end to end

Take a plain case of verb phrase ellipsis:

(3) Kim will read the report, and Lee will too.

Step one: locate the licensor. English verb phrase ellipsis is licensed by a finite auxiliary or by infinitival to. Here it is will. If you delete the auxiliary as well the sentence collapses: *Kim will read the report, and Lee too is not verb phrase ellipsis, it is a different construction with different properties.

Step two: identify the antecedent constituent. Not the antecedent string. The tests of Lesson 2 tell you which node is a constituent, and here the antecedent is the VP read the report, not the string will read.

Step three: copy it into the gap and check the fit. The gap must be the same category as the antecedent and must sit in a position where that category is legal.

[TP [NP Lee] [T-bar [T will] [VP read the report]]]

Step four: check for a mismatch that the theory must license. Verb phrase ellipsis tolerates a good deal of inflectional mismatch: Kim slept badly, and Lee will too pairs a past tense antecedent with a bare infinitive in the gap. Whatever is copied, it is not the phonological string.

Four steps, one sentence, one reading. Now change one input.

One input changed, two readings out

Replace the object with a possessive phrase:

(4) Kim will read her report, and Lee will too.

Run step three and the copy is read her report. But her is a variable, and there are two ways to resolve it. On the strict reading, Lee reads Kim's report. On the sloppy reading, Lee reads Lee's own report. Both readings are available, and both are available to every speaker asked. This ambiguity does not arise from the words: (4) has one word string and two structures underlying the gap, which is exactly the signature you learned to look for in Lesson 2.

ReadingWhat is in the gapParaphrase
Strictread her report, where her is fixed to KimLee will read Kim's report
Sloppyread x's report, with x bound by the subjectLee will read Lee's report

Now change the input again, this time to a relative clause. Antecedent-contained deletion is the case where the procedure appears to break:

(5) Kim visited every town that Lee did.

The antecedent VP is visited every town that Lee did. Copy it into the gap and the gap now contains a copy of itself, which contains a gap, which needs a copy. The regress never terminates. Speakers nonetheless understand (5) instantly and there is nothing marginal about it. The standard resolution is that the quantified object raises out of the VP before the copy is taken, so that the antecedent is the smaller VP visited x and the regress never starts. What matters here is not the mechanism but the shape of the argument: a procedure that produces an infinite regress on real data is telling you that the structure at the point of copying is not the structure you see.

In short: Ellipsis resolution is not string matching. It is constituent copying under an identity condition, and the strict and sloppy ambiguity plus antecedent-contained deletion are the two facts that force that conclusion.

A case ending that gives the game away

Sluicing deletes everything in an embedded question except the wh-phrase:

(6) Someone left early, but I do not know who.

English will not settle whether the deleted material is really there, because English wh-words carry almost no morphology. German will. The verb schmeicheln, to flatter, assigns dative case to its object, and the dative of wer is wem, not the accusative wen.

GermanGlossStatus
Er will jemandem schmeicheln, aber sie wissen nicht, wem.he wants to flatter someone dative, but they do not know whom dativeGood
*Er will jemandem schmeicheln, aber sie wissen nicht, wen.he wants to flatter someone dative, but they do not know whom accusativeBad

The remnant is inflected for a case that only schmeicheln can assign, and schmeicheln is not pronounced anywhere in the second clause. Either the case is assigned by a verb that is present without being pronounced, or case in German can be assigned by a verb in a different clause, which nothing else in the language suggests. Jason Merchant collected this and much more in The Syntax of Silence in 2001, and added a second argument that works across languages rather than within one. English allows a preposition to be stranded, and English allows a sluicing remnant with no preposition:

(7) Kim was arguing with someone, but I do not know who.
(8) Who was Kim arguing with?

Greek allows neither. Its counterpart of (8) is ungrammatical, and its counterpart of (7) requires the preposition to appear on the remnant. Merchant checked this pairing across a substantial sample and found it held: a language permits a bare sluicing remnant only if it permits preposition stranding in ordinary questions. That correlation makes no sense unless the sluice contains an ordinary question that has undergone ordinary wh-movement, with the whole clause then deleted.

The family, and what each member deletes

ConstructionExampleWhat goes missing
Verb phrase ellipsisKim read the report and Lee did tooThe VP, licensed by an auxiliary
SluicingSomeone called, but I forget whoEverything in the embedded question but the wh-phrase
GappingKim ordered pasta, and Lee risottoThe verb, in a coordinate structure, leaving two remnants
PseudogappingKim read more books than Lee did articlesThe verb only, leaving the object behind
StrippingKim read the report, and Lee tooEverything but one remnant and often a polarity word
Null complement anaphoraKim asked Lee to leave, but Lee refusedThe complement of a small class of verbs

The last row is not like the others. In 1976 Jorge Hankamer and Ivan Sag drew a line between what they called surface anaphora and deep anaphora. Surface anaphora, which includes verb phrase ellipsis and sluicing, requires a linguistic antecedent: you cannot walk into a room, watch someone lifting a piano, and say I do not think you can. Deep anaphora can be controlled by the situation: watching the same scene you can say I do not think you can do it. The pro-form is available to a hearer who saw the event; the gap is available only to a hearer who heard the sentence.

The strongest counterargument, and what it explains

A rival family of theories holds that the gap is an atom, a silent pro-form whose meaning is recovered from the semantics of the antecedent with no hidden syntax at all. It is not a fringe position, and it has a genuine result behind it. Consider extraction from an island:

(9) They want to hire someone who speaks a Balkan language, but I do not remember which.
(10) *They want to hire someone who speaks a Balkan language, but I do not remember which Balkan language they want to hire someone who speaks.

Sentence (10) violates the complex noun phrase constraint of the last lesson. Sentence (9) is fine. If (9) contains the structure of (10) and merely fails to pronounce it, why is the island violation not fatal in both? Ross noticed this in 1969 and it has been called island repair ever since. The pro-form theory has an easy answer: (9) contains no island because it contains no structure.

Set that against what the pro-form theory has to explain away: the missing antecedent in (1), the German dative in the sluicing table, the preposition stranding correlation, and the fact that a gap needs a linguistic antecedent while do it does not. Most current work keeps the structure and treats island repair as a fact about how deletion interacts with the offending nodes rather than as evidence that the nodes are absent. That is a bet rather than a proof, and you should hold it as one.

Common misconceptions

  • Ellipsis deletes repeated words. The antecedent and the gap routinely differ in tense, agreement and voice, so the identity condition cannot be phonological.
  • Any string can be elided if the context makes it recoverable. English verb phrase ellipsis needs a licensing auxiliary or infinitival to, and without one the sentence fails no matter how obvious the meaning.
  • The gap and a pro-form such as do it are two ways of saying the same thing. Hankamer and Sag showed they differ on whether a linguistic antecedent is required, and Grinder and Postal showed they differ on whether a pronoun can find an antecedent inside them.
  • Strict and sloppy readings come from an ambiguous pronoun. The pronoun is the same in both; what differs is whether it is a fixed reference or a variable bound by the subject.

The short version

  • Reading an ellipsis site is a four-step procedure: find the licensor, identify the antecedent constituent, copy it, and check what mismatches the copy tolerates.
  • The strict and sloppy ambiguity shows the gap holds structure with a variable in it; antecedent-contained deletion shows the copy cannot be taken from the surface structure.
  • German case on a sluicing remnant is assigned by a verb that is never pronounced, which is direct morphological evidence for structure in the gap.
  • Merchant's preposition stranding generalisation ties the shape of a sluice in a language to whether that language permits stranding in ordinary questions.
  • Verb phrase ellipsis, sluicing, gapping, pseudogapping and stripping delete different constituents under different licensing conditions, and null complement anaphora is a pro-form rather than a gap.
  • Island repair under sluicing is the best argument on the other side, and it is unresolved.

Sources

  1. Wikipedia. (n.d.). Ellipsis (linguistics). Wikimedia Foundation. en.wikipedia.org
  2. Wikipedia. (n.d.). Sluicing. Wikimedia Foundation. en.wikipedia.org
  3. Merchant, J. (2001). The syntax of silence: Sluicing, islands, and the theory of ellipsis. Oxford University Press.
  4. Hankamer, J., and Sag, I. (1976). Deep and surface anaphora. Linguistic Inquiry, 7(3), 391-428.
  5. Grinder, J., and Postal, P. (1971). Missing antecedents. Linguistic Inquiry, 2(3), 269-312.
Key terms
Verb phrase ellipsis
Deletion of a VP under identity with an antecedent, licensed in English by a finite auxiliary or by infinitival to.
Sluicing
Deletion of everything in an embedded question except the wh-phrase, as in someone left but I do not know who.
Strict and sloppy readings
The two interpretations of a pronoun inside an ellipsis site, one fixed to the original referent and one bound by the new subject.
Antecedent-contained deletion
An ellipsis site contained within its own antecedent, which produces an infinite regress unless the object leaves the VP first.
Missing antecedent
A pronoun whose antecedent occurs only inside an ellipsis site, which a silent pro-form cannot supply.
Preposition stranding generalisation
Merchant's finding that a language permits a bare sluicing remnant only if it permits preposition stranding in ordinary questions.
Surface anaphora
Anaphora requiring a linguistic antecedent, such as ellipsis, as against deep anaphora such as do it, which the situation can control.
Island repair
The observation that sluicing tolerates configurations whose unelided counterparts violate island constraints.

Module 5: Reference, Variation and Learning

How a grammar fixes what a pronoun may point at, how much of the machinery varies from language to language, and what a child could possibly have used to work any of it out.

Binding: Three Principles, and the Sentences That Broke Them

  • State Principles A, B and C in terms of binding domains and c-command, and apply them to a sentence you have not seen before.
  • Show that the relevant relation is c-command rather than linear precedence, using minimal pairs with possessors and with prepositional phrases.
  • Name a paradigm each principle fails on, and describe one reformulation the failures produced.

A rule that is nearly right

Start with the rule most people would write on a first attempt. A reflexive needs an antecedent in the same sentence; an ordinary pronoun needs one somewhere else. Here is the data it was written for:

(1) John likes himself. (himself is John)
(2) *John likes him. (on the reading where him is John)
(3) John said that Mary likes him. (him can be John)
(4) *John said that Mary likes himself.

The rule handles all four. Now break it. Every sentence below has an antecedent in the same sentence, and half of them are dead:

(5) *Himself likes John.
(6) John's mother likes him. (him can be John)
(7) *John's mother likes himself. (himself cannot be John)
(8) *He said that John left. (on the reading where he is John)

This lesson is the repair. We will fix the rule three times, and each repair is a piece of binding theory as it was stated in Chomsky's Lectures on Government and Binding in 1981. Then we will break the repaired version too, because the sentences that broke it are the reason the theory looks the way it does now.

First repair: the relation is c-command, not order

Sentence (5) fails and (1) succeeds, and the only difference is which noun phrase comes first. So the naive fix is a rule about order: the antecedent must precede the reflexive. That fix survives about ten seconds. Sentence (7) has John preceding himself and is still dead. What is wrong with (7) is that John is buried inside the subject noun phrase John's mother, so it is not high enough in the structure to bind anything outside that phrase.

[TP [NP [NP John's] [N mother]] [VP [V likes] [NP himself]]]

The relation you need is c-command: a node c-commands its sisters and everything inside them. In the bracketing above, the whole subject John's mother c-commands the object; John's alone does not, because its sister is mother and nothing else. In (1) the subject John is the whole subject, so it c-commands the object and binds it.

Sentence (6) is the confirming case. John does not c-command him, so him is not bound by it, so nothing rules the coreference out. The pronoun and the name simply happen to pick out the same person, which the grammar permits whenever binding is not involved.

Key idea: Binding is a structural relation, not a fact about word order or about who is mentioned first. Every argument in this lesson depends on that distinction, and the possessor cases are how you feel it.

Second repair: the domain has an edge

Sentence (4) has John c-commanding himself and is still ungrammatical, so c-command cannot be sufficient. The reflexive is too far away: it is in a different clause. Sentence (3) shows the mirror image, a pronoun that is fine at exactly the distance the reflexive is not. That is the complementarity that gives the theory its shape. Define the binding domain of a noun phrase as roughly the smallest clause containing it, its governor and a subject, and the three principles fall out:

PrincipleApplies toRequirementBroken by
AAnaphors: himself, themselves, each otherMust be bound inside its binding domain(4), (5), (7)
BPronouns: him, her, themMust be free inside its binding domain(2)
CR-expressions: John, the doctorMust be free everywhere(8)

Test the package on a sentence built to be awkward:

(9) John told Bill about himself.

Both John and Bill c-command himself and both are in its binding domain, so Principle A licenses either, and the sentence is genuinely ambiguous. Now:

(10) John believes himself to be competent.
(11) *John believes himself is competent.

The contrast follows from Lesson 9. In (10) the embedded clause is non-finite and has no subject of its own left in place, so the binding domain extends up to include John. In (11) the embedded clause is finite, which closes the domain below John, and the reflexive is stranded with nothing to bind it.

Third repair: Principle C is not about pronouns at all

Principle C is the strangest of the three because it has no positive requirement. It does not say what a name must be bound by; it says a name must never be bound. That single statement covers (8), and also:

(12) *He thinks John is a genius. (he is John)
(13) His mother thinks John is a genius. (his can be John)
(14) John thinks John is a genius. (marginal, and only with two different Johns)

In (13) the pronoun is inside the subject, so it does not c-command the name, so no violation arises. In (14) the higher John does c-command the lower one, which is why the sentence forces you to imagine two people with the same name. Principle C is thus the strongest evidence that binding cares about structure: the sentence is not incomprehensible, and the coreference is not implausible, and it is still unavailable.

So what?: Principle C predicts the unavailability of a reading, not the ungrammaticality of a string. That is a different and sharper kind of prediction, and it is the one that makes binding theory testable on languages with no reflexive morphology to speak of.

Where the repaired theory breaks

Now the sentences that did the damage. Each one is grammatical and each one violates a principle as stated.

Picture noun phrases. Principle A says an anaphor must be bound in its domain. Consider:

(15) John thought that a picture of himself would be on the cover.
(16) Pictures of himself in the Times had upset John for years.
(17) John saw a picture of him. (him can be John)

In (15) the antecedent is in a higher clause, past the edge of the domain. In (16) John does not c-command himself at all, since it is inside the object of a preposition inside the verb phrase and the reflexive is inside the subject. In (17) a reflexive and a pronoun are both available in one position, which the complementarity of Principles A and B says cannot happen. Carl Pollard and Ivan Sag argued in 1992 that anaphors in these positions are exempt from Principle A entirely, and are licensed by discourse conditions instead: the antecedent must be the person whose point of view the sentence reports.

Logophors. That discourse condition is visible where the antecedent is not even in the sentence:

(18) John was furious. The photograph of himself on the front page had been taken without permission.

No syntactic account can reach across a full stop. The phenomenon has a name, logophoricity, and languages such as Ewe mark it with a dedicated pronoun series used only for the person whose speech or thought is being reported.

Long-distance anaphors. Icelandic sig, Japanese zibun and Mandarin ziji are all anaphors by any morphological test and none of them respects an English-sized binding domain.

LanguageFormHow far it reachesExtra condition
IcelandicsigOut of the local clause, unboundedly in principleLong-distance readings require subjunctive mood in the embedded clause
JapanesezibunOut of the local clause, across several boundariesMust be bound by a subject, never by an object
MandarinzijiOut of the local clauseA first or second person subject in between blocks a higher antecedent

The subject orientation is the important part. English himself in (9) can take an object as antecedent; zibun cannot, ever. So the domain is not the only parameter: what counts as a possible binder varies too.

Principle B leaks. Locative and spatial phrases let a pronoun be bound where the theory says it may not:

(19) John pulled the blanket over him.
(20) John saw a snake near him.

Both allow him to be John, and swapping in himself is also fine, so this is another complementarity failure.

Principle C is not universal. Howard Lasnik reported in work collected in 1989 that Thai permits a name to be bound by a c-commanding name in exactly the configuration English rules out, so the counterpart of (12) with a repeated name is available. If Principle C were an inviolable universal, that language should not exist.

What the failures produced

Two reformulations matter. Tanya Reinhart argued in 1983 that the syntax should only handle genuinely bound variable readings, with accidental coreference left to a pragmatic principle that compares what the speaker said with what a reflexive would have said. That move removes the burden of explaining (17) and (19) from the syntax. Then in 1993 Reinhart and Eric Reuland restated the whole theory over predicates rather than over domains: a reflexive-marked predicate must be reflexive, and a reflexive predicate must be reflexive-marked. The picture noun phrase cases fall outside that statement automatically, because a picture of himself is not an argument of the same predicate as its antecedent. Neither reformulation is uncontested, and both are recognisably descendants of the three principles you fixed at the top of this lesson.

Common misconceptions

  • An anaphor needs an antecedent earlier in the sentence. Order is not the relation. In John's mother likes himself the antecedent is earlier and the sentence is dead; in pictures of himself upset John it is later and the sentence is fine.
  • Principle B says a pronoun cannot refer to anything mentioned nearby. It bans binding inside the domain only. John said Mary likes him is fine, and so is John's mother likes him.
  • Principle C is about pronouns preceding names. It is about c-command: his mother thinks John is a genius has the pronoun first and is perfectly good.
  • Reflexives and pronouns are always in complementary distribution. Picture noun phrases and locative phrases both allow either, which is one of the facts the classical theory could not state.

Where this leaves us

  • Binding is defined over c-command and a local domain, not over linear order or over plausibility of reference.
  • Principle A requires an anaphor to be bound in its domain, Principle B requires a pronoun to be free in its domain, and Principle C requires a name to be free everywhere.
  • The finite and non-finite contrast in believes himself to be competent against the version with a finite clause is the cleanest demonstration that the domain has an edge.
  • Picture noun phrases break Principle A three ways: antecedent too far, antecedent not c-commanding, and reflexive and pronoun both available.
  • Icelandic sig, Japanese zibun and Mandarin ziji reach outside the local clause and are subject-oriented, so both the domain and the class of possible binders vary.
  • Reinhart's bound variable account and Reinhart and Reuland's predicate-based reflexivity are the two reformulations those failures produced.

Sources

  1. Wikipedia. (n.d.). Binding (linguistics). Wikimedia Foundation. en.wikipedia.org
  2. Wikipedia. (n.d.). C-command. Wikimedia Foundation. en.wikipedia.org
  3. Chomsky, N. (1981). Lectures on government and binding, chapter 3. Foris Publications.
  4. Pollard, C., and Sag, I. (1992). Anaphors in English and the scope of binding theory. Linguistic Inquiry, 23(2), 261-303.
  5. Reinhart, T., and Reuland, E. (1993). Reflexivity. Linguistic Inquiry, 24(4), 657-720.
Key terms
C-command
The structural relation holding between a node and its sisters together with everything contained in them; the relation binding is defined over.
Binding domain
Roughly the smallest clause containing a noun phrase, its governor and a subject, inside which Principles A and B are evaluated.
Principle A
An anaphor must be bound within its binding domain.
Principle B
A pronoun must be free within its binding domain, which leaves it free to corefer outside that domain.
Principle C
An R-expression such as a name must not be bound at all, which predicts the loss of a reading rather than the loss of a string.
Exempt anaphor
A reflexive in a position where Principle A does not apply, licensed instead by the point of view the sentence reports.
Logophoricity
Marking of reference to the person whose speech or thought is being reported, which some languages express with a dedicated pronoun series.
Subject orientation
The restriction, found with zibun and sig, that an anaphor may only be bound by a subject and never by an object.

Parameters: One Switch, Several Consequences

  • Read a head direction table and predict the order of verb and object, adposition and noun, and complementiser and clause in a language you have not studied.
  • State the null subject cluster Rizzi proposed for Italian and name a language that has null subjects without the rest of it.
  • Derive German main clause word order from verb movement to C plus one phrase to the specifier, and use the subordinate clause facts as the argument.

Four words, and everything in the wrong place

Here is a Japanese sentence with a word-for-word gloss underneath it.

(1) Taroo ga hon o yonda.
Taroo nominative book accusative read-past.
Taro read a book.

The verb is last. Now watch the pattern repeat everywhere else in the language. Where English puts a preposition before its object, Japanese puts a postposition after it: Tookyoo kara is Tokyo-from. Where English puts a complementiser before its clause, Japanese puts one after: Taroo ga kita to omou is Taro nominative came that think, meaning I think that Taro came. Where English puts a relative clause after the noun, Japanese puts it before: Taroo ga yonda hon is Taro nominative read book, meaning the book that Taro read.

Four differences, or one? The X-bar theory of Lesson 5 gives you the vocabulary to say it is one. Every phrase has a head, and the head either precedes its complement or follows it. English chose one setting, Japanese chose the other, and the four facts are the same choice seen from four angles.

PhraseHead-initial: EnglishHead-final: Japanese
VPread a bookhon o yonda (book accusative read)
PPfrom TokyoTookyoo kara (Tokyo from)
CPthat Taro cameTaroo ga kita to (Taro came that)
Relative clause and nounthe book that Taro readTaroo ga yonda hon (Taro read book)

[VP [V read] [NP a book]] against [VP [NP hon o] [V yonda]]

Why this matters: A parameter is worth having only if it buys more than it costs. One switch that predicts four independent word order facts is a bargain; a switch invented for each fact separately explains nothing.

Does the bargain hold across the world's languages?

This is checkable, and the World Atlas of Language Structures has checked it. Chapter 81 classifies 1376 languages by the order of subject, object and verb: 564 are verb-final, 488 are verb-medial, 95 are verb-initial, and 189 have no dominant order. Chapter 95 is the one that tests the parameter, because it cross-tabulates the order of object and verb against the order of adposition and noun phrase.

CombinationLanguagesPredicted by a head direction parameter?
Object before verb, postpositions472Yes, consistently head-final
Verb before object, prepositions456Yes, consistently head-initial
Verb before object, postpositions42No
Object before verb, prepositions14No

Of the 984 languages that fall into one of these four cells, 928 take a harmonic combination. That is 94 percent, against the 50 percent you would expect if the two orders were independent. The parameter is doing real work. It is also not doing all the work, because 56 languages sit in the cells it forbids, and one of them is German, which puts its non-finite verb at the end of the clause and its prepositions at the front of theirs.

The null subject cluster

Italian permits a sentence with no subject at all.

(2) Parla italiano. (speaks Italian, meaning he or she speaks Italian)
(3) *Speaks Italian. (English, on the same reading)
(4) Piove. (rains, meaning it is raining)
(5) *Rains. (English)

Sentence (5) is the interesting half. English needs it in (5) even though the word refers to nothing whatever, which is the EPP of Lesson 6 demanding that the subject position be filled. Italian does not. Rizzi's 1982 proposal was that this single difference drags three more along with it, and the strength of the claim is that the other three are not obviously about subjects being absent.

PropertyItalianEnglish
Referential subject may be nullParla italianoNot available
Expletive subject may be nullPioveNot available
Subject may appear after the verb freelyHa telefonato GianniVery restricted
Subject may be extracted across a complementiserChi credi che abbia telefonato?Blocked, the that-trace effect of Lesson 11

The fourth row is the payoff. The that-trace effect was a stubborn, apparently arbitrary fact about English; the null subject parameter explains its absence in Italian by letting the subject be extracted from the postverbal position instead of from the position next to the complementiser. One switch, four consequences, one of which nobody would have connected to the others.

WALS chapter 101 puts English in a small minority: of 711 languages surveyed, only 82 require an independent subject pronoun in subject position the way English does, while 437 express pronominal subjects as affixes on the verb. Requiring an overt subject is the unusual setting, not the default.

Now the complication. The traditional story is that rich agreement morphology licenses the null subject, because the verb ending tells you who the subject was. Mandarin has no agreement morphology at all and drops subjects freely, as do Japanese and Korean. So either agreement is not the licensor, or there are two different ways to have a null subject, which is what most current work assumes: a pro licensed by agreement, and a radical or discourse-based dropping that leans on context instead. The cluster in the table also does not travel intact: not every null subject language shows free inversion.

German, and a verb that has to be second

German is the case that shows a parameter setting can be hidden under movement. Count the constituents before the finite verb.

(6) Maria hat gestern das Buch gelesen. (Maria has yesterday the book read)
(7) Gestern hat Maria das Buch gelesen. (Yesterday has Maria the book read)
(8) Das Buch hat Maria gestern gelesen. (The book has Maria yesterday read)
(9) *Gestern Maria hat das Buch gelesen.
(10) *Hat Maria gestern das Buch gelesen. (fine as a question, dead as a statement)

Exactly one constituent may precede the finite verb, and it can be almost any constituent. This is verb-second order, and the analysis uses the two positions Lesson 6 built: the finite verb moves to C, and one phrase moves to the specifier of CP above it.

[CP [NP Das Buch] [C hat] [TP Maria gestern gelesen]]

The argument for that analysis is not the main clause data at all. It is the subordinate clause:

(11) ..., dass Maria gestern das Buch gelesen hat. (that Maria yesterday the book read has)

When the complementiser dass is present, the finite verb goes to the end of the clause and verb-second disappears completely. A complementiser and a finite verb are in complementary distribution in German, which is exactly what you expect if they compete for one position. The clause-final order in (11) also shows the underlying head-final VP, so the head direction and the verb-second setting are two different switches operating on the same sentence.

Bottom line: Surface word order is the product of the base setting and whatever movement the language requires. German looks inconsistent until you separate the two, and then it is head-final in the VP with an obligatory verb-second rule on top.

How much of this survived

Parameters were proposed as a solution to a learning problem, which Lesson 15 takes up. If a child has only to set a few dozen switches rather than induce a grammar from nothing, acquisition looks tractable. The programme ran into three difficulties.

PositionWhat variesBest evidence for it
Macroparameters, as in Mark Baker's workA small number of high level switches with wide consequencesThe head direction correlations and the polysynthesis cluster
Microparameters, as in Richard Kayne's dialect studiesMany small properties of individual functional headsNorthern Italian dialects differing on dozens of tiny points while agreeing on the large ones
No parameters, as in Frederick Newmeyer's 2005 bookNothing; the clusters are typological tendencies with functional causesClusters that leak, and languages that take one property without the others

The clusters do leak, and honest textbooks now present the null subject cluster as a first approximation rather than a law. What has survived best is the Borer and Chomsky conjecture: that syntactic variation lives in the lexicon, specifically in the features of functional heads such as C and T, rather than in the rules that combine them. On that view the head direction parameter is a fact about how the phonology linearises a structure, and verb-second is a fact about the features of the German C. It is a considerable retreat from the switchboard picture, and it is where the field currently sits.

Common misconceptions

  • A parameter is a rule of the language. It is a choice point in the theory, and its interest lies entirely in what else the choice predicts. A parameter with one consequence is a restatement of the fact it was invented for.
  • Head-final languages are backwards versions of English. Japanese keeps the same hierarchical structure and the same constituent relations; only the linear order of head and complement differs, which is why the constituency tests of Lesson 2 give the same answers.
  • German is an SVO language with exceptions. The subordinate clause shows the base order is verb-final, and the main clause order is what verb movement to C produces.
  • Rich agreement is what makes null subjects possible. Mandarin, Japanese and Korean drop subjects with no agreement morphology at all.

Putting it together

  • A head direction setting predicts the order of verb and object, adposition and noun, complementiser and clause, and relative clause and noun together, and WALS chapter 95 shows 928 of 984 languages taking a harmonic combination.
  • Italian permits null referential and expletive subjects, free subject inversion, and subject extraction across a complementiser, and Rizzi tied all four to one setting.
  • Mandarin drops subjects without agreement morphology, so the licensing story cannot be agreement alone.
  • German verb-second follows from movement of the finite verb to C plus one phrase to the specifier of CP, and the subordinate clause with dass is the argument, because complementiser and finite verb never co-occur in that position.
  • Only 82 of the 711 languages in WALS chapter 101 require an independent subject pronoun the way English does.
  • Macroparameters, microparameters and the rejection of parameters are three live positions, and current work locates variation in the features of functional heads.

Sources

  1. Dryer, M. S. (2013). Order of subject, object and verb. In The World Atlas of Language Structures Online. Max Planck Institute for Evolutionary Anthropology. wals.info
  2. Dryer, M. S. (2013). Relationship between the order of object and verb and the order of adposition and noun phrase. In The World Atlas of Language Structures Online. Max Planck Institute for Evolutionary Anthropology. wals.info
  3. Dryer, M. S. (2013). Expression of pronominal subjects. In The World Atlas of Language Structures Online. Max Planck Institute for Evolutionary Anthropology. wals.info
  4. Baker, M. (2001). The atoms of language: The mind's hidden rules of grammar. Basic Books.
  5. Newmeyer, F. J. (2005). Possible and probable languages: A generative perspective on linguistic typology. Oxford University Press.
Key terms
Head direction parameter
The choice of whether heads precede or follow their complements, which fixes the order of verb and object, adposition and noun, and complementiser and clause together.
Harmonic word order
A combination of orders consistent with one head direction setting, such as object before verb together with postpositions.
Null subject parameter
The setting that permits a finite clause to lack an overt subject, associated by Rizzi with three further properties of Italian.
Free subject inversion
The option of placing a subject after the verb without special intonation or focus marking, as in Ha telefonato Gianni.
Verb-second
The requirement that the finite verb occupy the second position of a main clause, with exactly one constituent before it.
Radical pro-drop
Subject omission in a language with no agreement morphology, as in Mandarin, Japanese and Korean.
Microparameter
A small variation in the features of one functional head, of the kind that distinguishes closely related dialects.
Borer and Chomsky conjecture
The proposal that all syntactic variation resides in the lexical features of functional heads rather than in the combinatory rules.

The Poverty of the Stimulus, and What the Machines Changed

  • State the structure dependence argument in its auxiliary inversion form, and say precisely which premise each critic attacks.
  • Summarise the Pullum and Scholz corpus critique and the statistical learning replies, with the results each rests on.
  • Say what large language models do and do not settle about the learnability of syntax, and what evidence would move the dispute.

Ask Jabba a question

In a laboratory experiment reported in 1987, Stephen Crain and Mineharu Nakayama sat three to five year old children in front of a Star Wars puppet and asked them to put questions to it. One instruction was this: ask Jabba if the boy who is watching Mickey Mouse is happy. The target is a question with two instances of is in the input, and a child has to choose which one to move.

(1) Is the boy who is watching Mickey Mouse happy?
(2) *Is the boy who watching Mickey Mouse is happy?

Sentence (2) is what you get from the simplest rule that fits every simple question a child has ever heard: find the first auxiliary in the string and move it to the front. It is a rule about linear order, it is easy to state, and it is wrong. Sentence (1) comes from a rule about structure: move the auxiliary of the main clause, wherever in the string it happens to sit.

Crain and Nakayama's children made plenty of errors. They repeated auxiliaries, they restarted, they produced tangles. What they essentially never produced was (2). That result is one leg of an argument that has run for sixty years, and this lesson lays out the argument, the two most serious attacks on it, and the state of play now that machines can be trained on text and asked the same question.

The core of it: The dispute is not about whether children learn language from experience. It is about whether the experience contains enough information to select the structural rule over the linear one without a prior bias toward structure.

The argument, stated so it can be attacked

The poverty of the stimulus argument has three premises and a conclusion, and each premise is a separate target.

PremiseWhat it claimsHow you would test it
OneAdults know a structure dependent rule, not a linear oneGrammaticality judgments and elicited production
TwoChildren never go through a linear-rule stageCorpus study of child speech and elicitation experiments
ThreeThe crucial sentences are absent, or too rare, in what children hearCorpus study of child-directed speech
ConclusionThe bias toward structure is not learned from the dataFollows only if all three premises hold

The crucial sentence is one that distinguishes the two rules: a question whose subject contains an auxiliary of its own, as in (1). A child who hears only Is the boy happy? and Is Mummy coming? has no basis for choosing, because the first auxiliary and the main clause auxiliary are the same word.

The corpus attack

In 2002 Geoffrey Pullum and Barbara Scholz published an examination of the argument in The Linguistic Review that took premise three seriously enough to go and count. Their finding was that the crucial sentence type is not absent. Questions of the relevant shape occur in the Wall Street Journal corpus, and forms such as Is the boy who is holding the plate crying? and polar questions with relative clauses in the subject do appear in speech addressed to children. They also pressed a second point that is easy to miss: the argument requires the data to be absent, and showing that it is merely rare is not the same claim, because rare evidence can still be decisive if the learner is looking for it.

Julie Anne Legate and Charles Yang replied in the same journal that year with a quantitative standard. They asked how frequent a cue has to be before a child can be expected to use it, calibrating against phenomena where the acquisition timetable is known, and concluded that the auxiliary inversion cue falls well below the frequency of cues children demonstrably do use. Neither side disputes the corpus counts. They dispute what counts as enough.

What matters here: Pullum and Scholz did not show that children learn the structural rule from data. They showed that one premise of the standard argument was asserted rather than measured, which is a narrower and more damaging result than it is often reported to be.

The learning attack

The second attack accepts that the crucial sentences are rare and argues that a learner does not need them, because statistical structure elsewhere in the input does the work indirectly.

StudyModel and dataResult
Reali and Christiansen, 2005Word pair and word triple statistics from child-directed speechThe correct question form scored higher than the incorrect one
Kam and colleagues, 2008Reanalysis of the same models on a wider set of sentencesThe success rested on the frequency of the string who is, and vanished with other relative pronouns
Perfors, Tenenbaum and Regier, 2011Bayesian comparison of a linear grammar against a hierarchical one on child-directed speechThe hierarchical grammar wins on the corpus as a whole, without needing the crucial sentences

The Perfors result is the strongest thing the anti-nativist side has, and it is worth being precise about what it shows. The model was given both grammars and asked which better explains the corpus, trading off fit against simplicity. Hierarchy won. What the model was not asked to do was invent the hypothesis space: both grammars were supplied. A defender of the original argument can accept the result and say that the interesting question was always how the learner comes to entertain a hierarchical grammar at all.

What the language models did to the argument

From 2019 onward a new kind of evidence arrived. Large language models are trained on text with no grammar built in, and they can be probed with minimal pairs. On benchmarks of English grammatical contrasts, current models handle subject-verb agreement across relative clauses, island effects, binding contrasts and auxiliary inversion at high accuracy. Steven Piantadosi argued in 2023 that this refutes the whole approach: a system with no innate grammar learned the facts, so the facts are learnable.

Three things complicate that inference, and none of them is a defence of nativism as such.

The data budget. A child hears something on the order of ten million words by age three and perhaps a hundred million by adolescence. Large models are trained on hundreds of billions to trillions of tokens, three to five orders of magnitude more. A demonstration that a rule is learnable from a trillion words is not a demonstration that it is learnable from ten million. The BabyLM Challenge was set up precisely to test the developmentally plausible budgets of ten million and a hundred million words.

The direct test at child scale. In 2023 Aditya Yedetore, Tal Linzen, Robert Frank and Thomas McCoy trained recurrent and transformer models on CHILDES, a corpus of actual speech to children, and tested the auxiliary inversion generalisation directly. The models generalised linearly rather than hierarchically. That is the experiment the argument always wanted, run properly, and it came out on the nativist side.

The impossible languages control. If a model learns anything at all equally well, its success on English tells you nothing about what human grammars are like. Julie Kallini and colleagues tested this in 2024 by training the same architecture on systematically impossible languages, such as English with the words in a reversed or count-based order. The models learned them, but measurably less efficiently than they learned English. The architecture is therefore not indifferent to structure, which is a result neither side predicted cleanly.

ClaimSupported by the model results?
A statistical learner can represent structure dependent generalisationsYes, and this was not obvious in 1990
The generalisations are learnable from a child-sized corpusNot shown, and the direct test found the opposite
Human children learn the way these models doNot addressed by any of this work
Innate structural bias is unnecessary in principleWeakened, since the models are not neutral between possible and impossible languages

What would settle it

Ask each side what result would change its mind, and the dispute becomes tractable. A model trained on ten million words of transcribed child-directed speech, with no curriculum and no grammatical supervision, that generalises hierarchically on held-out constructions it has never seen, would be a serious problem for the nativist. A child, or a population of children, documented passing through a linear-rule stage would be a serious problem for the other side, and after four decades of looking nobody has found one. Both experiments are being run. Neither has finished.

One caution about the framing. Structure dependence is a single case study, chosen because it is crisp. The general argument covers binding, islands, control and much else, and it is entirely possible that the answer differs case by case: some structural facts falling out of the data, and others not. Treating the whole dispute as one yes or no question is how it stayed unresolved for sixty years.

Common misconceptions

  • The argument claims children never hear the crucial sentences. The strong version did, and Pullum and Scholz refuted it. The version worth defending is about frequency and about what a learner without a structural bias could extract.
  • Language models settle the question. They show the generalisations are learnable at internet scale. The learnability claim at issue concerns a child's input, which is smaller by several orders of magnitude.
  • Poverty of the stimulus is an argument for a dedicated grammar organ. It is an argument that some prior bias is needed. Whether that bias is specific to language or a general preference for hierarchical representation is a separate question the argument does not settle.
  • Corpus frequency is the whole issue. Legate and Yang's point is that raw counts mean nothing without a standard for how frequent a cue has to be, calibrated against cues children are known to use.

The takeaway

  • The auxiliary inversion case pits a linear rule against a structural one, and Crain and Nakayama found that children make many errors but not the linear one.
  • The argument has three premises, and Pullum and Scholz attacked the third by counting: the crucial sentences exist in corpora, including speech to children.
  • Legate and Yang replied with a frequency standard calibrated against known acquisition timetables, and the cue falls below it.
  • Statistical models can favour a hierarchical grammar over a linear one on child-directed speech, but in the strongest such study both grammars were supplied to the model in advance.
  • Models trained at child scale on CHILDES generalised linearly, not hierarchically, which is the direct test and it favours the nativist reading.
  • Models learn impossible languages less efficiently than possible ones, so their architecture is not neutral, and their success on English is weaker evidence than it first appears.

Sources

  1. Wikipedia. (n.d.). Poverty of the stimulus. Wikimedia Foundation. en.wikipedia.org
  2. Cowie, F. (2017). Innateness and language. In E. N. Zalta (Ed.), The Stanford Encyclopedia of Philosophy. Stanford University. plato.stanford.edu
  3. Pullum, G. K., and Scholz, B. C. (2002). Empirical assessment of stimulus poverty arguments. The Linguistic Review, 19(1-2), 9-50.
  4. Crain, S., and Nakayama, M. (1987). Structure dependence in grammar formation. Language, 63(3), 522-543.
  5. Perfors, A., Tenenbaum, J. B., and Regier, T. (2011). The learnability of abstract syntactic principles. Cognition, 118(3), 306-338.
Key terms
Poverty of the stimulus
The argument that a child's linguistic experience underdetermines the grammar acquired, so some prior bias must be doing part of the work.
Structure dependence
The property of grammatical rules that they refer to hierarchical constituents rather than to positions in a string.
Crucial sentence
An input sentence that distinguishes two candidate rules, here a polar question whose subject contains an auxiliary of its own.
Indirect statistical evidence
Information elsewhere in the input that lets a learner favour one grammar without ever meeting the crucial sentence.
CHILDES
A large archive of transcribed speech to and by children, used to estimate what a learner actually hears.
Data budget
The quantity of input a learner receives, roughly ten million words for a young child against hundreds of billions of tokens for a large model.
Impossible language control
Training the same architecture on a language no human grammar could have, to test whether the learner is neutral between systems.
Minimal pair benchmark
A test set of sentence pairs differing in one grammatical property, used to probe what a model has learned.

Module 6: Other Ways to Do It, and Three Sentences to Finish

Four frameworks that reject either the derivation or the autonomy of syntax, compared on the same data; then one last lesson in which three sentences are analysed end to end and each analysis is defended against a named rival.

Four Frameworks That Do Without a Derivation

  • Explain how Construction Grammar accounts for a sentence whose meaning is not contributed by any word in it.
  • Represent one sentence as a dependency structure and say what the representation gains and loses against a phrase structure tree.
  • State how HPSG and LFG handle long-distance dependencies without movement, and compare their coverage with the transformational account.

A verb that cannot take an object, taking an object

Here is a sentence Adele Goldberg put at the centre of a 1995 book:

(1) She sneezed the napkin off the table.

Everyone understands it and nobody has heard it before. Now check the verb. Sneeze is intransitive:

(2) She sneezed.
(3) *She sneezed the napkin.

Sentence (3) is dead, so the object in (1) cannot be licensed by sneeze. Neither can the meaning: nothing in the lexical entry for sneeze says anything about causing motion, and yet (1) unmistakably means that she caused the napkin to move. The theory you have been building for fifteen lessons says that a verb projects its arguments and the structure follows. Sentence (1) has an argument structure that its verb did not project.

Construction Grammar takes the obvious way out: the pattern itself carries meaning. There is a caused-motion construction, of the shape subject verb object oblique, whose meaning is X causes Y to move along path Z, and a verb that is compatible with it can be slotted in. On this view a grammar is an inventory of pairings of form with meaning, from morphemes through idioms up to fully schematic patterns, with nothing that is purely structural and no derivation anywhere.

Worth holding on to: The disagreement in this lesson is not about whether structure exists. All four frameworks build elaborate structures. It is about whether a sentence is derived by operations applying in sequence, and whether syntax is a module that meaning is read off afterwards.

Constructions all the way down

The strongest evidence for the constructional view comes from patterns that are neither fully idiomatic nor fully regular. Charles Fillmore and Paul Kay analysed one in 1999:

(4) What is this fly doing in my soup?

On the literal reading it asks about the fly's activity. On the reading everyone actually gets, it registers a complaint that something incongruous is present. That reading is available with almost any noun and any location, so it is productive, and it comes with a fixed frame: the auxiliary must be a form of be, the verb must be doing, and the complaint reading disappears if you change either. Neither a lexical entry nor a general rule captures it, because it is a rule that applies to exactly one shape.

The usage-based tradition, associated with Michael Tomasello and with usage-based linguistics more broadly, adds a developmental claim to the constructional one. Children start with concrete item-based patterns organised around particular verbs, generalise slowly, and never need an innate grammar because frequency and analogy do the work. The evidence is distributional: a two year old who uses one verb in three frames often uses another verb in only one, which is hard to state if the child has abstract argument structure rules.

Words linked to words

Dependency grammar makes a different cut. It keeps a hierarchy and throws away the phrase. In Lucien Tesniere's 1959 formulation there are no NP or VP nodes at all: there are words, and each word except one depends on exactly one other word.

put --nsubj--> Kim ; put --obj--> book ; book --det--> the ; put --obl--> table ; table --case--> on

Every relation in that structure is labelled, which is what phrase structure leaves implicit. Against that, the constituency tests of Lesson 2 have nothing to apply to. There is no node corresponding to on the table, so the fact that this string can be substituted, coordinated and clefted has to be stated some other way. Dependency structures nonetheless dominate practical parsing, and the Universal Dependencies project has produced over two hundred consistently annotated treebanks in more than a hundred and fifty languages, a scale no phrase structure annotation effort has matched. For languages with free word order and rich case marking, a representation that never has to decide what is adjacent to what is a considerable practical advantage.

Two frameworks that keep the structure and drop the derivation

Head-driven phrase structure grammar, set out by Carl Pollard and Ivan Sag in 1994, represents every linguistic object as a bundle of features, and builds sentences by unifying those bundles. There is one level of representation and nothing moves. Long-distance dependencies are handled by a feature named SLASH: a phrase with a gap in it is marked as such, the mark is passed up from the gap site to the filler, and the filler discharges it.

(5) Which book did Kim think Lee had read?

In HPSG the VP had read carries a SLASH value naming the missing object, the clause containing it inherits the value, and which book at the top cancels it. Every intermediate node bears the mark, which means the theory says something about intermediate positions, exactly as successive cyclic movement does. The Irish complementiser facts from Lesson 10 are as expressible here as they are in a movement account.

Lexical functional grammar, from Ronald Kaplan and Joan Bresnan in 1982, splits the job in two. A constituent structure records order and dominance; a functional structure records grammatical relations such as subject and object. Raising, which cost Lesson 9 a whole lesson, becomes a statement that one functional structure is shared between two predicates:

(6) Kim seems to be tired.

The subject of seem and the subject of be tired are the same object in the functional structure, so nothing has to move to make them the same. Passive is a lexical rule that changes which argument maps to which function, not an operation on a tree, which explains why passive is sensitive to individual verbs in a way syntactic operations usually are not.

The same data, four ways

FrameworkWhat replaces movementStrongest evidence for itWhere it is stretched
Construction GrammarNothing moves; patterns carry meaning directlySneeze the napkin off the table, and the whole idiom to schema continuumSystematic relations between constructions, which have to be stated as inheritance links
Usage-based approachesNothing moves; generalisation over stored exemplarsVerb-specific stages in early child speech, frequency effects on acceptabilityThe structure dependence facts of Lesson 15, which remain contested
Dependency grammarNothing moves; non-adjacent words are simply linkedFree word order languages, and treebanks in over a hundred languagesConstituency test results, coordination, and anything defined over phrases
HPSG and LFGFeature passing and structure sharingComputational implementability, no empty categories, lexical treatment of passiveIsland constraints, which still have to be stipulated as conditions on feature paths

Read the last column. Every framework in the table has to say something about islands, and none of them gets them for free. That is worth knowing, because the usual defence of the transformational account is that it derives constraints from locality, and the usual attack is that these are stipulations dressed as principles. The honest summary is that all five accounts, including the one this course has been building, stipulate something at that point.

In short: The choice between these frameworks is not settled by any single sentence. It is settled, to the extent it is settled at all, by which set of stipulations you find you can live with, and by what you want a grammar to be for.

What each one is for

That last clause is not a dodge. The frameworks were built for different purposes and they are good at different things. HPSG and LFG were designed to be computationally implementable, and grammars in both have been used in working parsers for decades. Dependency grammar was built for description and annotation and dominates practical natural language processing. Construction Grammar was built to handle the enormous middle ground between the fully idiomatic and the fully regular, which transformational syntax historically pushed to one side. The transformational tradition was built to explain why grammars have the shape they do rather than any other, which is why islands, binding and the poverty of the stimulus keep appearing in it and appear far less in the others.

A working syntactician in 2026 is expected to be able to read all of them. Papers in Glossa and in Language argue across the line regularly, and a claim stated in one framework can usually be restated in another with effort. Where restatement fails, you have found something worth writing about.

Common misconceptions

  • These frameworks reject the idea that syntax has structure. All four build structures at least as detailed as an X-bar tree. What they reject is a derivation, or the autonomy of syntax, or both.
  • Construction Grammar cannot state generalisations. It states them as inheritance relations between constructions, so the caused-motion pattern and the ditransitive pattern can share properties without either being derived from the other.
  • Dependency grammar is just phrase structure with the nodes rubbed out. The two are inter-translatable only for a restricted class of structures, and they disagree about which word heads a phrase in several well known cases, including determiners and auxiliaries.
  • HPSG and LFG are notational variants of transformational grammar. They make different predictions about lexical exceptions, and their treatment of passive as a lexical rule rather than a syntactic operation is a substantive claim with testable consequences.

Pulling it together

  • She sneezed the napkin off the table has an argument structure its verb cannot project, which is the central argument that patterns carry meaning independently of words.
  • The What is X doing Y construction is productive and yet tied to one fixed frame, which is the kind of case a rule-plus-lexicon architecture handles worst.
  • Dependency grammar keeps hierarchy and labelled relations while discarding phrases, which costs it the constituency tests and buys it annotation at scale.
  • HPSG passes a SLASH feature from gap to filler and LFG shares one functional structure between predicates, so both cover long-distance dependencies and raising without any movement.
  • Islands must be stipulated in every framework compared here, including the transformational one.
  • The frameworks were built for different purposes, and the choice among them is partly a choice about what a grammar is supposed to explain.

Sources

  1. Wikipedia. (n.d.). Construction grammar. Wikimedia Foundation. en.wikipedia.org
  2. Wikipedia. (n.d.). Dependency grammar. Wikimedia Foundation. en.wikipedia.org
  3. Goldberg, A. E. (1995). Constructions: A construction grammar approach to argument structure. University of Chicago Press.
  4. Pollard, C., and Sag, I. A. (1994). Head-driven phrase structure grammar. University of Chicago Press and CSLI Publications.
  5. Bresnan, J., Asudeh, A., Toivonen, I., and Wechsler, S. (2016). Lexical-functional syntax (2nd ed.). Wiley-Blackwell.
Key terms
Construction
A learned pairing of form with meaning, from a single morpheme up to a fully schematic argument structure pattern.
Caused-motion construction
The pattern subject verb object oblique, whose meaning of causing something to move is contributed by the pattern rather than by the verb.
Item-based pattern
An early child construction organised around one particular verb, generalised only gradually to others.
Dependency
A directed labelled link from a head word to a dependent word, with no phrasal node mediating between them.
SLASH feature
The HPSG device that records a missing element and passes the record from the gap site up to the filler that discharges it.
Functional structure
The LFG representation of grammatical relations such as subject and object, separate from the representation of order and dominance.
Structure sharing
One representation being the value of two different attributes at once, which does the work movement does in a raising sentence.
Lexical rule
A relation between two lexical entries, used in LFG and HPSG for passive, which predicts that passivisation can be idiosyncratic verb by verb.

Three Sentences, End to End, With the Objections Answered

  • Produce a complete analysis of a sentence: categories, constituency, structure, movement, case, and binding, written as a labelled bracketing.
  • Name the rival analysis of each sentence and give the diagnostic that decides between them.
  • Apply the same procedure to a sentence of your own, and state what evidence would refute your analysis.

Sentence one: The report seems to have been leaked to a journalist

Start with the string and end with a defence. The procedure has six steps and you will run it three times.

Categories. The is a determiner, report and journalist are nouns, seems is a finite verb, to is non-finite T, have and been are auxiliaries, leaked is a passive participle, and the second to is a preposition. The two instances of to are different words, which the distributional tests of Lesson 3 settle at once: only the second can be followed by a noun phrase and fronted with it.

Constituency. Substitution replaces the report with it; clefting isolates to a journalist; and seems to have been leaked to a journalist is a constituent because it survives ellipsis: The memo did too.

Structure.

[TP [NP The report] [T-bar [T -s] [VP [V seem] [TP [NP --] [T to] [VP have been [VP [V leaked] [NP --] [PP to a journalist]]]]]]]

Movement. Two applications of A-movement, in sequence. The report starts as the internal argument of leaked, where it receives the theme role. Passive morphology means the participle assigns no accusative case, so it must move; it lands in the subject position of the non-finite clause, which is the EPP requirement of Lesson 6. That position is caseless, because non-finite T assigns none. So it moves again, into the matrix subject position, where finite T assigns nominative.

Case and roles. One theta role for the report, assigned once at the bottom and carried along; one for the implicit agent, suppressed by the passive; one for a journalist, assigned by the preposition. Nominative on the report. This is Burzio's generalisation from Lesson 8 doing exactly what it was stated to do.

The rival, and the diagnostics. Someone could analyse seem as a control verb: the report is its argument, and it controls a PRO subject in the embedded clause. Three tests kill that.

TestPrediction if seem is a control verbWhat we find
Expletive subjectBlocked, since the subject bears a roleThere seems to have been a leak, which is fine
Idiom chunkIdiomatic reading lostThe cat seems to have been let out of the bag, idiom intact
Passive synonymyMeaning should shift when the embedded voice changesThe paper seems to have printed the leak and The leak seems to have been printed by the paper state the same fact

Remember: An analysis is not a diagram. It is a claim plus the evidence that would have refuted it, and the table above is the part that does the work.

Sentence two: Which of his own reports did the auditor say the committee had suppressed?

Categories and constituency. Which of his own reports is a noun phrase, established by substitution with what and by the fact that it can be answered with a bare noun phrase. Did is finite T in the C position. Say takes a clausal complement, optionally introduced by that.

Structure.

[CP [NP Which of his own reports] [C did] [TP the auditor [VP say [CP -- [TP the committee had [VP suppressed [NP --]]]]]]]

Movement. Two operations. T-to-C moves the finite auxiliary into C, which is why the subject follows it. A-bar movement takes the wh-phrase from the object position of suppressed to the specifier of the matrix CP, stopping at the embedded specifier of CP on the way, which is the dash inside the lower bracket.

Why there is a gap at all. Suppress is obligatorily transitive: *The committee had suppressed is out. The sentence is fine, so the frame is satisfied by something unpronounced.

Binding. His own needs an antecedent, and the intended one is the auditor, which does not c-command the front of the sentence. Lesson 13 says binding requires c-command, so the phrase must be evaluated in the position it came from, where the auditor does c-command it. Reconstruction is thus not a separate stipulation here; it is what makes the sentence interpretable at all.

The rival, and the diagnostics. The rival is base generation: the wh-phrase is generated at the front, and the object position holds a silent pronoun that it is merely construed with. Three facts weigh against it.

(1) *Which of his own reports did the auditor meet the journalist who suppressed?
(2) Which of his own reports did the auditor file without reading?
(3) Which of his own reports did the auditor say the committee had suppressed?

Sentence (1) is an island violation, and a base-generated construal has no reason to care about a relative clause boundary; a movement dependency does, by Lesson 11. Sentence (2) has a parasitic gap after reading, which is licensed only when a real gap runs past it. And the reconstruction fact above requires the phrase to be interpreted low, which is natural if it was there and awkward if it never was. Note also that adding that before the committee changes nothing, because this is object extraction; make it subject extraction and the that-trace effect appears.

Sentence three: Kim promised the committee to resign, and Lee did too

Who resigns. Kim. That is worth pausing on, because the nearest noun phrase to the infinitive is the committee, and for almost every other verb of this shape the nearest one wins:

(4) Kim persuaded the committee to resign. (the committee resigns)
(5) Kim promised the committee to resign. (Kim resigns)

Promise is the standard exception to the minimal distance principle, and Lesson 9 recorded that children over-regularise it until around eight years old.

Structure.

[TP [NP Kim] [VP [V promised] [NP the committee] [CP [TP [NP PRO] [T to] [VP resign]]]]]

Roles. Three for promise: an agent, a recipient, and a proposition. PRO bears the sole role of resign and is controlled by the matrix subject. This is control and not raising, and the same expletive test settles it: *There promised the committee to be a problem is dead, whereas the corresponding raising sentence with seem was fine.

The ellipsis. The second clause is verb phrase ellipsis, licensed by did, with the antecedent VP promised the committee to resign. Copy it in and something interesting happens: the PRO inside the copied material is controlled by the new subject, so the sentence means Lee promised the committee that Lee would resign, not that Kim would.

The rival, and the diagnostic. On a pro-form account the gap is an atom with no structure inside it, its content recovered from the meaning of the antecedent. That account has to explain why the controller switches. If the gap simply meant what the first clause meant, the second clause would say Lee promised that Kim would resign, and it does not. The controller switches because there is a PRO inside the gap looking for the nearest available subject, which is Lesson 12's sloppy reading arriving through a different door.

The upshot: Three sentences, three unpronounced elements: a trace of A-movement, a trace of A-bar movement, and a PRO inside an ellipsis site. None of them is audible and all three are detectable, which is the single claim this course has been making since Lesson 1.

The procedure, written out

StepWhat you doWhat tells you that you are wrong
1. CategoriesAssign each word a category on distribution, not meaningA word that fails the frames its assigned category licenses
2. ConstituencyRun substitution, coordination, movement, ellipsis, cleftingTests that disagree, which usually means the string is ambiguous
3. StructureWrite a labelled bracketing with binary branchingAn attachment the one-replacement test refuses
4. DependenciesIdentify every gap and say what filled itAn unsatisfied subcategorisation frame with no gap posited
5. Case and rolesCheck every noun phrase has one role and one caseA caseless noun phrase that stays put, or two roles on one argument
6. The rivalState the best competing analysis and the test that separates themNo test exists, in which case you have a notation, not a claim

Step six is the one that separates syntax from labelling. Anyone can draw a structure over a sentence. The question a reviewer asks is what the world would look like if you were wrong, and the tests you have collected across this course, expletives, idiom chunks, selectional restrictions, one-replacement, parasitic gaps, reconstruction and island sensitivity, are the accumulated answers to that question.

A last honest word about the framework. Everything above is stated in a broadly transformational idiom, and Lesson 16 showed you four traditions that would state these analyses differently. The facts are not in dispute. The report is understood as the thing leaked; the auditor binds the reflexive; Lee is the one who would resign. What varies is the machinery, and the mark of a good analysis in any framework is the same: it predicts a judgment you have not yet made.

Common misconceptions

  • The tree is the analysis. The structure is a summary of the analysis. The analysis is the set of tests that ruled out the alternatives, which is why step six exists.
  • Two sentences with the same meaning have the same structure. The report seems to have been leaked and It seems that the report was leaked mean the same and have different structures, as the position of the subject shows.
  • An ungrammatical string refutes a rule. It refutes the rule only if nothing else in the derivation could be responsible, which is why the diagnostic tables in this lesson vary one thing at a time.
  • A more complicated analysis is a worse one. It is worse only if it buys nothing. The two-step A-movement in sentence one is more complicated than a one-step account and it is what makes the case facts come out right.

What you now know

  • A full analysis runs six steps: categories, constituency, structure, dependencies, case and roles, and the rival with its diagnostic.
  • The report seems to have been leaked involves two applications of A-movement, forced by passive morphology and then by the absence of case on non-finite T.
  • Which of his own reports did the auditor say the committee had suppressed involves T-to-C, successive cyclic A-bar movement, and reconstruction of the fronted phrase for binding.
  • Island sensitivity, parasitic gap licensing and reconstruction together rule out base generation of the wh-phrase.
  • Kim promised the committee to resign is subject control, against the minimal distance principle, and the ellipsis in the second clause switches the controller, which a structureless gap cannot explain.
  • An analysis without a test that could refute it is a notation rather than a claim.

Sources

  1. Wikipedia. (n.d.). Transformational grammar. Wikimedia Foundation. en.wikipedia.org
  2. Wikipedia. (n.d.). Generative grammar. Wikimedia Foundation. en.wikipedia.org
  3. Carnie, A. (2021). Syntax: A generative introduction (4th ed.). Wiley-Blackwell.
  4. Sportiche, D., Koopman, H., and Stabler, E. (2014). An introduction to syntactic analysis and theory. Wiley-Blackwell.
  5. Adger, D. (2003). Core syntax: A minimalist approach. Oxford University Press.
Key terms
Labelled bracketing
A linear notation for a syntactic structure in which each pair of brackets carries the category of the node it encloses.
A-movement
Movement into an argument position, driven by case and the EPP, as in passive and raising.
A-bar movement
Movement into a non-argument position such as the specifier of CP, as in questions, relatives and clefts.
Reconstruction
Evaluation of a fronted phrase in its original position, required here so that the auditor can c-command the anaphor inside it.
PRO
The unpronounced subject of a control infinitive, which bears a theta role and is controlled by an argument of the higher verb.
Minimal distance principle
The default that the nearest noun phrase controls PRO, which promise violates and persuade obeys.
Diagnostic
A test whose outcome differs between two candidate analyses, which is what turns a structure into a claim.
Burzio's generalisation
The link between assigning an external theta role and assigning accusative case, which forces the passive object to move.

Open the interactive version with quizzes and progress →