๐ŸŽ Education · Graduate · EDUC 310

Learning Sciences & Instructional Design

A graduate survey of how human learning actually works and how to design instruction that respects it. You will move from the cognitive architecture of memory and cognitive load, through the strategies that research repeatedly shows produce durable learning, into the practical craft of writing objectives, designing assessments, giving feedback, building multimedia, and evaluating whether any ofโ€ฆ

Start the interactive course (quizzes, progress, videos) →

Free forever. No sign-up, no ads. 19 lessons. The full lesson text is below so you can read it right here.

Module 1: How Memory and Cognition Work

The cognitive architecture - sensory, working, and long-term memory - that every instructional decision must respect.

The Information-Processing Model

  • Describe the flow of information through sensory, working, and long-term memory.
  • Explain why attention is the gatekeeper of learning.
  • Distinguish encoding, storage, and retrieval as distinct processes.

Almost every defensible teaching decision rests on a simple picture of how the mind takes in and holds information. The dominant account in the learning sciences is the information-processing model, which describes learning as the flow of information through a series of memory systems, each with its own capacity and duration. It is a model, not a photograph of the brain, but it predicts an enormous amount about why instruction succeeds or fails.

The model has a history worth knowing. Richard Atkinson and Richard Shiffrin set out the canonical version in 1968, and it became so standard that psychologists still call it the modal model. Its power lies in being a functional account: it does not claim that sensory, working, and long-term memory sit in three separate boxes inside the skull, only that human performance behaves as though information passes through stages with sharply different capacities and time courses. Half a century of revision has changed the details - Alan Baddeley split working memory into components, Nelson Cowan reframed it as the currently activated portion of long-term memory - but the functional distinctions survived, and they are exactly what instructional design needs.

Key idea: The information-processing model is a functional account of stages with different capacities and durations, not a map of brain anatomy, and its instructional value survives the revisions to its details.

Three stores

Information enters through the senses and passes through three functionally distinct stores.

  • Sensory memory holds a raw, high-capacity impression of what the senses just registered - the fading echo of a sound, the brief visual afterimage - for a fraction of a second. Almost all of it decays unless attention selects it.
  • Working memory is where conscious thought happens: where you hold and manipulate the handful of items you are actively thinking about. It is severely limited in both capacity and duration. This is the bottleneck of learning, and Module 2 is devoted to its consequences.
  • Long-term memory is the vast, durable store of everything you know: facts, procedures, and organized knowledge structures. Its capacity is effectively unlimited and information there can last a lifetime.

Learning, in this model, is the process of building and reorganizing durable structures in long-term memory. Something has been learned when it can be retrieved and used later, not merely when it has passed through working memory once.

Key idea: Learning means building durable structures in long-term memory, so the test of learning is later retrieval and use, not momentary comprehension.

How much fits, and for how long

The stores differ by orders of magnitude, and the numbers matter for design. George Sperling's partial-report experiments in 1960 showed that visual sensory memory briefly holds far more than a person can report: participants could name any cued row of a flashed letter matrix, but the trace decayed within roughly half a second. George Miller's celebrated 1956 paper put the span of immediate memory at about seven items plus or minus two - though Miller was careful to say seven chunks, not seven raw pieces of information, a distinction the next lesson turns into a design tool. Later work using tasks that block rehearsal and grouping converged on a smaller figure: Nelson Cowan's reviews put the pure capacity limit at roughly four chunks in healthy young adults, and fewer still when items must be transformed rather than merely held.

Duration is equally lopsided. Unrehearsed material in working memory decays within seconds. Long-term memory, by contrast, has no demonstrated capacity ceiling and can hold material for decades. Harry Bahrick's studies of Spanish learned in school found substantial retention up to fifty years later among people who had never used the language since, with the steepest losses concentrated in the first three to six years and a long plateau afterward - what he called permastore. The asymmetry between a roughly four-chunk workspace and an effectively unlimited archive is the single most consequential fact in this course, and nearly every technique that follows is an attempt to exploit it.

Key idea: Working memory holds roughly four chunks for seconds while long-term memory has no known ceiling and can retain material for decades, and every design decision follows from that asymmetry.

Attention is the gatekeeper

Attention decides which slice of the flood of sensory input is admitted to working memory for processing. Because attention is limited, whatever the learner is not attending to is effectively not being learned. This is why a beautifully designed slide is useless if the learner is reading email, and why competing streams of information (a narrator talking over on-screen text that says something different) sabotage learning. The instructional implication is blunt: you cannot teach what the learner does not attend to, so managing attention is not a nicety, it is a precondition.

Two experimental literatures sharpen the point. Research on divided attention shows that adding a demanding secondary task during study sharply reduces later memory for the primary material even when comprehension feels intact at the time - the learner does not notice the loss, which is what makes it dangerous. Research on task switching shows that alternating between tasks carries a measurable reconfiguration cost on every switch, so what students experience as multitasking is really rapid alternation paid for in time and errors. Classroom studies of laptop note-taking point the same way: students who type tend to transcribe verbatim and do worse on conceptual questions than students writing longhand, and students merely seated within view of an off-task screen also do worse. Technology is not the villain here. Divided attention is, and technology is simply its most reliable supplier.

Attention in a real classroom

Consider a twenty-minute explanation of Ohm's law. A teacher who talks continuously while a slide deck cycles decorative photographs is competing with themselves for a single limited channel. A load-aware version of the same lesson signals the structure in advance ("there are three quantities and one relation - watch how each changes when another is held fixed"), strips out the photographs, pauses after each quantity for a ten-second written recall prompt, and poses one question that cannot be answered without having attended. The content is identical; what differs is that attention has been directed rather than assumed. Pacing and signaling are therefore not showmanship - they are the only mechanism by which content enters the system at all.

Key idea: Divided attention reliably degrades later memory even when comprehension feels fine, so instruction must direct and protect attention rather than assume it.

Three processes, not one

It helps to separate three things that casual talk of "memory" runs together.

  1. Encoding is getting information into long-term memory in the first place, which depends heavily on how deeply and meaningfully the learner processes it.
  2. Storage is the retention of that information over time.
  3. Retrieval is getting it back out when needed - and, as later lessons show, the act of retrieval is not a neutral readout but itself a powerful learning event.

Many teaching failures are really failures of one specific process. A student who "knew it last night but blanked on the test" often has an encoding or retrieval problem, not a storage problem. Diagnosing which process broke down tells you which fix to apply. Keep these three stores and three processes in mind; the rest of this course is, in a sense, an extended argument about how to move knowledge into long-term memory and get it back out reliably.

Depth of processing, not time on task

Encoding does not depend on how long material sits in working memory but on what is done with it there. Fergus Craik and Robert Lockhart proposed the levels of processing framework in 1972: material processed for meaning is retained far better than the same material processed for surface features such as typeface or rhyme. In the classic demonstration by Craik and Tulving, participants asked whether a briefly shown word fit a sentence frame later recognized many more of those words than participants asked whether the same word was printed in capitals, despite equal exposure time. The framework has been criticized for circularity, since depth is partly defined by the memory it produces, but the practical lesson is robust and returns in every strategy in Module 3: what the learner does with the material determines what gets encoded. Time on task is a proxy, and often a poor one.

Key idea: Encoding depends on the kind of processing rather than the duration of exposure, so tasks that force processing for meaning outperform tasks that merely keep material in view.

Where the model is contested

A graduate reader should hold the model loosely in three places. First, the strict three-box picture has been superseded: Baddeley's multicomponent model divides working memory into a phonological loop, a visuospatial sketchpad, a central executive, and an episodic buffer, while Cowan's embedded-processes account treats it as the currently activated subset of long-term memory plus a narrow focus of attention. Second, the boundary between working and long-term memory is porous - Ericsson and Kintsch argued that experts deploy a long-term working memory in which practised retrieval structures let them hold far more task-relevant information than the classical limit predicts. Third, the model is silent about motivation, emotion, and social context, all of which determine whether a learner engages the machinery at all, which is why Module 4 exists.

None of this undermines the design implications. Whatever the correct architecture turns out to be, the regularities instruction must respect - a severe momentary processing limit, a vast durable store, and a selective gate between them - are not in dispute.

Key idea: Competing accounts refine the architecture, but the design-relevant regularities of a narrow processing bottleneck and a vast durable store are not contested.

What this architecture rules out

Educational psychology is crowded with confident claims that sound cognitive and are not. Several are contradicted by the architecture just described, yet surveys find teachers endorse them at high rates.

  • "We only use ten percent of our brains." False. Imaging and lesion evidence show that essentially all of the brain is active across a day and that damage almost anywhere produces deficits. There is no dormant reserve waiting to be unlocked by a teaching method.
  • "People remember 10 percent of what they read, 20 percent of what they hear, and 90 percent of what they do." These percentages, usually attached to a triangle called Dale's Cone of Experience, are fabricated. Edgar Dale's 1946 cone was an unnumbered diagram of the concreteness of experiences; the figures were bolted on later by other people and trace back to no study at all.
  • "Some students are visual learners and must be taught visually." The learning-styles matching hypothesis has repeatedly failed the crossover experiment required to support it. Lesson 8 works through the evidence and the reason the idea persists.
  • "Students are left-brained or right-brained." Hemispheric specialization is real for some functions, notably language, but people do not come in dominant-hemisphere types, and no instructional benefit has been shown from sorting them that way.
  • "Digital natives process information differently." Growing up with technology does not produce a different cognitive architecture. Reviews find no evidence of a generation with different working-memory limits, and the divided-attention costs described above apply to everyone.

Dekker and colleagues surveyed teachers in the United Kingdom and the Netherlands and found roughly half endorsing the ten-percent claim and around 90 percent endorsing learning styles. Their most sobering finding was that greater general knowledge about the brain did not protect against these beliefs - if anything, teachers who read more popular neuroscience believed slightly more myths. Knowing the actual architecture, and knowing which claims it excludes, is the only reliable defence.

Recap

  • The information-processing model describes learning as a flow through sensory, working, and long-term memory, and it is a functional account rather than a brain map.
  • Working memory holds roughly four chunks for a matter of seconds; long-term memory has no known ceiling and can retain material for decades.
  • Attention is the selective gate, and divided attention degrades later memory even when in-the-moment comprehension feels intact.
  • Encoding, storage, and retrieval are distinct processes, and diagnosing which one failed tells you which fix to apply.
  • Depth of processing, not exposure time, determines what is encoded.
  • This architecture gives no support to the ten-percent myth, Dale's Cone percentages, learning-style matching, hemisphere types, or digital natives as a cognitive category.

Sources

  1. Miller, G. A. (1956). The magical number seven, plus or minus two: Some limits on our capacity for processing information. Psychological Review, 63(2), 81-97. Classics in the History of Psychology. psychclassics.yorku.ca
  2. Cowan, N. (2010). The magical mystery four: How is working memory capacity limited, and why? Current Directions in Psychological Science, 19(1), 51-57. pmc.ncbi.nlm.nih.gov
  3. Squire, L. R. (2004). Memory systems of the brain: A brief history and current perspective. Neurobiology of Learning and Memory, 82(3), 171-177. pubmed.ncbi.nlm.nih.gov
  4. Dekker, S., Lee, N. C., Howard-Jones, P., and Jolles, J. (2012). Neuromyths in education: Prevalence and predictors of misconceptions among teachers. Frontiers in Psychology, 3, 429. pmc.ncbi.nlm.nih.gov
  5. Sweller, J., van Merrienboer, J. J. G., and Paas, F. (2019). Cognitive architecture and instructional design: 20 years later. Educational Psychology Review, 31, 261-292. link.springer.com
  6. OpenStax. (2020). How memory functions. In Psychology 2e, Section 8.1. Rice University. CC BY 4.0. openstax.org
  7. Atkinson, R. C., and Shiffrin, R. M. (1968). Human memory: A proposed system and its control processes. In K. W. Spence and J. T. Spence (Eds.), The Psychology of Learning and Motivation (Vol. 2, pp. 89-195). Academic Press. find source โ†—
Key terms
Information-processing model
An account of learning as the flow of information through sensory, working, and long-term memory.
Sensory memory
A brief, high-capacity store holding raw sensory impressions for a fraction of a second.
Working memory
The limited-capacity system where conscious thinking and manipulation of information occur.
Long-term memory
The durable, effectively unlimited store of knowledge, skills, and organized structures.
Attention
The limited process that selects which sensory input reaches working memory.
Encoding
The process of getting information into long-term memory.

Working Memory, Schemas, and Chunking

  • Explain the capacity limits of working memory.
  • Define schemas and describe how they organize knowledge.
  • Explain how chunking and expertise expand effective capacity.

The single most important fact about the mind, for a teacher, is that working memory is small. When people are asked to hold unrelated items in mind, most can juggle only a handful at once, and that capacity shrinks further when the items must be manipulated rather than merely held. Working memory also empties quickly: without rehearsal, its contents fade within seconds. These limits are not a design flaw to be trained away; they are a fixed feature of human cognition that instruction must accommodate.

Key idea: Working-memory limits are a fixed feature of human cognition rather than a deficiency to be trained away, so instruction has to be designed around them.

Measuring the limit

How the limit is measured changes the number you get, which is why the literature contains several. A simple digit span task, where you repeat back a list, allows rehearsal and grouping and yields figures near Miller's seven. A complex span task interleaves storage with a processing demand - read a sentence, judge whether it makes sense, then remember the final word, and repeat - and yields spans of roughly three to five. Complex span is the better model of classroom cognition, because students almost never merely hold information; they hold it while doing something with it.

Complex span also predicts things instructors care about. Individual differences in complex span correlate substantially with reading comprehension, mathematical problem solving, and measures of fluid reasoning, which is one reason a lesson that overloads the workspace penalizes some students far more than others. Randall Engle's research program traces much of this variation to the ability to control attention and resist interference rather than to storage capacity as such - a useful reframing, because it means the practical problem in a classroom is often distraction and interference, not sheer shortage of slots.

Key idea: Complex span, which requires holding information while processing other information, is the classroom-relevant measure, and it varies across learners in ways that predict comprehension and problem solving.

The paradox of the expert

If working memory is so limited, how does a chess master glance at a board and reconstruct it perfectly, or a physicist read an equation as a single idea? The answer is schemas. A schema is an organized mental structure that groups many elements of information into a single, meaningful unit stored in long-term memory. The chess master does not see twenty separate pieces; they see a few familiar configurations - a castled king, a pawn chain - each of which is one chunk to working memory.

This is the resolution of the paradox: expertise works by building rich schemas in the effectively unlimited long-term store, which can then be pulled into working memory as single units. The expert has not enlarged their working memory; they have packed far more into each slot. Classic memory research made this vivid: experts recalled realistic board positions far better than novices, but their advantage largely collapsed for random positions, because scrambled boards match few stored schemas. Their edge was knowledge, not raw memory.

The chess evidence, stated carefully

The chess studies are cited so often that they deserve precision. Adriaan de Groot's 1946 work found that grandmasters and weaker players considered a similar number of moves; what differed was the quality of the moves that came to mind, suggesting perception rather than search. William Chase and Herbert Simon's 1973 experiments made the mechanism explicit: masters reconstructed realistic positions after a five-second glance far better than novices, but their advantage largely collapsed for randomly arranged pieces, and the pauses in their placements revealed that they were laying down groups of pieces at a time.

The graduate-level refinement matters. Fernand Gobet and Simon later reanalysed and extended these results and found that skilled players retain a small but reliable advantage even on random boards. The reason is that random arrangements still contain occasional familiar fragments, and masters have enough stored patterns to catch them. So the honest statement is not "experts have no advantage on random boards" but "the expert advantage shrinks dramatically once the material stops matching stored patterns, and what remains scales with the sheer number of patterns held." This is a stronger result for teaching, not a weaker one: it says expertise is a large, indexed library of situations, and that library is built one pattern at a time.

Key idea: Expert recall advantages depend on matching stored patterns, shrinking dramatically but not vanishing for random material, which shows expertise is an indexed library of situations rather than superior raw memory.

Chunking

Chunking is the process of grouping individual elements into larger meaningful units. The string F B I C I A N A S A is ten letters to someone who does not parse it, straining working memory. Grouped as FBI CIA NASA, it is three familiar chunks, trivially held. Chunking is why phone numbers are broken into groups, and why teaching students the underlying structure of a problem lets them hold more of it in mind at once.

SituationLoad on working memoryWhy
10 random letters, ungroupedOverloaded - exceeds capacityEach letter is a separate element
Same letters as 3 known acronymsEasily heldThree chunks backed by schemas
Novice reads a physics equation term by termHeavyNo schema; many separate symbols
Expert reads the same equationLightThe whole equation is one meaningful unit

Key idea: Chunking does not change the number of slots available; it changes how much meaning each slot can carry, and the meaning comes from schemas already in long-term memory.

Why you cannot simply train a bigger workspace

If working memory constrains so much, the obvious move is to expand it, and a commercial industry sells exactly that promise. The evidence does not support it. Monica Melby-Lervag and Charles Hulme's meta-analysis of working-memory training found reliable improvement on tasks resembling the trained ones, modest and short-lived gains on similar untrained tasks, and no convincing evidence of far transfer to reasoning, reading, or arithmetic. Their later meta-analysis with Thomas Redick reached the same conclusion with stricter controls: performance on measures of intelligence and other distant outcomes did not improve. A large consensus review of commercial brain-training programs by Daniel Simons and colleagues concluded that the marketing far outruns the evidence, with most improvements confined to the trained task itself.

The instructional consequence is decisive. Since the workspace cannot be enlarged, every gain must come from the other two directions: putting more meaning into each chunk (building schemas, the subject of this whole course) and wasting less of the workspace on avoidable demands (managing load, the subject of Module 2). Sweller's group makes the same point from the other end - domain-specific knowledge, not general cognitive skill, is what allows complex performance, which is why "teaching thinking skills" divorced from content reliably disappoints.

Key idea: Working-memory capacity itself resists training and shows no far transfer, so instruction must raise the meaning packed into each chunk and cut the demands that waste capacity.

Building chunks deliberately

What does chunk building look like in a real subject? In early reading, a child who decodes letter by letter has almost no capacity left for meaning; once frequent words are recognized as single units, comprehension becomes possible. The intervention is not "try harder to comprehend" but repeated practice to automate word recognition. In chemistry, a novice reading Fe2(SO4)3 processes eight separate symbols and subscripts, while an experienced student sees one compound, freeing capacity to reason about the reaction it participates in. In music, a beginner reads note by note, whereas a fluent sight reader recognizes scale fragments and chord shapes.

In each case the teaching move is identical in structure: identify the units that must become automatic, give distributed practice until they are, and only then pose problems that require holding several of those units at once. Sequencing this in the wrong order - demanding integration before the components are automatic - produces the familiar picture of a student who "understands it when we do it together" but collapses when working alone. Nothing about their motivation changed; their workspace simply filled.

Key idea: Automate the component units first, then require the integration that depends on them, because integration attempted before automation exhausts the workspace.

The instructional payoff

Two consequences follow immediately. First, what overwhelms a novice may be trivial for an expert, purely because of schemas, so instruction must be pitched to the learner's current knowledge, not the teacher's. This is the root of the "expert blind spot," where an instructor forgets how many separate pieces a beginner is juggling. Second, a central goal of teaching is deliberately schema construction: helping learners build organized structures so that what once required effortful assembly becomes a single automatic chunk. Much of what we call "understanding" is exactly this - having good schemas - and much of what we call practice is the slow work of building and automating them.

Common misconceptions

  • Working memory can be expanded with brain-training apps. Trained tasks improve; reasoning, reading, and mathematics do not. The limit is a design constraint, not a target.
  • Miller proved we hold seven items. Miller wrote about seven chunks, under conditions permitting rehearsal and grouping; strict measures put the limit nearer four.
  • Chess masters have better memories. Their advantage is domain-specific pattern knowledge, and it shrinks dramatically once the board stops resembling real play.
  • Struggling students are not trying. Overload looks exactly like disengagement from the outside, and the two require opposite responses.
  • Generic thinking skills can substitute for content. Complex performance rests on domain-specific schemas; skills taught without content rarely transfer.

Recap

  • Working memory holds only a few elements at once, and complex span - holding while processing - is the classroom-relevant measure.
  • Expertise resolves the paradox by packing more meaning into each slot through schemas stored in long-term memory.
  • The chess evidence shows a pattern-based advantage that shrinks sharply, but does not vanish, for random positions.
  • Capacity itself resists training and shows no far transfer, so gains must come from richer chunks and lower wasted load.
  • Teach components to automaticity before demanding tasks that require integrating several of them.
  • The expert blind spot arises because experts perceive as one chunk what novices must juggle as many.

Sources

  1. Miller, G. A. (1956). The magical number seven, plus or minus two: Some limits on our capacity for processing information. Psychological Review, 63(2), 81-97. Classics in the History of Psychology. psychclassics.yorku.ca
  2. Cowan, N. (2010). The magical mystery four: How is working memory capacity limited, and why? Current Directions in Psychological Science, 19(1), 51-57. pmc.ncbi.nlm.nih.gov
  3. Gobet, F., and Simon, H. A. (1996). Recall of random and distorted chess positions: Implications for the theory of expertise. Memory and Cognition, 24(4), 493-503. pubmed.ncbi.nlm.nih.gov
  4. Melby-Lervag, M., Redick, T. S., and Hulme, C. (2016). Working memory training does not improve performance on measures of intelligence or other measures of far transfer: Evidence from a meta-analytic review. Perspectives on Psychological Science, 11(4), 512-534. pubmed.ncbi.nlm.nih.gov
  5. Simons, D. J., Boot, W. R., Charness, N., Gathercole, S. E., Chabris, C. F., Hambrick, D. Z., and Stine-Morrow, E. A. L. (2016). Do "brain-training" programs work? Psychological Science in the Public Interest, 17(3), 103-186. pubmed.ncbi.nlm.nih.gov
  6. Tricot, A., and Sweller, J. (2014). Domain-specific knowledge and why teaching generic skills does not work. Educational Psychology Review, 26, 265-283. link.springer.com
  7. Chase, W. G., and Simon, H. A. (1973). Perception in chess. Cognitive Psychology, 4(1), 55-81. find source โ†—
Key terms
Schema
An organized mental structure that groups many elements of knowledge into a single meaningful unit.
Chunking
Grouping individual elements into larger meaningful units to reduce working-memory load.
Working-memory capacity
The small number of elements that can be held and manipulated in mind at once.
Expert blind spot
An expert's tendency to underestimate the difficulty a novice faces because the expert has schemas the novice lacks.
Schema construction
The instructional goal of helping learners build organized long-term knowledge structures.
Automaticity
The state in which retrieving or using a schema requires little or no conscious effort.

How Long-Term Memory Is Organized and Forgetting

  • Distinguish declarative and procedural long-term memory.
  • Explain forgetting as a failure of retrieval and the role of cues.
  • Describe how prior knowledge shapes what is remembered.

Long-term memory is not a single undifferentiated warehouse. Cognitive psychology distinguishes several kinds, and the two most useful for teaching are declarative and procedural memory. Declarative memory holds knowledge you can consciously state - facts (that Paris is the capital of France) and events you have experienced.

Procedural memory holds skills you perform, often without being able to articulate them: riding a bicycle, decoding print, executing a well-practiced algorithm. The two are built differently. Declarative knowledge can sometimes be acquired quickly through meaningful connection; procedural fluency is built slowly through practice until it becomes automatic. A student can know every rule of grammar declaratively and still write haltingly, because fluent writing is procedural.

Key idea: Long-term memory is not one store but several, and declarative knowledge and procedural fluency are built by different means and at different speeds.

Why we believe the systems are separate

The strongest evidence comes from cases where one system fails and the other survives. The patient known as H.M., described by William Scoville and Brenda Milner in 1957, underwent bilateral removal of medial temporal lobe tissue including most of the hippocampus and afterwards could form almost no new declarative memories: he would meet the same researcher repeatedly without recognition. Yet when he practised mirror drawing across days, his performance improved steadily - while he denied ever having done the task before. Skill was accumulating in a system that his damaged declarative memory could not report on.

Larry Squire's synthesis organizes these findings into a taxonomy that is worth carrying. Declarative memory subdivides into episodic memory for personally experienced events and semantic memory for general facts, both dependent on medial temporal structures. Nondeclarative memory covers procedural skills, priming, simple conditioning, and habituation, and it depends on other circuits including the basal ganglia and cerebellum. For a designer the operational lesson is blunt: telling someone how to do something adds to a different system than having them do it, and neither substitutes for the other. A lecture on pipetting technique and an hour at the bench are not interchangeable, and a student who can recite the procedure has not thereby practised it.

Key idea: Dissociations in amnesic patients show declarative and procedural memory depend on different systems, so explanation and practice build different things and cannot substitute for each other.

Memory as a connected network

Declarative knowledge is best pictured as a network of concepts linked by meaningful relationships, not a list of isolated items. When you learn that a whale is a mammal, you connect "whale" to a rich existing structure - warm-blooded, breathes air, feeds young with milk - and that web of connections is what makes the fact both memorable and useful. This is why meaningful material is far easier to remember than arbitrary material: it hooks into what you already know. It is also why the best single predictor of how easily someone learns new content is how much relevant prior knowledge they already have.

Key idea: Declarative knowledge is a network of meaningful connections, which is why meaningful material is easier to retain and why relevant prior knowledge predicts new learning better than almost any other variable.

The shape of forgetting

Hermann Ebbinghaus, experimenting on himself with nonsense syllables in the 1880s, produced the first quantitative forgetting curve: retention drops steeply within the first hours and days after learning, then flattens, so that the material still held after a week is lost far more slowly than the material lost on day one. The curve is well approximated by a decelerating function rather than a straight line, and the shape has been replicated many times, including in a careful modern replication of Ebbinghaus's own procedure.

But nonsense syllables are the worst case. Harry Bahrick tracked people who had studied Spanish in school and found that after an initial decline over the first three to six years, retention of vocabulary and grammar stabilized for decades among those who never used the language again - a plateau he called permastore. Meaningful, well-learned material behaves very differently from arbitrary material. The practical reading of the two findings together is that forgetting is steep and fast for weakly encoded, arbitrary content and shallow and slow for well-connected content that was learned to a high standard in the first place.

Key idea: Forgetting is steepest immediately after learning and flattens over time, and well-learned meaningful material can survive decades without use.

Forgetting is mostly a retrieval problem

Why do we forget? The intuitive picture - memories fading and crumbling like old photographs - is largely wrong for long-term memory. A more accurate and more useful view is that much forgetting is a failure of retrieval: the information is still there, but the path to it is weak or blocked. Evidence for this is everywhere.

A forgotten name suddenly surfaces given the right cue (its first letter, where you met the person). Information "lost" for years returns intact in the right context. The practical upshot is optimistic: if forgetting is often a retrieval-access problem, then the fix is to strengthen retrieval routes through practice - which is precisely what the strategies in Module 3 do.

Storage strength and retrieval strength

Robert and Elizabeth Bjork's new theory of disuse gives the retrieval account a precise form that is worth learning, because the rest of this course depends on it. They distinguish two independent properties of any memory. Storage strength reflects how thoroughly the item is learned and interconnected; on their account it only ever grows. Retrieval strength reflects how accessible the item is right now, and it rises with recent study or use and falls with time and competition from other memories.

Three consequences follow, and each one explains a phenomenon teachers see constantly. First, current performance measures retrieval strength, not storage strength, so a strong quiz score right after teaching says little about durable learning. Second, gains in storage strength are greatest when retrieval strength has been allowed to fall - which is exactly why spacing works, and why the review that feels effortful is the one that pays. Third, forgetting in this sense is functional rather than pathological: as retrieval strength drops, the opportunity for a big learning gain rises. Bjork and Bjork call this forgetting the friend of learning, and it reframes the classroom implication of the forgetting curve. The right response to forgetting is not to teach faster before it happens, but to schedule retrieval after it has begun.

Key idea: Storage strength and retrieval strength are independent, current performance reflects only retrieval strength, and letting retrieval strength decline before restudy produces the largest durable gains.

Retrieval cues and context

A retrieval cue is any stimulus that helps you access a stored memory. Memories are easier to retrieve when the cues present at recall match those present during encoding - a principle known as encoding specificity. This has real classroom consequences. If students only ever encounter a concept phrased one way, in one context, they may fail to retrieve it when it appears differently on a test or in real life. Teaching a concept in varied contexts builds multiple retrieval routes, so the knowledge is accessible from more directions.

Two classic studies anchor the principle. Godden and Baddeley had divers learn word lists on land or twenty feet underwater and recall them in the same or the other setting; free recall was substantially better when the environments matched. A related idea, transfer-appropriate processing, was demonstrated by Morris, Bransford and Franks: whether semantic or rhyme-based study was superior depended on whether the final test asked about meaning or about rhyme. Deep processing is not universally best; processing that matches the eventual demand is.

Two cautions keep this from being overapplied. Environmental context effects are reliable in free recall but small or absent in recognition, where the test item is itself a powerful cue, so students do not need to sit their exam in the room where they studied. And the design implication runs the opposite way from what people expect: rather than making study conditions match the test, vary them. Varied contexts build several independent routes to the same knowledge, which is what real-world transfer requires, since the situations that will demand the knowledge are not known in advance.

Key idea: Retrieval is easier when conditions at test match those at encoding, but because future demands are unknown, varying study contexts to build multiple routes serves transfer better than reproducing one context.

Prior knowledge shapes memory itself

Finally, remembering is reconstructive, not a literal playback. We rebuild memories using our existing schemas, which is why prior knowledge does not merely help us learn faster - it shapes what we encode and recall, sometimes distorting it toward what we expected. Two students in the same lecture can walk away with genuinely different memories because they filtered it through different schemas. This makes eliciting and building on accurate prior knowledge, the subject of Module 4, central rather than optional.

The experimental record on distortion is unusually vivid. Frederic Bartlett's 1932 studies had British participants recall an unfamiliar Native American folk tale, "The War of the Ghosts," repeatedly over months; their retellings progressively shortened, lost unfamiliar elements, and drifted toward the conventions of English storytelling - reconstruction toward a schema. In the Deese-Roediger-McDermott procedure, participants who study a list of associates such as bed, rest, awake, tired, dream frequently and confidently recall the unpresented word sleep. Elizabeth Loftus's misinformation experiments showed that a leading question asked after an event can alter what witnesses subsequently report seeing.

For a teacher the moral is not that memory is unreliable in general but that confidence is a poor guide to accuracy. A student who insists they were taught something incorrectly may be reporting a genuine, sincerely held reconstruction. This is one more argument for frequent low-stakes checks: the only way to discover what a learner's schema has actually built is to ask them to produce it, which is exactly what Module 3's retrieval practice does and what Module 6's formative assessment institutionalizes.

Common misconceptions

  • Memory works like a video recording. Recall reconstructs the past through current schemas, and the reconstruction can be confidently wrong.
  • Forgetting means the information is gone. Much of it is inaccessible rather than erased, which is why cues, context, and practice can restore it.
  • A single well-taught lesson secures the knowledge. Retrieval strength decays; only repeated retrieval over time raises storage strength.
  • Students should study in the room where they will be tested. Context effects are small in recognition, and varying study contexts serves transfer better.
  • Knowing the steps means being able to perform them. Declarative and procedural knowledge are built in different systems and by different means.

Recap

  • Long-term memory divides into declarative systems (episodic and semantic) and nondeclarative systems including procedural skill, as amnesic dissociations demonstrate.
  • Declarative knowledge is a network, so meaningful, well-connected material is far more durable than arbitrary material.
  • Forgetting is steep at first and then flattens, and well-learned meaningful content can survive decades.
  • Storage strength and retrieval strength are independent; performance now reflects only the latter.
  • Retrieval depends on cues, and varied encoding contexts build more routes back to the knowledge.
  • Recall is reconstructive, so learners can be confidently wrong and only production reveals what they hold.

Sources

  1. Squire, L. R. (2004). Memory systems of the brain: A brief history and current perspective. Neurobiology of Learning and Memory, 82(3), 171-177. pubmed.ncbi.nlm.nih.gov
  2. Ebbinghaus, H. (1913). Memory: A Contribution to Experimental Psychology (H. A. Ruger and C. E. Bussenius, Trans.; original work published 1885). Classics in the History of Psychology. psychclassics.yorku.ca
  3. Bahrick, H. P. (1984). Semantic memory content in permastore: Fifty years of memory for Spanish learned in school. Journal of Experimental Psychology: General, 113(1), 1-29. pubmed.ncbi.nlm.nih.gov
  4. Bjork, R. A., and Bjork, E. L. (2019). Forgetting as the friend of learning: Implications for teaching and self-regulated learning. Advances in Physiology Education, 43(2), 164-167. pubmed.ncbi.nlm.nih.gov
  5. Dudukovic, N., and Kuhl, B. (n.d.). Forgetting and amnesia. In R. Biswas-Diener and E. Diener (Eds.), Noba Textbook Series: Psychology. DEF Publishers. nobaproject.com
  6. McDermott, K. B., and Roediger, H. L. (n.d.). Memory (encoding, storage, retrieval). In R. Biswas-Diener and E. Diener (Eds.), Noba Textbook Series: Psychology. DEF Publishers. nobaproject.com
  7. OpenStax. (2020). Problems with memory. In Psychology 2e, Section 8.3. Rice University. CC BY 4.0. openstax.org
Key terms
Declarative memory
Memory for consciously stateable facts and events.
Procedural memory
Memory for skills and procedures performed, often without conscious articulation.
Retrieval cue
A stimulus that helps access a stored memory.
Encoding specificity
The principle that retrieval is easier when cues at recall match those present during encoding.
Reconstructive memory
The idea that recall rebuilds a memory using existing schemas rather than replaying it verbatim.
Prior knowledge
Relevant knowledge a learner already holds, the strongest predictor of new learning.

Module 2: Cognitive Load Theory

How to design instruction that fits within working-memory limits by managing three kinds of load.

Intrinsic, Extraneous, and Germane Load

  • Define the three types of cognitive load.
  • Distinguish load that is inherent to the material from load imposed by poor design.
  • Explain why reducing extraneous load is the designer's first job.

Cognitive load theory, developed within educational psychology, takes the fixed limits of working memory from Module 1 and turns them into a design discipline. Its central claim is simple: because working memory can process only a few elements at once, instruction must be designed so that the total demands placed on it do not exceed that capacity. When they do, learning stalls - not because the student is incapable, but because the channel is flooded. The theory distinguishes three sources of load, and telling them apart is the core diagnostic skill of instructional design.

Key idea: Cognitive load theory converts the fixed limits of working memory into a design discipline whose first task is telling apart the different sources of demand on that limit.

Where the theory came from

John Sweller arrived at the theory by way of a puzzle. In experiments during the 1980s, learners who successfully solved algebra and geometry problems often learned surprisingly little from doing so. Analysing their strategies, Sweller found that novices attack unfamiliar problems using means-ends analysis: hold the goal in mind, hold the current state in mind, find a difference, find an operator that reduces it, and repeat. That procedure is an effective way to reach an answer and a terrible way to learn, because it consumes the whole workspace with goal management and leaves nothing free for noticing the structure of the solution.

The prediction that follows is counterintuitive and was confirmed: give learners a goal-free version of the same problem ("calculate as many quantities as you can") and they learn more, because removing the specific goal removes the means-ends machinery. That result - problem solving success and learning coming apart - is the origin of cognitive load theory and remains its clearest demonstration that performance during instruction and learning from instruction are different things.

Key idea: The theory began with the finding that solving problems by means-ends analysis consumes working memory without building schemas, so success during practice and learning from practice can diverge.

The three loads

  • Intrinsic load is the inherent difficulty of the material itself, given the learner's current knowledge. It depends on element interactivity - how many pieces must be held in mind simultaneously because they interact. Learning isolated vocabulary is low in interactivity (each word is independent); learning to balance a chemical equation is high (many interacting constraints at once). Intrinsic load is real and cannot simply be wished away, though it can be managed.
  • Extraneous load is the load imposed by the way information is presented, not by the content. A confusing diagram, a cluttered slide, instructions split across two pages, or a search-heavy activity all consume working memory without contributing to learning. This is wasted load, and it is entirely the designer's responsibility.
  • Germane load is the productive mental effort devoted to building and automating schemas - the effort that actually constitutes learning. The goal is not to minimize all effort but to free up capacity so that more of it can be germane.

Element interactivity, counted

Intrinsic load is not a vague sense of difficulty; Sweller defines it through element interactivity, the number of elements that must be processed simultaneously because understanding any one depends on the others. The count is what makes the definition useful, and it always depends on the learner's schemas.

Compare two French tasks. Learning that chien means dog is a low-interactivity item: it can be learned entirely on its own, and learning fifty such pairs is fifty independent low-load acts. Learning to construct the French subjunctive is high-interactivity: the learner must simultaneously hold the trigger expression, the person and number, the stem, the ending, and the meaning being expressed, and getting any one wrong breaks the whole. Fifty vocabulary items may take longer in total but never overload; one subjunctive rule can overload immediately. The practical diagnostic question is therefore not "is this topic hard?" but "how many things must be true in the learner's head at once for this to make sense?"

The same count explains why prerequisites are not bureaucratic. If a learner has already automated the present-tense stem, that stem is one chunk rather than several, and the interactivity of the subjunctive task falls accordingly. Sequencing content so that today's interacting elements were yesterday's automated chunks is the single most reliable way to manage intrinsic load without diluting the content.

Key idea: Intrinsic load is set by element interactivity - how many pieces must be held simultaneously - and it falls as prerequisite elements become automated chunks.

They share one budget

The crucial insight is that all three loads draw on the same limited working memory. They add up. If intrinsic load is high and extraneous load is also high, their sum can exceed capacity, and there is no room left for the germane processing that builds understanding. This yields the designer's first commandment: ruthlessly reduce extraneous load, because every unit of working memory wasted on poor presentation is a unit stolen from learning. You usually cannot lower a topic's intrinsic difficulty, but you can almost always clean up the presentation.

Load typeSourceDesigner's move
IntrinsicInherent complexity of the contentSequence and segment it; build prerequisite schemas first
ExtraneousPoor presentation and designEliminate it - this is the top priority
GermaneEffortful schema buildingProtect and encourage it by freeing capacity

Key idea: The three loads draw on one shared budget, so eliminating extraneous load is the designer's first duty because it is the only component that is pure waste.

The contested third category

Graduate readers should know that germane load is the least settled part of the theory. Ton de Jong's widely cited critique argued that the three categories are difficult to distinguish empirically, that germane load in particular risks circularity - it is often inferred from the learning it is supposed to explain - and that the theory's measurement tools are too blunt to separate the components in practice. If any effortful processing that produces learning is germane by definition, the category cannot be falsified.

Sweller responded by substantially revising the construct. In his later formulation germane load is not a third independent source that adds to the total; it is the allocation of working-memory resources to dealing with intrinsic load - that is, to the element interactivity that constitutes the learning task. On this account there are two sources of load, intrinsic and extraneous, and germane processing describes where the available capacity is spent rather than adding to what is spent. Slava Kalyuga has argued for the same simplification on independent grounds. Newer work extends the framework to digital environments, where Skulmowski and Xu argue that the sheer act of navigating an interface adds a form of extraneous load that older accounts did not anticipate.

The practical guidance is unchanged by the dispute, which is why it is safe to teach the three-category version first: cut extraneous demands, manage intrinsic demands by sequencing and segmenting, and make sure the capacity you free is actually spent on schema construction rather than on nothing at all. What changes is your confidence in claims that a design "increased germane load" - that is a hypothesis about where capacity went, not a measurement.

Key idea: Whether germane load is a third additive source or simply the allocation of capacity to intrinsic processing is genuinely contested, and claims that a design raised germane load are hypotheses rather than measurements.

How load is measured

Because load is invisible, measurement is the theory's weakest joint, and knowing the options keeps you honest about what any study has shown. Three families are in use. Subjective ratings, most often Fred Paas's nine-point scale of invested mental effort administered right after a task, are cheap, surprisingly reliable, and used in the majority of published studies, but they conflate the sources of load. Dual-task methods ask learners to respond to a secondary probe, such as a colour change, during the primary task, and use their reaction time as an index of spare capacity; these separate load from performance more cleanly at the cost of intruding on the task. Physiological measures including pupil dilation, eye fixation patterns, and heart-rate variability are increasingly used and are sensitive, but they index arousal and effort in general rather than any particular category of load.

The upshot for reading the literature is that a study reporting "lower cognitive load" almost always means a lower rating on a self-report scale. That is real evidence, but it does not by itself tell you which component fell.

Key idea: Load is usually measured by self-reported mental effort, sometimes by secondary-task probes or physiological indices, and none of these cleanly separates intrinsic from extraneous from germane processing.

Load depends on the learner

One subtlety ties this back to Module 1: intrinsic load is not fixed by the topic alone but by the topic relative to the learner's schemas. The same equation is high-load for a novice and low-load for an expert. This means a design that is well-calibrated for a beginner can actually harm an expert by belaboring what they already chunk automatically - an effect explored in the next lessons. Cognitive load, in other words, is always a relationship between material and mind, never a property of the material by itself.

What the theory does not claim

One boundary condition is worth stating because it prevents a common overreach. David Geary distinguishes biologically primary knowledge, which humans evolved to acquire effortlessly through immersion - speaking a first language, recognizing faces, basic social cognition - from biologically secondary knowledge, which is culturally invented and acquired only through deliberate effort, such as reading, algebra, or chemical notation. Cognitive load theory is a theory of secondary knowledge. Nobody needs worked examples to learn to speak their mother tongue, and the fact that immersion works there is not evidence that immersion works for calculus.

Two further clarifications. Reducing extraneous load is not the same as reducing challenge; the whole point is to spend the freed capacity on genuinely difficult thinking. And low load is not automatically good - a task so undemanding that it requires no schema construction produces no learning, which is why Module 3's desirable difficulties are not a contradiction of this module but its complement.

Common misconceptions

  • Cognitive load theory says make everything easy. It says remove waste so that capacity is available for the hard part.
  • Germane load is a measurable third quantity. Its status is contested, and in Sweller's later formulation it describes where capacity is spent rather than an additional demand.
  • Intrinsic load is fixed by the topic. It is fixed by element interactivity relative to the learner's schemas, so prerequisites genuinely lower it.
  • Success on practice problems shows learning. Means-ends problem solving can produce correct answers with almost no schema construction.
  • The theory explains all learning. It addresses biologically secondary, culturally invented knowledge, not the primary abilities acquired through immersion.

Recap

  • Cognitive load theory turns working-memory limits into design rules and began with the finding that solving problems can teach very little.
  • Intrinsic load is set by element interactivity relative to the learner; extraneous load comes from presentation; germane processing is the effort that builds schemas.
  • All demands draw on one budget, so cutting extraneous load is the first and most reliable move.
  • The status of germane load as a separate additive category is genuinely contested, and measurement of load is imperfect.
  • Load is a relationship between material and learner, so no single design is optimal for everyone.
  • The theory concerns biologically secondary knowledge and does not claim that easier is always better.

Sources

  1. Sweller, J., van Merrienboer, J. J. G., and Paas, F. (2019). Cognitive architecture and instructional design: 20 years later. Educational Psychology Review, 31, 261-292. link.springer.com
  2. Sweller, J. (2010). Element interactivity and intrinsic, extraneous, and germane cognitive load. Educational Psychology Review, 22, 123-138. link.springer.com
  3. de Jong, T. (2010). Cognitive load theory, educational research, and instructional design: Some food for thought. Instructional Science, 38, 105-134. link.springer.com
  4. van Merrienboer, J. J. G., and Sweller, J. (2005). Cognitive load theory and complex learning: Recent developments and future directions. Educational Psychology Review, 17, 147-177. link.springer.com
  5. Skulmowski, A., and Xu, K. M. (2022). Understanding cognitive load in digital and online learning: A new perspective on extraneous cognitive load. Educational Psychology Review, 34, 171-196. link.springer.com
  6. Sweller, J. (1988). Cognitive load during problem solving: Effects on learning. Cognitive Science, 12(2), 257-285. find source โ†—
  7. Paas, F., and van Merrienboer, J. J. G. (2020). Cognitive-load theory: Methods to manage working memory load in the learning of complex tasks. Current Directions in Psychological Science, 29(4), 394-398. find source โ†—
Key terms
Cognitive load theory
A theory holding that instruction must fit within working-memory limits by managing total load.
Intrinsic load
The inherent difficulty of material, driven by how many elements interact, relative to the learner.
Extraneous load
Working-memory demand caused by poor presentation rather than by the content itself.
Germane load
Productive effort devoted to building and automating schemas.
Element interactivity
The number of interacting pieces that must be held in mind at once to understand something.
Working-memory budget
The shared, limited capacity that all three load types draw upon.

Cognitive Load Effects: Worked Examples, Split Attention, and Redundancy

  • Apply the worked-example and completion effects for novices.
  • Fix split-attention and redundancy problems in materials.
  • Explain the expertise-reversal effect and its design consequence.

Cognitive load theory is valued because it yields concrete, testable design effects, each replicated many times in experimental studies. This lesson covers the four most useful.

Each of these effects has the same logical form: a design choice that looks helpful, or at least harmless, turns out to change what learners can hold in mind, and therefore what they learn. Each was established by experiments that held content constant and varied only the presentation, which is what makes them design evidence rather than opinion. And each comes with boundary conditions, because an effect that helps one learner at one stage of expertise can hurt another.

Key idea: The cognitive load effects are replicated experimental findings in which changing only the presentation of identical content changes how much is learned, and each carries boundary conditions.

The worked-example effect

For novices learning a procedure, studying worked examples - fully solved problems with each step shown - typically produces better and faster learning than solving the equivalent problems unaided. The reason is load. A novice thrown at a hard problem often flails through weak, high-load search strategies (means-ends analysis) that consume working memory without building a schema.

A worked example lets them devote that capacity to understanding the solution's structure, which is what forms the schema. The practical pattern is to start with worked examples and gradually fade the support - first give complete examples, then completion problems (partly worked, with steps for the learner to finish), then full problems - as competence grows. This gradual withdrawal of support is sometimes called scaffolding and fading.

The founding experiment is worth knowing precisely. Sweller and Cooper, in 1985, gave secondary students algebra practice in one of two formats: conventional problems, or alternating worked example and problem pairs covering the same content. The worked-example group spent less time in acquisition and then outperformed the conventional group on similar problems afterwards. Less time, better learning - the reverse of what a simple practice-makes-perfect account predicts.

Two boundary conditions have emerged since, and both matter in practice. First, a worked example only helps if the learner processes it. Studying an example passively produces little; prompting learners to self-explain each step, or asking them to identify which principle justifies a line, converts viewing into schema construction. Second, the examples must be structured so the sub-goals are visible - a solution written as an undifferentiated run of algebra teaches less than the same solution grouped into labelled steps. Van Gog and Rummel add a further nuance from the social-cognitive side: worked examples can be delivered as static products or as observed modelling of the process, and models who visibly struggle and correct themselves can be more useful to some learners than flawless demonstrations, because they show how errors are detected.

Key idea: Worked examples outperform equivalent unaided practice for novices, but only when learners actively process them and the solution structure makes its sub-goals visible.

Fading the support: completion problems and backward fading

The natural question is when to stop giving examples, and the answer is not "all at once." The completion strategy, developed by van Merrienboer, replaces the abrupt jump from full example to full problem with a graded sequence in which progressively more of the solution is left blank. Backward fading, studied by Renkl and Atkinson, removes the last step first, then the last two, and so on, so the learner always finishes what they have just seen modelled - which keeps the load low at exactly the moment they are working out something new.

A concrete sequence for teaching titration calculations might run: a fully worked example with each step labelled; a second example with the final unit conversion left blank; a third with the last two steps blank; a fourth giving only the balanced equation; then an unaided problem. Each step introduces one new demand rather than all of them at once, and the learner's success rate stays high enough that the practice remains informative rather than demoralizing.

Key idea: Support should be withdrawn gradually through completion problems and backward fading, so that each step adds one new demand rather than transferring the entire task at once.

The split-attention effect

When understanding requires integrating two sources that cannot be understood in isolation - a diagram and a separate block of text explaining it - forcing the learner to hold one while searching for the other imposes needless extraneous load. This is the split-attention effect. The fix is physical integration: place each label directly on the diagram, put the explanation next to the step it explains, so the eye does not have to bridge a gap and working memory does not have to hold a placeholder. Any time a learner must mentally connect two separated but mutually dependent things, suspect split attention.

Noah Schroeder and Ada Cenkci's meta-analysis of spatial contiguity and split-attention studies found a consistent advantage for integrated over separated formats across a wide range of materials and age groups, and reported that the benefit tends to be larger when the material is complex and when learners have less prior knowledge - exactly the pattern the theory predicts, since integration matters most when the workspace is already close to full. Split attention also has a temporal form: narration that describes a step which appeared on screen ten seconds earlier forces the learner to hold the visual in memory while the words arrive, which is why simultaneous presentation of corresponding words and pictures outperforms successive presentation.

A common classroom instance is the laboratory handout that lists numbered procedures on page one and shows the apparatus diagram on page two. Students flip back and forth, holding "step 4" in mind while hunting for the stopcock. Reprinting the diagram with the relevant step printed beside each labelled part costs nothing and removes the shuttling entirely.

Key idea: Integrated formats reliably beat separated ones, with the largest benefits for complex material and low-prior-knowledge learners, and the principle applies to timing as well as to physical layout.

The redundancy effect

It is tempting to think more explanation is always safer, but the redundancy effect shows the opposite: presenting the same information in multiple simultaneous forms can hurt learning. The classic case is on-screen text that duplicates, word for word, what a narrator is saying. Learners cannot help trying to reconcile the two streams, which wastes capacity. If a diagram is self-explanatory, adding a redundant paragraph describing it can lower performance. The lesson: eliminate genuinely redundant information rather than piling it on. (Note the contrast with split attention, where the two sources are not redundant but are each necessary and should be integrated.)

Paul Chandler and Sweller established the effect in a series of experiments where an integrated diagram-and-text format beat a separated one, but a diagram alone beat the integrated format whenever the diagram was intelligible without the text. The decision rule follows directly: if the two sources are each unintelligible alone, integrate them; if either is self-sufficient, delete the other. Designers get into trouble by applying the split-attention fix to genuinely redundant material, producing a tidy, well-integrated presentation of information the learner did not need.

Key idea: Integrate two sources when neither is intelligible alone, but delete one when either is self-sufficient, because redundancy and split attention call for opposite remedies.

The expertise-reversal effect

Here is the twist that keeps design honest. The very techniques that help novices - worked examples, extra scaffolding, integrated explanations - can slow down or harm experts. This is the expertise-reversal effect. Once a learner has the relevant schema, guidance they no longer need becomes redundant information that adds extraneous load; for them, solving problems directly is better than studying worked examples.

The consequence is that no design is universally optimal. Support must be faded as expertise grows, and the "right" amount of guidance is a moving target that depends on the learner's current level. This is why adaptive sequencing and ongoing assessment (Module 6) are not luxuries but load-management necessities.

Slava Kalyuga's review makes the mechanism explicit and gives it an experimental signature. The effect appears as a crossover interaction: plot performance against instructional format for novices and for more knowledgeable learners, and the two lines cross. The mechanism is redundancy. Guidance that a learner no longer needs is information they must nonetheless process and reconcile with the schema they already hold, and that reconciliation costs working memory. Julian Roelle and Kirsten Berthold demonstrated the same reversal for prompts that focus processing of instructional explanations: helpful for those without the relevant knowledge, counterproductive for those with it.

Two practical consequences follow. First, any claim that a method "works" is incomplete without saying for whom; a study showing worked examples beat problem solving in a sample of beginners tells you nothing about a class of near-experts. Second, because expertise changes within a unit and not only between courses, you need a cheap, frequent way to tell where learners are - which is what rapid diagnostic tasks and low-stakes formative checks provide.

Key idea: Expertise reversal shows up as a crossover interaction and is caused by redundancy, so guidance must be faded on evidence of the learner's current knowledge rather than on a fixed schedule.

The boundary case: when struggle first is better

A genuine complication deserves naming, because it is often misused. A line of research on productive failure, associated with Manu Kapur, finds that having learners attempt to invent solutions to a problem before receiving instruction can outperform instruction-first sequences, particularly on conceptual understanding and transfer rather than on procedural fluency. Katharina Loibl, Ido Roll and Nikol Rummel's review argues that the benefit depends on specific conditions being met: the initial activity must generate a range of learner solutions, the subsequent instruction must explicitly compare those attempts with the canonical solution, and learners must have enough prior knowledge to generate something worth comparing.

This does not license unguided discovery, and Lesson 11 returns to why. The failed attempts do not teach on their own; they prepare the learner to notice what the subsequent explanation is about. Read carefully, productive failure is a claim about sequencing guidance, not about withholding it.

Key idea: Attempting a problem before instruction can beat instruction-first sequences for conceptual outcomes, but only when the follow-up instruction explicitly contrasts learners' attempts with the correct solution.

Common misconceptions

  • Practice is always the best way to learn a procedure. For novices, studying structured worked examples usually teaches more in less time.
  • Adding an explanation can only help. When either source is self-sufficient, the extra explanation is redundant and can lower performance.
  • Split attention and redundancy call for the same fix. One calls for integration, the other for deletion, and confusing them produces well-organized clutter.
  • Support should be removed when the unit ends. Fading should track evidence of expertise, since guidance a learner has outgrown becomes load.
  • Productive failure vindicates discovery learning. Its benefit depends on explicit instruction afterwards that contrasts learner attempts with the canonical solution.

Recap

  • Worked examples beat equivalent unaided practice for novices, provided learners actively process them and the sub-goal structure is visible.
  • Support should be faded gradually through completion problems and backward fading rather than withdrawn at once.
  • Integrated formats beat separated ones, most strongly for complex material and low-prior-knowledge learners, in space and in time.
  • Redundant duplication harms learning; integrate only sources that are unintelligible alone.
  • Expertise reversal is a crossover interaction driven by redundancy, so guidance must track the learner's current knowledge.
  • Problem solving before instruction can help conceptual learning when followed by instruction that contrasts learner attempts with the correct solution.

Sources

  1. Kalyuga, S. (2007). Expertise reversal effect and its implications for learner-tailored instruction. Educational Psychology Review, 19, 509-539. link.springer.com
  2. Schroeder, N. L., and Cenkci, A. T. (2018). Spatial contiguity and spatial split-attention effects in multimedia learning environments: A meta-analysis. Educational Psychology Review, 30, 679-701. link.springer.com
  3. van Gog, T., and Rummel, N. (2010). Example-based learning: Integrating cognitive and social-cognitive research perspectives. Educational Psychology Review, 22, 155-174. link.springer.com
  4. Roelle, J., and Berthold, K. (2013). The expertise reversal effect in prompting focused processing of instructional explanations. Instructional Science, 41, 635-656. link.springer.com
  5. Loibl, K., Roll, I., and Rummel, N. (2017). Towards a theory of when and how problem solving followed by instruction supports learning. Educational Psychology Review, 29, 693-715. link.springer.com
  6. Sweller, J., van Merrienboer, J. J. G., and Paas, F. (2019). Cognitive architecture and instructional design: 20 years later. Educational Psychology Review, 31, 261-292. link.springer.com
  7. Sweller, J., and Cooper, G. A. (1985). The use of worked examples as a substitute for problem solving in learning algebra. Cognition and Instruction, 2(1), 59-89. find source โ†—
Key terms
Worked example
A fully solved problem, with all steps shown, studied to build a schema with low load.
Completion problem
A partly worked example with remaining steps left for the learner, used to fade support.
Split-attention effect
The extra load caused when learners must integrate two separated but mutually dependent sources.
Physical integration
Placing related information together (labels on a diagram) to remove split attention.
Redundancy effect
The finding that presenting the same information in multiple simultaneous forms can harm learning.
Expertise-reversal effect
The reversal in which support that helps novices hinders experts, requiring guidance to be faded.

Module 3: Evidence-Based Learning Strategies

The study and teaching techniques that research repeatedly shows produce durable, transferable learning.

Spacing and the Testing Effect (Retrieval Practice)

  • Explain the spacing effect and why massed practice underperforms.
  • Explain the testing effect and design retrieval practice.
  • Distinguish learning from the feeling of fluency.

Two of the most robust findings in the entire science of learning are the spacing effect and the testing effect. Both have been replicated across ages, subjects, and decades, and both are widely underused because they feel less effective in the moment than the popular alternatives. Understanding why is the key to teaching them.

Key idea: Spacing and retrieval practice are among the best-established findings in the science of learning, and they are underused precisely because they feel less effective while you are doing them.

The spacing effect

The spacing effect is the finding that learning is more durable when study of the same material is distributed over time rather than massed into a single block. Studying a topic for one hour spread across three days beats studying it for three straight hours, when measured by retention later.

The mechanism is instructive: when you return to material after a delay, you have partly forgotten it, so retrieving and relearning it requires effort - and that effortful reconstruction strengthens the memory far more than smooth, uninterrupted rereading. Cramming (massed practice) can produce good performance on an immediate test and then rapid forgetting; spacing produces slightly worse immediate performance but dramatically better long-term retention.

The evidence base is unusually deep. Nicholas Cepeda and colleagues located 839 assessments of distributed practice across 317 experiments in their 2006 quantitative synthesis, and the benefit of spacing over massing was pervasive. Their more important contribution, though, was showing that the question "how long should the gap be?" has no fixed answer: the interval between study sessions and the interval before the final test operate jointly, so the optimal gap depends on when you need the knowledge.

Their 2008 experiment made this precise. Over 1,300 participants learned obscure facts in two sessions separated by gaps ranging from minutes to months, then were tested after delays of up to a year. The optimal gap grew with the retention interval, tracing what the authors called a temporal ridgeline, and settled at roughly ten to twenty percent of it. To remember something for a week, review it after about a day; to remember it for six months, review it after several weeks. Two practical corollaries follow. Short gaps are better than none, so a designer constrained to a single course should still separate exposures by days rather than minutes. And an interval that is too short costs more than one that is somewhat too long, because the curve falls away gently on the long side.

Alice Latimier, Hugo Peyre and Franck Ramus extended the analysis specifically to spaced retrieval episodes and again found a reliable advantage for distributing practice attempts over time rather than concentrating them, with the caveat that the studies vary widely in how they operationalize the gap.

Key idea: Spacing is one of the most replicated findings in memory research, and the optimal gap scales with how long you need the material to last, at roughly ten to twenty percent of the retention interval.

The testing effect

The testing effect - also called retrieval practice - is the finding that actively retrieving information from memory strengthens learning more than restudying the same information for the same time. A test is not merely a measurement; the act of pulling knowledge out of memory is itself one of the most powerful ways to make it durable. In controlled comparisons, learners who practiced recalling material vastly outperformed those who reread it, especially on delayed tests, even though the rereaders felt more confident. Retrieval works partly by strengthening the retrieval routes discussed in Module 1: every successful recall makes the next recall easier.

Two studies define the shape of the finding. Henry Roediger and Jeffrey Karpicke had students read prose passages and then either restudy them repeatedly or take repeated recall tests on them. Five minutes later, the repeated-study group did better, recalling about 83 percent against 71 percent for the repeated-testing group - exactly the outcome that makes rereading feel productive. One week later the ordering had reversed decisively: about 61 percent for the tested group against 40 percent for the restudy group. The advantage of restudy is real and it is temporary; the advantage of retrieval is smaller at first and grows.

Their 2008 experiment in Science isolated which component does the work. Students learned Swahili-English pairs under four schedules that varied whether learned items were dropped from further study or from further testing. Dropping items from restudy barely mattered; dropping them from testing was catastrophic. The two conditions that kept testing all items produced final recall around 80 percent, while the conditions that stopped testing them fell to roughly a third of that. Repeated retrieval, not repeated exposure, was carrying the learning.

The effect scales up. Chunliang Yang and colleagues synthesized 222 classroom studies covering more than 48,000 students and found that quizzing raised achievement by a medium amount, g = 0.499, with the benefit larger when corrective feedback was given, when the practice test format matched the final test, and when practice was repeated. Christopher Rowland's earlier meta-analysis added a mechanistic finding: recall tests produce larger benefits than recognition tests, which supports effortful retrieval rather than mere re-exposure as the active ingredient, and found only limited support for the idea that the benefit comes from semantic elaboration during retrieval.

Key idea: Retrieval practice loses to restudy on immediate tests and beats it decisively on delayed ones, and classroom syntheses put the benefit at about half a standard deviation, larger with feedback and repetition.

Why retrieval works

No single mechanism is agreed, and a graduate reader should know the main contenders. The retrieval effort account holds that difficult but successful retrievals produce the largest gains, which fits Rowland's finding that recall beats recognition and predicts the interaction with spacing: a longer gap makes retrieval harder and therefore more valuable, up to the point where it fails. The episodic context account proposes that retrieving an item reinstates and updates its contextual trace, giving later searches a fresher and more differentiated target. The bifurcation model explains the crossover between immediate and delayed tests statistically: testing splits items into a successfully retrieved set that is strongly boosted and a failed set that is not, so the average benefit only becomes visible once the weaker items have been forgotten anyway.

These are complementary rather than rival in practice. What they agree on is the design implication: the retrieval must be effortful and it must succeed often enough to be worth attempting, which is why difficulty needs calibrating rather than maximizing.

Key idea: Retrieval effort, episodic context updating, and the bifurcation of items into retrieved and failed sets all contribute, and all imply that retrieval should be effortful but usually successful.

Limits, cautions, and feedback

Four qualifications keep this from becoming a slogan. First, feedback matters. Retrieval without correction can consolidate errors, and Yang's synthesis found the classroom benefit substantially larger when corrective feedback followed. Second, failed retrieval is not automatically wasted, but it is much less useful than successful retrieval, so items that learners cannot yet produce need restudy before they need testing. Third, stakes change the picture: the benefits documented here come overwhelmingly from low-stakes or no-stakes retrieval, and high-stakes testing introduces anxiety effects that can offset the cognitive gain, particularly for students who are already anxious. Fourth, the effect is best documented for factual and conceptual recall; benefits for complex transfer and problem solving exist but are more variable, and depend on the practice questions actually demanding the target reasoning.

Key idea: Retrieval practice needs corrective feedback, should follow rather than replace initial learning, and delivers its documented benefits at low stakes.

The fluency trap

Why are these strategies underused? Because they violate our intuition about what learning feels like. Rereading and cramming produce a comforting sense of fluency - the material looks familiar, so we judge that we know it. But familiarity is not the same as retrievability. Spacing and retrieval practice deliberately introduce difficulty; they feel harder and less productive precisely because they are doing more work.

These are examples of what researchers call desirable difficulties: conditions that slow apparent progress but improve real, lasting learning. The teacher's job is partly to protect learners from their own metacognitive illusions by building spacing and retrieval into the design, since students left to their own devices reliably choose the fluent, less effective methods.

Putting them together

The two combine naturally into spaced retrieval practice: low-stakes quizzing on material, revisited at increasing intervals. Practically, this means frequent short recall quizzes rather than one big review; cumulative questions that revisit older material, not just the latest topic; and treating quizzes as learning events, not just grades. Flashcard systems that schedule reviews by difficulty are one popular implementation. None of this requires new technology - a teacher who opens class by asking students to write down everything they remember from last week is already using both effects.

A concrete twelve-week course schedule makes it tangible. Each session opens with a five-question written recall covering one topic from the previous week, one from three weeks ago, and one from the start of term, followed immediately by the answers with brief explanations. Homework sets always include two problems from earlier units alongside the current ones. A cumulative low-stakes quiz falls in weeks four, eight, and twelve, each weighted lightly enough that students treat it as practice. Nothing here adds content or removes teaching time; it redistributes review that was going to happen anyway, and it converts review from rereading into retrieval.

Note what the design protects against. Left to themselves, students overwhelmingly reread and highlight - the two techniques Dunlosky and colleagues rated as having low utility in their review of ten common study strategies, in the same paper that rated practice testing and distributed practice as high utility. Building spacing and retrieval into the course structure means learners get the benefit whether or not they have been persuaded of it.

Key idea: Spaced retrieval is implemented by redistributing review that already exists, using cumulative low-stakes recall with immediate feedback rather than adding content.

Common misconceptions

  • Tests only measure learning. Retrieval is itself among the most powerful learning events available.
  • Cramming does not work. It works well for an imminent test and poorly for anything later, which is why students keep doing it.
  • Longer gaps are always better. The optimal gap scales with the retention interval; too long a gap yields failed retrievals with little gain.
  • Any quizzing will do. Feedback, repetition, and matching the practice format to the target task substantially change the size of the benefit.
  • Feeling fluent means being ready. Fluency tracks recent exposure, not retrievability, and is the main reason effective strategies are rejected.

Recap

  • Distributed practice beats massed practice for retention, and the optimal gap is roughly ten to twenty percent of the interval before you need the knowledge.
  • Retrieval practice loses to rereading on immediate tests and wins decisively on delayed ones, with classroom effects near g = 0.5.
  • Repeated retrieval, not repeated exposure, carries the benefit, as the Swahili word-pair experiment showed.
  • Retrieval effort, context updating, and item bifurcation all plausibly contribute, and all imply calibrated rather than maximal difficulty.
  • Corrective feedback, low stakes, and initial learning before testing are the practical conditions for the effect.
  • The fluency illusion explains why learners reliably choose the weaker methods, so the schedule must be designed rather than left to preference.

Sources

  1. Cepeda, N. J., Pashler, H., Vul, E., Wixted, J. T., and Rohrer, D. (2006). Distributed practice in verbal recall tasks: A review and quantitative synthesis. Psychological Bulletin, 132(3), 354-380. pubmed.ncbi.nlm.nih.gov
  2. Cepeda, N. J., Vul, E., Rohrer, D., Wixted, J. T., and Pashler, H. (2008). Spacing effects in learning: A temporal ridgeline of optimal retention. Psychological Science, 19(11), 1095-1102. pubmed.ncbi.nlm.nih.gov
  3. Roediger, H. L., and Karpicke, J. D. (2006). Test-enhanced learning: Taking memory tests improves long-term retention. Psychological Science, 17(3), 249-255. pubmed.ncbi.nlm.nih.gov
  4. Karpicke, J. D., and Roediger, H. L. (2008). The critical importance of retrieval for learning. Science, 319(5865), 966-968. pubmed.ncbi.nlm.nih.gov
  5. Rowland, C. A. (2014). The effect of testing versus restudy on retention: A meta-analytic review of the testing effect. Psychological Bulletin, 140(6), 1432-1463. pubmed.ncbi.nlm.nih.gov
  6. Yang, C., Luo, L., Vadillo, M. A., Yu, R., and Shanks, D. R. (2021). Testing (quizzing) boosts classroom learning: A systematic and meta-analytic review. Psychological Bulletin, 147(4), 399-435. pubmed.ncbi.nlm.nih.gov
  7. Dunlosky, J., Rawson, K. A., Marsh, E. J., Nathan, M. J., and Willingham, D. T. (2013). Improving students' learning with effective learning techniques: Promising directions from cognitive and educational psychology. Psychological Science in the Public Interest, 14(1), 4-58. pubmed.ncbi.nlm.nih.gov
Key terms
Spacing effect
Distributed practice over time produces more durable learning than the same time massed together.
Massed practice
Concentrating study into a single block (cramming); good for immediate tests, poor for retention.
Testing effect
Retrieving information from memory strengthens it more than restudying for the same time.
Retrieval practice
Deliberately recalling information as a study method, another name for using the testing effect.
Fluency illusion
Mistaking the familiarity produced by rereading for genuine ability to retrieve.
Desirable difficulty
A condition that slows apparent progress but improves long-term learning.

Interleaving and Elaboration

  • Explain interleaving and contrast it with blocked practice.
  • Apply elaborative interrogation and self-explanation.
  • Explain how these strategies support discrimination and transfer.

Spacing and retrieval govern when and how you practice; this lesson adds two strategies about how the practice is arranged and how deeply you process: interleaving and elaboration.

Interleaving

Interleaving means mixing different but related problem types or topics within a practice session, rather than doing all of one type before moving to the next (blocked practice). A blocked math set does twenty problems of type A, then twenty of type B; an interleaved set shuffles A, B, C, D. Blocking feels smoother and produces better performance during practice, which is exactly why it is so popular and so misleading. Interleaving usually produces worse practice-session performance but better performance on a later mixed test.

The reason is discrimination. In a blocked set, you already know every problem is type A, so you never practice the crucial real-world skill of figuring out which kind of problem you are facing and thus which method to apply. Interleaving forces that judgment on every item. This is why interleaving especially helps in domains where the challenge is choosing the right approach - distinguishing similar species, categories of art, or which formula a word problem calls for. It is a close cousin of spacing (mixing topics naturally spaces each one) and another desirable difficulty.

The classroom evidence is direct. Doug Rohrer and Kelli Taylor gave college students practice at computing the volumes of four obscure solids, either blocked by solid type or shuffled. During practice the blocked group was more accurate; on a test a week later the interleaved group was far more accurate. In a later classroom study, Rohrer, Robert Dedrick and Sandra Stershic randomly assigned seventh-grade mathematics classes to interleaved or blocked versions of the same practice assignments across a term and found a substantial interleaving advantage on an unannounced test a month afterwards. The practice sets contained identical problems; only the order differed.

Key idea: Interleaving lowers performance during practice and raises it on delayed tests, and classroom trials that hold the problems constant and vary only their order reproduce the effect.

When interleaving helps, and when it does not

Interleaving is a genuine finding with genuine limits, and a graduate reader should hold both. Matthias Brunmair and Tobias Richter's meta-analysis of 59 studies found a moderate overall effect, Hedges' g = 0.42, but the moderators are as informative as the average. The effect was largest for visual category learning such as identifying painters' styles (g = 0.67), moderate for mathematics tasks (g = 0.34), ambiguous and nonsignificant for expository texts, and actually negative for word learning (g = -0.39), where blocking was better.

The pattern is not random. Interleaving helped most when the categories being learned were similar to each other and varied within themselves, and least when they were already easy to tell apart. This supports an attentional account: juxtaposing items from different categories draws attention to the features that distinguish them, while blocking draws attention to the features that items within a category share. Which of those two jobs your learners need determines whether to interleave.

So the design rule is conditional rather than universal. Interleave when the difficulty is telling confusable things apart - similar clinical presentations, statistical tests with overlapping assumptions, problem types that look alike but need different methods. Block when the difficulty is extracting the common structure of a single new category, and when items are already obviously distinct, as with unrelated vocabulary pairs. Note also that interleaving and spacing are partly confounded: mixing topics automatically spaces each one, so some of the reported benefit is spacing under another name, which is a reason to be cautious about attributing everything to discrimination.

Key idea: Interleaving helps most when categories are confusable and can hurt when they are already distinct, so it should be applied where discrimination is the bottleneck rather than as a blanket rule.

Elaboration

Elaboration means connecting new information to what you already know and to itself, processing it for meaning rather than surface. From Module 1 we know that meaningful, well-connected material is far more memorable, so elaboration is retrieval-route construction by design. Two well-studied techniques implement it:

  • Elaborative interrogation: asking and answering why a stated fact is true. Rather than memorizing that a particular adaptation exists, the learner asks "why would that adaptation help this animal survive?" and reasons out the connection, which anchors the fact in a causal web.
  • Self-explanation: explaining a process, a solution step, or a new idea to yourself in your own words, including why each step follows. Learners who self-explain while studying worked examples understand more deeply and transfer better, because they are actively building the schema rather than passively reading it.

Both work by directing capacity toward schema building rather than toward surface processing, and the connections they create give new knowledge more handles for later retrieval.

The evidence differs in strength between the two. Kiran Bisra and colleagues meta-analysed 69 experiments on prompted self-explanation and found a moderate overall benefit, g = 0.55, across subjects from mathematics to biology and across age groups, with the benefit appearing whether prompts were open-ended or structured. Elaborative interrogation has a thinner base: Dunlosky and colleagues rated it moderate rather than high utility, noting that most studies use factual material, that the technique depends on learners having enough prior knowledge to generate a plausible explanation, and that its effects on complex material are not well established. Gregory Donoghue and John Hattie's meta-analysis of ten learning techniques reaches similar conclusions, and adds the useful caution that the reported effects for most techniques vary enormously across studies, so the ordering of techniques is less stable than league tables imply.

Key idea: Prompted self-explanation has a solid moderate evidence base at about g = 0.55, while elaborative interrogation is promising but less well established, particularly for complex material.

Prompting elaboration well

Elaboration prompts fail in predictable ways, and the failures are instructive. A prompt that can be answered by copying a sentence from the text ("what does the passage say about osmosis?") produces no elaboration. A prompt that asks for justification the learner cannot possibly supply ("why did evolution select this mechanism?") produces guessing. The productive middle asks for a connection the learner has the material to make: "why would a cell in a hypotonic solution swell, given what you know about water movement?" or, during a worked example, "why is this step legitimate here when it was not legitimate in the previous problem?"

Two implementation details matter. Prompts placed during study interrupt reading and cost time, so they must earn their place; the meta-analytic evidence suggests they usually do. And self-explanation competes with instructional explanation - if the text already explains why, asking the learner to explain why is partly redundant, which is one reason the technique pairs so naturally with worked examples, where the reasoning behind each step is typically left implicit.

Key idea: Effective elaboration prompts ask for a connection the learner has enough knowledge to construct, and they add most where the reasoning is left implicit, as in worked examples.

Toward transfer

The deepest payoff of interleaving and elaboration is transfer: the ability to apply what was learned to new situations, which is the real goal of education. Blocked, shallow practice can produce knowledge that works only in the exact form it was drilled. Interleaving builds the discrimination needed to recognize when a method applies; elaboration builds the flexible, connected understanding that lets knowledge move to novel contexts. Neither guarantees transfer - transfer is genuinely hard and often fails - but both stack the odds in its favor far better than the fluent methods learners instinctively prefer.

It helps to be precise about what transfer means, because the word covers a wide range. Susan Barnett and Stephen Ceci proposed a taxonomy that distinguishes transfer along several dimensions at once: how different the knowledge domain is, how much time has elapsed, how different the physical and social context is, and how different the required functional use is. Near transfer - applying a method to a new problem of the same type next week - is routinely achieved. Far transfer - applying a principle learned in statistics to a decision at work two years later - is rare, and the historical record of programmes promising it is discouraging.

The mechanism that does support transfer is knowing the deep structure rather than the surface features. Learners who have only encountered a principle in one guise typically encode the guise. Interleaving attacks this by forcing attention onto the features that distinguish problem types; elaboration attacks it by connecting the principle to reasons that hold across instances; varied practice contexts, from Lesson 3, attack it by building multiple routes to the same knowledge. None of these is a transfer guarantee, and honest instructional design states its transfer claims modestly and tests them rather than asserting them.

Key idea: Near transfer is routine and far transfer is rare, and what supports transfer is encoding deep structure rather than surface features, which is exactly what interleaving and elaboration promote.

Common misconceptions

  • Interleaving always beats blocking. It depends on the material: strong for confusable visual categories, weaker for mathematics, and negative for unrelated word pairs.
  • Smooth practice means effective practice. Blocked practice produces higher accuracy during the session and lower retention afterwards.
  • Elaboration means adding more detail. It means building connections to prior knowledge and to reasons, not accumulating facts.
  • Self-explanation and re-reading the explanation are equivalent. The benefit comes from the learner generating the account, not from receiving it.
  • Good teaching produces far transfer routinely. Far transfer is difficult and often fails, and claims of it need evidence rather than assertion.

Recap

  • Interleaving mixes related problem types within a session and trains the discrimination that blocked practice never requires.
  • Its meta-analytic effect is moderate overall (g = 0.42) but strongly moderated by material and by how confusable the categories are.
  • Interleaving and spacing are partly confounded, since mixing topics automatically distributes practice on each.
  • Prompted self-explanation is well supported at about g = 0.55; elaborative interrogation is promising but less securely established.
  • Good prompts ask for connections the learner has the knowledge to construct, and add most where reasoning is left implicit.
  • Transfer depends on encoding deep structure, near transfer is common, and far transfer is rare and should be claimed cautiously.

Sources

  1. Brunmair, M., and Richter, T. (2019). Similarity matters: A meta-analysis of interleaved learning and its moderators. Psychological Bulletin, 145(11), 1029-1052. pubmed.ncbi.nlm.nih.gov
  2. Rohrer, D., and Taylor, K. (2007). The shuffling of mathematics problems improves learning. Instructional Science, 35, 481-498. link.springer.com
  3. Bisra, K., Liu, Q., Nesbit, J. C., Salimi, F., and Winne, P. H. (2018). Inducing self-explanation: A meta-analysis. Educational Psychology Review, 30, 703-725. link.springer.com
  4. Dunlosky, J., Rawson, K. A., Marsh, E. J., Nathan, M. J., and Willingham, D. T. (2013). Improving students' learning with effective learning techniques. Psychological Science in the Public Interest, 14(1), 4-58. pubmed.ncbi.nlm.nih.gov
  5. Donoghue, G. M., and Hattie, J. A. C. (2021). A meta-analysis of ten learning techniques. Frontiers in Education, 6, 581216. frontiersin.org
  6. Pashler, H., Bain, P., Bottge, B., Graesser, A., Koedinger, K., McDaniel, M., and Metcalfe, J. (2007). Organizing Instruction and Study to Improve Student Learning (NCER 2007-2004). Institute of Education Sciences, What Works Clearinghouse. ies.ed.gov
  7. Rohrer, D., Dedrick, R. F., and Stershic, S. (2015). Interleaved practice improves mathematics learning. Journal of Educational Psychology, 107(3), 900-908. find source โ†—
Key terms
Interleaving
Mixing different problem types or topics within a practice session.
Blocked practice
Practicing one type of problem fully before moving to the next; smooth but less durable.
Discrimination
The skill of identifying which type of problem one faces and thus which method to apply.
Elaboration
Connecting new information to prior knowledge and processing it for meaning.
Elaborative interrogation
Asking and answering why a stated fact is true to anchor it in reasoning.
Self-explanation
Explaining a process or solution to oneself in one's own words while studying.
Transfer
Applying learned knowledge or skills to new situations.

Dual Coding and Multimedia Learning Basics

  • Explain dual coding theory and the two processing channels.
  • Use words-plus-pictures to reduce load and deepen encoding.
  • Distinguish dual coding from the debunked 'learning styles' idea.

The final core strategy concerns the form in which information is represented. Dual coding theory proposes that the mind processes verbal information (words, spoken or written) and visual information (images, diagrams, spatial arrangements) through two partially separate channels. When information is coded in both a verbal and a visual form, it is encoded twice, through complementary routes, giving two independent ways to retrieve it later. This is why a well-chosen diagram paired with an explanation is remembered better than either alone.

The theory is Allan Paivio's, developed from the 1970s and summarized for educators by James Clark and Paivio in 1991. Its core claim is that the two systems are functionally independent but interconnected: they can operate in parallel, and activity in one can trigger activity in the other. Several long-standing findings support it. The picture superiority effect is that pictures are remembered better than their names in free recall. The concreteness effect is that concrete words such as hammer are remembered better than abstract words such as justice, and the standard explanation is that concrete words readily recruit an image code as well as a verbal one, while abstract words get only the verbal code.

Key idea: Dual coding theory holds that verbal and imaginal systems are separate but interconnected, and the picture superiority and concreteness effects are its main empirical support.

Two channels, one benefit

The dual-channel idea, which also underlies the multimedia principles in Module 6, has two consequences. First, it offers a capacity benefit: because the visual and verbal channels are partly independent, splitting information sensibly across them can ease the working-memory bottleneck compared with cramming everything into one channel (for example, a wall of text). Second, it offers an encoding benefit: dual-coded material builds richer, doubly connected representations. A process learned as both a labeled flow diagram and a verbal explanation has more retrieval handles than the same process learned as prose only.

Practically, dual coding means pairing words with relevant visuals: diagrams, timelines, graphs, concept maps, simple sketches, and gesture. It also means encouraging learners to generate their own visual representations - drawing a diagram of a process is a potent form of elaboration, because it forces them to decide how the parts relate spatially. The emphasis is on relevant visuals that carry the structure of the idea, not decorative images, which add extraneous load (recall the seductive-details problem from Module 2).

The distinction between relevant and decorative is sharper than it sounds, and it is worth applying a test. A visual earns its place if removing it would make some relation in the content harder to grasp. A flow diagram of a feedback loop passes, because the arrows carry causal structure that prose expresses only sequentially. A photograph of a laboratory beside a paragraph about titration fails, because nothing in the content is easier to understand with it there. Timelines, cross-sections, phase diagrams, tree structures, and graphs generally pass; stock photography, clip art, and thematic borders generally fail. Applying that test to an existing set of slides is often the fastest single improvement available to a designer, because it usually removes material rather than requiring new material to be made.

Learners can also be asked to build the visual representation themselves. Logan Fiorella and Richard Mayer group generative drawing with summarizing, mapping, self-explaining, teaching, and enacting as strategies that work by prompting learners to select, organize, and integrate material rather than receive it. Drawing is demanding precisely because it forces decisions - what is above what, what connects to what, what is a cause and what a consequence - that prose lets a reader skip. The evidence is positive but conditional: learners usually need support such as a partial diagram or a set of elements to place, since unsupported drawing can consume so much capacity on the drawing itself that little is left for the content.

Key idea: Having learners generate their own diagrams promotes selection, organization, and integration, but usually needs scaffolding so the act of drawing does not consume the capacity meant for learning.

A crucial clarification: this is not learning styles

Dual coding is often confused with the popular but unsupported idea of learning styles - the claim that each person is a "visual," "auditory," or "kinesthetic" learner and should be taught only in their preferred mode. Careful reviews have repeatedly failed to find evidence that matching instruction to a supposed style improves learning, and the belief persists mainly as a myth. Dual coding says something entirely different and evidence-based: everyone benefits from well-integrated words and pictures, because everyone has both channels. The recommendation is not to sort learners into modes but to combine modes for all of them.

Because this is the most widespread false belief in education, it is worth seeing exactly why it fails rather than accepting the verdict on authority. Harold Pashler, Mark McDaniel, Doug Rohrer and Robert Bjork set out the logic in their 2008 review. The claim at issue is not that people have preferences - they plainly do - but the meshing hypothesis: that matching instruction to a learner's preferred modality improves their learning. Testing that claim requires a specific design. You must classify learners by style, randomly assign them to at least two instructional modes, and test everyone on the same outcome. Support requires a crossover interaction: visual learners must do better with visual instruction and auditory learners must do better with auditory instruction. A main effect of one method being better for everyone, or of one group scoring higher overall, proves nothing about meshing.

Pashler and colleagues searched the literature and found that the overwhelming majority of studies never used this design. Of the handful that did, almost all produced results contradicting the meshing hypothesis. Later direct tests reached the same conclusion. Beth Rogowsky, Barbara Calhoun and Paula Tallal classified adults as auditory or visual learners, randomly assigned them to listen to an audiobook or read the same text, and tested comprehension immediately and later; there was no interaction between preferred style and instructional mode. Studies in anatomy and medical education have repeatedly found that students who study in line with their measured VARK preference do not outperform those who do not.

Key idea: The meshing hypothesis requires a crossover interaction between learner style and instructional mode, and the few studies using the correct design have failed to find one.

What is true in the neighbourhood of the myth

Rejecting learning styles does not mean rejecting everything nearby, and a careless dismissal loses real findings. Four things are true. Learners genuinely do have preferences, and preferences affect enjoyment, engagement, and choice even when they do not affect learning. Individuals genuinely differ in modality-specific abilities - some people have unusually strong spatial visualization or verbal fluency - and ability is not the same variable as preference; the two correlate weakly at best. Prior knowledge in the domain, which varies enormously between learners, is a far stronger moderator of what a given design does than any style measure. And most importantly, the modality that best suits a message depends on the content, not the recipient: a molecular structure needs a diagram for everyone, and a chronological argument usually needs prose for everyone.

That last point is the constructive replacement for style-matching. Instead of asking "what kind of learner is this?", ask "what representation does this idea require, and what does the learner already know?" Those two questions have evidence behind them; the style question does not.

Key idea: Preferences, modality-specific abilities, and prior knowledge are all real, but the representation should be matched to the content and the learner's knowledge, not to a supposed sensory type.

Why the myth survives

Learning styles remain endorsed by large majorities of educators. Sanne Dekker and colleagues found around ninety percent agreement among teachers in the United Kingdom and the Netherlands. Philip Newton's surveys of the published literature found that most papers mentioning learning styles assume their validity rather than examine it, and later work with colleagues found the belief still dominant in medical education despite a decade of rebuttals.

Several forces sustain it. It is intuitively plausible, because we notice our preferences. It is flattering and inclusive, because it reframes failure as mismatch rather than deficit. It is commercially supported by inventories and training. And it is confounded with genuinely good practice: teachers who "teach to all styles" end up presenting material in multiple representations, which does help - but it helps for the dual-coding reason, not the style reason, and it helps everyone rather than each person in their own way. The cost is not merely intellectual. Time spent administering inventories and building parallel materials is time not spent on spacing, retrieval, feedback, or worked examples, all of which have evidence behind them.

Key idea: The learning-styles belief persists because it is intuitive, flattering, commercially promoted, and confounded with the genuinely useful practice of presenting material in multiple representations.

Fitting the strategies together

Step back and see how Module 3's strategies interlock. Spacing and retrieval practice make memory durable and accessible. Interleaving builds the discrimination needed to apply knowledge. Elaboration and dual coding deepen and connect encoding so that knowledge is meaningful and richly cued. None is a gimmick; each is grounded in the cognitive architecture of Modules 1 and 2, and each shares a family resemblance - they all trade a comfortable feeling of fluency for genuine, transferable learning. A designer who builds these into instruction is not adding decoration; they are aligning the design with how memory actually works.

Common misconceptions

  • Matching teaching to a learner's style improves outcomes. The crossover interaction this requires has not been found by studies designed to detect it.
  • Dual coding is just learning styles with better branding. Dual coding says combine modes for everyone; styles says give each learner one mode.
  • Adding any image helps. Only images carrying the structure of the idea help; decorative ones add extraneous load.
  • Preferences are meaningless. They are real and affect engagement; they simply do not predict which instruction teaches a person more.
  • Drawing is always generative. Unsupported drawing can consume the very capacity it was meant to free, so it needs scaffolding.

Recap

  • Dual coding theory posits separate but interconnected verbal and imaginal systems, supported by picture superiority and concreteness effects.
  • Words plus relevant pictures ease the single-channel bottleneck and build doubly connected representations.
  • Generative drawing promotes selection, organization, and integration, but usually needs scaffolding.
  • The learning-styles meshing hypothesis requires a crossover interaction that well-designed studies have failed to find.
  • Preferences, modality abilities, and prior knowledge are real, but representation should match the content, not a supposed learner type.
  • The myth survives on intuition, flattery, commerce, and confusion with the genuinely useful practice of multiple representations.

Sources

  1. Pashler, H., McDaniel, M., Rohrer, D., and Bjork, R. (2008). Learning styles: Concepts and evidence. Psychological Science in the Public Interest, 9(3), 105-119. pubmed.ncbi.nlm.nih.gov
  2. Rogowsky, B. A., Calhoun, B. M., and Tallal, P. (2020). Providing instruction based on students' learning style preferences does not improve learning. Frontiers in Psychology, 11, 164. pmc.ncbi.nlm.nih.gov
  3. Newton, P. M. (2015). The learning styles myth is thriving in higher education. Frontiers in Psychology, 6, 1908. pmc.ncbi.nlm.nih.gov
  4. Newton, P. M., Najabat-Lattif, H. F., Santiago, G., and Salvi, A. (2021). The learning styles neuromyth is still thriving in medical education. Frontiers in Human Neuroscience, 15, 708540. pmc.ncbi.nlm.nih.gov
  5. Dekker, S., Lee, N. C., Howard-Jones, P., and Jolles, J. (2012). Neuromyths in education: Prevalence and predictors of misconceptions among teachers. Frontiers in Psychology, 3, 429. pmc.ncbi.nlm.nih.gov
  6. Clark, J. M., and Paivio, A. (1991). Dual coding theory and education. Educational Psychology Review, 3, 149-210. link.springer.com
  7. Fiorella, L., and Mayer, R. E. (2016). Eight ways to promote generative learning. Educational Psychology Review, 28, 717-741. link.springer.com
Key terms
Dual coding theory
The theory that verbal and visual information are processed through two partially separate channels.
Verbal channel
The processing route for spoken and written words.
Visual channel
The processing route for images, diagrams, and spatial information.
Dual-coded representation
Material encoded in both verbal and visual form, giving two retrieval routes.
Learning styles myth
The unsupported belief that matching instruction to a person's preferred sensory mode improves learning.
Generative drawing
Having learners create their own diagrams, a strong form of elaboration.

Module 4: Motivation, Mindset, and Prior Knowledge

The learner-side factors - motivation, beliefs about ability, and existing knowledge - that determine whether good design takes root.

Motivation: Self-Determination and Expectancy-Value

  • Distinguish intrinsic and extrinsic motivation.
  • Explain autonomy, competence, and relatedness as drivers of motivation.
  • Use the expectancy-value framework to diagnose disengagement.

A perfectly designed lesson teaches no one who will not engage with it. Motivation - what moves a learner to start, persist, and invest effort - is therefore not separate from instructional design but part of it. Two well-supported frameworks give designers practical leverage.

Key idea: Motivation is part of instructional design rather than a precondition for it, and two well-supported frameworks turn vague concern about engagement into specific diagnoses and moves.

Intrinsic and extrinsic motivation

Intrinsic motivation is doing an activity for its own sake - because it is interesting, satisfying, or meaningful. Extrinsic motivation is doing it for a separable outcome - a grade, a reward, avoiding a penalty. Both can drive learning, but they behave differently. Intrinsic motivation tends to produce deeper engagement and persistence, while purely extrinsic pressure can produce compliance that stops the moment the reward does.

A notable caution from research is the overjustification effect: offering strong external rewards for something a learner already enjoys can sometimes reduce their intrinsic interest, as the activity gets reframed as work done for pay. Rewards are not evil, but they are a blunt instrument, best used carefully.

The numbers behind that caution are worth knowing, because the claim is often stated too strongly or dismissed too quickly. Edward Deci, Richard Koestner and Richard Ryan meta-analysed 128 experiments and found that tangible rewards significantly undermined free-choice intrinsic motivation, with effect sizes of about d = -0.40 for rewards contingent on merely engaging in the task, -0.36 for completing it, and -0.28 for performing it well. The same analysis found that positive verbal feedback enhanced intrinsic motivation, at roughly d = +0.33, and that undermining was more pronounced in children than in college students.

This has been a genuinely contested literature. Judy Cameron and David Pierce argued from their own meta-analyses that the undermining effect is narrow and largely confined to particular reward arrangements, and the exchange between the two camps ran through several rounds in the same journal. Where the field has settled is roughly here: expected, tangible, salient rewards offered for engaging in an activity that the learner already finds interesting reliably reduce later voluntary engagement in it; unexpected rewards, rewards for genuinely uninteresting tasks, and informational praise do not show the same effect and may help. The useful design question is therefore not "rewards or no rewards" but "does this reward read to the learner as control, or as information about their competence?"

Key idea: Expected tangible rewards for already-interesting activities reliably reduce later voluntary engagement, while informational praise tends to increase it, so what matters is whether a reward is experienced as controlling or as competence information.

Extrinsic motivation is not one thing

The intrinsic-extrinsic split is too coarse for classroom use, and self-determination theory offers a more useful continuum of internalization. At one end, external regulation means acting purely to obtain a reward or avoid a punishment: the student writes the essay because failing to would cost marks. Introjected regulation means the pressure has been taken inside but not accepted: the student writes it to avoid guilt or shame. Identified regulation means the student personally endorses the goal: they write it because being able to argue in writing matters to them. Integrated regulation means the goal fits coherently with their other values and sense of self.

The practical point is that identified and integrated regulation behave much more like intrinsic motivation - persistent, higher quality, better wellbeing - even though the activity is still done for a separable outcome. Most academic work is never going to be intrinsically fascinating, so the realistic aim is not to make every task delightful but to move learners along this continuum by supplying rationales, acknowledging that a task is tedious, and connecting it to goals learners already hold. Giving a meaningful reason for a dull requirement is one of the cheapest motivational interventions available.

Key idea: Extrinsic motivation ranges from external regulation to fully integrated regulation, and the realistic instructional aim is internalization through rationales and connection to learners' own goals rather than making every task intrinsically interesting.

Self-determination theory

Self-determination theory proposes that intrinsic motivation flourishes when three basic psychological needs are met.

  • Autonomy: a sense of volition and choice, of acting from one's own interests rather than being coerced. Even small, genuine choices - which topic, which format, which order - raise engagement.
  • Competence: a sense of being effective and making progress. Tasks pitched at the right challenge, with clear feedback, feed this need; tasks that are hopelessly hard or trivially easy starve it.
  • Relatedness: a sense of connection to and being valued by others - teachers and peers. Learners invest more when they feel they belong and that someone cares whether they succeed.

The design implication is direct: build in meaningful choice, calibrate difficulty so learners experience earned progress, and cultivate a supportive, connected climate. These are not soft extras; they are levers on the engine that powers all the cognitive work in Modules 1 through 3.

Key idea: Autonomy, competence, and relatedness are the conditions under which motivation becomes self-sustaining, and each maps onto concrete, low-cost design choices.

Expectancy-value theory

Expectancy-value theory offers a complementary diagnostic. It holds that a learner's motivation for a task depends on two beliefs multiplied together: expectancy ("Can I succeed at this?") and value ("Do I care about this - is it useful, interesting, or important to me?"). Because they multiply, if either is near zero, motivation collapses. A student who sees no value in a task will not try even if success is certain; a student who values a task but is convinced they will fail will disengage to protect themselves.

This immediately guides diagnosis. When a learner is unmotivated, ask which term is missing: do they doubt they can succeed (raise expectancy with scaffolding, achievable steps, and evidence of progress) or do they not see the point (raise value by connecting the task to their goals, interests, or real-world use)? The right intervention depends entirely on which belief has failed.

Unpacking value, including cost

Jacquelynne Eccles and Allan Wigfield's formulation subdivides value into four components, and the fourth is the one instructors most often overlook. Attainment value is how much doing the task well matters to the learner's sense of who they are. Intrinsic value is the enjoyment of doing it. Utility value is its usefulness for the learner's other goals, present or future. Cost covers everything the task takes away: effort required, opportunities forgone, emotional price such as anxiety or fear of looking foolish.

Adding cost changes diagnosis substantially. A student who finds the material interesting and believes they can succeed may still disengage because the perceived cost is too high - the assignment competes with paid work, or an earlier humiliation makes participation feel dangerous. In their later restatement, Eccles and Wigfield emphasize that these judgments are situated: they are formed in a specific setting, at a specific moment, in light of a learner's identity and their reading of whether people like them belong there. Motivation is not a stable personal quantity you can measure once.

Key idea: Value has four components - attainment, intrinsic, utility, and cost - and perceived cost, including emotional and opportunity cost, is the most frequently missed source of disengagement.

Interventions that work, and how much

Utility value is the component most amenable to direct intervention, and there is good field evidence for it. Chris Hulleman and Judith Harackiewicz ran a randomized field experiment in high-school science in which one group periodically wrote short pieces connecting the course material to their own lives, while a control group wrote summaries. The relevance writing raised interest in science and course grades - but only for students who had low expectations of success. Students already confident of success showed no benefit and, in some analyses of related studies, slight harm, apparently because being told why something matters is unhelpful when you already believe it does.

Two lessons follow, and both generalize. First, motivational interventions are typically targeted, not universal: an intervention aimed at a belief a learner already holds is at best wasted. Second, having learners generate the connection themselves works better than being told the connection, which is elaboration in Lesson 7 wearing different clothes. Later work in this line has extended utility-value writing to undergraduate STEM courses with effects on persistence, though as with all field interventions, effect sizes are modest and vary by context.

Key idea: Getting learners to write their own connections between material and their lives raises interest and grades, but chiefly for those with low expectations of success, so motivational interventions must be targeted.

Two cautions

First, causation runs in both directions. It is comfortable to assume motivation produces achievement, but achievement also produces motivation: a learner who succeeds at something develops competence beliefs and interest in it. This is why fixing the instruction sometimes fixes the motivation, and why an unmotivated class is not automatically a motivation problem - it may be a comprehension problem, a load problem, or a prerequisite problem showing up as disengagement.

Second, be sceptical of large claims. Brief motivational interventions in education have a mixed replication record, and effects reported from small enthusiastic pilots often shrink in large preregistered trials, a pattern examined closely in the next lesson. Treat the frameworks here as reliable diagnostic tools - they tell you what to look at - and treat specific intervention effect sizes as provisional.

Key idea: Achievement causes motivation as well as the reverse, and brief motivational interventions have a mixed replication record, so these frameworks are best used diagnostically.

Common misconceptions

  • All rewards destroy intrinsic motivation. Expected tangible rewards for already-interesting tasks do; informational praise and rewards for genuinely dull tasks generally do not.
  • Extrinsic motivation is uniformly bad. Identified and integrated regulation are extrinsic yet behave much like intrinsic motivation.
  • Choice always raises motivation. Choices must be meaningful; trivial or overwhelming options add load without supporting autonomy.
  • An unmotivated student needs motivating. Disengagement often signals overload, missing prerequisites, or high perceived cost.
  • Telling students why a topic matters builds value. Learners generating the connection themselves works better, and helps mainly those with low expectations of success.

Recap

  • Intrinsic and extrinsic motivation differ in quality of engagement, and the overjustification effect is real but bounded.
  • Deci, Koestner and Ryan found tangible contingent rewards undermined free-choice motivation at about d = -0.3 to -0.4, while positive feedback enhanced it.
  • Extrinsic motivation runs along a continuum of internalization, and rationales move learners along it.
  • Self-determination theory identifies autonomy, competence, and relatedness as the conditions for self-sustaining motivation.
  • Expectancy multiplied by value predicts effort, and value includes attainment, intrinsic, utility, and cost components.
  • Utility-value writing raises interest and grades chiefly for students with low success expectations, so interventions must be targeted.

Sources

  1. Ryan, R. M., and Deci, E. L. (2000). Self-determination theory and the facilitation of intrinsic motivation, social development, and well-being. American Psychologist, 55(1), 68-78. selfdeterminationtheory.org
  2. Deci, E. L., Koestner, R., and Ryan, R. M. (1999). A meta-analytic review of experiments examining the effects of extrinsic rewards on intrinsic motivation. Psychological Bulletin, 125(6), 627-668. pubmed.ncbi.nlm.nih.gov
  3. Eisenberger, R., Pierce, W. D., and Cameron, J. (1999). Effects of reward on intrinsic motivation - negative, neutral, and positive: Comment on Deci, Koestner, and Ryan (1999). Psychological Bulletin, 125(6), 677-691. pubmed.ncbi.nlm.nih.gov
  4. Hulleman, C. S., and Harackiewicz, J. M. (2009). Promoting interest and performance in high school science classes. Science, 326(5958), 1410-1412. pubmed.ncbi.nlm.nih.gov
  5. Asher, M. W., Harackiewicz, J. M., Beymer, P. N., Hecht, C. A., Lamont, L. B., Else-Quest, N. M., Priniski, S. J., and Thoman, D. B. (2023). Utility-value intervention promotes persistence and diversity in STEM. Proceedings of the National Academy of Sciences, 120(19), e2300463120. pubmed.ncbi.nlm.nih.gov
  6. American Psychological Association, Coalition for Psychology in Schools and Education. (2015). Top 20 Principles From Psychology for PreK-12 Teaching and Learning. APA. apa.org
  7. Eccles, J. S., and Wigfield, A. (2020). From expectancy-value theory to situated expectancy-value theory: A developmental, social cognitive, and sociocultural perspective on motivation. Contemporary Educational Psychology, 61, 101859. find source โ†—
Key terms
Intrinsic motivation
Doing an activity for its own sake, because it is interesting or satisfying.
Extrinsic motivation
Doing an activity for a separable outcome such as a reward or grade.
Overjustification effect
The reduction of intrinsic interest that can follow strong external rewards for an already-enjoyed activity.
Self-determination theory
A theory that intrinsic motivation grows when autonomy, competence, and relatedness needs are met.
Expectancy-value theory
A theory that motivation depends on believing one can succeed (expectancy) and caring about the task (value).
Relatedness
The need to feel connected to and valued by others in a learning setting.

Mindset, Beliefs, and Self-Regulation

  • Explain growth and fixed mindsets and the evidence around them.
  • Describe how attributions affect persistence.
  • Define self-regulated learning and metacognition.

Beyond momentary motivation lie a learner's durable beliefs about ability and their capacity to manage their own learning. Both strongly shape how they respond to difficulty.

Growth and fixed mindsets

A fixed mindset is the belief that intelligence and ability are largely stable traits you either have or lack. A growth mindset is the belief that ability can develop through effort, good strategies, and help. These beliefs matter most in the face of setbacks. A learner who reads a failure as evidence of a fixed limit ("I'm not smart enough") tends to withdraw effort to protect their self-image; a learner who reads the same failure as information about what to work on tends to persist and adjust.

The research literature here deserves an honest, graduate-level caveat: while the underlying belief-behavior link is real, the average effect of brief mindset interventions is modest and varies by context, and early enthusiasm sometimes outran the evidence. Larger, well-run studies suggest such interventions can help, especially for students who are struggling or from disadvantaged backgrounds, but they are not a magic switch. The sound takeaway is not to slap "you can do it" posters on the wall, but to cultivate genuine growth-oriented conditions: valuing effort and strategy, framing errors as part of learning, and giving process-focused feedback.

What the evidence actually shows

This is a case where a graduate student should be able to state the numbers, because the gap between the popular story and the research record is wide and instructive.

Victoria Sisk and colleagues ran two meta-analyses in 2018. The first, covering 273 studies and more than 365,000 participants, examined the correlation between holding a growth mindset and academic achievement; the association was weak. The second, covering 43 studies and more than 57,000 participants, examined whether mindset interventions raise achievement; the overall effect was also weak. But their moderator analysis supported a specific version of the theory: students of low socioeconomic status and students academically at risk were the groups where benefits appeared.

Brooke Macnamara and Alexander Burgoyne's 2023 review went further, assessing 63 studies with nearly 98,000 participants against a checklist of design standards. They found widespread shortcomings in design, analysis, and reporting, and a signal of bias: authors with a financial interest in positive findings published significantly larger effects than authors without one. The overall effect was d = 0.05, and it was not significant after correcting for likely publication bias. Restricting to the studies that demonstrably changed students' mindsets as intended gave d = 0.04; restricting to the highest-quality evidence gave d = 0.02.

Against that, the strongest positive study is also the most rigorous. David Yeager and colleagues' National Study of Learning Mindsets randomized a short online intervention of under an hour across a nationally representative sample of US secondary students, with preregistered analyses, independent data processing, and blinded Bayesian corroboration. It improved grades among lower-achieving students and increased enrolment in advanced mathematics - and the effect on grades appeared only where peer norms in the school were consistent with the intervention's message. A one-hour, near-free intervention producing a small but real effect in the students who need it most, conditional on the surrounding culture, is a genuine finding. It is not the finding that mindset changes everything for everyone.

Key idea: Across meta-analyses the average effect of mindset interventions is close to zero, while the largest rigorous trial finds small, real benefits confined to lower-achieving students in supportive school contexts.

Reading the disagreement responsibly

The apparent conflict resolves once you notice that the three findings are answering different questions. Sisk and Macnamara ask what the average published mindset intervention does across all learners; the answer is very little. Yeager asks what a carefully constructed intervention does for a targeted subgroup in a supportive context; the answer is a small amount. Both can be true, and together they yield practical guidance that is neither hype nor dismissal.

What survives is this. Beliefs about ability do relate to how learners respond to difficulty, though the correlation is modest. Brief, generic mindset messaging aimed at everyone is unlikely to change achievement. Targeted intervention at transitions, for struggling students, embedded in a school culture that treats difficulty as normal, can help a little. And crucially, a mindset message that is contradicted by the surrounding practice - being told that ability grows while being ranked publicly by fixed test scores - is worse than useless, because it asks students to believe something their environment disconfirms. Design the environment first; the belief is downstream of what learners observe.

Key idea: Generic mindset messaging changes little, targeted intervention in a supportive context helps modestly, and messaging contradicted by the surrounding practice is counterproductive.

Attributions

Closely related is attribution - the causes a learner assigns to their outcomes. Attributing a poor result to something stable and uncontrollable ("I have no talent for this") predicts giving up. Attributing it to something controllable ("I didn't use an effective strategy" or "I didn't practice retrieval") predicts renewed effort, because the cause can be changed. Teachers influence attributions constantly through how they praise and console. Praising a student's process and strategy ("your approach of checking each step paid off") supports adaptive attributions better than praising fixed ability ("you're so smart"), which can ironically make learners more fragile when they later struggle.

Bernard Weiner's attribution theory organizes the causes learners cite along three dimensions rather than one, and using all three sharpens diagnosis. Locus asks whether the cause is internal (my effort, my ability) or external (the test was unfair, the teacher is unclear). Stability asks whether it persists over time (ability, difficulty of the subject) or fluctuates (effort on this occasion, luck, illness). Controllability asks whether the learner can do anything about it. The corrosive combination is internal, stable, and uncontrollable - "I am not a maths person" - because it implies that nothing the learner does will change the outcome, and the rational response to that belief is to stop trying and protect one's self-image. The productive combination is internal, unstable, and controllable - "I did not practise retrieval, I reread" - because it names something that can be different next time.

Notice that the productive attribution is not merely more optimistic; it is more specific. Telling a student "you just need to try harder" is internal and controllable but useless, because effort without a better strategy usually reproduces the same result and then confirms the fixed-ability story. The strongest attributional move available to a teacher is to identify the actual strategy failure - not enough spacing, no self-testing, worked examples read passively - so the learner has something concrete to change.

Carol Mueller and Carol Dweck's experiments with fifth-graders demonstrated the mechanism on the praise side. Children praised for intelligence after early success, compared with children praised for effort, subsequently preferred easier tasks, showed less persistence and enjoyment after a difficult set, attributed failure to low ability, and performed worse. Ability praise is not neutral encouragement; it teaches a theory about what success means.

Key idea: Attributions vary in locus, stability, and controllability, and the useful teaching move is to identify the specific strategy failure rather than to urge more effort or praise ability.

Self-regulated learning

The capstone learner skill is self-regulated learning: the ability to plan, monitor, and adjust one's own learning. It rests on metacognition - thinking about one's own thinking, including accurate judgments of what one does and does not understand. Self-regulated learners set goals, choose strategies (ideally the effective ones from Module 3), monitor their progress honestly, and change course when something is not working.

The importance of accurate monitoring returns us to the fluency illusion: poor self-regulators feel they understand after rereading and stop too early, while strong self-regulators test themselves, notice the gaps, and keep going. This is why teaching students about effective strategies and their own metacognitive biases is itself a high-value intervention. A major goal of good instruction is to gradually hand over the regulation of learning to the learner, so that the scaffolds an expert teacher provides are eventually internalized.

Barry Zimmerman's model describes self-regulation as a cycle with three phases. In forethought, the learner analyses the task, sets goals, plans strategies, and forms beliefs about efficacy and value. In the performance phase they execute, apply strategies, and monitor themselves. In self-reflection they judge the outcome, attribute it to causes, and react - and those judgments and attributions feed directly into the forethought of the next cycle. This is why a single instance of poor attribution is not a one-off: it becomes the planning assumption for the next attempt.

Key idea: Self-regulation is cyclical, so the way a learner explains one outcome becomes the planning assumption for the next attempt.

Why monitoring fails

Robert Bjork, John Dunlosky and Nate Kornell catalogue the systematic errors in learners' judgments of their own learning, and each has a design remedy. Learners mistake fluency for competence, judging that smoothly read material is known. They show foresight bias, predicting future performance while looking at the answer, which makes retrieval look easy. They show a stability bias, assuming that what they know now is what they will know at the exam, and therefore underestimating forgetting. And they systematically prefer massed, blocked, fluent study, which is the direct consequence of judging learning by how it feels.

The remedies are structural rather than motivational. Delayed judgments of learning, made without the material in front of you, are much better calibrated than immediate ones. A practice test taken without notes gives a learner data their intuition cannot supply. And explicitly teaching students the phenomenon - that fluency is misleading and that difficulty during practice is often a good sign - changes how they interpret the discomfort of effective strategies, which is what makes them willing to keep using them.

Guidance syntheses put this work among the more promising areas in education. The Education Endowment Foundation's review of metacognition and self-regulated learning concludes that explicitly teaching pupils to plan, monitor, and evaluate their learning is a high-impact, comparatively low-cost approach, with the crucial proviso that it must be taught within subject content rather than as a standalone study-skills course - which is precisely the domain-specificity point from Lesson 2.

Key idea: Metacognitive monitoring fails predictably through fluency, foresight, and stability biases, and the remedies are delayed judgments, unaided practice tests, and teaching learners about the illusions themselves.

Common misconceptions

  • Growth mindset interventions reliably raise achievement. Meta-analytic averages are near zero; benefits are small, targeted, and context-dependent.
  • The mindset research is worthless. The largest preregistered trial found real, if small, gains for lower-achieving students where school norms supported the message.
  • Praising intelligence builds confidence. Ability praise reduced persistence, enjoyment, and performance after difficulty in controlled experiments.
  • Telling students to try harder is good attributional feedback. Effort without a better strategy reproduces the result and confirms the fixed-ability story.
  • Study skills can be taught as a separate course. Self-regulation transfers best when taught inside the subject content it applies to.

Recap

  • Fixed and growth mindsets describe beliefs about the malleability of ability, and they matter most in the face of setbacks.
  • Sisk and colleagues and Macnamara and Burgoyne find weak average effects, with d near 0.05 in the most stringent analysis.
  • Yeager and colleagues' national experiment found small real gains for lower-achieving students where peer norms aligned with the message.
  • Attributions vary in locus, stability, and controllability, and the useful move is to name a specific strategy failure.
  • Ability praise measurably reduces persistence and performance after difficulty compared with process praise.
  • Self-regulation is cyclical, monitoring fails through predictable biases, and it is best taught inside subject content.

Sources

  1. Sisk, V. F., Burgoyne, A. P., Sun, J., Butler, J. L., and Macnamara, B. N. (2018). To what extent and under which circumstances are growth mind-sets important to academic achievement? Two meta-analyses. Psychological Science, 29(4), 549-571. pubmed.ncbi.nlm.nih.gov
  2. Macnamara, B. N., and Burgoyne, A. P. (2023). Do growth mindset interventions impact students' academic achievement? A systematic review and meta-analysis with recommendations for best practices. Psychological Bulletin, 149(3-4), 133-173. pubmed.ncbi.nlm.nih.gov
  3. Yeager, D. S., Hanselman, P., Walton, G. M., Murray, J. S., Crosnoe, R., Muller, C., et al. (2019). A national experiment reveals where a growth mindset improves achievement. Nature, 573, 364-369. pmc.ncbi.nlm.nih.gov
  4. Mueller, C. M., and Dweck, C. S. (1998). Praise for intelligence can undermine children's motivation and performance. Journal of Personality and Social Psychology, 75(1), 33-52. pubmed.ncbi.nlm.nih.gov
  5. Bjork, R. A., Dunlosky, J., and Kornell, N. (2013). Self-regulated learning: Beliefs, techniques, and illusions. Annual Review of Psychology, 64, 417-444. pubmed.ncbi.nlm.nih.gov
  6. Education Endowment Foundation. (2021). Metacognition and Self-Regulated Learning: Guidance Report. EEF. educationendowmentfoundation.org.uk
  7. Schunk, D. H. (2008). Metacognition, self-regulation, and self-regulated learning: Research recommendations. Educational Psychology Review, 20, 463-467. link.springer.com
Key terms
Fixed mindset
The belief that ability is a largely stable trait one has or lacks.
Growth mindset
The belief that ability can develop through effort, strategy, and help.
Attribution
The cause a learner assigns to an outcome, which shapes future persistence.
Process praise
Praise focused on effort and strategy rather than fixed ability.
Self-regulated learning
Planning, monitoring, and adjusting one's own learning.
Metacognition
Thinking about one's own thinking, including judging what one understands.

Constructivism, Prior Knowledge, and Misconceptions

  • Explain constructivism and the role of prior knowledge in learning.
  • Distinguish constructivism as a theory of learning from guidance in teaching.
  • Describe how to surface and address misconceptions.

Every learner arrives with a head full of existing ideas, and those ideas are not a blank slate to be written on but a structure that new learning must connect to, build on, or fight against. This is the core insight of constructivism, and it ties directly to the reconstructive, schema-based memory of Module 1.

Constructivism as a theory of learning

Constructivism holds that learners actively construct new understanding by integrating incoming information with their existing schemas, rather than passively receiving transmitted knowledge. Two classic ideas describe how this happens. Assimilation is fitting new information into an existing schema (a child who knows "dog" calls a new breed a dog).

Accommodation is changing the schema when the new information will not fit (adjusting the concept when they meet a cat). Learning, on this view, is a continual cycle of assimilating what fits and accommodating what does not. The single strongest predictor of new learning, as noted throughout this course, is relevant prior knowledge - because there is more for new material to connect to.

Assimilation and accommodation are Jean Piaget's terms, and his account added an important driver: disequilibrium. Learning is provoked when the existing schema fails to handle what the learner encounters, producing a discomfort the mind works to resolve. That is the theoretical ancestor of every "confront the misconception" strategy later in this lesson.

Lev Vygotsky's contribution runs alongside and is often misdescribed. His zone of proximal development is the range between what a learner can do alone and what they can do with the help of a more capable other. The zone is not a synonym for "slightly hard"; it is defined by the difference that assistance makes. Its instructional consequence, developed by David Wood, Jerome Bruner and Gail Ross as scaffolding, is that a helper takes over the parts of a task the learner cannot yet manage, keeps the goal in view, marks critical features, and withdraws that support as competence grows. This is the same fading logic as Module 2's worked examples, arrived at from a social rather than a cognitive-load direction, which is worth noticing: two independent traditions converge on graduated support.

Key idea: Piaget explains learning as the resolution of disequilibrium between schema and evidence, while Vygotsky's zone of proximal development and scaffolding describe graduated support that converges with cognitive load theory's fading.

A vital distinction: learning theory versus teaching method

Here the graduate learner must be careful, because "constructivism" is used in two senses that are often conflated. As a theory of how learning works - meaning is actively constructed on prior knowledge - constructivism is widely accepted and consistent with cognitive science. As a prescription for teaching method - specifically, that instruction should therefore be minimally guided, with learners left to discover principles largely on their own - it is contested and, for novices, not well supported.

A substantial body of evidence indicates that minimally guided instruction tends to be less effective and less efficient than guided instruction for learners who lack the prior knowledge to guide their own search, precisely because of the cognitive-load problems in Module 2: novices set loose flounder and overload. The resolution is that accepting the constructivist learning premise does not commit you to discovery-based teaching. Learners can construct knowledge actively while receiving strong guidance - worked examples, clear explanation, and structured practice are all fully compatible with, and arguably better for, active construction in novices.

The sharpest statement of this position is Paul Kirschner, John Sweller and Richard Clark's 2006 paper, whose title says that minimal guidance during instruction does not work and whose argument is the cognitive architecture of Modules 1 and 2: novices lack the schemas needed to guide a search, so unguided search loads working memory with problem-solving activity that builds nothing. Richard Mayer's earlier review reaches the same conclusion from the history of the field, noting that pure discovery methods have been proposed, tested, and found wanting in several successive waves - hence his proposal of a three-strikes rule against pure discovery learning.

The most useful evidence is a distinction rather than a verdict. Louis Alfieri and colleagues meta-analysed discovery-based instruction and found that unassisted discovery generally underperformed explicit instruction, while enhanced discovery - guided activity that includes scaffolds, worked examples, feedback, or elicited explanations - outperformed other forms of instruction. The active ingredient turns out to be guided activity, not discovery. This also explains the productive-failure results in Lesson 5: exploration works when it is followed by instruction that resolves it, and fails when it is left to stand alone.

One further caution about how this debate is often conducted. Critics of "constructivist teaching" and its defenders frequently talk past each other because neither side means anything precise by the term, which in practice covers everything from unstructured project work to carefully sequenced inquiry with heavy teacher support. When you read a study, ignore the label and ask what the teacher actually did, how much support the learners actually had, and what they already knew.

Key idea: Meta-analytic evidence favours guided activity over unassisted discovery, so the active ingredient is guidance rather than discovery, and arguments about constructivist teaching are usually arguments about undefined labels.

Misconceptions and conceptual change

The dark side of prior knowledge is that some of it is wrong. Misconceptions are intuitive but incorrect prior beliefs that are often deeply held and resistant to change - that heavier objects fall faster, that the seasons are caused by the earth's distance from the sun, that summing fractions means summing numerators and denominators. Crucially, simply telling a learner the correct fact frequently does not dislodge a misconception; the old belief persists alongside the new one and reasserts itself.

Genuine conceptual change usually requires more: first eliciting the learner's existing idea and making it explicit, then confronting it with an experience or argument it cannot explain (creating productive dissatisfaction), and then offering the correct conception as a more powerful, intelligible alternative and giving practice applying it. The instructional lesson is unavoidable: you must find out what learners already think, especially where they are likely to be wrong, and design directly against those errors. Assessment (Module 6) is one of the best tools for surfacing them.

Why misconceptions are so durable

Andrew Shtulman and Joshua Valcarcel provide the most striking evidence that correction does not erase. They gave university students and science-trained adults statements that are true scientifically but conflict with naive intuitions - that the earth revolves around the sun, that air has mass - alongside statements where science and intuition agree. Participants were slower and more error-prone on the conflict items, and the pattern persisted in people with substantial scientific training. Scientific knowledge, they concluded, suppresses earlier intuitions rather than supplanting them: the naive theory is still there, still generating an answer, and still having to be overridden.

Theoretical accounts of what a misconception is differ, and the difference matters for teaching. George Posner and colleagues described conceptual change as requiring four conditions - dissatisfaction with the current conception, and a new one that is intelligible, plausible, and fruitful - which frames the job as replacing one coherent theory with another. Andrea diSessa's knowledge-in-pieces view argues that novice intuitions are not coherent theories but loose collections of fragments activated by context, so the job is reorganizing and reweighting fragments rather than swapping theories. Michelene Chi's account holds that some misconceptions arise from assigning a concept to the wrong ontological category - treating heat as a substance rather than a process - and that these are the hardest of all, because no amount of within-category correction helps.

The practical upshot is shared across the accounts: expect the old idea to persist, expect it to resurface under time pressure or in unfamiliar contexts, and plan for repeated encounters rather than a single correction.

Key idea: Scientific knowledge suppresses rather than replaces naive intuitions, which continue to generate answers that must be overridden, so conceptual change requires repeated encounters rather than one correction.

Diagnosing and confronting misconceptions in practice

The physics education community built the most developed toolkit here, and it generalizes. The Force Concept Inventory, developed by David Hestenes and colleagues, is a multiple-choice instrument whose wrong answers are not random but are the specific naive beliefs students actually hold, so the pattern of errors diagnoses the misconception rather than merely counting failures. Administering it before and after instruction revealed something uncomfortable: traditional lecture courses produced small gains on conceptual understanding even when students passed the exams, while courses using interactive engagement produced substantially larger ones.

That finding scaled. Scott Freeman and colleagues meta-analysed 225 studies comparing traditional lecturing with active learning in undergraduate science, engineering, and mathematics. Examination and concept-inventory performance rose by 0.47 standard deviations under active learning, and the odds of failing were 1.95 times higher under traditional lecturing. Note carefully what "active learning" meant in these studies: mostly structured questioning, peer discussion of conceptual problems, and immediate feedback - guided activity, not discovery. The result is entirely consistent with the guidance evidence above.

Two design tools follow directly. Refutation texts state the common misconception explicitly, say plainly that it is wrong and why, and then give the correct account; they outperform ordinary expository text that simply presents the correct account without naming the rival. And distractor-driven items - questions whose wrong options encode known misconceptions - turn assessment into diagnosis, which is what makes formative assessment worth the time.

Key idea: Diagnostic instruments with misconception-based distractors, refutation texts that name the wrong idea explicitly, and guided active learning are the practical tools of conceptual change.

Common misconceptions

  • Constructivism as a learning theory implies discovery teaching. Learners construct meaning under strong guidance too, and novices do so more successfully.
  • Once corrected, a misconception is gone. The intuition is suppressed rather than erased and resurfaces under pressure.
  • Active learning means unguided activity. In the studies showing the largest gains it meant structured questioning, peer discussion, and immediate feedback.
  • Naming a wrong idea will reinforce it. Refutation texts that state and rebut the misconception outperform texts that ignore it.
  • The zone of proximal development just means moderately difficult work. It is defined by what assistance makes possible, not by difficulty alone.

Recap

  • Learners construct understanding by assimilating what fits and accommodating what does not, with disequilibrium as the driver.
  • Vygotsky's zone of proximal development and scaffolding converge with cognitive load theory on graduated, fading support.
  • Constructivism as a learning theory is well supported; minimally guided teaching is not, especially for novices.
  • Meta-analysis separates unassisted discovery, which underperforms, from enhanced guided discovery, which helps.
  • Naive intuitions are suppressed rather than replaced and continue to compete with correct knowledge.
  • Diagnostic inventories, refutation texts, and guided active learning are the practical instruments of conceptual change.

Sources

  1. Mayer, R. E. (2004). Should there be a three-strikes rule against pure discovery learning? The case for guided methods of instruction. American Psychologist, 59(1), 14-19. pubmed.ncbi.nlm.nih.gov
  2. Shtulman, A., and Valcarcel, J. (2012). Scientific knowledge suppresses but does not supplant earlier intuitions. Cognition, 124(2), 209-215. pubmed.ncbi.nlm.nih.gov
  3. Freeman, S., Eddy, S. L., McDonough, M., Smith, M. K., Okoroafor, N., Jordt, H., and Wenderoth, M. P. (2014). Active learning increases student performance in science, engineering, and mathematics. Proceedings of the National Academy of Sciences, 111(23), 8410-8415. pmc.ncbi.nlm.nih.gov
  4. Loibl, K., Roll, I., and Rummel, N. (2017). Towards a theory of when and how problem solving followed by instruction supports learning. Educational Psychology Review, 29, 693-715. link.springer.com
  5. Seifert, K., and Sutton, R. (2009). Educational Psychology (2nd ed.). Open Textbook Library. CC BY. open.umn.edu
  6. Kirschner, P. A., Sweller, J., and Clark, R. E. (2006). Why minimal guidance during instruction does not work: An analysis of the failure of constructivist, discovery, problem-based, experiential, and inquiry-based teaching. Educational Psychologist, 41(2), 75-86. find source โ†—
  7. Alfieri, L., Brooks, P. J., Aldrich, N. J., and Tenenbaum, H. R. (2011). Does discovery-based instruction enhance learning? Journal of Educational Psychology, 103(1), 1-18. find source โ†—
Key terms
Constructivism
The view that learners actively build understanding by integrating new information with existing schemas.
Assimilation
Fitting new information into an existing schema.
Accommodation
Changing a schema when new information does not fit it.
Guided instruction
Teaching that provides substantial structure and support, generally superior for novices.
Minimally guided instruction
Teaching that leaves learners to discover principles largely on their own; weak for novices.
Misconception
An intuitive but incorrect prior belief that resists simple correction.
Conceptual change
The process of restructuring a misconception into a correct understanding.

Module 5: Instructional Design Models and Objectives

Systematic frameworks - ADDIE and backward design - and the craft of writing measurable objectives with Bloom's taxonomy.

The ADDIE Model

  • Describe the five phases of ADDIE.
  • Explain the role of analysis and evaluation in the cycle.
  • Compare linear and iterative uses of the model.

Good instruction is rarely produced by improvisation; it is designed through a repeatable process. ADDIE is the most widely used generic framework for instructional design, named for its five phases: Analysis, Design, Development, Implementation, and Evaluation. It is less a rigid recipe than a checklist of the questions any serious design effort must answer, and it organizes the whole practical craft of this course.

Key idea: ADDIE is best understood as a checklist of questions any serious design effort must answer rather than as a prescriptive sequence.

Where ADDIE came from, and why it matters

Here is a fact that surprises most people who have used the framework for years: ADDIE has no original source document and no single author. Michael Molenda went looking for one and found that the acronym emerged as an informal, after-the-fact label for a family of instructional systems development procedures that had grown up in military and industrial training from the 1960s onward. The closest thing to an ancestor is the 1975 Interservice Procedures for Instructional Systems Development, produced at Florida State University for the US armed services, which laid out a five-phase process in enormous procedural detail. ADDIE became the shorthand for that whole family, and by the time it was widely taught, nobody could point to the document it came from.

Two consequences follow, and both should shape how you use it. First, ADDIE is not a research finding and carries no empirical claim of its own; it is an organizing convention, and evidence for the effectiveness of instruction lives in the specific choices made inside each phase, not in the phases themselves. Second, because it was never authored, there is no orthodox version to violate. Dee Andrews and Ludwika Goodson catalogued dozens of instructional design models even by 1980, most of them variations on the same phases with different emphases. Treat ADDIE as a vocabulary that lets designers talk to each other, and put your intellectual weight on what happens within the phases.

Key idea: ADDIE has no original source or single author but grew as a label for military and industrial instructional systems development, so it is an organizing convention rather than an evidence claim.

The five phases

  1. Analysis asks the prior questions: Who are the learners, and what prior knowledge do they have (Module 4)? What exactly do they need to be able to do? What is the gap between their current and desired state, and what constraints (time, resources, context) apply? Skipping analysis is the most common and most expensive design error, because it risks building a polished solution to the wrong problem.
  2. Design is the blueprint stage: writing measurable learning objectives, planning assessments that match them, sequencing content, and choosing strategies (ideally the evidence-based ones from Modules 2 and 3). No content is built yet; this is the plan.
  3. Development is where the actual materials are produced: the slides, readings, activities, videos, and assessments specified in the design.
  4. Implementation is delivery - the instruction is put into action with real learners, whether in a classroom, a workshop, or an online course.
  5. Evaluation asks whether it worked and how to improve it, using the methods of Module 7. Evaluation is not only a final step; formative evaluation runs throughout to catch problems early.

Analysis is four analyses, not one

Because analysis is the phase most often skipped, it repays being explicit about what it contains. Four distinct investigations hide behind the single word.

  1. Needs or gap analysis establishes the difference between current and required performance, and - crucially - whether instruction is the right remedy at all. Robert Mager's diagnostic question is the sharpest tool here: if the person's life depended on it, could they do it right now? If the answer is yes, the problem is not a lack of knowledge, and training will not fix it. Missing tools, unclear expectations, conflicting incentives, and absent feedback all produce performance gaps that look exactly like training needs and are immune to training.
  2. Learner analysis establishes prior knowledge above all, since Modules 1 and 2 have shown that prior knowledge determines what any design will do. It also covers language, accessibility requirements, and motivational starting points from Module 4.
  3. Task analysis decomposes the target performance into its components and their prerequisite relations. Cognitive task analysis, which elicits the reasoning and decision points that experts have automated and can no longer articulate, is the demanding version - and it matters precisely because of the expert blind spot from Lesson 2.
  4. Context analysis covers where learning will happen and where the performance must eventually happen, since the distance between the two is what transfer has to bridge.

A worked instance: a hospital asks for training because staff are not documenting a safety check. Needs analysis reveals that staff know the procedure perfectly and can demonstrate it on request, but the documentation form is on a different floor. No course will fix that. The output of analysis, in this case, is the recommendation not to build a course - which is a successful analysis, however unwelcome.

Key idea: Analysis comprises needs, learner, task, and context analysis, and its most valuable output is sometimes the finding that instruction is not the right remedy.

Linear on paper, iterative in practice

ADDIE is often drawn as a straight line, but skilled designers treat it as a cycle. Evaluation feeds back into analysis and design for the next version; a problem found during development sends you back to reconsider the design. Modern practice frequently uses rapid, iterative versions - building a rough prototype, testing it with a few real learners, and revising quickly - rather than perfecting each phase before the next. The value of ADDIE is not the illusion that design is linear; it is the guarantee that no essential question (especially analysis and evaluation, the two most often shortchanged) gets skipped.

PhaseCentral questionKey output
AnalysisWho, what gap, what constraints?Learner and needs analysis
DesignWhat objectives, assessments, sequence?Design blueprint
DevelopmentHow do we build it?Finished materials
ImplementationHow is it delivered?Live instruction
EvaluationDid it work; how to improve?Findings and revisions

Notice how the phases map onto this whole course: analysis draws on Module 4, design on objectives and Modules 2 and 3, evaluation on Module 7. ADDIE is the scaffold that holds the rest together.

What the middle phases actually produce

Design and development are often collapsed in practice, which is a mistake, because they fail in different ways. The output of design is a document that anyone else could build from: measurable objectives, the assessment that will demonstrate each one, a sequence with a stated rationale, the strategies chosen for each segment, and the media decisions with reasons. If the design document does not say why a topic comes third rather than first, or why a video was chosen over a text, those decisions will be made later by whoever happens to be building, on grounds of convenience.

The output of development is the artefacts themselves, and the discipline here is to build the cheapest thing that can be tested. Producing a polished animation before anyone has confirmed that learners understand the explanation is the classic and costly error, because polished artefacts resist revision - both technically and psychologically, since nobody wants to discard expensive work.

Implementation is not merely delivery either. It includes preparing whoever teaches or facilitates, arranging the environment, and anticipating the practical failures that sink otherwise sound designs: the room without the required software, the facilitator who has not seen the materials, the schedule that puts the practice session three weeks after the explanation. A design that assumes ideal conditions has not really been designed.

Key idea: Design produces a rationale-bearing blueprint, development should produce the cheapest testable artefact, and implementation must anticipate the practical conditions that sink sound designs.

Criticisms and alternatives

ADDIE attracts three recurring criticisms and each has produced an alternative worth knowing.

The first is that it encourages a waterfall in which each phase is completed before the next begins, so problems surface only at the end when they are most expensive to fix. Steven Tripp and Barbara Bichelmeyer proposed rapid prototyping as the remedy: build a rough working version early, put it in front of real learners, and let their difficulties drive design and analysis simultaneously. Modern variants including successive-approximation and agile approaches share this logic, and the empirical rationale is simple - learner behaviour reveals problems that no amount of front-end analysis anticipates.

The second criticism is that ADDIE is procedurally detailed and instructionally empty: it tells you to design, but not what good design looks like. David Merrill's first principles of instruction address that gap by extracting what a wide range of design theories agree on. Learning is promoted when learners tackle real-world problems; when existing knowledge is activated as a foundation; when new knowledge is demonstrated rather than only described; when learners apply it; and when they integrate it into their own world. Those five principles are a useful audit for any design that has passed through ADDIE's phases and still feels hollow.

The third criticism is that ADDIE says nothing about complex, integrated skills, which cannot be taught by decomposing them into parts and teaching each in isolation. Jeroen van Merrienboer's four-component instructional design model responds by organizing whole courses around sequences of increasingly complex whole tasks, supported by just-in-time procedural information and part-task practice for the components that must become automatic.

Key idea: ADDIE is criticized for encouraging a waterfall, for being instructionally empty, and for handling complex skills poorly, and rapid prototyping, Merrill's first principles, and 4C/ID respond to those criticisms.

Common misconceptions

  • ADDIE is a validated model. It is an organizing convention with no original source; the evidence lies in the choices made inside each phase.
  • Analysis means writing a needs survey. It comprises needs, learner, task, and context analysis, and may conclude that no course is warranted.
  • Every performance gap is a training gap. If people could do it if their life depended on it, the barrier is not knowledge.
  • Evaluation is the last phase. Formative evaluation runs throughout; leaving it to the end guarantees expensive surprises.
  • Following the phases produces good instruction. The phases guarantee only that no question is skipped, not that the answers are any good.

Recap

  • ADDIE names five phases - analysis, design, development, implementation, evaluation - and functions as a checklist rather than a theory.
  • It has no single author or founding document, having grown out of military and industrial instructional systems development.
  • Analysis contains four distinct investigations and may correctly conclude that instruction is not the answer.
  • Skilled designers treat the phases as a cycle with rapid prototyping and early learner testing.
  • Merrill's first principles supply the instructional substance ADDIE lacks: problem, activation, demonstration, application, integration.
  • Complex integrated skills need whole-task designs such as 4C/ID rather than decomposition into isolated parts.

Sources

  1. Merrill, M. D. (2002). First principles of instruction. Educational Technology Research and Development, 50(3), 43-59. link.springer.com
  2. Tripp, S. D., and Bichelmeyer, B. (1990). Rapid prototyping: An alternative instructional design strategy. Educational Technology Research and Development, 38(1), 31-44. link.springer.com
  3. Andrews, D. H., and Goodson, L. A. (1980). A comparative analysis of models of instructional design. Journal of Instructional Development, 3, 2-16. link.springer.com
  4. van Merrienboer, J. J. G., and Sweller, J. (2005). Cognitive load theory and complex learning: Recent developments and future directions. Educational Psychology Review, 17, 147-177. link.springer.com
  5. Wikipedia contributors. (n.d.). ADDIE model. Wikipedia. en.wikipedia.org
  6. Molenda, M. (2003). In search of the elusive ADDIE model. Performance Improvement, 42(5), 34-36. find source โ†—
  7. Branson, R. K., Rayner, G. T., Cox, J. L., Furman, J. P., King, F. J., and Hannum, W. H. (1975). Interservice Procedures for Instructional Systems Development (5 vols., TRADOC Pam 350-30). US Army Training and Doctrine Command. find source โ†—
Key terms
ADDIE
A five-phase instructional design framework: Analysis, Design, Development, Implementation, Evaluation.
Analysis phase
Determining learners, needs, the performance gap, and constraints before designing.
Design phase
Planning objectives, assessments, sequence, and strategies as a blueprint.
Development phase
Producing the actual instructional materials from the design.
Formative evaluation
Evaluation conducted during development to catch and fix problems early.
Iterative design
Building, testing, and revising in rapid cycles rather than perfecting each phase in sequence.

Backward Design and Alignment

  • Explain the three stages of backward design.
  • Define constructive alignment among objectives, assessment, and instruction.
  • Diagnose misalignment in a course.

Where ADDIE governs the overall process, backward design (also called Understanding by Design) governs the logic of planning a course or unit. Its insight is that most planning happens in the wrong order. Teachers often start from activities and content ("what will we cover, what will we do?") and only later ask how to test it. Backward design insists you start from the end and work backward, in three stages.

Key idea: Backward design reverses the usual planning order so that assessment is decided immediately after goals and before any activity is chosen.

The two failure modes it is designed to prevent

Grant Wiggins and Jay McTighe named the two habits their framework exists to correct, and recognizing them in your own planning is more useful than memorizing the stages. Activity-oriented design starts from things that would be engaging to do - the simulation, the debate, the field trip - and works outward, so the lesson is busy and enjoyable and it is never quite clear what learners were supposed to end up able to do. Coverage-oriented design starts from the material that must be got through - chapters one to fourteen by December - so the syllabus is complete and nobody has asked what understanding is supposed to survive it.

Both failures share a structure: the means have replaced the ends. Backward design's contribution is procedural rather than theoretical. By forcing the question "what evidence would convince me they got there?" before any activity is selected, it makes the ends impossible to skip.

Key idea: Backward design exists to prevent activity-oriented and coverage-oriented planning, in both of which the means silently replace the ends.

The three stages

  1. Identify desired results. First decide what learners should know and be able to do by the end - the goals and enduring understandings. This is where you write clear learning objectives (the next lesson's focus).
  2. Determine acceptable evidence. Before planning any lessons, decide how you will know the results were achieved: what assessment or performance would demonstrate the objective has been met? Designing the assessment second, not last, keeps the whole course honest.
  3. Plan learning experiences and instruction. Only now do you design the activities, readings, and practice - chosen specifically to prepare learners to succeed on that evidence and reach those results.

The reversal seems small but changes everything. When assessment is an afterthought, courses drift into "covering material" with tests that measure whatever happened to be emphasized. When assessment is designed right after objectives, every subsequent choice can be checked against a clear target.

Enduring understandings and essential questions

Stage one contains a distinction that does real work. Wiggins and McTighe sort content into three rings: what is worth being familiar with, what is important to know and do, and the small core of enduring understandings - the transferable big ideas that should still be usable years later. Most syllabi treat everything as ring two, which is why they are unteachable in the time available. Deciding explicitly that some material is ring one, to be encountered but not mastered, is what creates room for the depth that ring three requires.

Paired with these are essential questions: open, recurring, arguable questions that organize a unit and cannot be answered with a single fact. "What were the causes of the First World War?" is a topic. "How do we decide who is responsible when many parties contribute to a catastrophe?" is an essential question, and it makes the historical content into evidence for something. The related idea of the six facets of understanding - being able to explain, interpret, apply, take perspective, empathize, and self-assess - is a checklist for assessment design, because a unit that only ever asks for explanation has evidence of just one facet.

Key idea: Distinguishing enduring understandings from material merely worth being familiar with is what creates the time for depth, and essential questions turn topics into evidence for a transferable idea.

Constructive alignment

The principle underlying backward design is alignment: the learning objectives, the assessments, and the instructional activities should all point at the same target. In an aligned course, the objective states a capability, the assessment directly measures that capability, and the activities give learners practice in that capability. When these three fall out of alignment, instruction fails in predictable ways.

The term is John Biggs's, and his rationale is worth stating because it explains why alignment is more than tidiness. Biggs starts from the constructivist premise of Lesson 11: students construct meaning from what they do, not from what the teacher intends. The teacher therefore controls learning only indirectly, through the activities and assessments that determine what students actually spend their time doing. Alignment is the mechanism by which intentions reach behaviour. In an aligned system, the rational student who is optimizing for the assessment is thereby doing exactly what the objectives require - which is the only sustainable arrangement, because students will optimize for the assessment regardless.

That last point deserves emphasis, because it is the empirical engine of the whole idea. Research on student approaches to learning has repeatedly found that assessment demands drive study behaviour more powerfully than stated aims, syllabi, or exhortation - an effect sometimes called backwash. Benjamin Bloom's collaborator-era observation and later work on the hidden curriculum both converge on it: the assessment is the curriculum from the learner's point of view. A course whose stated objective is critical analysis and whose exam rewards memorized summaries has not failed to teach critical analysis by accident; it has actively taught students that memorized summaries are what counts.

Biggs pairs alignment with the SOLO taxonomy, which classifies the structural complexity of a response - from missing the point, to one relevant aspect, to several unrelated aspects, to an integrated relational account, to an extended abstract account that generalizes beyond the case. Its practical use is in writing outcome statements and rubric levels that describe qualitative differences in understanding rather than quantities of content covered.

Key idea: Biggs grounds alignment in the fact that students construct meaning from what they do, so assessment demands drive study behaviour more powerfully than stated aims and effectively define the curriculum.

Objective saysAssessment doesResult
Learners will analyze and critique argumentsMultiple-choice recall of definitionsMisaligned - the test rewards memorizing, not analyzing
Learners will solve novel problemsReproduce three memorized worked solutionsMisaligned - measures recall, not problem solving
Learners will write a persuasive essayWrite and revise a persuasive essay to a rubricAligned - the assessment is the capability

Why alignment is the master check

Alignment is the single most useful lens for critiquing any course, including your own. If the objective is at a high cognitive level (analyze, create) but the assessment only asks for recall, the objective is not really being taught or measured, whatever the syllabus claims. If activities drill one skill but the exam demands another, learners will be blindsided and will, rationally, study for the test rather than the stated goals.

Diagnosing a struggling course therefore starts with three questions asked side by side: What are the objectives? What do the assessments actually require? What do the activities actually practice? Wherever those three diverge is where the design is broken - and it explains why the next lesson, on writing precise objectives, is the foundation the entire structure rests on.

A worked audit

Consider a research methods course whose stated objective is that students will critique the design of a published study. The lecture schedule covers sampling, validity threats, and statistical tests. The weekly activities are multiple-choice quizzes on definitions. The final assessment is a closed-book exam with short-answer questions asking students to define terms and identify the test appropriate to a described dataset.

Laid side by side, the divergence is obvious. The objective is at the evaluate level; the activities practise recall; the assessment measures recall plus one application step. Nothing in the course ever asks a student to read a real study and judge it, so nothing in the course develops or measures the stated capability, and students who prepare rationally will never attempt it. The repair is not to add a lecture on critique. It is to make the assessment a critique - give students a published paper, ask for a structured evaluation against criteria - and then to rebuild the weekly activities as progressively less scaffolded practice at exactly that, starting from a worked critique modelled by the instructor. Notice how much of Modules 2 and 3 this pulls in: the worked example, the fading, the spaced retrieval of criteria.

Key idea: Auditing a course means laying objectives, activities, and assessment side by side and repairing from the assessment outward, since that is what students will optimize for.

Where backward design is criticized

Three objections are worth taking seriously. The first is that it presumes outcomes can be fully specified in advance, which fits procedural and conceptual goals better than open-ended, exploratory, or creative work where valuable learning is genuinely unforeseeable. The reasonable response is that even open work has some intended qualities - originality, defensibility, technical control - and those can be named without scripting the product.

The second is that rigorous alignment can slide into narrow teaching to the test. This is a real risk and Lesson 19 treats it as a validity failure: alignment is a virtue when the assessment is a faithful sample of the capability, and a vice when the assessment is a narrow proxy that can be gamed. Aligning to a bad assessment makes a course worse, not better.

The third is practical: full backward design is labour-intensive, and applied superficially it becomes a template-filling exercise producing documentation nobody uses. Practitioner accounts of implementing the framework consistently report that the value lies in the arguments the process forces - about what is genuinely enduring, and about what evidence would actually convince you - rather than in the completed planning documents.

Key idea: Backward design is weakest for genuinely open-ended learning, dangerous when aligned to a poor assessment, and valuable mainly for the arguments it forces rather than the documents it produces.

Common misconceptions

  • Backward design means writing the exam first and teaching to it. It means deciding what evidence would demonstrate the goal, then designing instruction that genuinely produces it.
  • Alignment is a paperwork exercise. It is the mechanism by which teacher intentions reach student behaviour, because students optimize for assessment.
  • Everything in a syllabus deserves equal depth. Distinguishing enduring understandings from familiarity-level content is what makes depth possible.
  • A misaligned course can be fixed by adding a lecture. The repair starts with the assessment, because that is what determines student effort.
  • Tighter alignment is always better. Aligning to a narrow or gameable assessment degrades the course.

Recap

  • Backward design proceeds from desired results to acceptable evidence to learning experiences, in that order.
  • It exists to prevent activity-oriented and coverage-oriented planning, where means displace ends.
  • Enduring understandings and essential questions concentrate limited time on transferable ideas.
  • Biggs grounds constructive alignment in students constructing meaning from what they do, which is driven by assessment.
  • Auditing a course means comparing objectives, activities, and assessment and repairing from the assessment outward.
  • The framework is weakest for open-ended learning and harmful when aligned to a poor assessment.

Sources

  1. Biggs, J. (1996). Enhancing teaching through constructive alignment. Higher Education, 32, 347-364. link.springer.com
  2. Newell, A. D., Foldes, C. A., Haddock, A. J., Ismail, N., and Moreno, N. P. (2024). Twelve tips for using the Understanding by Design curriculum planning framework. Medical Teacher, 46(1), 34-39. pmc.ncbi.nlm.nih.gov
  3. Crowe, A., Dirks, C., and Wenderoth, M. P. (2008). Biology in Bloom: Implementing Bloom's taxonomy to enhance student learning in biology. CBE-Life Sciences Education, 7(4), 368-381. pmc.ncbi.nlm.nih.gov
  4. Merrill, M. D. (2002). First principles of instruction. Educational Technology Research and Development, 50(3), 43-59. link.springer.com
  5. Wikipedia contributors. (n.d.). Backward design. Wikipedia. en.wikipedia.org
  6. Wiggins, G., and McTighe, J. (2005). Understanding by Design (2nd ed.). Association for Supervision and Curriculum Development. find source โ†—
  7. Cohen, S. A. (1987). Instructional alignment: Searching for a magic bullet. Educational Researcher, 16(8), 16-20. find source โ†—
Key terms
Backward design
Planning that starts from desired results, then evidence, then instruction (Understanding by Design).
Desired results
The goals and understandings learners should reach, defined first in backward design.
Acceptable evidence
The assessment or performance that would show an objective was met, designed before activities.
Constructive alignment
The matching of objectives, assessments, and activities to the same learning target.
Misalignment
A mismatch among objectives, assessment, and instruction that undermines learning.
Enduring understanding
A central, transferable idea worth retaining long after the course ends.

Writing Learning Objectives with Bloom's Taxonomy

  • Write measurable objectives using observable verbs.
  • Place objectives on the levels of Bloom's revised taxonomy.
  • Match objective level to appropriate assessment.

Everything in Modules 5 and 6 depends on one deceptively hard skill: writing a learning objective that states, in measurable terms, what a learner will be able to do. A good objective is the target that assessment measures and instruction serves. A vague objective makes alignment impossible, because you cannot align to a blur.

Key idea: A learning objective is the target that assessment measures and instruction serves, so vagueness in the objective propagates into every downstream decision.

Measurable, observable verbs

The classic mistake is to write objectives around invisible mental states: "students will understand photosynthesis" or "students will appreciate poetry." The trouble is that you cannot directly observe understanding or appreciation, so you cannot tell whether the objective was met. A strong objective uses an observable, measurable verb that names something the learner will actually do: explain, calculate, classify, compare, design, critique.

Compare "understand the causes of the war" with "list and explain three causes of the war and rank them by importance." The second can be assessed; the first cannot. A useful format specifies the audience, the behavior (observable verb), and often the conditions and criteria: "Given a data set, the learner will construct a correctly labeled histogram."

The three-part format comes from Robert Mager, whose 1962 handbook argued that an objective is useful only if it specifies the performance (what the learner will do), the conditions under which they will do it, and the criterion of acceptable performance. The conditions and criterion are what most drafts omit, and they carry more weight than they look. "Solve quadratic equations" is a different objective depending on whether a formula sheet is available, whether a calculator is permitted, and whether three out of five correct is enough. Those choices determine the assessment, the practice, and often the amount of instruction required.

Mager's approach has been criticized, fairly, for encouraging trivial objectives: performances that are easy to observe and specify are not always the ones that matter, and a course written entirely in behavioural terms can end up optimizing for what is measurable. The workable compromise, and the one most institutions use, is to state a small number of substantial objectives with observable verbs while accepting that some valuable outcomes - a taste for the subject, intellectual humility - are aims rather than objectives, and should be named as such rather than either faked into behavioural language or quietly dropped.

Key idea: Mager's format specifies performance, conditions, and criterion, and the frequently omitted conditions and criterion largely determine what the assessment and practice must look like.

Bloom's revised taxonomy

Bloom's taxonomy classifies cognitive objectives by complexity, and its widely used revised version arranges six levels from lower-order to higher-order thinking. Each is paired with characteristic verbs.

LevelWhat the learner doesSample verbs
RememberRecall facts and basic conceptslist, define, name, recall
UnderstandExplain ideas in their own wordsexplain, summarize, paraphrase, classify
ApplyUse knowledge in new situationssolve, use, calculate, demonstrate
AnalyzeBreak into parts and see relationshipscompare, contrast, differentiate, examine
EvaluateJudge based on criteriacritique, justify, assess, defend
CreateCombine parts into something newdesign, compose, construct, propose

Two cautions keep the taxonomy from being misused. First, the levels are not a strict ladder that must be climbed in fixed order, nor is "higher" always better - a course genuinely needs some remember-level objectives, because you cannot analyze what you cannot recall, and fluent lower-order knowledge frees working memory for higher-order work. The point is to be intentional about the level you are targeting, not to chase the top of the pyramid. Second, the level of an objective must match the level of its assessment (the alignment principle): an "evaluate" objective demands an assessment that requires judgment, not a recall quiz.

What the 2001 revision actually changed

Because the revised taxonomy is usually presented as a relabelled pyramid, it is worth knowing what David Krathwohl and Lorin Anderson's revision of Benjamin Bloom's 1956 handbook actually did. Three changes matter. The categories became verbs rather than nouns, reflecting that they name cognitive processes rather than types of content. The top two categories were reordered and renamed, so that synthesis became create and moved above evaluate. And most importantly, a single ordered list became a two-dimensional table: one axis for the cognitive process (remember through create) and a second for the knowledge dimension - factual, conceptual, procedural, and metacognitive knowledge.

That second axis is the part most often ignored and the part with the most practical value. It makes visible that "apply procedural knowledge" and "apply conceptual knowledge" are different objectives requiring different instruction, and it gives metacognitive knowledge - knowing about strategies, about the task, and about oneself as a learner - a formal place in a curriculum. Everything in Lesson 10 about self-regulation sits in that cell, and courses that never write an objective there tend never to teach it.

Two further cautions come from research on how the taxonomy is used in practice. Studies of instructor coding show that assigning a level to a question is much less reliable than it appears, because the same question sits at different levels depending on what the learner has already been taught: a problem that requires analysis the first time is a remember-level task once the solution has been rehearsed. Anne Crowe and colleagues, developing a discipline-specific version for biology, found that agreement improved substantially when the taxonomy was operationalized with subject-specific examples rather than generic verb lists. So treat verb tables as prompts, not classifiers, and always ask what level the task is for this learner given what they have already done.

Key idea: The 2001 revision turned Bloom's single list into a two-dimensional table crossing cognitive process with factual, conceptual, procedural, and metacognitive knowledge, and level assignments depend on what the learner has already been taught.

The other two domains

Bloom's committee described three domains, and the cognitive one has so dominated practice that the others are often forgotten. The affective domain, set out by Krathwohl, Bloom and Masia in 1964, orders dispositions from receiving and responding, through valuing, to organizing values and finally characterizing oneself by them. It matters for exactly the outcomes that resist behavioural phrasing - professional ethics, willingness to revise one's view, care for accuracy - and it at least gives a vocabulary for naming them and for noticing when a curriculum claims them and never addresses them. The psychomotor domain, developed by others including Simpson and Dave, orders physical skill from imitation through manipulation and precision to naturalization, and it remains the natural language for clinical, laboratory, athletic, and craft instruction.

Neither is as well developed as the cognitive taxonomy, and neither should be forced. But a nursing curriculum whose objectives are entirely cognitive has failed to say anything about sterile technique or about how a student should respond to a distressed patient, and the omission is not a formatting problem.

Key idea: Affective and psychomotor taxonomies give a vocabulary for dispositions and physical skills that a purely cognitive objective set silently omits.

Four recurring errors

Most weak objectives fail in one of four ways, and each has a quick repair.

  • Describing teaching rather than learning. "This unit will cover three theories of motivation" states what the instructor does. Objectives are always about what the learner will be able to do.
  • Bundling several objectives into one. "Students will analyse, evaluate, and design experiments" contains three objectives at three levels, and no single assessment can validly measure it. Split it.
  • Using a high-level verb for a low-level task. Writing "evaluate" when the assessment asks students to recall a textbook's evaluation is more damaging than writing "remember" honestly, because it disguises the misalignment from everyone including the designer.
  • Writing more objectives than the course can serve. Twenty objectives for a twelve-week course guarantees that most receive no real assessment or practice, which returns you to coverage-oriented design by another route.

Key idea: Objectives commonly fail by describing teaching, bundling multiple targets, inflating the verb above the actual task, or proliferating beyond what the course can genuinely assess.

From verb to assessment

Because the verb encodes the cognitive level, it also tells you what kind of assessment fits. "Recall the parts of a cell" is honestly assessed by a labeling task; "design an experiment to test a hypothesis" cannot be - it needs a task where the learner actually designs something. This is the practical bridge into Module 6: choosing an objective's verb is simultaneously choosing, in outline, how you will assess it. Write the verb carelessly and every downstream decision inherits the confusion; write it precisely and the assessment almost designs itself.

A short worked sequence shows the chain running end to end. Start with a vague aim: "students will understand experimental control." Ask what a person who understood it could do that a person who did not could not. One answer: given a described study, they could identify the uncontrolled variable that threatens its conclusion. That yields the objective - "given a one-paragraph description of a study, the learner will identify the principal uncontrolled variable and explain the threat it poses to the study's conclusion, for four of five cases" - which specifies performance, conditions, and criterion, and sits at the analyze level.

The assessment is now almost written: five short study descriptions with a two-part response. The instruction follows as well: worked examples of the analysis, faded to completion problems where the variable is identified but the threat must be explained, then interleaved practice mixing threat types so learners must discriminate among them (Lesson 7), spaced across the term with cumulative retrieval (Lesson 6). This is what it means to say the objective is the foundation - it is the only decision that constrains all the others.

Common misconceptions

  • Higher levels of the taxonomy are better. Fluent lower-order knowledge is a prerequisite for higher-order work and frees the capacity it requires.
  • The taxonomy is a strict ladder. It classifies processes; it does not prescribe a fixed teaching order.
  • A question has an inherent Bloom level. The level depends on what the learner has already done; a rehearsed analysis becomes recall.
  • The revision merely renamed the categories. It added a whole knowledge dimension, including metacognitive knowledge.
  • Everything valuable can be written as a measurable objective. Some outcomes are legitimately aims, and should be named as aims rather than faked into behavioural language.

Recap

  • Objectives must use observable verbs, because invisible states such as understanding cannot be assessed directly.
  • Mager's format adds conditions and criterion, which largely determine what the assessment and practice must be.
  • Bloom's revised taxonomy classifies six cognitive processes from remember to create.
  • The 2001 revision added a knowledge dimension covering factual, conceptual, procedural, and metacognitive knowledge.
  • Level assignments are relative to the learner's prior experience, and generic verb tables classify unreliably.
  • Choosing the verb effectively chooses the assessment, which is why the objective constrains every downstream decision.

Sources

  1. Adams, N. E. (2015). Bloom's taxonomy of cognitive learning objectives. Journal of the Medical Library Association, 103(3), 152-153. pmc.ncbi.nlm.nih.gov
  2. Crowe, A., Dirks, C., and Wenderoth, M. P. (2008). Biology in Bloom: Implementing Bloom's taxonomy to enhance student learning in biology. CBE-Life Sciences Education, 7(4), 368-381. pmc.ncbi.nlm.nih.gov
  3. Biggs, J. (1996). Enhancing teaching through constructive alignment. Higher Education, 32, 347-364. link.springer.com
  4. Wikipedia contributors. (n.d.). Bloom's taxonomy. Wikipedia. en.wikipedia.org
  5. Newell, A. D., Foldes, C. A., Haddock, A. J., Ismail, N., and Moreno, N. P. (2024). Twelve tips for using the Understanding by Design curriculum planning framework. Medical Teacher, 46(1), 34-39. pmc.ncbi.nlm.nih.gov
  6. Anderson, L. W., and Krathwohl, D. R. (Eds.). (2001). A Taxonomy for Learning, Teaching, and Assessing: A Revision of Bloom's Taxonomy of Educational Objectives. Longman. find source โ†—
  7. Mager, R. F. (1962). Preparing Instructional Objectives. Fearon Publishers. find source โ†—
Key terms
Learning objective
A statement of what a learner will be able to do, in measurable, observable terms.
Observable verb
An action verb naming something a learner does that can be seen and assessed.
Bloom's taxonomy
A classification of cognitive objectives by complexity, from remembering to creating.
Lower-order thinking
Remember and understand levels: recall and basic comprehension.
Higher-order thinking
Analyze, evaluate, and create levels: complex reasoning and production.
Criterion
The standard of performance a learner must meet, often stated in an objective.

Module 6: Assessment, Feedback, and Multimedia Design

Designing assessments that measure the right thing, feedback that changes performance, and multimedia that respects cognition.

Designing Assessments: Formative, Summative, Validity, and Reliability

  • Distinguish formative and summative assessment by purpose.
  • Define validity and reliability and why both matter.
  • Choose assessment formats aligned to objectives.

Assessment is how we find out what learners have actually learned - and, done well, it is also one of the most powerful ways to produce learning, as the testing effect showed. Designing it well means being clear about its purpose and its quality.

Key idea: Assessment is both the means of finding out what has been learned and, done well, one of the most powerful ways of producing learning.

Formative versus summative

The most important distinction is by purpose. Formative assessment is assessment for learning: low-stakes checks used during instruction to reveal where learners are and to guide the next teaching move - a quick quiz, a show of hands, a one-minute written summary, a problem worked on the board. Its results feed back into instruction and into the learner's own studying (recall self-regulation).

Summative assessment is assessment of learning: higher-stakes judgments of achievement at the end of a unit or course - a final exam, a capstone project, a certification test. The same instrument can serve either purpose depending on how it is used; a practice test used to guide study is formative, the same test used to assign a grade is summative. Strong instruction is rich in formative assessment, because frequent low-stakes feedback both improves learning directly and prevents nasty surprises at the summative stage.

How strong is the formative assessment evidence?

Formative assessment is one of the most confidently promoted ideas in education, and it is worth knowing where the confidence comes from and how well it has held up. Paul Black and Dylan Wiliam's 1998 review of some 250 studies concluded that improving classroom formative assessment produced substantial learning gains, and the effect sizes they mentioned - roughly 0.4 to 0.7 - have been quoted in policy documents ever since.

Two subsequent developments complicate that picture and are essential for a graduate reader. Neal Kingston and Brooke Nash conducted a formal meta-analysis restricted to studies with adequate designs and found a much smaller average effect, around 0.20, with wide variation by subject and by what "formative assessment" actually meant in each study. Randy Bennett's critique made the conceptual point underneath: formative assessment is not a single intervention but a loose family of practices ranging from a two-second show of hands to a full diagnostic system, and averaging across them produces a number that describes nothing in particular.

The defensible position is therefore narrower and more useful than the slogan. What has good support is the specific mechanism: frequent low-stakes retrieval with corrective feedback improves retention (Module 3), and diagnostic information that actually changes the next teaching decision improves outcomes. What is not supported is the idea that adding assessment events, by themselves, produces large gains. The question to ask of any formative practice is not "are we assessing formatively?" but "what did we learn from this, and what will we do differently as a result?" If there is no answer, the assessment was summative in disguise and merely cost class time.

Key idea: Black and Wiliam's widely quoted large effects for formative assessment shrink to roughly 0.20 in stricter meta-analysis, and what is well supported is the mechanism - retrieval with feedback and diagnostic information that changes the next decision - not the label.

Validity and reliability

Two technical criteria determine whether an assessment is any good. Validity asks whether the assessment measures what it claims to measure. A test of "scientific reasoning" that mostly rewards reading speed has a validity problem. Validity is bound up with alignment (Module 5): an assessment is valid for an objective only if it actually taps that objective's capability. Reliability asks whether the assessment yields consistent results - across occasions, versions, and graders.

A rubric that different graders apply to give wildly different scores is unreliable; a test whose score depends on which random form a student got is unreliable. The two are distinct and both necessary. An assessment can be reliable but not valid (it consistently measures the wrong thing, like a precise scale that is miscalibrated) and it cannot be truly valid without being reasonably reliable (random, inconsistent results cannot faithfully reflect a real capability). A good assessment must be both.

PropertyQuestion it answersFailure looks like
ValidityDoes it measure the intended capability?A reasoning test that mostly measures reading speed
ReliabilityAre the results consistent?Graders disagree; scores swing by test version

Validity is a property of interpretations, not of tests

The everyday phrasing above - "a valid test" - is a useful simplification that professional practice abandoned decades ago, and the correction changes how you work. Samuel Messick argued, and the Standards for Educational and Psychological Testing now codify, that validity is a property of the interpretations and uses of scores rather than of the instrument. The same twenty-item quiz may support a valid inference that a student has learned this week's vocabulary and an invalid inference that they are ready to practise as a translator. Asking "is this test valid?" is therefore incomplete; the question is always "valid for what claim, about whom, for what decision?"

It follows that the older habit of listing separate types of validity - content validity, criterion validity, construct validity, as though a test could possess some and lack others - has been replaced by a single unified concept supported by several kinds of evidence: evidence based on test content, on response processes, on internal structure, on relations to other variables, and on the consequences of use. David Cook and Rose Hatala's primer sets this out as a practical workflow: state the intended interpretation, list the assumptions it depends on, and gather the evidence that most threatens the weakest assumption.

Two threats to validity are worth naming because they recur in ordinary courses. Construct-irrelevant variance is anything that affects scores but is not part of the target capability - reading load in a mathematics word problem, time pressure in an assessment of accuracy, familiarity with an interface. Construct underrepresentation is failing to sample the capability adequately, as when a course on scientific reasoning tests only vocabulary. Most misalignment from Lesson 13 is construct underrepresentation wearing another name.

Key idea: Validity belongs to score interpretations rather than to instruments, is supported by several kinds of evidence rather than divided into types, and is most often threatened by construct-irrelevant variance and construct underrepresentation.

Where unreliability comes from

Reliability is easier to compute than to think about, and three of its sources behave differently. Internal consistency, usually reported as coefficient alpha, indexes how far items measure the same thing; it rises mechanically with the number of items and with item homogeneity, so a long test of one narrow skill will look highly reliable regardless of whether that skill is the one you wanted. Test-retest reliability indexes stability over occasions, which matters when a decision depends on the assumption that performance is not fluctuating. Inter-rater reliability indexes agreement between judges, and it is the dominant source of error in essays, projects, and performance tasks.

Two practical implications follow. First, short assessments are unreliable almost by construction: a single item is a coin flip, and a four-item quiz cannot support any individual judgement, however carefully written. This is a strong argument for many small low-stakes assessments rather than few large ones - the aggregate is more dependable and, from Module 3, the frequent retrieval is itself instructive. Second, inter-rater reliability is fixable by design: explicit criteria, worked exemplars at each performance level, calibration sessions where markers score the same scripts and compare, and double marking of a sample.

Key idea: Reliability has distinct sources - internal consistency, stability, and rater agreement - and the practical remedies are more assessment points and explicit criteria with marker calibration.

Choosing the format

No format is best in the abstract; the right choice follows from the objective's cognitive level. Selected-response items (multiple choice, true/false) can efficiently and reliably sample broad knowledge and even some higher-order reasoning if carefully written, but they cannot directly measure whether a learner can produce or create. Constructed-response and performance tasks (essays, projects, designs, demonstrations) can measure higher-order capabilities directly, at the cost of more time and lower grading reliability unless a clear rubric is used.

The design rule is the alignment rule once more: pick the format that validly measures the specific objective. A "recall" objective is fine to assess with multiple choice; a "design an experiment" objective requires a task where the learner designs an experiment, scored against explicit criteria. Match the tool to the target, and use rubrics to shore up the reliability of open-ended judgments.

One widespread belief deserves correcting: that multiple-choice items can only test recall. They can test considerably more, but only if they are built for it. Items that present a novel scenario and ask which interpretation follows, which next step is indicated, or which of four proposed explanations is inconsistent with given data require analysis and evaluation, not retrieval of a memorized sentence. Maria Xiromeriti and Philip Newton found that structured guidance for writing higher-order questions - centring items on a problem to be solved rather than a fact to be recalled - measurably improved the cognitive level of the items produced.

What degrades multiple-choice assessment is well documented. Item-writing flaws such as implausible distractors, options of very unequal length, grammatical cues, "all of the above," negatively worded stems, and overlapping options let test-wise students score without knowing the content, which is construct-irrelevant variance in its purest form. Studies analysing flawed items in real examinations find that they distort item difficulty and discrimination, so the flaws are not cosmetic. Writing four genuinely plausible distractors - ideally drawn from the misconceptions in Lesson 11 - is the hardest and most valuable part of the work.

Key idea: Well-constructed multiple-choice items can assess analysis and evaluation, while common item-writing flaws let test-wise students succeed without knowing the content.

Common misconceptions

  • Formative assessment reliably produces large gains. Stricter meta-analysis puts the average nearer 0.20; the mechanism matters more than the label.
  • A test is either valid or not. Validity attaches to a specific interpretation and use of scores, not to the instrument.
  • There are several separate types of validity. Modern practice treats validity as one concept supported by different kinds of evidence.
  • High reliability means a good test. A long, homogeneous test of the wrong construct is highly reliable and useless.
  • Multiple-choice items can only assess recall. Scenario-based items can require analysis and evaluation if written deliberately.

Recap

  • Formative and summative assessment differ by purpose and use, not by instrument.
  • The strong claims for formative assessment shrink under stricter analysis, but retrieval with feedback and decision-changing diagnosis are well supported.
  • Validity is a property of score interpretations and uses, supported by evidence about content, response processes, internal structure, relations to other variables, and consequences.
  • Construct-irrelevant variance and construct underrepresentation are the two threats most often seen in ordinary courses.
  • Reliability arises from internal consistency, stability, and rater agreement, and short assessments are unreliable by construction.
  • Format should follow the objective's cognitive level, and well-written selected-response items can reach beyond recall.

Sources

  1. Cook, D. A., and Hatala, R. (2016). Validation of educational assessments: A primer for simulation and beyond. Advances in Simulation, 1, 31. pmc.ncbi.nlm.nih.gov
  2. Xiromeriti, M., and Newton, P. M. (2024). Solving not answering: Validation of guidance for writing higher-order multiple-choice questions in medical science education. Medical Science Educator, 34, 1469-1477. pmc.ncbi.nlm.nih.gov
  3. Pham, H., Besanko, J., and Devitt, P. (2018). Examining the impact of specific types of item-writing flaws on student performance and psychometric properties of the multiple choice question. MedEdPublish, 7, 225. pmc.ncbi.nlm.nih.gov
  4. Yang, C., Luo, L., Vadillo, M. A., Yu, R., and Shanks, D. R. (2021). Testing (quizzing) boosts classroom learning: A systematic and meta-analytic review. Psychological Bulletin, 147(4), 399-435. pubmed.ncbi.nlm.nih.gov
  5. Hamilton, L., Halverson, R., Jackson, S., Mandinach, E., Supovitz, J., and Wayman, J. (2009). Using Student Achievement Data to Support Instructional Decision Making (NCEE 2009-4067). Institute of Education Sciences, What Works Clearinghouse. ies.ed.gov
  6. Black, P., and Wiliam, D. (1998). Assessment and classroom learning. Assessment in Education: Principles, Policy and Practice, 5(1), 7-74. find source โ†—
  7. Kingston, N., and Nash, B. (2011). Formative assessment: A meta-analysis and a call for research. Educational Measurement: Issues and Practice, 30(4), 28-37. find source โ†—
Key terms
Assessment
The process of gathering evidence about what learners have learned.
Formative assessment
Low-stakes assessment for learning, used during instruction to guide next steps.
Summative assessment
Higher-stakes assessment of learning, judging achievement at the end.
Validity
The degree to which an assessment measures what it claims to measure.
Reliability
The degree to which an assessment yields consistent results.
Rubric
An explicit set of criteria and performance levels used to score open-ended work reliably.

Feedback That Improves Learning

  • Explain the three questions effective feedback answers.
  • Distinguish feedback about the task, process, and self.
  • Time and frame feedback to change performance.

Feedback is repeatedly identified as among the most powerful influences on learning - but the same research shows its effects are highly variable: some feedback helps a great deal, some does nothing, and some actively harms. The difference lies in what the feedback addresses and how it is delivered. Simply giving more feedback is not the goal; giving the right kind is.

The variability is not a hedge; it is the central finding, and the numbers are worth carrying. Avraham Kluger and Angelo DeNisi's meta-analysis of 131 controlled studies yielding 607 effect sizes found an average effect on performance of about d = 0.41 - and that over a third of the effects were negative, meaning the feedback made performance worse than no feedback at all. John Hattie and Helen Timperley's later synthesis reported a higher average, around d = 0.79, which is the figure most often quoted in staff development. Benedikt Wisniewski, Klaus Zierer and Hattie revisited the question with a fresh meta-analysis in 2020 and found an overall effect of d = 0.48 once extreme values were removed, with 17 percent of effects still negative and enormous heterogeneity remaining.

Take those three numbers together and the honest summary is this: feedback is on average one of the more powerful interventions available, its average conceals effects ranging from strongly helpful to actively harmful, and the published averages themselves vary by a factor of nearly two depending on which studies are included. Anyone who tells you feedback has an effect size of 0.79 and stops there has told you the least useful part of the finding. The design question is which kind, delivered how, to whom.

Key idea: Meta-analytic averages for feedback range from about d = 0.4 to 0.8, and roughly a fifth to a third of measured effects are negative, so the variability rather than the average is the finding that matters.

Three questions good feedback answers

A widely used model holds that effective feedback answers three questions from the learner's point of view.

  1. Where am I going? (Feed up) - what is the goal or standard I am aiming for? Feedback only makes sense against a clear objective, which is why Module 5's work pays off here.
  2. How am I doing? (Feed back) - what is the gap between my current performance and that goal?
  3. Where to next? (Feed forward) - what specific action will close the gap? This third question is the one weak feedback most often omits, yet it is what actually changes future performance.

Feedback that only reports a score answers none of these well. Feedback that says "your thesis states a claim but two of your three body paragraphs do not support it - revise them to give evidence for the thesis" answers all three and tells the learner exactly what to do.

Royce Sadler stated the underlying requirement even more starkly in 1989, and his formulation is the most useful single test of any feedback system. For feedback to improve work, the learner must (1) possess a concept of the standard being aimed for, (2) be able to compare their current performance with that standard, and (3) be able to take action to close the gap. If any of the three is missing, the feedback cannot function, however well written. This explains a familiar frustration: students who receive detailed comments and do not improve frequently fail at step one - they cannot see what the comment refers to because they do not yet hold the standard - or at step three, because the work is finished and there is nothing left to act on.

The practical corollaries are unusually concrete. Show exemplars at several quality levels and have learners judge them before they write, so the standard exists before the feedback arrives. Have learners self-assess against the criteria before submitting, so the comparison in step two is one they have attempted themselves. And always leave something to act on - a revision, a resubmission, a next task of the same type - so step three is available.

Key idea: Sadler's three conditions - holding the standard, comparing performance to it, and being able to act - must all be met, and feedback on finished work with no further action fails the third by construction.

Task, process, and self

Feedback can be aimed at different levels, and the level matters. Feedback about the task ("this answer is incorrect because...") corrects specific work. Feedback about the process ("your strategy of checking units would catch this class of error") is often the most powerful, because it transfers to future tasks and builds the self-regulation of Module 4. Feedback about self-regulation ("you caught your own mistake by rereading - keep doing that") strengthens the learner's monitoring.

In contrast, feedback about the self as a person - unfocused praise like "great job, you're so clever" - is generally the least effective for learning and, as the mindset lesson warned, can even backfire by directing attention to the ego rather than the work. The practical guidance is to favor specific task and process feedback over global praise or blame.

Kluger and DeNisi's feedback intervention theory explains why self-level feedback can be actively harmful, and the mechanism is worth understanding rather than memorizing. Attention is limited and hierarchically organized. Feedback that draws attention to the task details prompts task-relevant processing. Feedback that draws attention upward to the self - "you are the sort of person who does badly at this" - prompts self-relevant processing such as rumination, comparison with others, or affect regulation, which competes for the same limited capacity that the task requires. This is Module 1 and Module 2 reappearing inside a social interaction: feedback that redirects working memory away from the work degrades the work.

Two related findings follow the same logic. Grades attached to comments suppress the effect of the comments, because the mark is a self-level signal that captures attention first and the comments then go unread. And normative feedback that positions a learner relative to peers reliably produces worse outcomes than criterion-referenced feedback that positions them relative to the standard, for the same reason.

Key idea: Feedback that directs attention to the self rather than the task diverts limited capacity into self-relevant processing, which is why person-level praise, marks attached to comments, and normative comparisons can all reduce performance.

Timing, specificity, and load

Three further principles sharpen feedback. Specificity: actionable, focused feedback beats vague comments ("be clearer") that give the learner nothing to act on. Timing: feedback generally helps most when it still connects to the effort and can be used before the learner moves on, though for fostering durable retrieval a slight delay can sometimes be beneficial - the key is that the learner has a real opportunity to act on it.

Manageability: a page bleeding with thirty corrections overwhelms working memory (Module 2) and gets ignored; a few prioritized, high-leverage points are more likely to be understood and used. Finally, feedback only helps if it is received and acted upon. Building in a step where learners must respond to feedback - revise the draft, redo the problem, explain what they will change - turns feedback from a comment into a cause of improvement, and closes the loop that the three questions opened.

Timing, and the evidence for it

The timing question is genuinely unsettled, and the apparent contradictions resolve once you separate the situations. For simple factual retrieval, delaying corrective feedback slightly can improve later retention, plausibly because the delay adds spacing and the learner must retrieve the item again when the answer arrives. For complex tasks where the learner is building a procedure, immediate feedback prevents the practice of errors and stops learners flailing. And for anything where the learner will not otherwise return to the material, feedback delayed past the point of use is not delayed feedback but no feedback.

The dependable rule is therefore not "immediate" or "delayed" but in time to be used. A comment returned after the module has ended, on work that will not be revisited, cannot satisfy Sadler's third condition no matter how good it is. Where resources are scarce, moving effort from end-of-module comments to mid-module feedback on drafts is usually the single highest-value change available.

Key idea: Slight delays can help simple retrieval and immediacy helps complex procedural practice, but the governing rule is that feedback must arrive in time to be acted upon.

Building a feedback system rather than writing comments

The Education Endowment Foundation's guidance on teacher feedback makes a point that individual comment-writing advice misses: feedback is a system-level design problem. Their recommendations begin before any feedback is given - lay the groundwork with high-quality initial instruction, since feedback cannot repair an explanation that never worked - and then deliver feedback that focuses on moving learning forward, plan how learners will use it, and implement it in a way that is sustainable for teachers. That last criterion is not an afterthought. A feedback policy requiring lengthy written comments on every piece of work reliably degrades into ritual compliance, and the ritual is what produces the null and negative effects in the literature.

A worked redesign: instead of marking thirty essays with paragraphs of individual comment, mark them for two or three recurring problems, teach a fifteen-minute whole-class session addressing the most common one with exemplars, give each student a one-line personal note identifying which of the recurring problems is theirs, and require a targeted revision of one paragraph. Total teacher time falls; the standard becomes visible; every learner acts on the feedback. Each of Sadler's three conditions is met, which the paragraphs of individual comment did not manage.

Key idea: Feedback is a system design problem, and approaches that are unsustainable for teachers decay into ritual, which is a major source of the null and negative effects in the research.

Common misconceptions

  • More feedback is better. A third of measured feedback effects are negative, and volume is not the active ingredient.
  • Praise is harmless encouragement. Person-level praise redirects attention to the self and competes with task processing.
  • A grade plus comments gives students both. The mark captures attention first and the comments typically go unread.
  • Immediate feedback is always best. The rule is in time to be used; slight delays can help simple retrieval.
  • Detailed comments on final work close the loop. Without an opportunity to act, Sadler's third condition fails and the comments cannot function.

Recap

  • Feedback averages between about d = 0.4 and 0.8 across meta-analyses, with a substantial minority of negative effects.
  • Effective feedback answers where the learner is going, how they are doing, and what to do next.
  • Sadler's three conditions require the learner to hold the standard, compare against it, and be able to act.
  • Task and process feedback outperform person-level feedback, which diverts capacity into self-relevant processing.
  • Timing should be governed by usability rather than by a fixed immediate-or-delayed rule.
  • Sustainable, system-level feedback design beats individually heroic commenting that decays into ritual.

Sources

  1. Wisniewski, B., Zierer, K., and Hattie, J. (2020). The power of feedback revisited: A meta-analysis of educational feedback research. Frontiers in Psychology, 10, 3087. pmc.ncbi.nlm.nih.gov
  2. Sadler, D. R. (1989). Formative assessment and the design of instructional systems. Instructional Science, 18, 119-144. link.springer.com
  3. Education Endowment Foundation. (2021). Teacher Feedback to Improve Pupil Learning: Guidance Report. EEF. educationendowmentfoundation.org.uk
  4. Education Endowment Foundation. (n.d.). Feedback. Teaching and Learning Toolkit. EEF. educationendowmentfoundation.org.uk
  5. Yang, C., Luo, L., Vadillo, M. A., Yu, R., and Shanks, D. R. (2021). Testing (quizzing) boosts classroom learning: A systematic and meta-analytic review. Psychological Bulletin, 147(4), 399-435. pubmed.ncbi.nlm.nih.gov
  6. Kluger, A. N., and DeNisi, A. (1996). The effects of feedback interventions on performance: A historical review, a meta-analysis, and a preliminary feedback intervention theory. Psychological Bulletin, 119(2), 254-284. find source โ†—
  7. Hattie, J., and Timperley, H. (2007). The power of feedback. Review of Educational Research, 77(1), 81-112. find source โ†—
Key terms
Feed up
Feedback clarifying the goal or standard the learner is aiming for.
Feed back
Feedback on the gap between current performance and the goal.
Feed forward
Feedback specifying the next action to close the gap; often the most neglected part.
Process feedback
Feedback about the strategy used, which tends to transfer to future tasks.
Self-level feedback
Unfocused praise or blame about the person, generally least effective for learning.
Actionable feedback
Specific feedback the learner can directly act on to improve.

Multimedia Learning Principles

  • State the core assumptions behind multimedia learning.
  • Apply key multimedia principles to design materials.
  • Explain each principle in terms of cognitive load.

Much modern instruction is multimedia - combinations of words and graphics in slides, videos, and e-learning. A well-developed research program has produced a set of evidence-based multimedia principles for designing such materials, and the striking thing is that every one of them follows directly from the cognitive architecture in Modules 1 through 3. They rest on three assumptions you already know: information travels through two channels (dual coding), each channel has limited capacity (working memory), and meaningful learning requires active processing. The principles are simply the design rules that keep material inside those limits.

Key idea: Every multimedia principle is a consequence of dual channels, limited capacity, and the need for active processing, so the principles are derivations rather than a list to memorize.

Five processes and three kinds of demand

Richard Mayer's cognitive theory of multimedia learning specifies what a learner must actually do with a narrated diagram, and the specification is what makes the principles predictable rather than arbitrary. Meaningful learning requires five processes: selecting relevant words from the narration, selecting relevant images from the graphic, organizing the selected words into a verbal model, organizing the selected images into a pictorial model, and integrating the two models with each other and with prior knowledge. The last step is where understanding happens, and it is the step most often starved of capacity.

Against those processes stand three kinds of demand on the same limited workspace. Extraneous processing is effort spent on things that do not serve the goal, and it comes from the design. Essential processing is the effort the material's own complexity requires. Generative processing is the effort of organizing and integrating - the productive kind. The principles group into three families accordingly: reduce extraneous processing, manage essential processing, and foster generative processing. That is precisely the intrinsic, extraneous, and germane distinction from Module 2 restated for message design, and Mayer's grouping is a good example of theory earning its keep by generating rules rather than merely labelling findings.

Key idea: Meaningful multimedia learning requires selecting, organizing, and integrating verbal and pictorial material, and the principles are organized by whether they cut extraneous demand, manage essential demand, or promote generative processing.

Reducing extraneous processing

  • Coherence principle: exclude material that does not serve the goal. Interesting but irrelevant stories, decorative images, and background music (seductive details) add extraneous load and depress learning. Less is more.
  • Signaling principle: highlight the essential material and its structure (headings, arrows, bold key terms, cues) so attention lands in the right place without a costly search.
  • Redundancy principle: do not add on-screen text that merely duplicates the narration (the redundancy effect from Module 2). Narration plus graphic usually beats narration plus graphic plus identical text.
  • Spatial and temporal contiguity: place corresponding words and pictures near each other on the page and present them at the same time rather than separated (the split-attention fix). A label beside the part it names beats a label in a distant key.

Two of these have been examined meta-analytically, which is worth knowing because it separates the well-evidenced principles from the merely plausible ones. NarayanKripa Sundararajan and Olusola Adesope synthesized the seductive details literature and found that adding interesting but irrelevant material reliably depressed both retention and transfer, with the harm greater for transfer - the outcome that depends most on the integration step above. Noah Schroeder and Ada Cenkci's meta-analysis of spatial contiguity, discussed in Lesson 5, found integrated layouts consistently superior to separated ones, with larger benefits for complex material and lower-prior-knowledge learners.

The coherence principle is the hardest of these to accept, because it contradicts a strong professional instinct. A designer asked to improve a dull explainer will reach for a story, a striking image, or music. The evidence says that when those additions do not carry the content's structure, they reduce learning even while raising ratings of enjoyment - a divergence between satisfaction and learning that Lesson 18 treats as a systematic hazard of evaluation.

Key idea: Seductive details reliably depress retention and especially transfer, and integrated layouts reliably beat separated ones, so the coherence and contiguity principles are among the best-evidenced in the set.

Managing essential processing

  • Segmenting principle: break a complex lesson into learner-paced segments rather than one continuous torrent, so working memory can process each chunk before the next arrives.
  • Pre-training principle: teach the names and characteristics of key components before explaining the whole process, so learners are not decoding vocabulary and following the logic at the same time. This directly lowers the peak load of the main explanation.
  • Modality principle: when presenting a graphic that needs verbal explanation, spoken narration often works better than on-screen text, because it uses the auditory channel and frees the visual channel for the graphic - splitting the load across both channels rather than overloading the visual one.

Note that the modality principle is the one most often overstated, and its boundary conditions are well documented. Narration frees the visual channel only if the words can be processed as they pass. When the material contains unfamiliar technical terms, when the sentences are long or syntactically complex, when learners are not native speakers of the language, or when the learner cannot control the pace, written text becomes better than speech, because text can be reread and speech cannot. The safe formulation is that spoken narration helps when the verbal load is light and transient processing is feasible; otherwise on-screen text, or both together, may be preferable. Blanket advice to "narrate rather than write" is not what the research supports.

Key idea: The modality advantage for narration holds when the verbal material is short, familiar, and learner-paced, and reverses for technical, complex, or non-native language where text can be reread.

Fostering generative processing

  • Multimedia principle: people generally learn better from words and relevant pictures than from words alone - the foundational dual-coding payoff, provided the pictures carry meaning rather than decoration.
  • Personalization principle: a conversational, human tone tends to engage learners more than stiff, formal wording, prompting more active processing.

Notice that these principles can pull against naive instincts. The urge to make a slide "richer" with music, extra text, and decorative images is exactly what the coherence and redundancy principles warn against. The discipline of multimedia design is largely subtractive: remove what does not serve learning, integrate what does, split load sensibly across channels, and let the learner's limited working memory spend its budget on building the schema. Every principle here is a specific application of one idea you have carried since Module 1 - respect the bottleneck.

Boundary conditions the literature makes clear

A graduate reader should know where this evidence base is thin, because the principles are frequently applied far outside the conditions that established them.

  • Novices, mostly. The great majority of experiments used learners with low prior knowledge on the topic. The expertise-reversal effect from Lesson 5 applies here directly: redundancy, signalling, and integrated formats help those who need the support and can hinder those who do not.
  • Short lessons, immediate tests. Typical materials run a few minutes and outcomes are measured within the hour. Whether an advantage in a five-minute animation persists across a semester is usually unknown, and Module 3 warns that immediate advantages can reverse.
  • System-paced presentation. Many effects, notably modality and segmenting, depend on the learner being unable to control the flow. Give learners a pause button and some of the differences shrink, because they can self-segment.
  • Laboratory conditions. Effects established with undivided attention in a quiet room may behave differently in a classroom or on a phone. Alexander Skulmowski and Kate Man Xu argue that in genuinely digital settings, the demands of navigating an interface constitute a further source of extraneous load that the classical experiments never included.

Key idea: Most multimedia evidence comes from novices learning short, system-paced lessons in laboratory conditions with immediate tests, so the principles should be applied as strong defaults rather than laws.

Applying the set in priority order

Faced with an existing video or deck, the principles are most useful applied in a fixed order, because they are not independent. Start by subtracting: remove anything that duplicates the narration word for word, then remove music, decorative imagery, and interesting-but-irrelevant anecdotes. This costs nothing, requires no new production, and frees the capacity that everything else depends on. Next reposition: move each label beside the part it names, and align the timing of words with the visuals they describe. Then segment: break a continuous explanation into chunks with a learner-controlled boundary between them, and pre-teach any vocabulary the main explanation assumes. Only then add: insert prompts that require the learner to explain, predict, or draw, which is generative processing and depends on there being capacity left for it.

The order matters because adding generative activity to an overloaded presentation makes it worse, not better. A retrieval prompt inserted into a cluttered, redundant, badly timed video is competing with the clutter for the same workspace. This is the most common failure in well-intentioned redesign: designers reach for the additive principles first because they feel like teaching, when the subtractive ones do more and cost less.

Key idea: Apply the principles in the order subtract, reposition, segment, then add, because generative activity inserted into an overloaded presentation competes with the clutter rather than replacing it.

A contested extension: emotional design

The coherence principle appears to forbid anything decorative, which raises an obvious question: does making material visually appealing ever help? A line of work on emotional design tests whether warm colours, rounded anthropomorphic shapes, and other affect-inducing features improve learning by increasing engagement rather than distracting from content. Rachel Wong and Adesope's replication and extension meta-analysis found the picture mixed - some benefits, considerable heterogeneity, and effects that depend heavily on the outcome measured and on the design used.

The honest position is that emotional design is an active research question rather than a settled principle, and that the distinction from seductive details is subtle but real: emotional design modifies how the essential graphic looks, while seductive details add non-essential material. Colouring the relevant parts of a diagram warmly is not the same as putting a sunset behind it. Until the evidence firms up, treat coherence as the default and emotional design as a hypothesis worth testing in your own context.

Key idea: Emotional design modifies how essential graphics look rather than adding non-essential material, and its benefits remain heterogeneous and unsettled, so coherence stays the default.

Common misconceptions

  • Richer slides teach more. Coherence and redundancy both predict the opposite, and seductive details reliably depress transfer.
  • Always narrate rather than use on-screen text. Modality reverses for technical, complex, or non-native verbal material that needs rereading.
  • These principles are universal laws. They come mostly from novices, short lessons, system-paced delivery, and immediate tests.
  • Any attractive design is a seductive detail. Emotional design alters the essential graphic; seductive details add irrelevant material.
  • Higher learner ratings mean better multimedia. Added detail often raises enjoyment while lowering learning.

Recap

  • Multimedia learning requires selecting, organizing, and integrating verbal and pictorial material through two limited channels.
  • The principles group into reducing extraneous, managing essential, and fostering generative processing.
  • Coherence and spatial contiguity are among the best-evidenced principles, with seductive details harming transfer most.
  • Segmenting, pre-training, and modality manage essential load, and modality has real boundary conditions.
  • The evidence base rests largely on novices learning short, system-paced lessons under laboratory conditions.
  • Emotional design is an unsettled extension, distinct from seductive details but not yet a dependable principle.

Sources

  1. Sundararajan, N., and Adesope, O. (2020). Keep it coherent: A meta-analysis of the seductive details effect. Educational Psychology Review, 32, 707-734. link.springer.com
  2. Schroeder, N. L., and Cenkci, A. T. (2018). Spatial contiguity and spatial split-attention effects in multimedia learning environments: A meta-analysis. Educational Psychology Review, 30, 679-701. link.springer.com
  3. Wong, R. M., and Adesope, O. O. (2021). Meta-analysis of emotional designs in multimedia learning: A replication and extension study. Educational Psychology Review, 33, 357-385. link.springer.com
  4. Fiorella, L., and Mayer, R. E. (2016). Eight ways to promote generative learning. Educational Psychology Review, 28, 717-741. link.springer.com
  5. Skulmowski, A., and Xu, K. M. (2022). Understanding cognitive load in digital and online learning: A new perspective on extraneous cognitive load. Educational Psychology Review, 34, 171-196. link.springer.com
  6. Sweller, J., van Merrienboer, J. J. G., and Paas, F. (2019). Cognitive architecture and instructional design: 20 years later. Educational Psychology Review, 31, 261-292. link.springer.com
  7. Mayer, R. E., and Moreno, R. (2003). Nine ways to reduce cognitive load in multimedia learning. Educational Psychologist, 38(1), 43-52. find source โ†—
Key terms
Multimedia principle
People learn better from relevant words and pictures together than from words alone.
Coherence principle
Excluding extraneous material improves learning; seductive details hurt it.
Signaling principle
Cueing the essential material and its structure directs attention efficiently.
Contiguity principle
Corresponding words and pictures should be placed near and presented together.
Modality principle
Explaining a graphic with narration can beat on-screen text by splitting load across channels.
Segmenting principle
Breaking a lesson into learner-paced chunks manages essential load.
Pre-training principle
Teaching key components' names and features before the full process lowers peak load.

Module 7: Evaluating Instruction

How to judge whether instruction worked, using structured models and evidence, and how to iterate.

Levels of Evaluation and Program Evaluation

  • Distinguish reaction, learning, behavior, and results levels of evaluation.
  • Explain why satisfaction is a weak proxy for learning.
  • Choose appropriate evidence for each evaluation level.

The final phase of ADDIE and the whole point of design is evaluation: determining whether instruction achieved its goals and how to improve it. Doing this well requires being clear about what kind of success you are measuring, because "did it work?" hides several very different questions.

Key idea: The question "did it work?" hides several different questions, and evaluation is defensible only when the evidence gathered matches the specific claim being made.

Four levels of evaluation

A classic and widely used framework distinguishes four levels of training evaluation, each harder to measure and more meaningful than the last.

  1. Reaction: did learners find it engaging, relevant, and satisfying? Usually measured by surveys ("smile sheets"). Easy to collect, but the weakest evidence of actual learning.
  2. Learning: did learners actually acquire the intended knowledge or skills? Measured by valid assessments (Module 6), ideally comparing before and after.
  3. Behavior: are learners applying what they learned in the real setting - the classroom, the job, life - some time later? This is transfer (Module 3), and it is measured by observation or performance in context, not in the training room.
  4. Results: did the application produce the intended outcomes that motivated the training in the first place - better student achievement, fewer errors, improved organizational metrics?

The levels rise in value and in difficulty. Most evaluation stops at Level 1 because it is cheap, which is precisely the problem.

The framework is Donald Kirkpatrick's, set out in a series of articles for the American Society for Training Directors beginning in 1959 and refined by him and his family over the following half-century. Its longevity is remarkable and largely deserved, because it names distinctions practitioners genuinely confuse. But two of its implied claims have been examined and are not supported, and knowing which is what separates competent use from ritual use.

The causal chain is not established. The model is often read as implying that reaction leads to learning, which leads to behaviour, which leads to results. George Alliger and Elizabeth Janak examined this thirty years after its publication and found the assumed causal linkage between levels to be an assumption rather than a finding. A later meta-analysis by Alliger and colleagues examined the actual correlations among training criteria and found the relationships between reaction measures and learning measures to be weak - close to negligible for affective reactions and immediate learning. Reaction is not a leading indicator of learning; it is a different variable.

The hierarchy of value is a simplification. Higher levels are not uniformly more meaningful. A results measure that is confounded by everything else happening in an organization can be less informative than a well-designed learning measure. And levels three and four are frequently unmeasurable within the resources available, which in practice means that insisting on them produces no evaluation at all rather than better evaluation.

Key idea: Kirkpatrick's levels usefully distinguish different claims, but the assumed causal chain from reaction to results is not empirically supported and higher levels are not automatically more informative.

The satisfaction trap

The most important lesson from this framework is that reaction is a weak proxy for learning. Learners can love a session and learn little, or dislike a demanding one and learn a great deal. Recall the fluency illusion and desirable difficulties from Module 3: the very techniques that produce durable learning (spacing, retrieval practice, interleaving) often feel harder and less enjoyable in the moment, so they can lower satisfaction scores while raising real learning.

An instructor who optimizes for smile sheets may systematically drift toward comfortable, fluent, less effective instruction. Positive reactions are worth having - a hated course will not be completed - but they must never be mistaken for evidence that learning occurred. To claim learning, measure learning.

The single most instructive study here is Louis Deslauriers and colleagues' 2019 experiment in a large-enrolment physics course. Students were randomly assigned, within the same course and covering the same content, to either a passive lecture by an experienced instructor or an actively engaged classroom. Two things were measured: how much students actually learned, and how much they felt they had learned. The active-learning students learned substantially more and reported feeling that they had learned less. They also rated the instructor less favourably. The mechanism the authors propose is the fluency illusion of Module 3 operating in real time: a fluent expert lecture is easy to follow, and ease is misread as learning, while the effortful struggle of active engagement is misread as poor teaching.

The consequences for evaluation are severe rather than academic. An instructor rewarded on satisfaction scores faces a direct incentive to teach in the way that produces less learning, and will receive confirming feedback for doing so. Timothy Uttl and colleagues' meta-analysis of the student-evaluation-of-teaching literature reaches a compatible conclusion: once prior-ability confounds and small-study effects are handled properly, ratings and actual learning are not meaningfully related. Reaction data is worth collecting - it flags material problems, and a course learners refuse to complete teaches nobody - but it is not a proxy for learning, and using it as one systematically selects against effective instruction.

Scott Freeman and colleagues' meta-analysis, discussed in Lesson 11, supplies the other half of the picture: across 225 studies, active learning raised examination performance by 0.47 standard deviations and reduced failure rates substantially. Combine that with the Deslauriers finding and the position is uncomfortable but clear - the more effective approach can be the less popular one.

Key idea: In a randomized within-course experiment, active learning produced more learning while students felt they learned less and rated the instructor lower, so optimizing for satisfaction scores selects against effective teaching.

Matching evidence to the question

Sound evaluation chooses evidence that fits the level being claimed. If you want to know whether learners enjoyed it, a survey is fine. If you want to know whether they learned it, you need a valid assessment, and a pre/post comparison is far stronger than a single post-test (which cannot separate what the instruction added from what learners already knew). If you want to know whether it transferred, you must look at real performance later, in the actual context.

And if you want to know whether it mattered, you track the downstream results that justified the effort. A common and costly error is claiming a high-level success (learners now perform better on the job) on the strength of low-level evidence (they rated the workshop highly). Match the evidence to the claim, and be honest about which level you have actually reached.

Why a pre/post comparison is still weak

A pre/post design is much better than a single post-test, but calling it strong evidence overstates it, and the specific threats are worth naming because each has a cheap partial remedy.

  • Maturation and history. Learners change and encounter other influences over the interval. A comparison group experiencing the same interval without the instruction is the remedy.
  • Testing effects. The pre-test is itself a learning event - that is the whole of Module 3 - so some of the gain may be caused by having taken the pre-test. Using different but equivalent forms, or pre-testing only a random subset, addresses this.
  • Regression to the mean. If learners were selected because they scored low, their scores will tend to rise on retest whether or not anything was taught. Selecting on a variable other than the pre-test, or comparing with an equivalently selected control group, addresses it.
  • Instrumentation. If the post-test is easier, or graded by someone who knows which group is which, the gain may be in the measurement rather than the learners. Blind marking and equated forms address it.

None of this argues for abandoning modest designs. It argues for stating conclusions at the strength the design supports. "Scores rose by twelve points, in a design without a comparison group, so we cannot separate the instruction from maturation and the pre-test itself" is a perfectly respectable sentence, and far more useful than an unqualified claim that the course worked.

Key idea: Pre/post designs are threatened by maturation, the learning effect of the pre-test itself, regression to the mean, and instrumentation, so conclusions should be stated at the strength the design supports.

Alternatives and complements to the four levels

Three other frameworks are worth knowing, since each answers a question Kirkpatrick does not.

Daniel Stufflebeam's CIPP model evaluates context (what needs exist), input (is the plan sound and are the resources adequate), process (was it implemented as designed), and product (what were the outcomes). Its distinctive contribution is the process question, which Kirkpatrick omits entirely and which explains a large share of null results: an intervention that was never delivered as designed has not been tested. Before concluding that a design does not work, establish that it happened.

Robert Brinkerhoff's success case method deliberately samples the most and least successful cases rather than averaging, on the grounds that when transfer to real practice is patchy, an average conceals the conditions that made the difference. It answers "what makes this work when it works?", which is often the more actionable question.

Logic models make the assumed chain from inputs through activities and outputs to outcomes explicit before evaluation begins. Their value is diagnostic: writing the chain down usually reveals which link is being assumed rather than tested, and that link is where the evaluation should look first.

Key idea: CIPP adds the implementation-fidelity question that explains many null results, the success case method identifies the conditions under which an intervention works, and logic models expose which link in the causal chain is merely assumed.

Common misconceptions

  • Positive reactions indicate learning. The correlation is weak, and effective effortful instruction can lower satisfaction.
  • Kirkpatrick's levels form a validated causal chain. The linkage between levels is an assumption, examined and not supported.
  • Higher levels are always better evidence. A confounded results measure can be less informative than a good learning measure.
  • A pre/post gain shows the instruction worked. Maturation, the pre-test itself, regression to the mean, and instrumentation all produce gains.
  • A null result means the design failed. It may never have been implemented as designed, which is why process evaluation matters.

Recap

  • Kirkpatrick's four levels - reaction, learning, behaviour, results - distinguish claims that are routinely confused.
  • The assumed causal chain between levels is not supported, and reaction correlates weakly with learning.
  • Deslauriers and colleagues found active learning produced more learning and lower perceived learning and instructor ratings.
  • Freeman and colleagues found active learning raised performance by 0.47 standard deviations across 225 studies.
  • Pre/post designs face maturation, testing, regression, and instrumentation threats and should be reported accordingly.
  • CIPP, the success case method, and logic models supply the implementation and mechanism questions Kirkpatrick omits.

Sources

  1. Deslauriers, L., McCarty, L. S., Miller, K., Callaghan, K., and Kestin, G. (2019). Measuring actual learning versus feeling of learning in response to being actively engaged in the classroom. Proceedings of the National Academy of Sciences, 116(39), 19251-19257. pmc.ncbi.nlm.nih.gov
  2. Freeman, S., Eddy, S. L., McDonough, M., Smith, M. K., Okoroafor, N., Jordt, H., and Wenderoth, M. P. (2014). Active learning increases student performance in science, engineering, and mathematics. Proceedings of the National Academy of Sciences, 111(23), 8410-8415. pmc.ncbi.nlm.nih.gov
  3. Sitzmann, T., Brown, K. G., Casper, W. J., Ely, K., and Zimmerman, R. D. (2008). A review and meta-analysis of the nomological network of trainee reactions. Journal of Applied Psychology, 93(2), 280-295. pubmed.ncbi.nlm.nih.gov
  4. What Works Clearinghouse. (n.d.). Handbooks and Other Resources (standards for assessing study designs). Institute of Education Sciences. ies.ed.gov
  5. Hamilton, L., Halverson, R., Jackson, S., Mandinach, E., Supovitz, J., and Wayman, J. (2009). Using Student Achievement Data to Support Instructional Decision Making (NCEE 2009-4067). Institute of Education Sciences. ies.ed.gov
  6. Alliger, G. M., and Janak, E. A. (1989). Kirkpatrick's levels of training criteria: Thirty years later. Personnel Psychology, 42(2), 331-342. find source โ†—
  7. Uttl, B., White, C. A., and Gonzalez, D. W. (2017). Meta-analysis of faculty's teaching effectiveness: Student evaluation of teaching ratings and student learning are not related. Studies in Educational Evaluation, 54, 22-42. find source โ†—
Key terms
Program evaluation
The systematic determination of whether instruction achieved its goals and how to improve it.
Reaction level
Evaluation of learners' satisfaction and perceived relevance; the weakest evidence of learning.
Learning level
Evaluation of the knowledge or skills actually acquired, via valid assessment.
Behavior level
Evaluation of whether learning is applied in the real setting later; that is, transfer.
Results level
Evaluation of the downstream outcomes the instruction was meant to produce.
Pre/post comparison
Measuring before and after instruction to isolate what the instruction added.

Data, Iteration, and Improving Instruction

  • Use assessment and process data to locate design problems.
  • Apply iterative improvement and simple comparisons responsibly.
  • Guard against common pitfalls in interpreting educational data.

Evaluation is only worthwhile if it feeds improvement. This closing lesson is about turning evidence into better instruction through disciplined iteration - the loop that makes ADDIE a cycle rather than a line.

Key idea: Evaluation earns its cost only when it changes a design decision, which means reading data for signals about the instruction rather than verdicts about the learners.

Reading the data for design signals

Assessment results are not just grades; they are a map of where instruction succeeded and failed. An item analysis - looking at performance question by question - can reveal a specific topic everyone missed (a lesson to redesign), a question that even strong students got wrong (possibly a flawed item or a hidden misconception), or a concept that a pre-test shows learners already knew (time to reallocate).

Process data help too: where do learners drop out of an online module, where do they slow down, which activities do they skip? Patterns across many learners point to design problems more reliably than any single learner's experience. The mindset is diagnostic - treat a bad result as information about the design, not merely about the learners, and ask what specifically to change.

Item analysis has a small technical vocabulary that repays learning, because each statistic answers a different design question. Item difficulty, conventionally the proportion answering correctly, tells you whether an item is informative: items answered correctly by everyone or by no one discriminate nothing, whatever else they do. Item discrimination, often computed as the correlation between success on the item and total score, tells you whether the item behaves like the rest of the assessment. A negative discrimination is the single most valuable signal available, because it means students who did well overall got this item wrong - which points either to a flawed item or to a genuine misconception that stronger students hold more confidently.

Distractor analysis is where the diagnosis lives. If a wrong option attracts forty percent of learners, that option is telling you what they believe, and if the distractors were built from the misconceptions in Lesson 11, it names the misconception directly. If a distractor attracts nobody, it is doing no work and the item is effectively a three-option question.

A worked read: an item on interpreting a confidence interval shows difficulty 0.35, discrimination near zero, and one distractor attracting 45 percent of responses. Difficulty alone might suggest the topic is hard and needs more time. The near-zero discrimination and the dominant distractor together say something different - that the whole class, strong and weak alike, holds one specific wrong interpretation. That is not a time problem; it is a conceptual-change problem, and the remedy is a refutation-style treatment naming the wrong interpretation explicitly, not another practice set.

Key idea: Difficulty tells you whether an item is informative, discrimination flags flawed items or shared misconceptions, and distractor analysis names what learners actually believe.

Iterate deliberately

Improvement works best as small, repeated cycles: identify a likely weak point from the evidence, change one thing, and check whether the next cohort or the next run does better. This mirrors the rapid, iterative view of ADDIE. Changing many things at once feels efficient but makes it impossible to learn which change helped.

Where feasible, a simple comparison - running two versions of a lesson and comparing valid learning outcomes (an A/B comparison) - can settle a design debate that opinion never would. Even informally, the habit of forming a specific hypothesis ("moving the worked examples earlier will raise the quiz scores"), making one change, and checking the result is what separates genuine improvement from mere tinkering.

Two disciplines from outside education sharpen this habit. Improvement science, as developed for education by Anthony Bryk and colleagues, organizes work into short plan-do-study-act cycles run by the people doing the teaching, on problems specified precisely enough to test - "our learners fail the unit conversion step," not "our learners struggle with chemistry." Its central discipline is defining the problem and the measure before changing anything. Design-based research works at a longer timescale, iterating a design in a real setting while deliberately developing the theory that explains why it works, so that what transfers to other settings is the principle rather than the materials.

Both share a commitment that is easy to state and hard to keep: write down what you expect to happen before you make the change. A prediction recorded in advance can be wrong, which is what makes it informative. A judgement formed after seeing the results can always be made to fit them.

Key idea: Improvement science and design-based research both require the problem, the measure, and the expected result to be specified before the change is made, because only a prediction made in advance can be disconfirmed.

Pitfalls in interpreting educational data

Evidence is only as good as its interpretation, and several traps recur.

  • Correlation is not causation. That high performers also used the optional review tool does not prove the tool caused their performance; motivated students may simply do both. Only a fair comparison can support a causal claim.
  • Beware tiny samples and noise. One class doing better could easily be chance. Small differences from few learners are unreliable; look for consistent patterns across enough cases.
  • Do not optimize the measure at the expense of the goal. If a test becomes the target, instruction can drift toward teaching to the test in a narrow way that raises scores without raising real understanding - a validity failure in disguise. Keep checking that the assessment still reflects the true objective.
  • Immediate performance can mislead. As the whole course has stressed, performance during or right after instruction is a poor guide to durable learning; a change that improves the end-of-lesson quiz but is never checked weeks later may be optimizing fluency, not retention.
  • Regression to the mean manufactures gains. Any group selected for scoring badly will tend to score better on retest, with no intervention at all. Interventions targeted at the lowest performers are therefore almost guaranteed to look successful unless there is a comparison group selected the same way.
  • Effect sizes in education are small. Matthew Kraft's analysis of the field's intervention literature argues that the benchmarks borrowed from psychology - 0.2 small, 0.5 medium, 0.8 large - badly mislead when applied to real programmes affecting broad achievement measures, where 0.2 is in fact a large effect. Treat any classroom claim above about 0.5 as requiring scrutiny of the outcome measure rather than celebration.
  • Published effects shrink. Macnamara and Burgoyne's mindset review, discussed in Lesson 10, found larger published effects among authors with a financial interest, and a headline effect that vanished after correcting for publication bias. This pattern is general rather than particular to mindset work, and it is a reason to weight large preregistered trials far above collections of small enthusiastic studies.

Closing the loop

Put together, the discipline of good instruction is a cycle you have now traced end to end: analyze learners and needs, design aligned objectives and assessments, develop materials that respect cognitive load and multimedia principles, implement them, evaluate honestly at the right level, read the data for design signals, change one thing, and go around again. Every element rests on the same foundation laid in Module 1 - a clear-eyed understanding of how memory and cognition actually work. Instruction improves not by intuition or fashion but by this patient, evidence-driven loop, and you are now equipped to run it.

Evaluating other people's evidence

Most design decisions will rest on research you did not run, so the last skill this course can give you is a way of reading it. Five questions do most of the work.

  1. Compared with what? An intervention beating no instruction says almost nothing. An intervention beating a serious business-as-usual alternative says a great deal. Studies whose control condition was doing nothing are the most common source of inflated claims.
  2. Measured when? An immediate post-test rewards fluency. A delayed test measures what this course means by learning. If the interval is not stated, assume it was short.
  3. Measured how? Was the outcome the target capability, or a proxy? Was it a researcher-designed instrument scored by the researchers, or an independent measure? Construct underrepresentation from Lesson 15 is rampant.
  4. For whom? Prior knowledge moderates nearly everything in this course, and expertise reversal means a result from novices can invert for experienced learners.
  5. How large, and how certain? A significant result is not a large one, and a single study is not a literature. Prefer preregistered trials and meta-analyses that report heterogeneity and check for publication bias.

The What Works Clearinghouse handbooks formalize a version of this by rating studies on design quality before allowing them to support a claim, and its standards are a useful external reference when a colleague cites a study you find unconvincing. The point is not scepticism for its own sake. It is that the failure mode in instructional design is not too little enthusiasm for evidence but too much enthusiasm for weak evidence.

Key idea: Read research by asking what the comparison was, when and how outcomes were measured, for which learners, and how large and how certain the effect is.

Common misconceptions

  • A hard question means the topic needs more time. Near-zero discrimination with one dominant distractor indicates a shared misconception, not insufficient practice.
  • More data produces better decisions. Data changes nothing unless a specific decision was waiting on it.
  • Gains among the lowest performers show the intervention worked. Regression to the mean produces exactly that pattern with no intervention at all.
  • An effect size of 0.3 is small. For a real programme affecting broad achievement, effects that size are substantial; psychology's benchmarks mislead here.
  • Improvement means changing several things at once. Simultaneous changes make attribution impossible and prevent learning from the iteration.

Recap

  • Assessment data are design signals: difficulty, discrimination, and distractor analysis each answer a different question.
  • Negative or near-zero discrimination flags a flawed item or a misconception held even by strong students.
  • Iterate in small cycles, changing one thing and recording the prediction before the change.
  • Improvement science and design-based research supply disciplined versions of this loop at different timescales.
  • Correlation, small samples, regression to the mean, measure-gaming, and publication bias are the recurring interpretive traps.
  • Effect sizes in real educational settings are small, so claims above roughly 0.5 warrant scrutiny of the outcome measure.

Sources

  1. Hamilton, L., Halverson, R., Jackson, S., Mandinach, E., Supovitz, J., and Wayman, J. (2009). Using Student Achievement Data to Support Instructional Decision Making (NCEE 2009-4067). Institute of Education Sciences, What Works Clearinghouse. ies.ed.gov
  2. What Works Clearinghouse. (n.d.). Handbooks and Other Resources (standards and procedures for reviewing study designs). Institute of Education Sciences. ies.ed.gov
  3. Macnamara, B. N., and Burgoyne, A. P. (2023). Do growth mindset interventions impact students' academic achievement? A systematic review and meta-analysis with recommendations for best practices. Psychological Bulletin, 149(3-4), 133-173. pubmed.ncbi.nlm.nih.gov
  4. Freeman, S., Eddy, S. L., McDonough, M., Smith, M. K., Okoroafor, N., Jordt, H., and Wenderoth, M. P. (2014). Active learning increases student performance in science, engineering, and mathematics. Proceedings of the National Academy of Sciences, 111(23), 8410-8415. pmc.ncbi.nlm.nih.gov
  5. Yang, C., Luo, L., Vadillo, M. A., Yu, R., and Shanks, D. R. (2021). Testing (quizzing) boosts classroom learning: A systematic and meta-analytic review. Psychological Bulletin, 147(4), 399-435. pubmed.ncbi.nlm.nih.gov
  6. Kraft, M. A. (2020). Interpreting effect sizes of education interventions. Educational Researcher, 49(4), 241-253. find source โ†—
  7. Bryk, A. S., Gomez, L. M., Grunow, A., and LeMahieu, P. G. (2015). Learning to Improve: How America's Schools Can Get Better at Getting Better. Harvard Education Press. find source โ†—
Key terms
Item analysis
Examining assessment performance question by question to locate design problems.
Iterative improvement
Refining instruction through small, repeated change-and-check cycles.
A/B comparison
Comparing two versions of instruction by valid outcomes to test which works better.
Correlation versus causation
The caution that an association between two variables does not prove one caused the other.
Teaching to the test
Narrowly optimizing scores in ways that raise the measure without raising real understanding.
Improvement loop
The ongoing cycle of designing, evaluating, and revising instruction based on evidence.

Open the interactive version with quizzes and progress →