💼 Business & Management · Undergraduate · MGMT 320

Operations & Supply Chain Management

Every organization you deal with runs on operations, and almost none of it is visible from the outside. This course opens it up and teaches you to compute it. You will analyze processes and find bottlenecks, apply Little's Law in both directions, and see why a queue forms at a coffee counter that is only eighty percent busy. You will build a control chart from raw measurements, calculate whether a…

Start the interactive course (quizzes, progress, videos) →

Free forever. No sign-up, no ads. 16 lessons. The full lesson text is below so you can read it right here.

Module 1: Processes, Capacity, and Flow

What operations is, why the process is the unit of analysis, and the three pieces of arithmetic that describe any process: capacity, bottlenecks, and Little's Law.

What Operations Actually Is

  • Describe the transformation model and apply it to both a factory and a service organization.
  • Contrast goods and services on the dimensions that change how operations are managed.
  • Compare the four competitive priorities and explain why trade-offs among them are real but not absolute.
  • Compute single-factor and multifactor productivity and explain what each measure hides.

The big picture

Stand in a coffee shop at 8:15 on a Tuesday and watch. Someone takes an order. Someone else pulls a shot while steaming milk for the order before it. A third person is restocking cups from a back room whose shelves were filled by a delivery that arrived at 5:40 that morning, from a warehouse forty miles away, holding beans that were roasted eleven days ago in a plant that bought them from an importer who bought them from a cooperative in Colombia. The whole thing has to produce a drink you will accept, in under four minutes, at a price you will pay, roughly two hundred times before ten o'clock. Nobody in the shop is thinking about any of this. It works because somebody designed it to.

That design work is operations management: the design, operation, and improvement of the systems that create and deliver an organization's products and services. It is the part of a business that actually makes the thing. Finance decides how to fund it, marketing decides what to promise about it, but if the drink is cold, late, or wrong, none of that matters.

A word about this course before we start. Principles of Management (MGMT 301) surveys operations in a single lesson: the control cycle, a first look at bottlenecks, lean at Toyota, and supply chain fragility. That lesson tells you the ideas exist. This course makes you compute them. Almost every lesson here ends with a number you worked out yourself, because operations is a quantitative discipline pretending to be common sense. Common sense will tell you that a process running at 90 percent of capacity is efficient. The arithmetic in Module 2 will tell you it has a queue nine times longer than the same process at 50 percent, and the arithmetic is right.

Be honest with yourself about one limit. A text course can teach you to draw a process, find its constraint, and price a buffer. It cannot teach you what it feels like to stand on a line at 2 a.m. when a machine is down and the shift supervisor is looking at you. Every operations person will tell you that walking the floor, what Toyota calls going to see for yourself, is where the real learning happens. Treat this course as the arithmetic and vocabulary you should already have in your head when you get there.

The transformation model

The simplest useful model of any operation has three parts and a loop. Inputs go in: materials, labor, equipment, energy, information, capital, and in services, the customer. A transformation process changes them. Outputs come out: goods, services, or nearly always some mixture. A feedback loop measures the outputs and adjusts the process.

The transformation itself can be physical (a steel coil becomes a fender), locational (a package moves from Memphis to Denver), exchange-based (a share of stock changes hands), physiological (a patient's infection is cured), psychological (a student understands compound interest), or informational (raw claims data becomes an approved reimbursement). Once you see the model you cannot unsee it. A university transforms uninformed students into informed ones using faculty time, libraries, and buildings, and its capacity constraint on a Tuesday morning is classrooms, which is why your 8 a.m. section exists.

Key idea: Every operation takes inputs, transforms them, produces outputs, and measures the result. Naming the inputs, the transformation, and the output precisely is the first analytical move in this entire course.

Goods and services are not two categories

Textbooks used to split the world into manufacturing and services. The split is real but it is a continuum, not a wall. A restaurant meal is a manufactured good delivered inside a service experience. A software subscription is a service delivered by a manufactured artifact. What matters is not the label but the handful of properties that genuinely change how you manage the operation.

PropertyGoods-heavy operationService-heavy operationWhy it changes management
StorabilityOutput can be inventoriedOutput cannot be inventoriedA factory can build in slow months and sell in fast ones. An empty hotel room tonight is gone forever, which is why capacity and demand must be matched in real time.
TimingProduction separate from consumptionProduced and consumed at onceYou can inspect a fender before shipping it. You cannot inspect a haircut before the customer receives it.
Customer contactLow; customer rarely in the processHigh; customer is inside the processThe customer is an uncontrolled input. A patient who forgets to fast before a blood test injects variability no factory ever faces.
Quality judgmentMeasurable against specificationsPartly perceptual and comparativeTolerances work for a shaft diameter. They do not work for whether the nurse seemed rushed.
LocationCan be distant from customersUsually near customersFenders travel. Haircuts do not, which is why service networks have many small sites and factories have few large ones.

Two consequences follow immediately. First, service capacity is perishable, so services live or die on how well they match capacity to a demand pattern they only partly control. That is why Module 2 spends a whole lesson on waiting lines. Second, high customer contact means high variability, and variability is the enemy that every technique in this course is ultimately fighting.

How big is each side? In the United States, private service-providing industries employ roughly five out of every six nonfarm workers, while manufacturing employs on the order of 12 to 13 million people, about eight percent of nonfarm payrolls. Manufacturing's share of output is larger than its share of jobs, contributing roughly a tenth of gross domestic product, because productivity per manufacturing worker is high. Transportation and warehousing alone employ more than six million people. Operations is not a niche; it is where most of the workforce actually is.

Key idea: Goods and services differ most in storability, simultaneity, and customer contact. Those three properties, not the label, determine which operations tools apply.

Competitive priorities and the reality of trade-offs

An operation cannot be designed until you know what it is supposed to be good at. Four competitive priorities cover almost everything a customer might value:

  • Cost: low unit cost, achieved through scale, high utilization, standardization, and process discipline.
  • Quality: split into two very different things. Performance quality means how good the design is (a luxury sedan versus an economy car). Conformance quality means how reliably the thing matches its design. A cheap car built exactly to spec has high conformance quality and modest performance quality, and those are separate management problems.
  • Delivery: also two things. Speed is how fast, reliability is how predictably. Reliability is usually worth more than speed, because a supplier who is always eight days is easier to plan around than one who averages five days but sometimes takes twenty.
  • Flexibility: product flexibility (how easily you can make something different) and volume flexibility (how easily you can make more or less).

Wickham Skinner argued in 1969 that a factory trying to be excellent at everything ends up mediocre at all of it, and proposed the focused factory: pick the priorities, then design consistently for them. His observation is still the most useful sentence in the field. A plant tooled for high-volume low-cost production has long changeovers and large batches, which is exactly what makes it inflexible. Those are not two flaws; they are one design choice seen from two angles.

But the trade-off is not absolute, and this is where the field got more interesting. Japanese manufacturers in the 1970s and 1980s demonstrably improved quality and cost and flexibility at the same time, which the trade-off view said was impossible. The resolution is the sand cone model proposed by Ferdows and De Meyer in 1990: capabilities build in a sequence. Quality comes first, because defects waste everything downstream. Dependability builds on quality, because you cannot promise dates when a third of output needs rework. Speed builds on dependability. Cost improvement comes last and rests on all three, because the cheapest thing you can do is stop wasting material, time, and rework. Trade-offs are real at the frontier; below the frontier, improvement in one dimension often helps the others.

Terry Hill's vocabulary is worth keeping. Order qualifiers are the levels you must meet to be considered at all. Order winners are what actually make the customer choose you. Airline safety is a qualifier: nobody chooses an airline for being safe, but they will not choose an unsafe one. Schedule and price are the winners. The trap is spending improvement budget on qualifiers you already meet.

Key idea: Choose priorities explicitly. Trade-offs bind at the frontier, capabilities build in the order quality, dependability, speed, cost, and improving a qualifier you already meet buys nothing.

Productivity, computed two ways

Productivity is output divided by input. That sounds too simple to argue about, until you pick which input.

Take a small bakery. On Monday it produced 1,200 loaves using 40 labor hours.

Labor productivity = 1,200 loaves / 40 hours = 30 loaves per labor hour. That is a single-factor measure. It is easy to compute and easy to game: buy a machine that does half the work, and labor productivity soars while total cost rises.

Multifactor productivity fixes that by putting several inputs in the denominator, valued in money. Monday's numbers:

ItemAmount
Output: 1,200 loaves at 4.00 dollars4,800 dollars
Labor600 dollars
Materials1,100 dollars
Energy150 dollars
Overhead and capital250 dollars
Total input2,100 dollars

Multifactor productivity = 4,800 / 2,100 = 2.29. The number is unitless and meaningless on its own; it is only useful as a comparison across time or across sites.

Now the bakery installs a larger oven. Output rises to 1,320 loaves, labor falls to 480 dollars, energy rises to 190 dollars, materials rise proportionally to 1,210 dollars, and overhead rises to 320 dollars with the new equipment.

  • New revenue: 1,320 x 4.00 = 5,280 dollars.
  • New input total: 480 + 1,210 + 190 + 320 = 2,200 dollars.
  • New multifactor productivity: 5,280 / 2,200 = 2.40, up 4.8 percent.
  • New labor productivity, if labor hours fell to 32: 1,320 / 32 = 41.25 loaves per hour, up 37.5 percent.

Look at the gap between those two percentages. Labor productivity rose 37.5 percent; the honest overall gain was 4.8 percent. The difference is the oven, the energy it burns, and the capital it consumed. Whenever someone quotes a productivity gain, your first question is which inputs are in the denominator.

Key idea: Single-factor productivity is easy to compute and easy to inflate by substituting one input for another. Multifactor productivity is the honest version, and it is the one to demand.

Why the process is the unit of analysis

Here is the shift in perspective that this whole course rests on. Most organizations are drawn as boxes: purchasing, engineering, production, quality, shipping, customer service. Work does not flow that way. Work flows sideways, across the boxes, through a process: a set of activities that transforms inputs into an output for a customer.

Consider a hospital discharge. The physician writes the order at 9:10. The nurse cannot act until she finishes rounds at 10:30. The pharmacy must fill discharge medications, which takes ninety minutes because it is batched. Transport must be called and averages forty minutes. Housekeeping cannot clean the room until the patient leaves, and cannot start until a request appears in their queue. The patient leaves at 3:40. Ask each department how it performed and every one will report acceptable numbers. Ask the process how it performed and the answer is six and a half hours to move someone who was medically ready at 9:10, during which the bed was unavailable to a patient boarding in the emergency department.

Nobody in that story did a bad job. The process is bad. That is the operations insight in one sentence, and it is why the next two lessons teach you to draw a process, compute its capacity, and find the one place where fixing something actually matters.

Where the field came from

A short lineage helps, because you will meet these names for the rest of the course. Interchangeable parts, developed in armories in the late 1700s and 1800s, made assembly possible without individual fitting. Frederick Taylor's scientific management, from roughly 1880 to 1915, brought measurement and standardization along with real human costs that MGMT 301 examines. Ford's moving assembly line at Highland Park in 1913 cut chassis assembly from over twelve hours to about ninety minutes and cut the Model T's price along with it. Ford W. Harris published the economic order quantity in 1913, which you will derive in Module 3. Walter Shewhart drew the first control chart at Bell Labs in 1924, which you will build in Module 2. World War II created operations research, applying mathematics to convoy routing and search. Postwar Toyota, under Taiichi Ohno, built the production system that Module 5 covers properly. Computers brought material requirements planning in the 1970s, which you will run by hand in Module 4. Containerization, retail logistics, and offshoring built the global chains of Module 6, and the disruptions of 2020 to 2023 tested them in public.

Each of those developments is an answer to a question about flow, variability, or cost. That is all operations has ever been about.

Common misconceptions

  • "Operations means factories." Most operations jobs are in services: hospitals, banks, airlines, distribution centers, restaurants, call centers, universities. The tools are the same; the perishability of capacity makes them matter more.
  • "Operations is just cost cutting." Cost is one of four priorities. An operation designed purely for cost will be slow, rigid, and fragile, and Module 6 shows what that looked like in 2021.
  • "We should be excellent at everything." Below the frontier, improvement is often free in multiple dimensions. At the frontier, it is not, and pretending otherwise produces a plant that is expensive, slow, and inflexible at once.
  • "Higher productivity means people working harder." Almost all durable productivity gains come from changing the process, the equipment, or the product design, not from asking for more effort.

Recap

  • Operations management designs, runs, and improves the systems that create and deliver goods and services; every operation is inputs, a transformation, outputs, and feedback.
  • Goods and services differ most in storability, simultaneity of production and consumption, and customer contact, and those differences drive the tools you use.
  • The four competitive priorities are cost, quality, delivery, and flexibility; trade-offs bind at the frontier, and the sand cone model says capabilities build in the order quality, dependability, speed, cost.
  • Single-factor productivity is easy to inflate by input substitution; multifactor productivity, worked here from 2.29 to 2.40 for a bakery, is the honest measure.
  • Work flows across departments through processes, so the process, not the department, is the unit of analysis for the rest of this course.

Sources

  1. Association for Supply Chain Management. (2025). Supply chain and operations resources. ASCM. ascm.org
  2. OpenStax. (2018). Achieving world-class operations management. In Introduction to Business. Rice University. openstax.org
  3. U.S. Bureau of Labor Statistics. (2025). Employment by major industry sector. Employment Projections program. bls.gov
  4. Britannica. (2025). Mass production. Encyclopaedia Britannica. britannica.com
  5. MIT OpenCourseWare. (2013). Introduction to operations management. Massachusetts Institute of Technology. ocw.mit.edu
Key terms
Operations management
The design, operation, and improvement of the systems that create and deliver an organization's goods and services.
Transformation model
The view of any operation as inputs, a transformation process, outputs, and a feedback loop that measures and corrects.
Process
A set of activities that transforms inputs into an output for a customer, usually running sideways across departments.
Competitive priorities
Cost, quality, delivery, and flexibility: the dimensions on which an operation chooses to compete.
Conformance quality
How reliably output matches its design specification, as distinct from how good the design itself is.
Order qualifier
A performance level a firm must meet to be considered by customers at all, as opposed to what wins the order.
Sand cone model
The proposition that operating capabilities build cumulatively in the order quality, dependability, speed, then cost.
Multifactor productivity
Output value divided by the combined value of several inputs, such as labor, materials, energy, and capital.

Process Flow, Capacity, and Bottlenecks

  • Draw a process flow diagram and define its flow unit precisely.
  • Compute the capacity of each resource and identify the bottleneck of a multi-stage process.
  • Calculate utilization, implied utilization, cycle time, and takt time, and use them to size a line.
  • Explain how setup time, batch size, and capacity cushions change effective capacity.

The big picture

A sandwich shop near campus has a line out the door at noon and an empty counter at two. The owner is convinced she needs more staff. She hires two more people, puts them behind the counter, and the noon line is exactly as long as before. She has just learned the most expensive lesson in operations at retail price: capacity is not the sum of your people. It is set by one place in the process, and if you add people anywhere else, you have bought nothing.

This lesson gives you the arithmetic to find that place before you spend the money. There are only four quantities, and you will use them constantly for the rest of the course: capacity, flow rate, utilization, and cycle time. Everything else in process analysis is bookkeeping on top of those four.

Step one: define the flow unit

Before you compute anything, you have to say what is flowing. The flow unit is the thing you are tracking through the process: a sandwich, a patient, a mortgage application, a pallet, a phone call. This sounds like pedantry. It is not. Half of all botched process analyses come from mixing units.

An emergency department that measures capacity in patients per hour and staffs in nurse-hours per patient will get the wrong answer if half of its patients are quick suture cases and half are chest pain workups. If your flow unit varies wildly in the work it demands, you either define separate flow units and analyze separate flows, or you convert to a common unit such as standard minutes of work. Pick the unit, write it at the top of the page, and never switch mid-analysis.

Key idea: Every process analysis begins by naming the flow unit. All capacities, rates, and times must then be expressed per that unit.

Capacity of a resource

A resource is anything that does work: a person, a machine, an oven, a room. Its activity time (also called unit load or processing time) is how long it takes to handle one flow unit. Then:

Capacity of a resource = (number of parallel units) divided by (activity time).

Our sandwich shop has four steps. Watch the arithmetic carefully, because the whole lesson turns on it.

StepActivity timeParallel unitsCapacity per hour
Take order30 seconds1 cashier3,600 / 30 = 120
Assemble sandwich75 seconds1 maker3,600 / 75 = 48
Toast90 secondsoven holds 33 x (3,600 / 90) = 120
Pay and hand off40 seconds1 register3,600 / 40 = 90

The bottleneck is the resource with the smallest capacity: assembly, at 48 sandwiches per hour. And here is the rule that governs everything: the capacity of the whole process equals the capacity of its bottleneck. This shop can make 48 sandwiches an hour, no matter that its cashier could take 120 orders and its oven could toast 120 sandwiches.

Now put the owner's decision through the arithmetic. She hires a second sandwich maker. Assembly capacity becomes 2 x 48 = 96 per hour. Process capacity is now limited by the register at 90 per hour. She has bought 42 extra sandwiches an hour for one hire, an excellent trade. She hires a third maker. Assembly goes to 144 per hour; process capacity stays at 90, because the register is now the bottleneck. That third hire bought exactly zero. The bottleneck moved, and she did not notice.

Key idea: Process capacity equals bottleneck capacity. Adding capacity at a non-bottleneck changes nothing, and adding capacity at the bottleneck helps only until the bottleneck moves somewhere else.

Flow rate, utilization, and implied utilization

Flow rate (or throughput) is what the process actually produces per unit time. It is the smaller of demand and capacity. You cannot produce more than customers want, and you cannot produce more than the bottleneck allows.

Flow rate = minimum of (demand rate, process capacity).

Utilization of a resource = flow rate divided by that resource's capacity. With two makers and demand of 70 sandwiches per hour, capacity is 90 and demand is 70, so flow rate is 70. Then:

  • Cashier: 70 / 120 = 58 percent utilized.
  • Assembly: 70 / 96 = 73 percent.
  • Oven: 70 / 120 = 58 percent.
  • Register: 70 / 90 = 78 percent.

Utilization can never exceed 100 percent, because flow rate can never exceed capacity. That makes it useless for diagnosing an overloaded process, which is why we also compute implied utilization = demand rate divided by capacity. Implied utilization can and should exceed 100 percent when demand outruns capacity, and the amount by which it does tells you how badly you are short. With only one sandwich maker and demand of 70, assembly's implied utilization is 70 / 48 = 146 percent. That is not a busy station; that is a station whose queue grows without limit until the lunch rush ends.

Cycle time is the time between successive units coming out of the process: cycle time = 1 / flow rate. At a flow rate of 90 per hour, a sandwich emerges every 3,600 / 90 = 40 seconds. Note carefully that cycle time is not how long one sandwich takes to get through the shop. That is flow time, and the next lesson separates them properly, because confusing the two is the single most common error in this material.

Takt time and line balancing

When you are designing a line rather than measuring one, you work backwards from demand. Takt time is the rhythm that demand requires:

Takt time = available production time divided by required output.

A small assembly cell works an eight-hour shift with two fifteen-minute breaks, so available time is 480 - 30 = 450 minutes = 27,000 seconds. Today's requirement is 540 units. Takt time = 27,000 / 540 = 50 seconds per unit. To meet demand, a finished unit must come off the line every 50 seconds. Not faster, which would build inventory nobody ordered, and not slower, which misses the schedule.

Now suppose the total work content of one unit is 235 seconds, divided into tasks of 45, 30, 50, 25, 40, and 45 seconds. How many stations do you need?

Start with the floor. Theoretical minimum stations = total work content / takt time = 235 / 50 = 4.7, which rounds up to 5 stations. That is a lower bound, not a plan. Now try to build it, assigning tasks in precedence order without letting any station exceed 50 seconds.

StationTasks assignedStation timeIdle time
14545 s5 s (adding the 30 would give 75)
23030 s20 s (adding the 50 would give 80)
35050 s0 s
42525 s25 s (adding the 40 would give 65)
54040 s10 s (adding the 45 would give 85)
64545 s5 s

Six stations, not five. The theoretical minimum was unreachable because tasks are indivisible lumps that do not tile neatly into 50-second slots, and precedence forbids the rearrangements that might have helped. That failure is the lesson. Line balancing efficiency = total work content divided by (number of stations x cycle time). With six stations at a 50-second cycle: 235 / (6 x 50) = 78 percent. Twenty-two percent of paid station time is idle. Real line design attacks this by splitting or combining tasks, moving work between stations, and sometimes redesigning the product so a 45-second task becomes two 22-second ones. The bottleneck station, at 50 seconds, sets the line's actual pace no matter what the average station does.

Key idea: Takt time is the beat demand requires; the slowest station sets the beat the line can deliver. The gap between them, measured as balance efficiency, is paid idle time.

Design capacity, effective capacity, and the honest numbers

The capacities you compute from activity times are design capacity: the output under ideal conditions. Nothing runs under ideal conditions. Effective capacity subtracts planned losses: changeovers, preventive maintenance, breaks, meetings, cleaning, startup and shutdown. Actual output is what you really got, after unplanned losses: breakdowns, absences, material shortages, quality holds.

A press line has a design capacity of 1,000 units per week. Planned maintenance and changeovers consume ten percent, so effective capacity is 900. Last week it produced 810.

  • Utilization = actual output / design capacity = 810 / 1,000 = 81 percent.
  • Efficiency = actual output / effective capacity = 810 / 900 = 90 percent.

Both numbers are honest and they answer different questions. Efficiency asks how well you ran the plan you had. Utilization asks how much of the asset you actually used. A plant reporting 90 percent efficiency and 81 percent utilization is telling you that the planned losses, not the unplanned ones, are the bigger opportunity, and that is the sort of finding that reorients an improvement program.

Setups, batches, and the capacity you throw away

Now the piece that quietly determines most factory capacity. A machine processes one unit in 2 minutes and requires a 60-minute setup whenever it changes over to a new part.

Batch sizeSetup + run timeTime per unitCapacity per hour
60 units60 + 120 = 180 min180 / 60 = 3.00 min20.0
120 units60 + 240 = 300 min300 / 120 = 2.50 min24.0
300 units60 + 600 = 660 min660 / 300 = 2.20 min27.3
600 units60 + 1,200 = 1,260 min1,260 / 600 = 2.10 min28.6

Bigger batches buy capacity, with diminishing returns: going from 60 to 300 gains 36 percent, going from 300 to 600 gains only another 5 percent. So why not run enormous batches? Because every unit in a batch of 600 sits in inventory waiting for the other 599, and because you cannot respond to a customer who wants the part you just changed away from. That tension between setup cost and inventory cost is exactly what Module 3 formalizes as the economic order quantity, and it is exactly what Module 5's setup-reduction work attacks from the other direction. Cut the setup from 60 minutes to 6, and a batch of 60 now takes 6 + 120 = 126 minutes, or 2.1 minutes per unit: the same capacity that used to require a batch of 600. Setup reduction does not just save an hour. It makes small batches affordable, and small batches are what make a process responsive.

Key idea: Batch size buys capacity when setups are long, at the price of inventory and responsiveness. Reducing setup time buys the capacity without paying that price.

Capacity strategy over time

Capacity decisions at the plant level come in three postures. A lead strategy builds capacity ahead of expected demand: you capture growth and deter entrants, and you eat the cost if the demand does not arrive. A lag strategy adds capacity only after demand is proven: high utilization, low risk of stranded assets, and lost sales during the gap. A match strategy adds in smaller increments as demand appears, trading some scale economy for less exposure.

Whichever posture you choose, plan a capacity cushion: the deliberate slack you keep, computed as 100 percent minus average utilization. Cushions of ten to twenty percent are common, and they exist because of the arithmetic in the next lesson but one. High-variability operations such as hospitals, fire departments, and job shops need large cushions. Low-variability, capital-intensive operations such as paper mills and refineries run at 90 percent and above precisely because their demand and processing are stable enough to permit it. A hospital that plans for 95 percent bed occupancy has not designed an efficient hospital; it has designed one with an ambulance diversion problem.

Common misconceptions

  • "The bottleneck is the slowest worker." The bottleneck is the resource with the lowest capacity, which accounts for parallel units. A slow oven that holds three items can have more capacity than a fast one that holds one.
  • "Every station should be busy." Non-bottleneck stations must have idle time; if they do not, they are starving the bottleneck or building inventory in front of it. Idle time at a non-constraint is free.
  • "Utilization above 100 percent means we are working hard." Utilization cannot exceed 100 percent by definition. What you mean is implied utilization, and above 100 percent it means the queue is growing.
  • "Capacity is a fixed number." It depends on product mix, batch size, setup practice, and staffing pattern. The same equipment can have very different capacity next month.

Recap

  • Name the flow unit first; every capacity and time is expressed per flow unit.
  • Resource capacity equals parallel units divided by activity time; the bottleneck has the lowest capacity, and process capacity equals bottleneck capacity.
  • Flow rate is the smaller of demand and capacity; utilization is flow rate over capacity, and implied utilization is demand over capacity and may exceed 100 percent.
  • Takt time is available time over required output; balance efficiency measures how much station time indivisible tasks leave idle.
  • Design, effective, and actual capacity differ by planned and unplanned losses; utilization and efficiency measure different gaps.
  • Long setups make large batches economical and processes sluggish; cutting setup time buys capacity without the inventory penalty.

Sources

  1. MIT OpenCourseWare. (2013). Introduction to operations management: process analysis. Massachusetts Institute of Technology. ocw.mit.edu
  2. Wikipedia. (2025). Bottleneck (production). en.wikipedia.org
  3. Association for Supply Chain Management. (2025). Capacity and process terminology. ASCM. ascm.org
  4. OpenStax. (2018). Production and operations management. In Introduction to Business. Rice University. openstax.org
Key terms
Flow unit
The item being tracked through a process, such as a patient, sandwich, or claim; all rates and times are expressed per flow unit.
Activity time
The time a resource takes to process one flow unit, sometimes called unit load or processing time.
Capacity
The maximum output rate of a resource or process; for a resource, parallel units divided by activity time.
Bottleneck
The resource with the lowest capacity in a process; it sets the capacity of the entire process.
Flow rate
Actual output per unit time, equal to the smaller of demand rate and process capacity.
Implied utilization
Demand rate divided by capacity; unlike utilization it can exceed 100 percent, showing how far demand outruns the resource.
Takt time
Available production time divided by required output; the rhythm at which finished units must emerge to meet demand.
Capacity cushion
Deliberate slack capacity, equal to 100 percent minus planned average utilization, held to absorb variability.

Little's Law, Flow Time, and Process Types

  • State Little's Law and apply it in all three directions to solve for inventory, flow rate, or flow time.
  • Distinguish flow time, cycle time, lead time, and takt time, and compute flow time efficiency.
  • Translate Little's Law into financial terms using inventory turns and the cash-to-cash cycle.
  • Classify processes from project through continuous flow and explain the product-process matrix.

The big picture

In 1961 an MIT professor named John Little published a proof of a relationship so simple that people had been using it informally for decades without knowing it was always true. It says that the average number of things in a system equals the rate at which things arrive multiplied by the average time each one spends there. Written out:

Inventory = Flow rate x Flow time, usually abbreviated I = R x T.

That is the whole thing. It looks trivial. It is the most useful equation in operations, for three reasons. First, it holds for any stable system regardless of the distribution of arrivals, the distribution of service times, the queue discipline, or the number of servers. It does not care whether your arrivals are Poisson or bunched or scheduled. It only requires that the system be stable over the measurement period, meaning what goes in eventually comes out and the average inventory is not trending. Second, it works at any level: one machine, one department, a whole factory, an entire supply chain. Third, and this is the practical payoff, you almost always know two of the three quantities and want the third.

Let us use it three ways.

Solving for flow time: how long is a patient in the department?

An emergency department treats 9 patients per hour on average. A census taken at random moments over several weeks finds an average of 12 patients physically in the department.

T = I / R = 12 patients / 9 patients per hour = 1.33 hours = 80 minutes.

The average patient spends eighty minutes there. Notice what you did not need: no timing study, no tracking individual patients, no chart review. Two numbers that any department already collects gave you the average length of stay. This is why Little's Law matters operationally. Flow time is expensive to measure directly and cheap to infer.

Solving for inventory: how many calls are in the system?

A call center handles 300 calls per hour. From the moment a caller connects to the moment they hang up, including the time on hold, the average is 6 minutes, which is 0.1 hours.

I = R x T = 300 x 0.1 = 30 calls in the system at any instant.

If you staff 22 agents, then 22 of those 30 are being talked to and about 8 are on hold at any moment. That number tells you whether your hold music budget is the problem or your staffing is. Run the same calculation for a supply chain: a distributor ships 500 units a day and ocean transit takes 21 days, so I = 500 x 21 = 10,500 units are floating on the water at any moment, financed, insured, and unavailable. Module 6 returns to that number when we price near-shoring.

Solving for flow rate: how fast is this line moving?

A bank branch has, on average, 8 customers waiting in line, and a customer who joins the line waits on average 4 minutes before reaching a teller.

R = I / T = 8 / (4/60 hours) = 8 / 0.0667 = 120 customers per hour.

Notice that this application used the queue alone, not the whole branch. Little's Law applies to any subsystem you can draw a boundary around, as long as you use the inventory, flow rate, and flow time of that same boundary. Draw the box, count what is in it, and the law holds.

Key idea: I = R x T holds for any stable system, whatever the distributions. Measure the two easy quantities and infer the hard one.

Four times that get confused, and one table that fixes it

Students lose more points on this vocabulary than on any calculation in the course. Learn it once, properly.

TermDefinitionSandwich shop example
Flow time (throughput time)Time one flow unit spends from entry to exit, including all waitingYou walk in at 12:04 and leave with your sandwich at 12:13: flow time is 9 minutes
Cycle timeTime between successive units leaving the process; the inverse of flow rateA sandwich comes out every 40 seconds
Takt timeAvailable time divided by required output; the beat demand requiresTo serve 90 customers in the noon hour, one sandwich must exit every 40 seconds
Lead timeTime from a customer placing an order to receiving it; often includes queueing before the process even startsA catering order placed Monday for Thursday has a 3-day lead time
Value-added timeTime actually spent transforming the flow unit145 seconds of assembling, toasting, and wrapping

The gap between flow time and value-added time is where operations lives. Flow time efficiency = value-added time / flow time. In the sandwich case: 145 seconds of work in a 9-minute (540-second) visit is 27 percent, which is unusually good because a sandwich shop is a short, simple process.

Now run the same calculation on an office process. A purchase requisition at a mid-sized company takes six business days from submission to approved purchase order. Someone times the actual work: 5 minutes of the requester filling the form, 8 minutes of a manager reviewing, 15 minutes of procurement checking contracts and pricing, 10 minutes of finance coding it, and 7 minutes of issuing the order. Total value-added time is 45 minutes. Flow time, at eight working hours per day, is 2,880 minutes.

Flow time efficiency = 45 / 2,880 = 1.6 percent.

For 98.4 percent of its life, the requisition is sitting in somebody's inbox. This ratio, computed for office and service processes, is routinely between one and five percent, and it explains why speeding up the work almost never speeds up the process. If you doubled the speed of every human in that chain, flow time would fall from six days to five days and twenty-two minutes. The queues are the process. Attacking them is where the time is.

Key idea: Flow time is dominated by waiting, not working. Flow time efficiency of one to five percent is normal in office processes, so improvement means removing queues and handoffs, not working faster.

Little's Law in dollars

Accountants have been using Little's Law for a century under different names. Inventory turns = annual cost of goods sold / average inventory value, and days of inventory = 365 / turns, which is just flow time for a dollar of goods.

A retailer with 3.6 billion dollars in annual cost of goods sold and 400 million dollars of average inventory turns 3.6 billion / 400 million = 9 times a year, which is 365 / 9 = 40.6 days of inventory. Every item sits, on average, forty days. For context, the U.S. Census Bureau's monthly inventories-to-sales ratio for total business has hovered around 1.3 to 1.4 in recent years, implying roughly forty days of inventory across the whole economy.

Stretch the boundary and you get the cash-to-cash cycle, the number of days your money is tied up:

Cash-to-cash = days of inventory + days sales outstanding - days payables outstanding.

A distributor with 45 days of inventory, 30 days to collect from customers, and 60 days before it pays suppliers has a cycle of 45 + 30 - 60 = 15 days. It finances fifteen days of operations. Dell in the late 1990s built to order, held roughly a week of inventory, collected from consumers immediately by credit card, and paid suppliers on terms, producing a famously negative cash-to-cash cycle: suppliers financed Dell's growth. That is not an accounting trick. It is process design, visible in the operations and only later in the balance sheet.

Process types: five families

Two variables, volume and variety, sort nearly every operation into one of five families. They move in opposite directions: as volume rises, variety must fall, because the equipment and layout that make high volume cheap are precisely the ones that make change expensive.

TypeVolume / varietyLayoutExamplesCost structure
ProjectOne of a kind, uniqueFixed position: resources come to the productA bridge, a film, a satellite, an auditVery high unit cost, low fixed investment in dedicated equipment
Job shopLow volume, very high varietyProcess layout: like machines grouped togetherMachine shop, print shop, emergency department, custom tailorHigh unit cost, flexible general-purpose equipment, long queues
BatchModerate volume, moderate varietyProcess or cellularBakery, apparel, paint, specialty chemicals, textbook printingSetups dominate; batch size drives cost
Repetitive / assembly lineHigh volume, low variety with optionsProduct layout: arranged in sequenceCars, appliances, fast food, phone assemblyLow unit cost, high fixed cost, changeovers expensive
Continuous flowVery high volume, one productProduct layout, often literally pipesOil refining, paper, sugar, electricity, float glassLowest unit cost, enormous fixed cost, runs 24 hours because stopping is ruinous

Robert Hayes and Steven Wheelwright drew this as the product-process matrix in 1979, with product variety on one axis and process type on the other. Their argument was that viable operations sit near the diagonal, and firms that drift off it get punished. Off the diagonal in one direction you have a custom furniture maker who installed a high-volume line and now has enormous fixed costs spread over tiny batches. Off it in the other direction you have a commodity producer using a job shop, whose unit costs cannot survive contact with a competitor running continuous flow.

The interesting exception is mass customization: high variety at near mass-production cost. It is real, and it works through specific mechanisms rather than by wishing. Modularity means the product is assembled from standard modules whose combinations create the variety. Postponement means the customizing step is pushed as late as possible: Benetton famously knitted sweaters in undyed yarn and dyed them once color demand was known, and paint stores stock white base and tint at the counter rather than stocking two thousand premixed colors. Late-stage configuration is why a Subway sandwich has millions of possible combinations built at line speed: the variety is added at the last station, by the customer, from standardized components. Note that in every case the upstream process is still high-volume and low-variety. Mass customization does not repeal the matrix; it moves the customization point.

Key idea: Volume and variety trade off, and the process type must match the product. Mass customization works by postponing variety to the last possible step, not by abolishing the trade-off.

Service processes have their own matrix

Roger Schmenner's service process matrix sorts services by labor intensity and by the degree of customization and interaction. Low labor intensity with low customization gives the service factory: airlines, trucking, hotels, where the challenge is capital utilization and demand smoothing. Low labor intensity with high customization gives the service shop: hospitals, auto repair, where the challenge is scheduling variable work through expensive equipment. High labor intensity with low customization gives mass service: retail, retail banking, schools, where the challenge is hiring, training, and managing a large dispersed workforce. High labor intensity with high customization gives professional service: law, medicine, consulting, architecture, where the challenge is cost control and the sheer variability of what each client needs. The tools you reach for change with the quadrant, which is why an airline and a law firm do not manage capacity the same way even though both sell time.

Common misconceptions

  • "Little's Law requires random arrivals or exponential service times." It requires nothing of the sort. It is distribution-free and holds exactly for any stable system over a long enough period.
  • "Cycle time and flow time are the same thing." Cycle time is the gap between successive outputs; flow time is one unit's journey. A line producing a car every 60 seconds may take 20 hours to build any individual car.
  • "Faster workers mean faster processes." When flow time efficiency is 1.6 percent, working faster touches 1.6 percent of the elapsed time. Remove a handoff or a batching rule instead.
  • "Mass customization proves the volume-variety trade-off is dead." It works by keeping the upstream process standardized and moving the customization point downstream. The trade-off is relocated, not repealed.

Recap

  • Little's Law says inventory equals flow rate times flow time, holds for any stable system, and can be solved in all three directions.
  • Flow time, cycle time, takt time, lead time, and value-added time are five different quantities; flow time efficiency is value-added time over flow time and is often under five percent.
  • Inventory turns, days of inventory, and the cash-to-cash cycle are Little's Law expressed in money.
  • Five process types run from project through job shop, batch, line, and continuous flow, trading variety for volume.
  • The product-process matrix says viable operations sit near the diagonal; mass customization succeeds by postponing variety, not by abolishing the trade-off.

Sources

  1. Wikipedia. (2025). Little's law. en.wikipedia.org
  2. U.S. Census Bureau. (2025). Manufacturing and trade inventories and sales. Economic Indicators. census.gov
  3. Hayes, R. H., & Wheelwright, S. C. (1979). Link manufacturing process and product life cycles. Harvard Business Review, 57(1), 133-140. hbr.org
  4. Wikipedia. (2025). Mass customization. en.wikipedia.org
  5. MIT OpenCourseWare. (2013). Introduction to operations management. Massachusetts Institute of Technology. ocw.mit.edu
Key terms
Little's Law
Inventory equals flow rate times flow time (I = R x T), true for any stable system regardless of distributions.
Flow time
The total time one flow unit spends in the process, including all waiting.
Flow time efficiency
Value-added time divided by flow time; commonly one to five percent in office and service processes.
Inventory turns
Annual cost of goods sold divided by average inventory value; Little's Law stated in financial terms.
Cash-to-cash cycle
Days of inventory plus days sales outstanding minus days payables outstanding; the days a firm finances its own operations.
Job shop
A low-volume, high-variety process with a functional layout in which similar machines are grouped together.
Continuous flow
A very high volume, single-product process with enormous fixed cost that runs without stopping.
Postponement
Delaying the step that creates product variety until as late as possible, the central mechanism of mass customization.

Module 2: Variability, Waiting, and Quality

Why queues form in processes that are not full, how to compute the cost of running hot, and how to measure and manage the variation that causes both waiting and defects.

Why Lines Form: Variability, Queues, and Service Design

  • Demonstrate that variability alone creates waiting even when average capacity exceeds average demand.
  • Use the utilization factor to estimate how waiting time explodes as a process approaches full capacity.
  • Compute the benefit of pooling servers and explain when pooling is and is not appropriate.
  • Apply variability reduction, capacity cushions, and queue psychology to service system design.

The big picture

Here is a puzzle that trips up almost everyone the first time. A barber can cut hair in exactly ten minutes. Customers arrive at an average of one every twelve minutes. Average demand, five per hour, is comfortably below average capacity, six per hour. Utilization is 10/12, about 83 percent. How long does the average customer wait?

The intuitive answer is zero, and if arrivals really were exactly every twelve minutes, the intuitive answer would be correct: customer one arrives at 0 and leaves at 10, customer two arrives at 12 to an empty chair, and no one ever waits. The barber is idle two minutes out of every twelve, forever. Perfect regularity, no queue.

Now keep every average identical and let the arrivals vary. Same mean gap of twelve minutes, same ten-minute service, just realistic spacing.

CustomerArrivesService startsService endsWait
100100
2410206
32424340
43034444
54848580
65358685
77272820

The gaps are 4, 20, 6, 18, 5, and 19 minutes, which average exactly 12. Service is exactly 10 minutes every time. Average capacity still exceeds average demand. And the average wait is (0 + 6 + 0 + 4 + 0 + 5 + 0) / 7 = 2.1 minutes, with the barber idle for 12 of the 82 minutes.

Nothing about the averages changed. Only the variability changed, and waiting appeared out of nothing. Work through why: when two customers arrive close together, the second one waits, and that wait cannot be repaid later. When a long gap follows, the barber sits idle, and idle time is lost forever. Queueing systems have a ratchet in them. Bunching creates waits that idleness cannot undo.

Key idea: Variability, not overload, is what creates waiting lines. A process with average capacity above average demand will still have queues if arrivals or service times vary at all.

The utilization factor, and why running hot is so expensive

The second force is utilization. Sir John Kingman published an approximation in 1961 that is close enough for management work and simple enough to hold in your head. The expected wait in queue is:

Wait = (variability factor) x (utilization factor) x (average service time)

where the utilization factor is u / (1 - u) with u the utilization, and the variability factor is the average of the squared coefficients of variation of arrivals and service, that is (CVa squared + CVs squared) / 2. A coefficient of variation is just the standard deviation divided by the mean, so a CV of 1 describes the completely random arrival pattern typical of walk-in customers, a CV of 0 describes perfect regularity, and a CV above 1 describes bunching.

Ignore the variability factor for a moment by setting both CVs to 1, which makes it exactly 1, and look only at what utilization does.

Utilization uu / (1 - u)Wait with 10-minute service
50 percent1.010 minutes
70 percent2.3323 minutes
80 percent4.040 minutes
90 percent9.090 minutes
95 percent19.0190 minutes
98 percent49.0490 minutes

Read that table twice, because it explains more about the world than any other table in this course. Waiting does not rise linearly with utilization; it rises hyperbolically and goes to infinity as utilization approaches one. Moving from 50 to 80 percent utilization multiplies the wait by four. Moving from 90 to 95 percent, an increase of five percentage points, doubles the wait again.

Turn it around and it becomes a management tool. A clinic at 95 percent utilization has a 190-minute average wait. Add enough capacity to drop utilization to 90 percent, roughly a five percent capacity increase, and the wait falls to 90 minutes. That is a 53 percent reduction in waiting for a five percent increase in cost. No amount of exhorting the staff to hurry produces anything close to that. Conversely, the executive who says "our machine is only 85 percent utilized, we can absorb more volume" is proposing to buy a small revenue increase with an enormous increase in delay.

This is also the honest answer to why hospitals, fire departments, and emergency rooms cannot run at the utilization levels a factory can. Their arrival variability is high and their service variability is enormous, so the variability factor is large, and the only lever left is to hold utilization down. A capacity cushion is not waste in these settings. It is the product.

Key idea: Waiting rises with u / (1 - u), which explodes near full utilization. Small capacity additions near the top of that curve buy enormous reductions in delay.

Pooling: the cheapest improvement in service operations

Suppose you run two service desks in the same building, each with its own line, each with one server. Service takes 10 minutes on average, so each server can handle 6 per hour. Each desk gets 5.4 customers per hour, so each runs at 90 percent utilization. Using standard single-server queueing results with random arrivals and service, the average number waiting at each desk is 8.1 people, and the average wait is 8.1 / 5.4 = 1.5 hours, or 90 minutes.

Now knock down the wall. One line feeds both servers. Total arrivals are 10.8 per hour against a combined capacity of 12 per hour, so utilization is still 90 percent. Nothing has been added: same two servers, same arrival rate, same service times. The standard two-server result gives an average of 7.67 people waiting, and the wait is 7.67 / 10.8 = 0.71 hours, or about 43 minutes.

The wait fell by more than half for free. The mechanism is easy to see once you look for it: in the separated system, one server can be idle while a customer waits in the other line. That is pure waste, and it happens constantly. Pooling eliminates it. This is why banks, airports, and pharmacies converted to single serpentine lines feeding multiple windows, and it is why the deli counter's number system beats the free-for-all at the bakery next door.

Pooling is not free of downsides, and a good operations manager knows all four. First, specialization is lost: if one desk handled only passport renewals and the other only visa applications, the pooled servers must know both, and average service time may rise. Second, the customer experience can worsen: one long line looks intimidating even when it moves faster, which is why some retailers keep separate lines despite knowing the arithmetic. Third, priority mixing: pooling a five-minute transaction with a forty-minute one makes the quick customers wait behind the slow ones, which is why supermarkets keep an express lane and hospitals triage. Fourth, physical distance: pooling call centers across three cities works because calls travel free; pooling emergency rooms across three cities does not.

The same mathematics reappears in Module 3 as inventory pooling and in Module 6 as network design. Whenever independent random demands are combined, the combined variability grows more slowly than the combined volume, and that gap is the pooling benefit.

Key idea: Combining separate queues into one line feeding several servers cuts waiting sharply at zero cost, because it prevents a server from being idle while someone waits elsewhere.

The three levers, in the order you should reach for them

Kingman's approximation is three factors multiplied together, so there are exactly three things you can do about a wait.

  1. Reduce variability. This is usually the cheapest and is almost always tried last. Appointment systems and reservations convert random arrivals into scheduled ones, dropping arrival CV toward zero. Standardized work and better training reduce service CV. Restricting scope, as urgent care does by refusing complex cases, reduces service variability by construction. Triage separates the fast from the slow so each stream has low internal variability. Requiring complete paperwork up front eliminates the enormous service-time outliers caused by missing information.
  2. Reduce utilization. Add capacity, or shift demand. Shifting demand is often free: off-peak pricing (matinee tickets, happy hour, off-peak transit fares, utility time-of-use rates), appointment slots spread across the day, and staggered class schedules all move demand from the steep part of the curve to the flat part. Flexible capacity does the same thing from the other side: part-time staff during the lunch rush, cross-trained employees who can move to whichever station is drowning, on-call arrangements, and self-service options that peel off the simple transactions.
  3. Reduce service time. The most obvious lever and often the weakest, because Kingman multiplies service time by the utilization factor, and cutting service time also cuts utilization, so improvements here do compound. But service time is bounded below by the work itself, while the utilization factor has no upper bound.

When you cannot remove the wait, manage the waiting

Sometimes the queue is unavoidable, and then the perception of waiting becomes the design problem. David Maister's 1985 propositions on the psychology of waiting lines have held up well:

  • Unoccupied time feels longer than occupied time. Mirrors by elevators, menus handed out in line, and status screens all fill the time.
  • Pre-process waits feel longer than in-process waits. Getting someone seated with a menu, or moving a patient from the waiting room into an exam room, converts one to the other even if nothing else changes.
  • Anxiety makes waits feel longer. "Did they forget me?" is worse than the wait itself.
  • Unexplained waits feel longer than explained ones. "The doctor is with an emergency" measurably shortens the felt wait.
  • Uncertain waits feel longer than known, finite ones. A countdown, a position in queue, or a delivery tracking map converts an uncertain wait into a bounded one.
  • Unfair waits are by far the worst. A single line exists partly to guarantee first-come first-served, which people care about intensely.

The often-told airport example illustrates the point. A large airport received persistent complaints about baggage claim waits. Adding baggage handlers cut the wait but the complaints continued. The airport then routed arriving passengers to a more distant carousel, so the walk took most of the previously idle waiting time. Complaints dropped sharply. The objective time to bag was similar; the unoccupied portion was nearly gone. Treat this as a lesson in perception management rather than an excuse for it: the honest version of this move is to occupy the wait usefully, not to hide it.

Key idea: When a wait cannot be removed, occupy it, explain it, bound it, and make it visibly fair. Perceived waiting time responds to design even when actual waiting time cannot.

Common misconceptions

  • "If capacity exceeds demand there should be no line." True only if nothing varies. Any variability in arrivals or service produces queues at any utilization above zero.
  • "High utilization is good management." High utilization is good for expensive, low-variability equipment and terrible for high-variability service systems. Ninety-five percent utilization in a clinic is a waiting-room crisis with a spreadsheet defense.
  • "You need exact distributions to use queueing theory." The approximation used here needs only means and coefficients of variation, both of which you can estimate from a morning of observation.
  • "Adding a server doubles throughput and halves the wait." It roughly doubles capacity, but the effect on waiting is far larger than proportional because of where you sit on the u / (1 - u) curve.

Recap

  • Variability alone creates waiting; a hand-simulated barber with identical averages produced a 2.1-minute average wait purely from irregular arrivals.
  • Kingman's approximation multiplies a variability factor, the utilization factor u / (1 - u), and the average service time.
  • The utilization factor explodes near capacity: 1.0 at 50 percent, 9.0 at 90 percent, 49.0 at 98 percent, so small capacity additions near the top buy large reductions in delay.
  • Pooling two 90 percent utilized single-server desks into one two-server queue cut the wait from about 90 minutes to about 43 with no added resources.
  • The three levers are reducing variability, reducing utilization, and reducing service time, in roughly that order of cost-effectiveness.
  • Where waits are unavoidable, occupy, explain, bound, and equalize them.

Sources

  1. Wikipedia. (2025). Kingman's formula. en.wikipedia.org
  2. Wikipedia. (2025). Queueing theory. en.wikipedia.org
  3. Maister, D. H. (1985). The psychology of waiting lines. In J. Czepiel et al. (Eds.), The Service Encounter. Lexington Books. hbr.org
  4. MIT OpenCourseWare. (2013). Introduction to operations management: variability and queues. Massachusetts Institute of Technology. ocw.mit.edu
Key terms
Coefficient of variation
Standard deviation divided by the mean; a unitless measure of variability, equal to 1 for completely random arrivals.
Utilization factor
The quantity u / (1 - u) in Kingman's approximation, which grows without bound as utilization approaches one.
Kingman's approximation
Expected queue wait equals the average of squared coefficients of variation, times u / (1 - u), times average service time.
Pooling
Serving one combined queue with several servers instead of maintaining separate queues, which sharply reduces waiting at no added cost.
Capacity cushion
Deliberate slack held below full utilization; in high-variability services it is the mechanism that keeps waits finite.
Demand shaping
Moving demand rather than adding capacity, using off-peak pricing, appointments, or staggered scheduling.
Pre-process wait
Waiting before service has visibly begun, which is perceived as longer than an equivalent wait after service starts.

Quality and Statistical Process Control

  • Distinguish common-cause from special-cause variation and explain why the distinction determines the correct response.
  • Build x-bar and R control charts from raw subgroup data and interpret out-of-control signals.
  • Construct a p-chart for attribute data and apply run rules without inviting false alarms.
  • Explain why control limits and specification limits are different things with different sources.

The big picture

A packaging line fills bags of coffee to a target of 340 grams. This morning a bag came off at 343 grams. The supervisor adjusts the filler down. Twenty minutes later a bag reads 337, and she adjusts up. By noon the line is producing bags scattered from 330 to 350, far worse than when she started, and she is certain that without her constant attention it would be worse still.

She has done something with a name: tampering. She has treated ordinary, expected variation as if it were a signal, and every correction added a new error on top of the one already there. W. Edwards Deming demonstrated this with a funnel and a marble in front of thousands of executives, showing that a funnel adjusted after each drop produces a wider scatter than a funnel left alone. The costliest quality decisions in most organizations are not failures to act. They are actions taken on noise.

Avoiding that mistake requires a way to tell noise from signal, and that is exactly what a control chart does. Principles of Management (MGMT 301) introduces Deming, total quality management, and the general idea that quality is built in rather than inspected in. This lesson takes the tool apart and builds one from raw numbers.

What quality actually means

Three definitions circulate, and they are not the same. Joseph Juran called quality fitness for use, which puts the customer's purpose at the center. Philip Crosby called it conformance to requirements, which puts the specification at the center and makes quality measurable. Deming emphasized predictable uniformity at low cost, suited to the market. In practice you need Crosby's definition to run a process and Juran's to decide what the specification should be in the first place.

David Garvin's eight dimensions are a useful checklist when you are deciding what to measure: performance, features, reliability, conformance, durability, serviceability, aesthetics, and perceived quality. A product can be excellent on some and poor on others, and the honest question is which ones the customer is actually paying for. For services, the widely used SERVQUAL framework offers five dimensions: reliability, assurance, tangibles, empathy, and responsiveness. Across many studies, reliability, meaning doing what you promised when you promised, dominates the others in driving customer judgment. That is a genuinely useful finding for a manager choosing where to spend.

Two kinds of variation, and only two

Walter Shewhart's central insight, published in 1931 and still the foundation of the field, is that variation comes in two flavors that require opposite responses.

  • Common-cause variation is the natural, ever-present scatter produced by the process as it is currently designed: slight differences in ambient temperature, material density, operator technique, and machine wear. It is stable and predictable within limits. A process showing only common-cause variation is in statistical control, which does not mean it is good, only that it is predictable. Responding to individual common-cause points is tampering, and it makes things worse.
  • Special-cause variation (also assignable-cause) comes from something specific that was not there before: a new lot of material, a tool that broke, an untrained substitute, a valve that stuck. It is not predictable from past data. Responding to it by finding and removing the cause is exactly right.

The management implication is sharp. If a process is in control and you do not like its output, do not chase individual data points and do not blame the operator. Change the process: different equipment, different material, different method. Deming estimated that the large majority of quality problems belong to the system, and therefore to management, rather than to the worker. His red bead experiment made the point brutally by having "willing workers" scoop beads from a container containing a fixed proportion of red ones, then praising and firing them based on results that were pure sampling noise.

Key idea: Common-cause variation calls for changing the system; special-cause variation calls for finding the specific cause. Reversing the two either tampers with a stable process or ignores a real problem.

Building an x-bar and R chart from raw data

Back to the coffee bags. Instead of reacting to single bags, the line takes a subgroup of five consecutive bags every half hour and records the net weight in grams.

SubgroupMeasurements (grams)Mean (x-bar)Range (R)
1341, 339, 342, 338, 3403404
2343, 340, 339, 341, 3423414
3338, 337, 340, 339, 3413394
4340, 342, 341, 339, 3383404
5342, 341, 343, 340, 3393414
6339, 338, 340, 341, 3373394

Two averages drive everything: the grand mean and the average range.

  • Grand mean = (340 + 341 + 339 + 340 + 341 + 339) / 6 = 2,040 / 6 = 340.0 grams.
  • Average range = (4 + 4 + 4 + 4 + 4 + 4) / 6 = 4.0 grams.

Control limits use published constants that depend on subgroup size. For a subgroup of 5, the standard values are A2 = 0.577, D3 = 0, D4 = 2.114, and d2 = 2.326. These come from the statistics of the range and are tabulated in every quality handbook; you look them up, you do not derive them.

X-bar chart limits:

  • Upper control limit = grand mean + A2 x average range = 340.0 + (0.577 x 4.0) = 340.0 + 2.31 = 342.31 grams.
  • Lower control limit = 340.0 - 2.31 = 337.69 grams.

R chart limits:

  • Upper control limit = D4 x average range = 2.114 x 4.0 = 8.46 grams.
  • Lower control limit = D3 x average range = 0 x 4.0 = 0.

Every subgroup mean (339 to 341) sits inside 337.69 to 342.31, and every range (4) sits inside 0 to 8.46. The process is in statistical control. The supervisor should leave the filler alone. Every one of her morning adjustments was tampering.

Always read the R chart first. If the spread is unstable, the x-bar limits computed from that spread are meaningless, because they were derived from the average range. Stability of dispersion is a precondition for interpreting the location.

Now take a seventh subgroup: 344, 345, 343, 346, 342. Mean = 344, range = 4.

The range is fine, so dispersion has not changed, but the mean of 344 is above the upper control limit of 342.31. That is a special-cause signal, and now action is correct. Note the diagnostic value of the pattern: the spread is unchanged and the location has shifted, which points at things that move the whole process at once (a new setting, a new material lot, a changed temperature) rather than things that add scatter (a worn bearing, an inconsistent operator).

One more number falls out for free. The process standard deviation can be estimated as average range / d2 = 4.0 / 2.326 = 1.72 grams. The next lesson uses that number to answer the question the control chart cannot: whether this process can actually hold the customer's tolerance.

Key idea: Control limits are computed from the process's own variation, using the grand mean and the average range. A point outside them is a signal; points inside them are noise, and acting on noise makes output worse.

Control limits are not specification limits

This is the single most important confusion in quality management, and it costs organizations real money every year.

Specification limits come from the customer, the engineer, or the regulator. They say what is acceptable. In our example, the customer might require 340 grams plus or minus 5, so 335 to 345.

Control limits come from the process itself. They say what the process actually does when nothing unusual is happening. Here, 337.69 to 342.31.

Nothing forces these to line up. A process can be in perfect statistical control and produce mostly scrap, if its natural spread is wider than the specification. A process can be wildly out of control and still produce parts within specification, if its natural spread is much narrower. Putting specification limits on a control chart, which people do constantly, destroys the chart, because you can no longer tell whether a point outside the line means "something changed" or "the customer would reject this one." They answer different questions and belong on different charts. The bridge between them is process capability, which is the first topic of the next lesson.

Attribute data: the p-chart

Not everything is measured on a scale. Sometimes each item is simply conforming or not: an invoice is correct or wrong, a flight is on time or late, a surgical site is infected or not. For proportions, use a p-chart.

An accounts payable team audits 200 invoices a week for twenty weeks and finds an overall defect proportion of 0.03, meaning 3 percent of invoices contain an error.

  • Standard error of the proportion = square root of [ p x (1 - p) / n ] = square root of [ 0.03 x 0.97 / 200 ] = square root of 0.0001455 = 0.01206.
  • Upper control limit = 0.03 + 3 x 0.01206 = 0.03 + 0.0362 = 0.0662, or 6.6 percent.
  • Lower control limit = 0.03 - 0.0362 = -0.0062, which is impossible, so it is set to 0.

So a week with 10 bad invoices out of 200 (5 percent) is noise; do not investigate. A week with 18 out of 200 (9 percent) is above the upper limit and worth a genuine root cause hunt. A week with zero errors is inside the limits and is not evidence that the new training worked, which is exactly the kind of premature celebration control charts exist to prevent.

For counts rather than proportions, such as defects per vehicle or complaints per thousand room-nights, the analogous tool is the c-chart, built on the Poisson distribution, with limits at the mean plus or minus three times the square root of the mean.

Run rules, and how to avoid drowning in false alarms

A point outside three sigma is not the only signal. The Western Electric rules add patterns that are individually unlikely inside the limits:

  • One point beyond three sigma from the center line.
  • Two of three consecutive points beyond two sigma on the same side.
  • Four of five consecutive points beyond one sigma on the same side.
  • Eight consecutive points on the same side of the center line.
  • Six consecutive points steadily increasing or decreasing.

These catch small sustained shifts that a three-sigma rule alone would miss for a long time. But there is an honest cost. Each rule carries its own false alarm rate, and applying all of them together roughly triples the chance of a false signal compared with the three-sigma rule alone. A chart that cries wolf weekly gets ignored within a month. The practical advice is to use the three-sigma rule plus one or two pattern rules that match the failure modes you actually have, and to write down in advance what you will do when a signal appears.

Key idea: Run rules increase sensitivity to small shifts at the cost of more false alarms. Choose a small set deliberately rather than switching on everything the software offers.

Inspection cannot create quality

Deming's third point for management was to cease dependence on mass inspection. The arithmetic behind that slogan is worth stating: inspection is itself a process with its own error rate, typically finding 80 to 90 percent of defects when done by a human on a repetitive visual task, and it adds cost without adding value. Two hundred percent inspection, where every item is checked twice, characteristically finds fewer defects per pass than one hundred percent inspection, because each inspector relaxes knowing another will look.

Where inspection does belong is at points where it prevents wasted downstream work, especially just before an expensive or constrained operation, and at the boundary where a defect would reach a customer. Module 5 shows the alternative that actually works: build detection into the process itself with mistake-proofing devices and a stop-the-line rule, so defects are caught at the moment and place they are created.

Note also that control charts apply far beyond factories. Hospitals chart infection rates and door-to-balloon times, call centers chart handle time and abandonment, airlines chart on-time departure, and public health agencies chart disease counts. The mathematics does not know or care what the flow unit is.

Common misconceptions

  • "A point outside the control limits means a defective product." It means the process changed. The item may be well within specification, and an in-control process may produce items outside specification.
  • "In control means good." In control means predictable. A process can be predictably terrible, which is precisely why capability analysis exists.
  • "More adjustments mean more care." Adjusting a stable process on individual readings is tampering, and it demonstrably widens the output distribution.
  • "SPC is for manufacturing." Any repeated process with measurable output can be charted, and healthcare quality improvement now leans on these charts heavily.

Recap

  • Quality is defined as fitness for use, conformance to requirements, or predictable uniformity; you need more than one definition to both set and hold a standard.
  • Variation is either common cause, which requires changing the system, or special cause, which requires finding a specific event; confusing them produces tampering.
  • An x-bar and R chart is built from a grand mean and an average range using tabulated constants; here the limits were 337.69 to 342.31 grams with a range limit of 8.46.
  • Read the R chart first, because unstable dispersion invalidates the x-bar limits computed from it.
  • Control limits describe the process; specification limits describe the customer's requirement, and the two must never be drawn on the same chart.
  • A p-chart handles proportion data, with limits at the mean plus or minus three standard errors, truncated at zero.
  • Run rules add sensitivity and false alarms in equal measure; pick a small set and define the response in advance.

Sources

  1. American Society for Quality. (2025). Statistical process control and quality tools. ASQ. asq.org
  2. Wikipedia. (2025). Control chart. en.wikipedia.org
  3. Britannica. (2025). W. Edwards Deming. Encyclopaedia Britannica. britannica.com
  4. National Institute of Standards and Technology. (2013). Process or product monitoring and control. In NIST/SEMATECH e-Handbook of Statistical Methods. nist.gov
Key terms
Common-cause variation
The natural, stable scatter inherent in a process as designed; responding to individual points of it is tampering.
Special-cause variation
Variation from a specific identifiable event not part of the normal process, which warrants investigation and removal.
Tampering
Adjusting a stable process in response to ordinary variation, which demonstrably increases the spread of output.
Subgroup
A small set of consecutive items measured together, forming one point on a control chart.
Control limits
Boundaries computed from the process's own variation, marking the range of ordinary output; not the same as specification limits.
Specification limits
Boundaries set by the customer, engineer, or regulator defining acceptable output, independent of what the process actually does.
p-chart
A control chart for the proportion of nonconforming items, with limits at the mean proportion plus or minus three standard errors.
Western Electric rules
A set of pattern-based signals, such as eight consecutive points on one side of the center line, that detect small sustained shifts.

Capability, Six Sigma, and the Cost of Quality

  • Compute Cp and Cpk from process data and explain what each one does and does not capture.
  • Explain the Six Sigma method, the 1.5 sigma shift convention, and the documented critiques of the program.
  • Categorize and compute the four costs of quality and evaluate a prevention investment.
  • Judge when statistical improvement methods fit a problem and when they do not.

The big picture

The control chart from the last lesson answered one question: is this process stable? It deliberately said nothing about whether the process is good enough. A filler can be beautifully stable at 340 grams with a spread of plus or minus 6 grams while the customer requires plus or minus 3. Stability and adequacy are separate properties, and the tool that connects them is process capability.

This lesson does capability, then Six Sigma, then the cost of quality, and it treats all three the way a practitioner should: as genuinely useful arithmetic wrapped in a great deal of consulting mythology that you should be able to strip off.

Cp: can the process fit inside the tolerance at all?

The customer's tolerance has a width: upper specification limit minus lower specification limit. The process has a width too, conventionally taken as six standard deviations, which covers about 99.73 percent of output for a roughly normal distribution. Divide one by the other:

Cp = (upper specification limit - lower specification limit) / (6 x standard deviation)

Continue with the coffee bags. The control chart gave a standard deviation estimate of 1.72 grams. Suppose the specification is 340 grams plus or minus 5, so the limits are 335 and 345.

Cp = (345 - 335) / (6 x 1.72) = 10 / 10.32 = 0.97

A Cp of 0.97 means the process spread is very slightly wider than the tolerance. Even perfectly centered, it will produce output beyond the specification. Interpretation is conventional and worth memorizing:

CpMeaningApproximate nonconforming rate if perfectly centered
Below 1.00Process cannot hold the toleranceMore than 2,700 parts per million
1.00Process exactly fills the toleranceAbout 2,700 per million (0.27 percent)
1.33Common minimum for an established processAbout 63 per million
1.67Common requirement for safety-critical workAbout 0.6 per million
2.00The "six sigma" levelAbout 0.002 per million

To reach a Cp of 1.33 on this line, you need the standard deviation to fall to 10 / (6 x 1.33) = 1.25 grams, a 27 percent reduction in variation. That is a concrete engineering target, which is exactly the point of computing the index.

Cpk: is the process actually centered?

Cp has a serious blind spot: it does not know where the process is sitting. A process with a narrow spread parked right against the upper limit will show an excellent Cp and produce scrap all day. Cpk fixes this by measuring the distance from the process mean to the nearer specification limit in units of three standard deviations:

Cpk = the smaller of [(upper limit - mean) / (3 x standard deviation)] and [(mean - lower limit) / (3 x standard deviation)]

Centered at 340 with standard deviation 1.72:

  • Upper side: (345 - 340) / (3 x 1.72) = 5 / 5.16 = 0.97.
  • Lower side: (340 - 335) / 5.16 = 0.97.
  • Cpk = 0.97, equal to Cp because the process is centered.

Now let the mean drift to 342 with the spread completely unchanged:

  • Upper side: (345 - 342) / 5.16 = 3 / 5.16 = 0.58.
  • Lower side: (342 - 335) / 5.16 = 7 / 5.16 = 1.36.
  • Cpk = 0.58, while Cp is still 0.97.

The variation did not change at all, and capability collapsed by 40 percent. Cp is the potential: what the process could do if perfectly centered. Cpk is the actual. The gap between them is entirely a centering problem, and centering is usually the cheapest fix in the whole quality toolkit, often a machine setting rather than a capital project. Whenever you see Cp much larger than Cpk, go adjust the mean before you buy anything.

Key idea: Cp measures whether the spread fits inside the tolerance; Cpk measures whether the process, where it actually sits, fits. A large gap between them is a centering problem and usually cheap to fix.

Six Sigma: the method, the number, and the convention

Six Sigma began at Motorola in 1986, developed by engineer Bill Smith and later formalized by Mikel Harry, and became famous when General Electric adopted it company-wide under Jack Welch in 1995. Two things travel under the name, and they should be kept apart.

The first is a statistical target. A six sigma process has specification limits six standard deviations from the mean, which is a Cp of 2.0. For a stable, centered, normal process that implies about two defects per billion opportunities. But the famous Six Sigma number is 3.4 defects per million opportunities, which is far worse than two per billion. Where does the difference come from?

It comes from the 1.5 sigma shift, a convention Motorola adopted to account for the fact that real process means drift over the long run: tools wear, materials change, seasons change. The convention assumes the mean may wander up to 1.5 standard deviations off center over time, leaving only 4.5 sigma of margin on the near side, which corresponds to 3.4 defects per million. It is important to be honest about what this is. It is an engineering allowance chosen by a company, not a law of nature and not a derived result. Some processes drift more, some drift far less. When you see 3.4 defects per million quoted as though it were a mathematical constant, you are looking at a convention that has been promoted to a fact.

The second thing that travels under the name is a project method, and this is the part with lasting value. DMAIC structures improvement in five phases:

  • Define the problem, the customer requirement, and the project scope, with a charter that names the financial stake.
  • Measure the current process, which includes validating that the measurement system itself is reliable, a step teams routinely skip and routinely regret.
  • Analyze the data to find root causes rather than plausible stories, using tools such as Pareto charts, cause-and-effect diagrams, hypothesis tests, and regression.
  • Improve by designing, piloting, and verifying a change.
  • Control by locking the gain in with control charts, updated standard work, and a handover to the process owner.

The Control phase is what distinguishes DMAIC from ordinary problem solving, and it is the phase most often abandoned. A version for new designs, DMADV or design for Six Sigma, substitutes Design and Verify for Improve and Control. The belt hierarchy of green belts, black belts, and master black belts provides trained internal capacity and a career ladder.

The critiques, taken seriously

An honest course has to report the other side, and there is a substantial one.

It is a poor fit for innovation. The most cited case is 3M. James McNerney arrived from GE in 2001 and pushed Six Sigma across the company, including into research and development. Efficiency and margins improved. But 3M's own new product vitality index, the share of sales from recently introduced products, fell markedly over that period, and when George Buckley succeeded him in 2005 he explicitly loosened the process discipline in the research labs, saying that invention is by nature a disorderly process that cannot be reduced to a defined sequence. The general point survives the specific case: DMAIC assumes a repeatable process with a measurable defect. Early-stage discovery has neither.

The reported savings are self-reported. GE announced billions of dollars in Six Sigma benefits, and those figures came from GE. Project savings are typically calculated by the team that ran the project, using assumptions that team chose, and frequently never reconcile to an actual reduction in the profit and loss statement. Finance sign-off on savings is the standard fix and is often skipped.

Firm-level evidence is weak. Studies looking at large publicly announced adopters have generally failed to find that they outperformed comparable firms, and a widely reported analysis in the mid-2000s found that a majority of large Six Sigma adopters had trailed the S and P 500 after adoption. Selection effects run both ways here, so this is not proof that the method fails; it is good evidence that adopting a program is not the same as improving.

It can become bureaucracy. Belt certifications, project charters, and tollgate reviews are overhead. When the certification count becomes the metric, you get projects invented to be certified rather than problems chosen because they matter, which is the Goodhart failure that MGMT 301 covers in detail.

The fair conclusion: DMAIC is an excellent discipline for a repeatable, measurable, defect-driven problem in a stable process, and using it there is one of the higher-return activities available to an operations manager. Applying it to exploratory, creative, or highly variable work is a category error, and treating 3.4 defects per million as a universal target is superstition.

Key idea: Six Sigma's project discipline is genuinely valuable for repeatable defect problems; its famous 3.4 per million figure is a drift convention, and its reported financial results deserve the same skepticism you would apply to any self-reported number.

The cost of quality, computed

Armand Feigenbaum and Joseph Juran gave managers the argument that actually moves budgets. Quality costs fall into four buckets, and only two of them are voluntary.

  • Prevention: design reviews, supplier qualification, training, mistake-proofing, process capability studies. Spent before anything goes wrong.
  • Appraisal: inspection, testing, audits, calibration, measurement systems. Spent to find problems that already exist.
  • Internal failure: scrap, rework, downtime, re-testing, downgrading. Discovered before the customer sees it.
  • External failure: warranty claims, returns, field service, recalls, liability, lost customers. Discovered by the customer.

A mid-sized manufacturer with 24 million dollars in annual sales tallies its current spending:

CategoryCurrent annual costAfter investing in prevention
Prevention120,000320,000
Appraisal180,000144,000
Internal failure460,000276,000
External failure940,000470,000
Total1,700,0001,210,000

Current quality costs are 1,700,000 / 24,000,000 = 7.1 percent of sales, which is squarely in the normal range for a firm that has not worked on this. Note the shape of the spending: 82 percent of it is failure cost, meaning money spent on things that already went wrong.

Now the proposal. Spend an extra 200,000 dollars a year on prevention: supplier qualification, operator training, mistake-proofing fixtures, and capability studies on the three worst processes. The projection assumes internal failure falls 40 percent, external failure falls 50 percent, and appraisal falls 20 percent because there is less to catch.

  • New total = 320,000 + 144,000 + 276,000 + 470,000 = 1,210,000 dollars.
  • Gross saving = 1,700,000 - 1,210,000 = 490,000 dollars.
  • Net of the 200,000 investment already included in that total, the firm is 490,000 dollars better off per year on a 200,000 dollar spend, a return of about 2.5 to 1.
  • Quality cost falls from 7.1 percent of sales to 5.0 percent.

The pattern behind these numbers is the widely used rule of thumb that a defect caught in design costs one unit, caught in production costs ten, and caught in the field costs a hundred. Treat the 1-10-100 pattern as a memorable ordering rather than a measured constant; the ratios vary enormously by industry, and in a regulated or safety-critical setting the external multiplier is far higher than a hundred.

Two honest cautions on cost of quality accounting. First, the biggest external failure cost, the customer who quietly never returns, does not appear in any ledger, so every cost of quality study understates the case for prevention. Second, projections like the one above are projections. The disciplined version is to measure the four buckets before and after, and to have finance rather than the project team confirm the after.

Key idea: Most quality spending is failure cost, and shifting money into prevention typically returns more than it costs. The strongest external failure cost, lost customers, is invisible in the accounts, so the true case for prevention is stronger than the numbers show.

Does higher quality cost more?

Split the question and the confusion disappears. Higher performance quality usually does cost more: a better design, better materials, more features. Higher conformance quality usually costs less, because scrap, rework, warranty, and expediting all fall. Crosby's slogan that quality is free was overstated as a slogan but correct about conformance. This distinction resolves most arguments in a room about whether a quality program will pay for itself: ask which kind of quality is being proposed.

A final note on standards. ISO 9001 certification demonstrates that an organization has documented its processes and follows its documentation. That is genuinely valuable for consistency and for supplier qualification, and it is often contractually required. It does not certify that the product is good, because a firm can document a mediocre process faithfully. The Baldrige Excellence Framework, administered by the National Institute of Standards and Technology, evaluates performance results as well as process, which makes it a stronger signal and a much heavier lift.

Common misconceptions

  • "Cp and Cpk mean the same thing." Cp assumes perfect centering. When Cp is much larger than Cpk, the process is off center and the fix is usually a setting, not a purchase.
  • "Six Sigma means 3.4 defects per million everywhere." The 3.4 figure encodes an assumed 1.5 sigma long-term drift. A truly centered six sigma process would run at about two defects per billion.
  • "Quality always costs more." Performance quality costs more; conformance quality generally pays for itself by removing failure costs.
  • "ISO 9001 certification means the products are good." It means the documented process is followed. Consistency, not excellence.

Recap

  • Cp compares tolerance width with six standard deviations of process spread; Cp of 1.33 is a common minimum, and 2.0 is the six sigma level.
  • Cpk measures the distance to the nearer specification limit, so a mean drifting from 340 to 342 grams cut capability from 0.97 to 0.58 with no change in spread.
  • DMAIC is the durable part of Six Sigma; the 3.4 defects per million figure rests on a 1.5 sigma drift convention adopted by Motorola.
  • Documented critiques include poor fit with innovation, as at 3M, self-reported savings, weak firm-level evidence, and certification bureaucracy.
  • Quality costs split into prevention, appraisal, internal failure, and external failure; the worked case moved 200,000 dollars into prevention and cut total quality cost from 1.7 million to 1.21 million.
  • Performance quality costs more; conformance quality usually pays for itself.

Sources

  1. Wikipedia. (2025). Process capability index. en.wikipedia.org
  2. Wikipedia. (2025). Six Sigma. en.wikipedia.org
  3. American Society for Quality. (2025). Cost of quality. ASQ. asq.org
  4. National Institute of Standards and Technology. (2025). Baldrige Excellence Framework. nist.gov
  5. International Organization for Standardization. (2015). ISO 9001: Quality management systems. iso.org
Key terms
Cp
Tolerance width divided by six process standard deviations; the potential capability of a perfectly centered process.
Cpk
The distance from the process mean to the nearer specification limit divided by three standard deviations; actual capability including centering.
1.5 sigma shift
Motorola's convention allowing the process mean to drift 1.5 standard deviations over the long run, which converts a six sigma process into 3.4 defects per million.
DMAIC
The Six Sigma improvement sequence: define, measure, analyze, improve, control.
Prevention cost
Money spent before defects occur, on training, design review, supplier qualification, and mistake-proofing.
External failure cost
Cost of defects discovered by the customer, including warranty, recalls, and the unmeasured cost of customers who never return.
Performance quality
How good the design itself is, which generally costs more to raise, as distinct from conformance quality.
ISO 9001
An international standard certifying that an organization documents its processes and follows that documentation, not that its products are good.

Module 3: Inventory

Why inventory exists, how much to order at once, how much to hold against uncertainty, and how to decide a single perishable order when there is no second chance.

Why Inventory Exists: EOQ, Discounts, and ABC

  • Classify inventory by type and by the function it performs, and identify the components of holding cost.
  • Derive and compute the economic order quantity, and demonstrate the flatness of its cost curve.
  • Evaluate a quantity discount schedule where holding cost varies with unit price.
  • Apply ABC classification to set different control policies for different items.

The big picture

Walk into the back room of any store and you are looking at a decision. Every case on those shelves represents money that was spent early, space that is being rented, and risk that the item will be obsolete, damaged, or stolen before it sells. Every empty spot on those shelves represents a different risk: a customer who wanted something and left without it.

Inventory is the buffer between two processes that do not run at the same rate. That is the whole concept. Because it is a buffer, the lean tradition treats it with suspicion, and Module 5 explains why with some force. But the suspicion is easy to misapply. Inventory is not waste; it is a purchase of insurance whose premium you should compute rather than assume. Sometimes it is a bargain, sometimes a rip-off, and the arithmetic in this module tells you which.

The scale is not small. United States business inventories run in the trillions of dollars, and the Census Bureau's inventories-to-sales ratio for total business has hovered around 1.3 to 1.4 in recent years, meaning the economy holds roughly forty days of goods at any moment. Getting the number right by a few percent moves real money.

Types and functions

Inventory is classified two ways at once, and both matter. By form: raw materials, work in process, finished goods, and MRO (maintenance, repair, and operating supplies, the gloves and filters and bearings that never appear in the product). By function, which is the more useful lens:

  • Cycle stock exists because you order or produce in batches rather than one unit at a time. It is the direct consequence of setup and ordering costs, and it is what the economic order quantity sizes.
  • Safety stock exists because demand and lead time are uncertain. The next lesson sizes it.
  • Pipeline or in-transit stock exists because movement takes time. Little's Law gives it directly: a distributor shipping 500 units a day with 21 days of ocean transit has 500 x 21 = 10,500 units on the water permanently, financed and unavailable.
  • Anticipation stock is built deliberately ahead of a known event: a seasonal peak, a plant shutdown, an announced price increase, a threatened strike.
  • Decoupling stock sits between process stages so that a stoppage at one does not immediately stop the other.

Naming the function tells you how to reduce it. Cycle stock falls when setup or ordering cost falls. Safety stock falls when variability or lead time falls. Pipeline stock falls only when transit time falls. Telling a plant to "cut inventory 20 percent" without saying which function you are attacking is how firms end up cutting safety stock and discovering the reason it was there.

Key idea: Classify inventory by the function it performs. Each function has a different cause, and therefore a different lever for reducing it.

The two costs that fight

Holding cost is what it costs to keep one unit for one year. It has more components than people expect: the cost of capital tied up (usually the largest piece), warehouse space and handling, insurance, property taxes, obsolescence, spoilage, damage, and shrinkage. Added up, holding cost typically runs 20 to 30 percent of the item's value per year, and higher for anything fashionable, perishable, or technological. A firm that assumes 10 percent because that is its interest rate is understating the cost by half or more, and will systematically order too much.

Ordering cost (or setup cost in production) is what it costs to place one order, regardless of size: the buyer's time, transmitting the order, receiving, inspecting, and processing the invoice. In a factory the parallel is the setup or changeover cost from Module 1, which was worth 60 minutes of a machine's capacity in that lesson's example.

These two pull in opposite directions. Order large amounts and you place few orders (low ordering cost) but hold a lot (high holding cost). Order small amounts and you flip it. Somewhere in between is a minimum, and finding it is a nineteenth-century-style calculus problem that Ford W. Harris solved in 1913 while working at Westinghouse.

Deriving the economic order quantity

Let D be annual demand, S the cost per order, H the annual holding cost per unit, and Q the order quantity. Assume for now that demand is constant and known, lead time is constant, the entire order arrives at once, and there are no discounts.

With constant demand, inventory sawtooths from Q down to 0, so average inventory is Q/2. And you place D/Q orders per year. So:

  • Annual holding cost = (Q/2) x H, a straight line rising with Q.
  • Annual ordering cost = (D/Q) x S, a curve falling with Q.
  • Total annual cost = (Q/2)H + (D/Q)S.

Set the derivative with respect to Q to zero: H/2 - DS/Q squared = 0, so Q squared = 2DS/H, giving

EOQ = square root of (2 x D x S / H)

There is a shortcut worth knowing. At the optimum, the rising line and the falling curve cross: holding cost exactly equals ordering cost. You can find the EOQ by setting (Q/2)H = (D/Q)S and solving, and you can check any EOQ calculation in two seconds by confirming the two costs match.

Work it. A hardware distributor sells 9,600 units a year of a fitting. Each order costs 60 dollars to place and receive. Holding costs 3 dollars per unit per year.

  • EOQ = square root of (2 x 9,600 x 60 / 3) = square root of (1,152,000 / 3) = square root of 384,000 = 620 units.
  • Orders per year = 9,600 / 620 = 15.5, so roughly one order every 365 / 15.5 = 23.5 days.
  • Annual ordering cost = 15.5 x 60 = 929 dollars.
  • Annual holding cost = (620 / 2) x 3 = 310 x 3 = 930 dollars.
  • Total = 1,859 dollars, and the two components match, as they must.

The most useful thing about EOQ is that it is forgiving

Every input to that formula is an estimate. Annual demand is a forecast. Holding cost involves a judgment call about the cost of capital and obsolescence. Ordering cost is an allocation. Does the imprecision matter?

Order quantityOrdering costHolding costTotalExcess over optimum
400 (35 percent below EOQ)1,4406002,0409.7 percent
5001,1527501,9022.3 percent
620 (the EOQ)9299301,8590
7507681,1251,8931.8 percent
900 (45 percent above EOQ)6401,3501,9907.0 percent

Being 35 percent below the optimum costs under 10 percent. Being 45 percent above costs 7 percent. The total cost curve is remarkably flat near its minimum, which has an enormously practical consequence: you can round the EOQ to something convenient at almost no cost. Round 620 up to a full pallet of 640, or down to a case-pack multiple of 600, or to whatever fills a truck, and the penalty is a rounding error. This flatness is why EOQ survives despite assumptions that are never quite true. It gives you the right order of magnitude, and the right order of magnitude is what you needed.

Key idea: The EOQ cost curve is flat near its optimum, so the formula is robust to imprecise inputs and you should round to practical quantities without guilt.

Attacking the formula instead of solving it

Notice that S, the ordering or setup cost, sits inside the square root. That makes it a lever. Suppose the distributor moves to electronic ordering and cuts the cost per order from 60 dollars to 6.

EOQ = square root of (2 x 9,600 x 6 / 3) = square root of 38,400 = 196 units.

Average inventory drops from 310 units to 98, and total annual cost falls from 1,859 dollars to about 588. Note that a tenfold reduction in ordering cost produced only a threefold reduction in order quantity, because of the square root, but the cost saving is large.

This is the mathematical bridge to Module 5. Toyota did not accept setup cost as given and then compute an optimal batch size. It attacked the setup, driving die changes on large presses from hours down to minutes through what Shigeo Shingo systematized as single-minute exchange of dies. The point of that work was never the saved hour. It was that when S collapses, small batches become economical, and small batches are what make a process responsive, low in inventory, and quick to surface defects.

Quantity discounts, worked properly

Suppliers offer lower unit prices for larger orders, and this breaks the basic EOQ in two ways. First, purchase cost now varies with the decision, so it must enter the comparison. Second, if holding cost is a percentage of unit value, then H itself changes with the price tier.

Same item: D = 9,600 per year, S = 60 dollars per order, and holding cost is 20 percent of unit value per year. The supplier quotes:

Order quantityUnit priceAnnual holding cost per unit (20 percent)
1 to 99920.004.00
1,000 to 2,49919.403.88
2,500 or more19.003.80

The procedure: compute the EOQ at each price, discard any that is not feasible in its own tier, and for those tiers use the smallest quantity that qualifies for the discount. Then total up purchase plus ordering plus holding for every candidate and pick the winner.

Tier 1, price 20.00, H = 4.00. EOQ = square root of (1,152,000 / 4) = square root of 288,000 = 537 units, which is feasible in the 1 to 999 range.

  • Purchase: 9,600 x 20.00 = 192,000
  • Ordering: (9,600 / 537) x 60 = 17.88 x 60 = 1,073
  • Holding: (537 / 2) x 4.00 = 1,074
  • Total: 194,147 dollars

Tier 2, price 19.40, H = 3.88. EOQ = square root of (1,152,000 / 3.88) = 545 units, which is not feasible because this tier begins at 1,000. Use Q = 1,000.

  • Purchase: 9,600 x 19.40 = 186,240
  • Ordering: (9,600 / 1,000) x 60 = 576
  • Holding: (1,000 / 2) x 3.88 = 1,940
  • Total: 188,756 dollars

Tier 3, price 19.00, H = 3.80. EOQ = square root of (1,152,000 / 3.80) = 551 units, also infeasible. Use Q = 2,500.

  • Purchase: 9,600 x 19.00 = 182,400
  • Ordering: (9,600 / 2,500) x 60 = 230
  • Holding: (2,500 / 2) x 3.80 = 4,750
  • Total: 187,380 dollars

Tier 3 wins by 1,376 dollars a year over tier 2 and 6,767 over tier 1. But before signing, look at what the model does not price. An order of 2,500 units is 95 days of supply. That means 47,500 dollars of cash committed at each order, a much larger exposure if the item is superseded by a new version, and a much bigger loss if the forecast that produced the 9,600 turns out to be optimistic. The model's holding cost of 20 percent captures obsolescence only as an average. If this is a fashion item, an electronic component, or anything with a version number, raise the holding percentage until the arithmetic reflects the actual risk, and watch the answer change.

Key idea: With discounts, compare total purchase plus ordering plus holding cost at each candidate quantity, remembering that holding cost itself falls with the price. Then check whether the winning quantity commits more cash and obsolescence risk than the model priced.

ABC classification: spending attention where it pays

A distribution center may stock 40,000 items. Nobody can manage 40,000 items with equal care, and nobody should try. ABC classification applies the Pareto pattern: rank items by annual dollar usage (annual demand times unit cost), then split them into classes.

ItemAnnual demandUnit costAnnual dollar usageShareClass
A-1011,20085.00102,00039.8 percentA
B-205400220.0088,00034.3 percentA
C-33030,0001.2036,00014.0 percentB
D-4109,0002.0018,0007.0 percentB
E-5122,5003.007,5002.9 percentC
F-62080,0000.064,8001.9 percentC
Total256,300100 percent

Two items out of six, 33 percent of the catalog, account for 74 percent of the dollars. In a real catalog of thousands the pattern is sharper: roughly 20 percent of items typically carry 70 to 80 percent of the value.

Notice what the ranking does to intuition. Item F-620 moves 80,000 units a year, by far the most physical activity, and it is a C item worth less than 5,000 dollars annually. Item B-205 moves only 400 units and is an A item. Volume is not value.

The policies then differ deliberately:

  • A items: tight control. Frequent review, accurate forecasts, close supplier relationships, cycle counting monthly or more often, low safety stock relative to demand because you watch them closely, and negotiation effort concentrated here because a one percent price reduction is worth real money.
  • B items: normal control. Periodic review, standard reorder points, quarterly counting.
  • C items: loose control and generous stock. Large orders, annual counting, and simple visual systems. The classic C item policy is the two-bin system: keep the item in two containers, and when the first empties, reorder and start on the second. It requires no computer, no forecast, and no analyst, and for a five-cent washer it is exactly right. Running out of a C item can shut a line just as effectively as running out of an A item, so C items get plenty of stock, they just do not get attention.

Related to all of this is cycle counting, the practice of counting a few items every day on a schedule driven by ABC class rather than shutting the warehouse once a year for a full physical count. Cycle counting finds and fixes record errors continuously, which matters enormously because Module 4's material requirements planning is worthless if inventory records are wrong. The common target is 95 to 99 percent record accuracy, and a plant below that is planning on fiction. Shrinkage is a real driver of the gap: the National Retail Federation's surveys have put retail shrink in the neighborhood of 1.5 percent of sales, which for a large chain is a very large number.

Common misconceptions

  • "Inventory is always waste." Cycle stock is the price of batching, safety stock is priced insurance, and pipeline stock is unavoidable while goods are moving. The lean point is that inventory hides problems, not that its correct level is zero.
  • "The EOQ formula is too crude to use." Its cost curve is flat near the optimum, so approximate inputs still give a good answer. Its real limitation is the constant-demand assumption, not input precision.
  • "Take the discount, it is free money." The discount is money; the extra holding cost, cash commitment, and obsolescence exposure are also money. Compute all of them.
  • "A items are the fast movers." A items are the high annual dollar usage items. A slow, expensive part outranks a fast, cheap one.

Recap

  • Inventory is classified by form and, more usefully, by function: cycle, safety, pipeline, anticipation, and decoupling stock, each with a different lever.
  • Holding cost typically runs 20 to 30 percent of item value per year and is usually understated; ordering and setup cost is the countervailing force.
  • EOQ is the square root of 2DS/H; in the worked case 620 units, with ordering and holding cost each about 930 dollars.
  • The cost curve is flat near the optimum, so rounding to practical quantities is nearly free.
  • Cutting setup cost from 60 to 6 dollars cut the EOQ to 196 units, which is the arithmetic behind lean's attack on changeovers.
  • With quantity discounts, compare purchase plus ordering plus holding across tiers; here the 2,500-unit tier won by 1,376 dollars but committed 95 days of supply.
  • ABC classification concentrates control on the 20 percent of items carrying most of the value, and cycle counting keeps records accurate enough for planning to work.

Sources

  1. Wikipedia. (2025). Economic order quantity. en.wikipedia.org
  2. Wikipedia. (2025). ABC analysis. en.wikipedia.org
  3. U.S. Census Bureau. (2025). Manufacturing and trade inventories and sales. Economic Indicators. census.gov
  4. Association for Supply Chain Management. (2025). Inventory management fundamentals. ASCM. ascm.org
  5. National Retail Federation. (2025). National retail security survey. NRF. nrf.com
Key terms
Cycle stock
Inventory that exists because items are ordered or produced in batches rather than one at a time.
Pipeline inventory
Goods in transit, equal by Little's Law to the shipment rate times the transit time.
Holding cost
The annual cost of keeping one unit in stock, including capital, space, insurance, obsolescence, and shrinkage; typically 20 to 30 percent of item value.
Ordering cost
The fixed cost of placing and receiving one order, independent of order size; the setup cost is its production equivalent.
Economic order quantity
The order size that minimizes the sum of annual ordering and holding cost, equal to the square root of 2DS/H.
Quantity discount analysis
Comparing total purchase, ordering, and holding cost at each feasible quantity when unit price and holding cost vary with order size.
ABC classification
Ranking items by annual dollar usage so that control effort is concentrated on the small share of items carrying most of the value.
Two-bin system
A visual reorder method for low-value items: when the first container empties, reorder and begin using the second.

Inventory Under Uncertainty: Safety Stock and the Newsvendor

  • Compute safety stock and reorder points from demand variability, lead time, and a target service level.
  • Show why lead-time variability usually dominates demand variability in setting safety stock.
  • Distinguish cycle service level from fill rate and compare continuous with periodic review.
  • Solve the single-period newsvendor problem using the critical ratio for both high-margin and low-margin items.

The big picture

The last lesson assumed demand was constant and known. It never is. A distributor whose average demand is 50 units a day will have days at 32 and days at 79, and a supplier who promises nine days will occasionally take fourteen. Everything expensive about inventory management lives in that gap between the average and what actually happened.

There are two distinct questions here, and confusing them causes real damage. The first is the repeating question: I sell this item continuously, so how much extra should I carry to absorb the surprises? That is safety stock. The second is the one-shot question: I get exactly one order, the item is worthless afterward, so how many should I buy? That is the newsvendor problem. Same subject, completely different mathematics, and this lesson does both.

Safety stock: the arithmetic

What matters is not demand variability by itself but variability over the lead time, because the lead time is the window in which you are exposed. Once you place an order you cannot place another that will arrive sooner, so the question is: how much could demand exceed my expectation during those days?

Start with the simple case. Daily demand averages 50 units with a standard deviation of 15. Lead time is a constant 9 days.

  • Expected demand during lead time = 50 x 9 = 450 units.
  • Standard deviation over the lead time = daily standard deviation x square root of lead time = 15 x square root of 9 = 15 x 3 = 45 units.

Note the square root. Variances add across independent days, and standard deviations are square roots of variances, so nine days of variability is three times one day's, not nine times. This square root shows up everywhere in this course, and it is always the same fact.

Now choose a cycle service level: the probability of not running out during a replenishment cycle. Each service level corresponds to a z-value from the standard normal table.

Safety stock = z x standard deviation of lead-time demand
Reorder point = expected lead-time demand + safety stock

Service levelzSafety stockReorder pointExtra units over previous row
90 percent1.281.28 x 45 = 58508-
95 percent1.651.65 x 45 = 7452416
98 percent2.052.05 x 45 = 9254218
99 percent2.332.33 x 45 = 10555513
99.9 percent3.093.09 x 45 = 13958934
100 percentinfiniteinfiniteinfiniteinfinite

Read the last two rows carefully. Going from 99 percent to 99.9 percent costs 34 more units, more than double what it cost to go from 95 to 98. And the final row is not a joke: with a normal distribution, a 100 percent service level requires infinite stock, because there is no demand level that cannot in principle be exceeded. Any executive who asks for "no stockouts, ever" is asking for a number that does not exist. The honest conversation is about which service level, at what cost, for which items, and that is exactly what ABC classification is for.

Key idea: Safety stock equals z times the standard deviation of demand over the lead time. Service level rises with sharply diminishing returns, and 100 percent is mathematically unreachable.

The part that actually matters: variable lead time

Lead time is rarely constant, and when it varies, the formula gets an extra term:

Standard deviation of lead-time demand = square root of [ (mean lead time x demand variance) + (mean demand squared x lead time variance) ]

Keep everything the same and let the supplier's lead time have a mean of 9 days with a standard deviation of 2 days:

  • First term: 9 x (15 squared) = 9 x 225 = 2,025.
  • Second term: (50 squared) x (2 squared) = 2,500 x 4 = 10,000.
  • Total variance = 12,025, so the standard deviation = square root of 12,025 = 110 units.
  • Safety stock at 95 percent = 1.65 x 110 = 181 units, against 74 before.
  • Reorder point = 450 + 181 = 631 units.

Safety stock more than doubled, and look at where it came from. Demand variability contributed 2,025 of the 12,025 total variance, about 17 percent. Lead time variability contributed 10,000, about 83 percent. In this example, and in a great many real ones, the supplier's inconsistency is the dominant driver of your inventory investment, not your customers' inconsistency.

That single finding reorients purchasing. Suppose you can either negotiate a shorter average lead time or a more consistent one. Push the supplier to cut lead time variability from 2 days to half a day:

  • Variance = 2,025 + (2,500 x 0.25) = 2,025 + 625 = 2,650.
  • Standard deviation = 51.5, so safety stock at 95 percent = 1.65 x 51.5 = 85 units.

Consistency alone cut safety stock from 181 units to 85, a 53 percent reduction, with no change in average lead time and no change in price. This is why a good buyer measures supplier on-time delivery variance and not just average lead time, and why "reliable" beats "fast" in Module 1's discussion of competitive priorities. Reliability is worth money you can compute.

Key idea: When lead time varies, its variance is usually the larger term. Reducing supplier lead-time variability is often the cheapest inventory reduction available.

Cycle service level is not fill rate

These two get used interchangeably and they are very different numbers. Cycle service level is the probability of no stockout during a cycle. Fill rate is the fraction of demand satisfied directly from stock. A 95 percent cycle service level says you run short in one cycle out of twenty; it says nothing about how short.

Work it for our item, ordering the EOQ of 620 units from the previous lesson, with lead-time demand standard deviation 45 and z = 1.65. The expected number of units short per cycle is the standard deviation times the unit normal loss function at z, which at z = 1.65 is 0.0206:

  • Expected shortage per cycle = 0.0206 x 45 = 0.93 units.
  • Fill rate = 1 - (0.93 / 620) = 1 - 0.0015 = 99.85 percent.

A 95 percent cycle service level delivered a 99.85 percent fill rate. Customers experience the fill rate; the planner's spreadsheet reports the cycle service level. Reporting the smaller number as though it were the customer experience makes managers buy inventory nobody needed. The gap is widest when the order quantity is large relative to the variability, which is most of the time.

Continuous versus periodic review

Everything so far assumed continuous review: you watch the position constantly and order Q units the moment it hits the reorder point. That is what a barcode system does.

Periodic review checks stock every T days and orders up to a target level. It is what you do when a delivery truck comes every Tuesday, when a supplier requires consolidated orders, or when counting is manual. Its cost is exposure: you must survive not just the lead time but the lead time plus the review period, because a shortage developing the day after a review goes unnoticed until the next one.

With a 7-day review period and a 9-day lead time, the protection interval is 16 days:

  • Standard deviation over 16 days = 15 x square root of 16 = 15 x 4 = 60 units.
  • Safety stock at 95 percent = 1.65 x 60 = 99 units, against 74 under continuous review.
  • Order-up-to level = (50 x 16) + 99 = 800 + 99 = 899 units.

Periodic review costs 25 more units of safety stock here, and buys simplicity and consolidated shipping. That is a real trade, not an error.

Pooling, again

Module 2 showed that combining queues reduces waiting. The same square root works on inventory. Suppose the distributor operates 9 regional warehouses, each facing independent demand with a lead-time demand standard deviation of 45 units, each holding 74 units of safety stock: 666 units in total.

Consolidate into one national warehouse. Independent variances add, so the combined standard deviation is 45 x square root of 9 = 135, and:

  • Pooled safety stock = 1.65 x 135 = 223 units.
  • Reduction: 666 to 223, a saving of 443 units, or two thirds.

This is the square root law of inventory pooling: consolidating n independent locations cuts safety stock by a factor of the square root of n. It is why e-commerce companies centralized aggressively in the 2000s. The countervailing forces are equally real: outbound freight rises, delivery time rises, and the single site becomes a single point of failure, which is why the same companies then rebuilt regional networks for speed. Module 6 prices that trade properly. Virtual pooling, where stock stays distributed but any location can be drawn on through a shared system, captures part of the benefit without the transport penalty.

Key idea: Consolidating n independent stocking locations reduces required safety stock by a factor of the square root of n, at the cost of freight, speed, and concentration risk.

The newsvendor problem: one order, no second chance

Now the other question entirely. A bakery makes croissants each morning. Unsold croissants cannot be carried to tomorrow at full value. There is no reorder point, no cycle, no safety stock. There is exactly one decision, made before demand is known.

The trade-off is between two mistakes. Bake too few and you lose profit on sales you could have made. Bake too many and you lose the cost of what you throw away.

  • Underage cost (Cu) = profit lost per unit of unmet demand = price - cost.
  • Overage cost (Co) = loss per unsold unit = cost - salvage value.

The optimal policy stocks up to the point where the probability of selling one more unit just equals the ratio of these costs. That ratio is the critical ratio:

Critical ratio = Cu / (Cu + Co), and you stock the quantity at which the cumulative probability of demand equals that ratio.

Worked case one: the croissants. They cost 1.10 dollars to produce and sell for 3.50. Day-old croissants go on a discount rack for 0.25. Daily demand is roughly normal with a mean of 180 and a standard deviation of 40.

  • Cu = 3.50 - 1.10 = 2.40 dollars.
  • Co = 1.10 - 0.25 = 0.85 dollars.
  • Critical ratio = 2.40 / (2.40 + 0.85) = 2.40 / 3.25 = 0.738.
  • The z-value with 73.8 percent of the normal distribution below it is about 0.64.
  • Optimal quantity = 180 + (0.64 x 40) = 180 + 25.6 = 206 croissants.

Bake 206, well above the average demand of 180, and accept that on a typical day some go to the discount rack. The arithmetic says so because running out costs nearly three times what an extra croissant costs. On about 74 percent of days you will have enough.

Worked case two: the florist. Same demand distribution, different economics. A fresh arrangement costs 38 dollars in flowers and labor and sells for 60. Unsold arrangements are worthless the next day.

  • Cu = 60 - 38 = 22 dollars.
  • Co = 38 - 0 = 38 dollars.
  • Critical ratio = 22 / (22 + 38) = 22 / 60 = 0.367.
  • The z-value at 36.7 percent is about -0.34.
  • Optimal quantity = 180 + (-0.34 x 40) = 180 - 13.6 = 166 arrangements.

Here you deliberately stock below average demand and plan to disappoint some customers, because each unsold arrangement destroys 38 dollars while each lost sale costs only 22. Same demand, opposite answer. This is the entire lesson of the newsvendor model: the optimal quantity depends on the ratio of the two costs, and the average demand is almost never the right order. Order the mean only in the special case where Cu happens to equal Co.

The model generalizes further than bakeries. It applies to a magazine print run, a concert's merchandise order, a hotel's overbooking level, a ski resort's seasonal staffing, a hospital's platelet order (platelets expire in about five days), and a flu vaccine order. In the vaccine case the underage cost is not a lost sale but illness and hospitalization, so the critical ratio is very close to one and the correct order is far above expected demand, which is exactly what public health agencies do. When the underage cost is a human cost rather than a margin, the model still works; you just have to be explicit about the number you are putting in the numerator, and that explicitness is a feature.

Key idea: For a single perishable order, stock to the critical ratio Cu / (Cu + Co). High-margin perishables should be overstocked, low-margin ones understocked, and the mean is rarely optimal.

What people actually do, and why it is wrong

Laboratory experiments on the newsvendor problem, beginning with work by Schweitzer and Cachon in 2000, found a robust and stubborn bias: people pull toward the mean. Given a high-margin product where the model says order well above average demand, subjects order too little. Given a low-margin product where the model says order below average, they order too much. They do it repeatedly, with feedback, with training, and with financial incentives.

The explanation most consistent with the data is a discomfort with visible leftovers combined with an anchor on the forecast. Leftovers are conspicuous and feel like a mistake; lost sales are invisible. That asymmetry in what you can see, rather than any failure of arithmetic, is what drives the bias. Knowing about it is genuinely useful: when your intuition and the critical ratio disagree, it is very likely your intuition is doing the pulling.

Common misconceptions

  • "Safety stock should cover the worst case." There is no worst case in a continuous distribution. Choose a service level, price it, and be explicit about what you are buying.
  • "Nine days of lead time means nine times the variability." Variances add, so standard deviation grows with the square root: three times, not nine.
  • "A 95 percent service level means 5 percent of customers go unserved." That confuses cycle service level with fill rate; the worked case had a 99.85 percent fill rate at a 95 percent cycle service level.
  • "For a perishable item, order the average." Only if the underage and overage costs happen to be equal. Otherwise the critical ratio moves you above or below the mean, sometimes a long way.

Recap

  • Safety stock equals z times the standard deviation of demand over the lead time; standard deviation grows with the square root of lead time.
  • Service level has sharply diminishing returns, and a 100 percent service level would require infinite stock.
  • With variable lead time, the lead-time variance term dominated at 83 percent of total variance; cutting supplier variability from 2 days to half a day cut safety stock from 181 units to 85.
  • Cycle service level and fill rate are different numbers; 95 percent cycle service produced a 99.85 percent fill rate here.
  • Periodic review must cover lead time plus the review period, costing 99 units of safety stock against 74 for continuous review.
  • Pooling 9 locations cut total safety stock from 666 units to 223, by a factor of the square root of 9.
  • The newsvendor critical ratio Cu / (Cu + Co) gave 206 croissants against a mean of 180, and 166 flower arrangements against the same mean.
  • Experimental subjects consistently pull their orders toward the mean in both directions, because leftovers are visible and lost sales are not.

Sources

  1. Wikipedia. (2025). Safety stock. en.wikipedia.org
  2. Wikipedia. (2025). Newsvendor model. en.wikipedia.org
  3. Association for Supply Chain Management. (2025). Inventory planning and service levels. ASCM. ascm.org
  4. MIT OpenCourseWare. (2013). Introduction to operations management: inventory management. Massachusetts Institute of Technology. ocw.mit.edu
Key terms
Safety stock
Extra inventory held to absorb variability in demand and lead time, equal to z times the standard deviation of lead-time demand.
Reorder point
The inventory position at which a replenishment order is placed: expected lead-time demand plus safety stock.
Cycle service level
The probability of not running out during a replenishment cycle; the source of the z-value in the safety stock formula.
Fill rate
The fraction of total demand met directly from stock, usually much higher than the cycle service level for the same safety stock.
Periodic review
Checking stock every T days and ordering up to a target, which must protect against lead time plus the review period.
Square root law of pooling
Consolidating n independent stocking locations reduces required safety stock by a factor of the square root of n.
Newsvendor problem
A single-period decision for a perishable item, solved by stocking to the critical ratio of underage to total mismatch cost.
Critical ratio
Cu divided by (Cu + Co); the cumulative demand probability at which the optimal single-period order quantity sits.

Module 4: Forecasting, Planning, and the Bullwhip

Predicting demand and measuring how wrong you were, converting a forecast into a workforce and a parts schedule, and understanding why order swings grow as they travel upstream.

Forecasting: Methods and Honest Limits

  • Compute moving average, weighted moving average, and exponential smoothing forecasts from a demand series.
  • Build seasonal indices and use them to disaggregate an annual forecast.
  • Measure forecast error with MAD, MAPE, and the tracking signal, and diagnose bias.
  • State what forecasting can and cannot deliver, and explain why shortening lead time often beats improving accuracy.

The big picture

Every plan downstream of this lesson depends on a number that nobody knows. How many units to build, how many nurses to schedule, how much steel to buy, whether to open the second line: all of it rests on a forecast of demand that has not happened yet.

Two temperaments fail here. The first treats the forecast as a fact, plans as if it were certain, and is astonished every quarter. The second concludes that forecasting is futile and plans on instinct, which is a forecast too, just an undocumented one. The professional position is in between: produce the best estimate you can, attach an honest measure of how wrong it is likely to be, and design the operation so that being wrong is survivable. That last clause is the part most textbooks underplay, and this lesson ends on it.

What is in a demand series

Any time series of demand decomposes into four things:

  • Level: the underlying average, the baseline.
  • Trend: a persistent drift up or down.
  • Seasonality: a repeating pattern tied to the calendar, whether hourly (a lunch rush), weekly (Saturday retail), or annual (garden centers in spring).
  • Random variation: noise.

Read that list again, because the fourth item defines the ceiling on everything you can do. Noise is by definition not forecastable. If half the movement in a series is random, then a perfect model of the other half still leaves you badly wrong on any given period. Much of the frustration in forecasting comes from teams treating irreducible noise as a modeling failure, then chasing it with more complexity and making things worse by fitting the noise.

Some texts add a fifth element, longer business cycles, which are real but rarely forecastable at the horizon operations planners work on.

Where there is no history at all, you are in qualitative territory: sales force composites, executive judgment, the Delphi method (repeated anonymous rounds with feedback, designed to stop the loudest voice from dominating), and market research. These have their own biases. Sales forecasts sag when quotas are set from them and swell when budget is allocated from them, which is a Goodhart problem, not a statistical one.

The benchmark you must beat

Before any method, compute the naive forecast: next period equals this period. It is free, it requires no software, and in many series it is surprisingly hard to beat. Any proposed method that does not beat the naive forecast on out-of-sample data should be rejected, and a shocking number of expensive systems fail this test. Always compute the naive benchmark first.

Moving averages

Take six months of demand: 120, 135, 128, 142, 150, 138.

A three-month moving average forecast for month 7 is (142 + 150 + 138) / 3 = 143.3. Simple and stable, but it weights a three-month-old observation exactly as heavily as last month's, and it responds sluggishly to change.

A weighted moving average fixes the weighting. With weights of 0.5, 0.3, and 0.2 on the most recent three months: (0.5 x 138) + (0.3 x 150) + (0.2 x 142) = 69 + 45 + 28.4 = 142.4. Better, but now you have to pick weights, and you must store the last n observations.

Exponential smoothing, worked all the way through

Exponential smoothing solves both problems with one line of arithmetic. The new forecast is the old forecast adjusted by a fraction of the error you just made:

New forecast = alpha x (latest actual) + (1 - alpha) x (previous forecast)

Equivalently, new forecast = previous forecast + alpha x (latest actual - previous forecast). The smoothing constant alpha runs from 0 to 1: high alpha reacts quickly and follows noise, low alpha is stable and lags real changes. You need to store exactly one number, the last forecast, which is why this method ran the world's inventory systems for decades on very little memory.

Run the same six months from a starting forecast of 120, at two values of alpha.

MonthActualForecast, alpha = 0.2Forecast, alpha = 0.5
1120120.00120.00
2135120.00120.00
3128123.00127.50
4142124.00127.75
5150127.60134.88
6138132.08142.44
7 (forecast)-133.26140.22

Trace one row so the mechanism is concrete. For month 5 at alpha = 0.2: the previous forecast was 124.00 and the month 4 actual was 142, so the new forecast is 0.2 x 142 + 0.8 x 124.00 = 28.4 + 99.2 = 127.60.

Now look at the pattern. This series is trending upward, and both columns sit below the actuals almost the whole way. That is not a coincidence or a bad alpha. Simple exponential smoothing is a weighted average of past values, and a weighted average of past values in a rising series will always lag behind it. The method is structurally biased on trended data.

The fix has a name. Holt's method, or double exponential smoothing, adds a second smoothed component for the trend itself and projects it forward. Holt-Winters, or triple exponential smoothing, adds a third component for seasonality. Both are the same idea applied repeatedly, and both are standard in any planning system.

Key idea: Exponential smoothing updates the forecast by a fraction of the last error, needing only one stored number. On trended data it lags systematically, which is what Holt's trend term exists to correct.

Seasonality, computed

A garden center's quarterly sales, in thousands of dollars:

QuarterYear 1Year 1 ratio to averageYear 2Year 2 ratio to averageSeasonal index
Q11800.7202200.7330.727
Q24201.6804801.6001.640
Q33001.2003401.1331.167
Q41000.4001600.5330.467
Total1,0001,2004.001

The procedure in three steps. First, compute each year's quarterly average: 1,000 / 4 = 250 for year 1, and 1,200 / 4 = 300 for year 2. Second, divide each quarter by its own year's average, which strips out the growth between years: 180 / 250 = 0.720, 220 / 300 = 0.733, and so on. Third, average the ratios for each quarter across years. The four indices should sum to the number of periods, and 4.001 confirms the arithmetic.

Now use them. Suppose a trend line or a management judgment puts year 3 at 1,400 thousand dollars. The quarterly average would be 1,400 / 4 = 350.

  • Q1: 350 x 0.727 = 254
  • Q2: 350 x 1.640 = 574
  • Q3: 350 x 1.167 = 408
  • Q4: 350 x 0.467 = 163
  • Sum: 1,399, which recovers the annual total.

This is the workhorse calculation behind aggregate planning in the next lesson, because a plant that plans to the annual average and ignores a Q2 index of 1.64 will be 64 percent short in spring and idle in December.

Measuring how wrong you were

A forecast without an error measure is an opinion. Take the alpha = 0.2 forecasts for months 3 through 6, where error is actual minus forecast.

MonthActualForecastErrorAbsolute errorAbsolute percent error
3128123.00+5.005.003.9 percent
4142124.00+18.0018.0012.7 percent
5150127.60+22.4022.4014.9 percent
6138132.08+5.925.924.3 percent
Total+51.3251.3235.8 percent
  • Mean absolute deviation (MAD) = 51.32 / 4 = 12.83 units. Easy to interpret, in the units of the thing being forecast.
  • Mean absolute percent error (MAPE) = 35.8 / 4 = 8.95 percent. Unitless, so it compares across products, but it misbehaves badly for items with near-zero demand because dividing by a tiny actual produces enormous percentages.
  • Running sum of forecast errors (RSFE) = +51.32. Signed, so it detects bias. In an unbiased forecast, positive and negative errors cancel and this hovers near zero.
  • Tracking signal = RSFE / MAD = 51.32 / 12.83 = +4.0.

The tracking signal is the diagnostic that matters, and the conventional action limit is plus or minus 4. At exactly +4.0 this forecast has tripped its limit: every single error was positive, meaning the method underforecast every month. That is not bad luck, it is the structural lag on trended data that we predicted above, and the correct response is to change the method rather than to change alpha.

For comparison, running the same calculation on the alpha = 0.5 column gives a MAD of 8.58, a MAPE of 5.9 percent, and a tracking signal of about +3.0. Higher responsiveness helped, as it should on a trending series, but the bias is still there and still points at the same fix.

Key idea: MAD and MAPE measure size of error; RSFE and the tracking signal measure bias. A tracking signal beyond plus or minus 4 means the method is systematically wrong, not merely imprecise.

The honest limits

This is the part to remember five years from now.

Simple methods are competitive. Spyros Makridakis has run a series of open forecasting competitions since 1982, comparing methods on thousands of real series. The findings have been remarkably stable across four decades: statistically sophisticated methods do not automatically beat simple ones; combining several forecasts usually beats any single method; and accuracy degrades as the horizon lengthens. In the M4 competition published in 2018, the winning entries were hybrids that combined statistical structure with machine learning, while several pure machine learning submissions performed worse than simple statistical benchmarks. Complexity is not free and is not automatically better.

Aggregate forecasts are far more accurate than detailed ones. This is the pooling square root once more. Forecasting total demand for a product family next quarter might carry a 5 to 10 percent error. Forecasting one specific item, in one specific color, at one specific store, for one specific week, routinely carries errors of 30 to 50 percent or worse, and for slow-moving items the notion of a percentage error nearly breaks down. Any planning process that pretends item-level weekly forecasts are reliable is building on sand. This is precisely why aggregate planning exists as a separate step, and why safety stock rather than forecast precision handles item-level uncertainty.

Judgmental overrides usually hurt. Studies of firms where planners adjust statistical forecasts find that small adjustments typically make accuracy worse, while large adjustments, made when the planner genuinely knows something the model cannot (a promotion, a lost customer, a competitor's plant fire), can help. The practical rule is to require a documented reason for any override and to track override accuracy separately, which by itself reduces the number of pointless adjustments.

Shortening the lead time beats improving the forecast. This is the single most valuable operations insight in the lesson. Forecast error grows with the horizon. If your supplier's lead time is 90 days, you must forecast 90 days out, where errors are large. Cut the lead time to 20 days and you are forecasting 20 days out, where errors are dramatically smaller, without changing your forecasting method at all. Zara built a business on this: by keeping much production close to its distribution center and running short cycles, it commits to far less inventory on a much shorter forecast horizon. When someone proposes buying a better forecasting system, always ask what it would cost instead to halve the lead time.

Every forecast needs a range. A single number cannot be used to set safety stock, because the previous lesson's formula needs a standard deviation. A forecast delivered without an error estimate has withheld the half of the information the operation actually needs.

Key idea: Forecast accuracy is bounded by noise, worse at detailed levels, and worse at longer horizons. Reducing lead time reduces the horizon you must forecast, which is often cheaper than improving the forecast itself.

Common misconceptions

  • "A better model would fix our forecast accuracy." Some of the error is irreducible noise, and forty years of competitions show sophistication buys much less than people expect.
  • "The forecast is the target." A forecast is an estimate of what will happen; a target is what you want to happen. Merging them corrupts the estimate, because nobody will forecast below quota.
  • "Exponential smoothing handles trends." Simple exponential smoothing systematically lags a trend. Use Holt's method when a trend is present, which the tracking signal will tell you.
  • "More granular forecasts are more useful." They are less accurate by a wide margin. Plan aggregate, buffer at the item level.

Recap

  • Demand series contain level, trend, seasonality, and noise; noise is not forecastable and sets the accuracy ceiling.
  • Always compute the naive forecast as a benchmark before adopting any method.
  • Exponential smoothing updates by alpha times the last error; at alpha 0.2 and 0.5 it lagged a rising series all the way, which Holt's trend term corrects.
  • Seasonal indices are built by dividing each period by its own year's average and averaging across years; here they were 0.727, 1.640, 1.167, and 0.467.
  • MAD was 12.83 units and MAPE 8.95 percent, while a tracking signal of +4.0 flagged systematic underforecasting rather than random error.
  • Simple methods compete well, combinations beat singles, aggregate beats detailed, and long horizons are always worse.
  • Cutting lead time shortens the forecast horizon and is often cheaper than improving forecast accuracy.

Sources

  1. Wikipedia. (2025). Exponential smoothing. en.wikipedia.org
  2. Wikipedia. (2025). Makridakis Competitions. en.wikipedia.org
  3. Association for Supply Chain Management. (2025). Demand planning and forecasting. ASCM. ascm.org
  4. U.S. Census Bureau. (2025). Advance monthly retail trade survey. Economic Indicators. census.gov
Key terms
Naive forecast
Using the most recent actual as the forecast for the next period; the free benchmark every method must beat.
Exponential smoothing
A forecast that equals alpha times the latest actual plus (1 - alpha) times the previous forecast, requiring only one stored value.
Smoothing constant
The weight alpha between 0 and 1 that sets responsiveness; high values track change quickly and also track noise.
Holt's method
Double exponential smoothing, which adds a separately smoothed trend component to correct the lag on trended series.
Seasonal index
The ratio of a period's demand to the average period, used to disaggregate an annual forecast into periods.
Mean absolute deviation
The average of absolute forecast errors, expressed in the units being forecast.
Tracking signal
Running sum of forecast errors divided by MAD; values beyond plus or minus 4 indicate systematic bias rather than random error.
Forecast horizon
How far ahead a forecast must reach, set largely by lead time; error grows with the horizon, so shortening lead time improves effective accuracy.

Aggregate Planning, Master Scheduling, and MRP

  • Compare level and chase aggregate plans by computing their full annual cost.
  • Describe the planning hierarchy from sales and operations planning down to shop floor scheduling.
  • Run a material requirements planning explosion through a two-level bill of materials with lead-time offsetting.
  • Explain the difference between independent and dependent demand and the known limits of MRP and ERP systems.

The big picture

You now have a forecast. It says demand will be 300 units in January, rising to 800 in March, and falling back to 500 by June. Nobody can act on that yet. A forecast is not a plan. Somebody has to decide how many people to employ, whether to work overtime, whether to build ahead into inventory, and then, at a much finer grain, exactly how many brackets to order in week two so that a bookshelf can ship in week six.

That translation runs through a hierarchy, and each level answers a different question at a different horizon.

LevelHorizonUnit of decisionQuestion
Strategic capacity planning1 to 5 yearsPlants, lines, distribution centersWhat capacity should exist at all?
Sales and operations planning / aggregate planning6 to 18 monthsProduct families, total units, headcountHow many people and how much inventory?
Master production scheduleWeeks to monthsSpecific end items by weekWhat exactly do we build, and when?
Material requirements planningWeeksComponents and raw materialsWhat do we order or make, and when do we start?
Shop floor schedulingDays to hoursIndividual jobs at machinesWhat runs next on this machine?

The reason for the hierarchy is the accuracy limit from the last lesson. Aggregate forecasts are decent; item-level weekly forecasts are poor. So the long-horizon decisions are made in aggregate units where forecasts hold up, and the detailed decisions are made close in, where they are needed and where information is better.

Sales and operations planning

Sales and operations planning is the monthly meeting where a demand plan and a supply plan get reconciled into one set of numbers that every function agrees to use. That sounds procedural and it is the most consequential meeting most manufacturers hold, because in its absence sales forecasts optimistically, finance budgets conservatively, and operations plans to something else entirely, and all three are surprised in the same month.

The discipline that makes it work is a single set of numbers: one demand plan, expressed in both units and dollars, owned jointly, with disagreements resolved in the room rather than in three separate spreadsheets. The standard failure mode is the meeting that reviews last month's variance without committing to next month's number.

Aggregate planning: level versus chase, costed out

Aggregate planning decides the total output rate, workforce size, and inventory level for a product family over the medium term. The levers are hiring and firing, overtime and undertime, subcontracting, part-time and temporary labor, building inventory ahead, accepting backorders, and shaping demand itself with pricing or promotion.

Two pure strategies anchor the choices. A chase strategy matches output to demand every period by varying the workforce. A level strategy holds output constant and absorbs the mismatch with inventory and backorders. Cost them out on a real six-month season.

Month123456Total
Demand3005008008006005003,500

Assumptions: each worker produces 100 units per month; the current workforce is 5; hiring costs 1,200 dollars per worker; layoff costs 2,000 dollars per worker; holding costs 8 dollars per unit of ending inventory per month; and labor costs 4,000 dollars per worker per month.

Chase plan. Staff exactly to demand: 3, 5, 8, 8, 6, and 5 workers.

  • Workforce changes from the starting 5: down 2, up 2, up 3, no change, down 2, down 1.
  • Hires: 2 + 3 = 5, at 1,200 dollars each = 6,000 dollars.
  • Layoffs: 2 + 2 + 1 = 5, at 2,000 dollars each = 10,000 dollars.
  • Inventory carried: zero, so no holding cost.
  • Labor: 3 + 5 + 8 + 8 + 6 + 5 = 35 worker-months at 4,000 = 140,000 dollars.
  • Total: 156,000 dollars.

Level plan. Hire one worker to reach 6, produce 600 units every month, and let inventory absorb the swings.

MonthProductionDemandChangeEnding inventory
1600300+300300
2600500+100400
3600800-200200
4600800-2000
560060000
6600500+100100
Total3,6003,5001,000 unit-months
  • Hiring: 1 worker = 1,200 dollars. No layoffs.
  • Holding: 1,000 unit-months at 8 dollars = 8,000 dollars.
  • Labor: 6 workers x 6 months = 36 worker-months at 4,000 = 144,000 dollars.
  • Total: 153,200 dollars.

The level plan wins by 2,800 dollars, a bit under two percent. Now change one number and watch the answer flip. If the product were bulkier or more perishable so that holding cost were 20 dollars per unit-month rather than 8, the level plan's holding cost becomes 20,000 dollars and its total becomes 165,200, while chase is unchanged at 156,000. Chase now wins by 9,200 dollars.

That flip is the lesson. The choice between level and chase is determined by the ratio of inventory cost to workforce change cost, and neither strategy is inherently superior. A hospital cannot inventory care and so must chase. A cement plant cannot cheaply hire and fire skilled operators and so tends to level. Most real plans are hybrids: hold a stable core workforce, cover the peak with overtime (typically 25 to 50 percent more per hour but with no hiring or layoff cost), and subcontract the remainder.

One more caution the arithmetic cannot show. The chase plan laid off five people and hired five people in six months. The model prices that at 16,000 dollars. It does not price the loss of experienced people who do not come back, the productivity of new hires who take three months to reach full speed, the safety record of a green workforce, or what the rest of the town concludes about working for you. Those costs are real and they are systematically absent from aggregate planning spreadsheets.

Key idea: Level and chase plans are chosen by the ratio of holding cost to workforce change cost. The worked case flipped from level to chase when holding cost rose from 8 to 20 dollars per unit-month.

From the aggregate plan to the master production schedule

The aggregate plan says 600 units of the family per month. The master production schedule disaggregates that into specific end items in specific weeks: 100 walnut bookshelves in week 6, 150 oak ones in week 8. It is a build plan, not a forecast, and it is what everything downstream explodes from.

Two features are worth naming. Available to promise is the uncommitted portion of a scheduled build, which lets a salesperson answer a delivery question without guessing. Time fences divide the schedule into a frozen zone where changes require executive approval because material is committed, a slushy zone where trade-offs are negotiated, and a liquid zone far enough out that changes are free. Without time fences, the schedule changes daily and nothing downstream can plan.

Material requirements planning, worked

Joseph Orlicky's insight in the 1960s was a distinction so clean it created an industry. Independent demand, meaning demand for finished goods, must be forecast, because it comes from customers you do not control. Dependent demand, meaning demand for the components inside those finished goods, should never be forecast, because it can be calculated exactly from the build schedule and the bill of materials. Forecasting bracket demand when you know you are building 100 bookshelves each needing 8 brackets is not just unnecessary, it is worse than the arithmetic.

Work a full explosion. The product is a bookshelf, item B, with this bill of materials:

  • Each B requires 2 side panels (S), 4 shelves (H), 1 back board (K), and 8 brackets (R).
  • Each S requires 1 panel blank (P). Each H also requires 1 panel blank (P). P is therefore a common component.
ItemLead timeOn hand
B (bookshelf, assembly)1 week0
S (side panel)2 weeks60
H (shelf)1 week100
K (back board)2 weeks40
R (bracket, purchased)3 weeks200
P (panel blank, purchased)2 weeks150

The master schedule calls for 100 bookshelves in week 6 and 150 in week 8. The logic at each level is always the same three steps: compute gross requirements from the parent's planned order releases, subtract on-hand to get net requirements, then offset by the lead time to get the planned order release date.

Level 0, item B. Gross 100 in week 6 and 150 in week 8. Nothing on hand, so net is the same. Lead time 1 week, so release 100 in week 5 and 150 in week 7. Those two releases are what every component below now responds to. Note that components are driven by the parent's release dates, not its due dates, which is where beginners go wrong.

ItemGross requirementOn handNet requirementLead timePlanned order release
S (2 per B)200 in wk 5; 300 in wk 760140 in wk 5; 300 in wk 72 wk140 in wk 3; 300 in wk 5
H (4 per B)400 in wk 5; 600 in wk 7100300 in wk 5; 600 in wk 71 wk300 in wk 4; 600 in wk 6
K (1 per B)100 in wk 5; 150 in wk 74060 in wk 5; 150 in wk 72 wk60 in wk 3; 150 in wk 5
R (8 per B)800 in wk 5; 1,200 in wk 7200600 in wk 5; 1,200 in wk 73 wk600 in wk 2; 1,200 in wk 4

Now the interesting one. Item P is required by both S and H, so its gross requirements are the sum of two parents' releases: 140 in week 3 and 300 in week 5 from S, plus 300 in week 4 and 600 in week 6 from H.

Week23456
Gross requirement for P0140300300600
On hand at start of week1501501000
Net requirement00290300600
Planned order release (2-week lead time)290300600--

Follow the on-hand line: 150 blanks cover the 140 needed in week 3, leaving 10, which covers part of week 4's 300 and leaves a net of 290. Offsetting each net requirement by the two-week lead time produces releases in weeks 2, 3, and 4.

Stand back and look at what this produced. To ship bookshelves in week 6, a purchase order for 600 brackets and 290 panel blanks has to be placed in week 2, four weeks earlier. No human intuition generates that date. This calculation, run over thousands of items and dozens of levels, is what a planning system exists to do, and it is why inventory record accuracy matters so intensely: an on-hand figure of 150 that is really 90 pushes a shortage four weeks downstream, where it will be discovered on the assembly floor.

Key idea: Dependent demand is calculated, not forecast. MRP explodes a master schedule through the bill of materials by netting against on-hand stock and offsetting by lead time, and a common component sums the requirements of all its parents.

Lot sizing, and the limits of MRP

The explosion above used lot-for-lot ordering: order exactly the net requirement. That minimizes inventory and is right when setup or ordering cost is low. When ordering cost is significant you use a fixed order quantity, often the EOQ, or a periodic order quantity that groups several weeks of requirements into one order. The trade-off is exactly Module 3's, applied week by week.

MRP has four documented weaknesses, and knowing them is what separates a user from a believer.

  • It assumes infinite capacity. Basic MRP will cheerfully schedule 900 units into a week when the plant can make 600. Capacity requirements planning was bolted on to check this, producing closed-loop MRP and eventually manufacturing resource planning, or MRP II.
  • It is only as good as its data. Wrong on-hand balances, wrong bills of materials, and wrong lead times all produce confidently wrong plans. Practitioners generally want inventory record accuracy above 95 percent and bill of materials accuracy above 98 percent before the output is trustworthy, which is why cycle counting from Module 3 is a prerequisite rather than a nicety.
  • It is nervous. A small change in the master schedule can cascade into large changes across hundreds of component orders, a phenomenon actually called system nervousness. Time fences and firm planned orders exist to damp it.
  • Its lead times are fixed inputs, and real lead times are mostly queue time. A 3-week purchased lead time and a 1-week assembly lead time are entered as constants, but real shop lead times expand when the plant is busy, exactly as Module 2's queueing arithmetic predicts. A system whose fixed lead times are optimistic will chronically start work too late, then expedite.

Enterprise resource planning systems extended MRP II across finance, human resources, sales, and procurement on one database, which is genuinely valuable and genuinely difficult. Implementation failures are well documented: Hershey's 1999 go-live coincided with an inability to fulfill orders during its critical Halloween season, and Nike attributed roughly 100 million dollars of lost sales in 2001 to problems with a demand planning implementation. In both cases the reported causes were the familiar ones: an aggressive schedule, insufficient testing, process changes bundled with a system change, and data that was not clean before conversion.

Common misconceptions

  • "Level production is more efficient." It is cheaper only when inventory is cheap relative to workforce changes. Change one cost and the answer reverses.
  • "MRP forecasts component demand." It calculates it. Forecasting dependent demand throws away exact information you already have.
  • "Components are scheduled from the parent's due date." They are driven by the parent's planned order release date, which is the due date offset by the parent's own lead time.
  • "A new ERP system will fix our operations." A system enforces a process; it does not design one. Implementations that fail generally fail on data quality, process definition, and schedule pressure rather than on software.

Recap

  • Planning runs as a hierarchy from strategic capacity through sales and operations planning, master scheduling, MRP, and shop floor scheduling, because forecast accuracy differs sharply by horizon and grain.
  • The worked chase plan cost 156,000 dollars and the level plan 153,200; raising holding cost to 20 dollars per unit-month reversed the ranking.
  • The master production schedule commits to specific end items by week, protected by time fences and reported through available-to-promise.
  • Independent demand is forecast; dependent demand is calculated by exploding the bill of materials, netting on-hand stock, and offsetting lead time.
  • In the worked explosion, shipping bookshelves in week 6 required releasing bracket and panel blank orders in week 2.
  • MRP assumes infinite capacity, depends on accurate records, is nervous under schedule change, and treats lead times as fixed when they are largely queue time.

Sources

  1. Wikipedia. (2025). Material requirements planning. en.wikipedia.org
  2. Association for Supply Chain Management. (2025). Sales and operations planning and CPIM body of knowledge. ASCM. ascm.org
  3. Wikipedia. (2025). Enterprise resource planning. en.wikipedia.org
  4. OpenStax. (2018). Production planning and scheduling. In Introduction to Business. Rice University. openstax.org
Key terms
Sales and operations planning
A monthly cross-functional process that reconciles demand and supply plans into a single agreed set of numbers.
Chase strategy
An aggregate plan that matches output to demand each period by varying the workforce, carrying no inventory.
Level strategy
An aggregate plan that holds output constant and absorbs demand swings with inventory and backorders.
Master production schedule
A committed build plan of specific end items by period, from which all component requirements are exploded.
Time fence
A boundary in the schedule separating frozen, negotiable, and freely changeable periods, which damps schedule churn.
Dependent demand
Demand for components, calculable exactly from the build schedule and bill of materials rather than forecast.
Planned order release
The date a component order must be started or placed, equal to the net requirement date minus that item's lead time.
System nervousness
The tendency of small master schedule changes to cascade into large changes in component orders throughout an MRP plan.

The Bullwhip Effect

  • Demonstrate arithmetically how a small demand change becomes a large order change as it moves upstream.
  • Identify the four operational causes of bullwhip amplification and match a countermeasure to each.
  • Explain what the beer distribution game reveals about structure versus individual judgment.
  • Evaluate the bullwhip ratio and the evidence on where amplification actually occurs.

The big picture

Babies are born at a steady rate. Diaper consumption is about as stable as consumer demand gets: no fashion cycle, no weather sensitivity, no weekend spike worth mentioning. In the early 1990s Procter and Gamble looked at the order patterns for Pampers along its own supply chain and found something that should not have been there. Retail sales were smooth. Retailer orders to distributors were noticeably choppier. Distributor orders to Procter and Gamble were choppier still. And Procter and Gamble's own orders to its material suppliers swung most violently of all.

Nothing about babies had changed. The variability was manufactured by the supply chain itself. Hau Lee, V. Padmanabhan, and Seungjin Whang named the phenomenon the bullwhip effect in a 1997 paper, after the way a small flick of the wrist produces an enormous crack at the far end of the whip.

This is the most important structural pathology in supply chain management, and the striking thing about it is how little human irrationality it requires. You can generate most of it with four lines of arithmetic and a perfectly sensible inventory policy, which is what we will do now.

The arithmetic of amplification

A retailer sells 100 units a week. Its supplier's lead time is two weeks, and it wants one extra week of demand as safety stock, so its target inventory position is three weeks of demand: 300 units. In steady state it holds 300 and orders 100 each week, exactly replacing what it sold.

Now demand rises to 110 units a week, a 10 percent increase. What does a sensible planner order this week? Two things: enough to cover the new demand, and enough to raise the inventory position to the new target.

  • New target inventory position = 3 x 110 = 330 units.
  • Gap between the old 300-unit position and the new target = 30 units.
  • Order = new demand + gap to target = 110 + 30 = 140 units.

A 10 percent rise in consumer demand just produced a 40 percent rise in the order. The planner did nothing foolish. Ordering only 110 would leave the retailer permanently under-buffered at the new demand level.

Now step upstream. The distributor does not see consumer demand. It sees an order that jumped from 100 to 140, a 40 percent increase, and it runs the identical policy with its own three-week target.

  • New target = 3 x 140 = 420 units, against a position of about 300.
  • Gap to target = 120 units.
  • Order to the factory = 140 + 120 = 260 units.

The factory now sees demand of 260 against a historical 100: a 160 percent increase caused by a 10 percent change in diaper consumption. Add a fourth echelon and it gets worse. Nobody lied, nobody panicked, nobody made an error.

EchelonDemand it observesOrder it placesIncrease over baseline
Consumer-11010 percent
Retailer11014040 percent
Distributor140260160 percent
Factory260Capacity crisis-

Then comes the second half, which hurts more. Demand settles at 110 and stays there. Inventories throughout the chain are now at their new targets, so the gap term disappears and everyone orders only 110. But the factory built to 260. For several weeks the chain orders below the new steady demand while it burns off the inventory it over-ordered. The factory, which just hired and ran overtime, now faces orders far under its baseline. That is the whipsaw: a boom followed by a bust, both manufactured internally, and it is why suppliers to long chains find capacity planning nearly impossible.

Key idea: Rational order-up-to policies amplify demand changes at every echelon, because each stage orders to cover both the new demand and the gap to a new inventory target. No irrationality is required.

The four causes, and how to tell them apart

Lee and colleagues identified four operational causes. They are worth separating carefully, because each has a different fix.

1. Demand signal processing. The arithmetic above. Each echelon re-forecasts from the orders it receives rather than from real consumption, so a signal gets re-interpreted and re-amplified at every stage. Longer lead times make it worse, because the target inventory is proportional to lead time, so the gap term is bigger.

2. Order batching. Ordering costs, full-truckload economics, and monthly planning runs all push firms to order in lumps rather than continuously. Consider twenty retailers each consuming 5 units a day and each ordering a 150-unit truckload monthly. If their order weeks are scattered, the factory sees a smooth 100 units a day. If they all order in the last week of the month, which is exactly what happens when sales quotas and month-end closes run on the calendar, the factory sees nothing for three weeks and 3,000 units in one. Same consumption, wildly different production requirement.

3. Price fluctuation and forward buying. Trade promotions invite buyers to purchase far more than they can sell, storing the surplus and skipping the next several periods. In the grocery industry at its peak, a large share of manufacturer volume was bought on deal rather than to meet demand. The purchase pattern then bears almost no resemblance to the consumption pattern, and the manufacturer's own promotion calendar is generating the variability its factories struggle with. Procter and Gamble's shift toward everyday low pricing in the early 1990s was explicitly aimed at this.

4. Rationing and shortage gaming. When supply is short and a supplier allocates proportionally to orders, the rational move for every buyer is to inflate the order. If you need 100 and expect 50 percent allocation, you order 200. When supply recovers, the phantom orders are cancelled and the supplier discovers that half its backlog was never real. This is not a theoretical curiosity. It was visible in personal protective equipment in early 2020 and in semiconductors in 2021 and 2022, where buyers placed duplicate orders across multiple suppliers and distributors, inflating apparent demand and making the shortage look worse than it was, then cancelling as supply recovered.

Key idea: Bullwhip has four distinguishable causes: forecast re-processing, order batching, promotional buying, and shortage gaming. Diagnose which one you have before choosing a fix.

Countermeasures matched to causes

CauseCountermeasureHow it works
Demand signal processingShare point-of-sale data upstream; vendor-managed inventory; collaborative planningEvery echelon forecasts from real consumption instead of re-forecasting a forecast. Walmart's supplier data sharing and the Procter and Gamble vendor-managed inventory arrangement are the standard examples.
Demand signal processingShorten lead timesTarget inventory is proportional to lead time, so a shorter lead time shrinks the gap term that does the amplifying.
Order batchingCut ordering cost with electronic ordering; mixed truckloads; third-party consolidationLowers the economic order quantity so orders can be smaller and more frequent, which is Module 3's arithmetic applied to the whole chain.
Order batchingStagger order cycles across customersUncorrelated ordering weeks smooth the aggregate even when each customer batches.
Price fluctuationEveryday low pricing; fewer, smaller trade dealsRemoves the incentive to buy on price rather than on need, so purchases track consumption.
Shortage gamingAllocate on historical sales, not current ordersRemoves the payoff for inflating orders, since a bigger order no longer earns a bigger share.
Shortage gamingNon-cancellable orders; capacity reservation contracts; visible capacity informationMakes phantom demand costly to place and reduces the fear that drives it.

Notice the pattern across the table: almost every fix either shares information or removes an incentive. Very few involve better forecasting, and that is the point. A supply chain that amplifies by construction cannot be forecast its way out of the problem.

The beer game, and what it actually proves

The beer distribution game, developed at MIT in the 1960s out of Jay Forrester's system dynamics work and used in classrooms ever since, puts four players in a chain: retailer, wholesaler, distributor, and factory. Each sees only the orders from the stage below and its own inventory, faces a fixed shipping delay, and pays a penalty for both holding stock and backorders. Consumer demand is typically a single small step, from 4 cases to 8, held constant thereafter.

The result is famously reliable. Order swings of ten or twenty times the demand change are ordinary, inventories oscillate wildly, and players end up alternating between mountains of stock and desperate shortages. Afterward, participants almost universally blame each other: the retailer panicked, the factory was asleep, the wholesaler hoarded.

Then the facilitator reveals the demand pattern, which was a single one-time step of four cases, and the room goes quiet.

John Sterman's experimental work identified the specific cognitive error that adds to the structural one. Players systematically underweight the supply line: they forget to account for the orders they have already placed but not yet received. Seeing low inventory, they order more, then more again, and when the earlier orders finally arrive all at once they are massively overstocked. This is a general failure in reasoning about systems with delays, and it shows up in hiring, in capital investment, and in public health responses as readily as it does in beer.

The teaching point is the combination. Most of the oscillation is structural, produced by delays, local information, and sensible reordering rules, so replacing the players changes little. A smaller but real part is cognitive, and it can be reduced by making the supply line visible on the screen where the decision is made. Both halves are design problems, not character problems.

Key idea: The beer game shows that a well-behaved demand step produces violent oscillation through structure alone, compounded by the human tendency to ignore orders already in the pipeline.

Measuring it, and an honest complication

The standard measure is the bullwhip ratio: the variance of a stage's orders divided by the variance of the demand it received. A ratio above 1 means amplification, below 1 means smoothing. A retailer whose weekly demand has a variance of 400 and whose orders have a variance of 1,600 has a bullwhip ratio of 4.

Here the honest complication. When Gerard Cachon, Taylor Randall, and Glen Schmidler examined United States industry data in work published in 2007, they found the picture was not uniform. Wholesalers and distributors did generally amplify, consistent with the theory. But many manufacturing industries actually smoothed demand rather than amplifying it, because production smoothing, seasonal build, and capacity constraints work in the opposite direction. Aggregation also hides the effect: a whole industry's data averages across firms and products in a way that cancels individual amplification.

What survives the complication is the mechanism, which is well documented at the level of individual products and firms, and the countermeasures, which work. What does not survive is the claim that every supply chain everywhere amplifies by a large factor. When someone asserts a bullwhip problem, ask for the ratio, computed on a specific product, at a specific stage.

Common misconceptions

  • "Bullwhip is caused by panicky or irrational buyers." The worked example used a completely rational order-up-to policy. Psychology adds to the effect; structure creates it.
  • "Better forecasting will fix it." Each echelon forecasting more accurately from distorted orders still amplifies. The fix is to give every stage the undistorted signal, not to model the distorted one better.
  • "It only hurts factories." It hurts everyone: retailers get stockouts then gluts, distributors carry excess, and the factory alternates between overtime and idleness.
  • "Bigger safety stock solves it." Larger inventory targets make the gap term larger, which amplifies more. Shorter lead times and shared data reduce it; more buffer at every stage does not.

Recap

  • The bullwhip effect is the growth of order variability as it moves upstream, documented at Procter and Gamble and named by Lee, Padmanabhan, and Whang in 1997.
  • A 10 percent demand rise produced a 40 percent retailer order and a 160 percent distributor order using an ordinary order-up-to policy, followed by a bust as the chain burned off the excess.
  • The four causes are demand signal processing, order batching, price fluctuation and forward buying, and rationing and shortage gaming.
  • Countermeasures share information or remove incentives: point-of-sale data, vendor-managed inventory, shorter lead times, cheaper ordering, staggered cycles, everyday low pricing, and allocation based on past sales.
  • The beer game demonstrates that structure, plus the human tendency to underweight the supply line, produces oscillation from a tiny demand step.
  • The bullwhip ratio measures amplification, and industry evidence shows wholesalers typically amplify while many manufacturers smooth.

Sources

  1. Wikipedia. (2025). Bullwhip effect. en.wikipedia.org
  2. Wikipedia. (2025). Beer distribution game. en.wikipedia.org
  3. Lee, H. L., Padmanabhan, V., & Whang, S. (1997). The bullwhip effect in supply chains. MIT Sloan Management Review, 38(3), 93-102. mitsloan.mit.edu
  4. Association for Supply Chain Management. (2025). Supply chain collaboration and demand distortion. ASCM. ascm.org
Key terms
Bullwhip effect
The amplification of order variability as demand information moves upstream through a supply chain.
Order-up-to policy
Ordering enough to cover current demand plus the gap between current inventory position and a target level.
Demand signal processing
Each echelon re-forecasting from the orders it receives rather than from actual consumption, which re-amplifies the signal at every stage.
Order batching
Placing infrequent large orders for economic or calendar reasons, which converts smooth consumption into lumpy demand upstream.
Shortage gaming
Inflating orders during an allocation because supply is rationed in proportion to orders, creating phantom demand that is later cancelled.
Vendor-managed inventory
An arrangement in which the supplier sees actual consumption and decides replenishment, removing one layer of order distortion.
Beer distribution game
A four-stage simulation that reliably produces violent order oscillation from a single small demand step.
Bullwhip ratio
The variance of a stage's orders divided by the variance of the demand it received; above 1 indicates amplification.

Module 5: Lean and Constraints

The Toyota Production System told from its origins, including the half most Western programs dropped, and the theory of constraints applied to factories, offices, and hospitals.

The Toyota Production System, Told Accurately

  • Explain the postwar constraints that produced the Toyota Production System and its two pillars.
  • Compute the number of kanban cards for a pull loop and explain what removing cards does.
  • Describe heijunka, standard work, kaizen, and the seven wastes with their Japanese counterparts muri and mura.
  • Assess the respect-for-people pillar using the NUMMI case and identify where lean fails when misapplied.

The big picture

Japan, 1950. Toyota is close to collapse. A financial crisis forces a restructuring that includes layoffs, a bitter labor dispute follows, and the president, Kiichiro Toyoda, resigns to take responsibility. The company that emerges has almost no capital, a domestic market that is small and wants many different vehicle types in small quantities, and a productivity gap with American manufacturers that was estimated at the time to be roughly nine to one.

Everything distinctive about the Toyota Production System comes from that position. Toyota could not afford Ford's approach of enormous dedicated machines running enormous batches, because it had neither the capital nor the volume. It could not afford large inventories, because it had no cash to tie up. It could not afford defects, because scrapping material it could barely buy was ruinous. What looks in hindsight like a philosophy was, at the time, a set of forced moves.

That origin matters for a practical reason. Firms that adopt the tools without the constraints get the appearance and not the result. Principles of Management (MGMT 301) introduces just-in-time, jidoka, and kaizen at survey level. This lesson gives you the mechanism, the arithmetic, the half of the system that usually gets dropped, and an honest account of where it fails.

The house: two pillars on a foundation

Toyota's own teaching diagram is a house. The roof is the goal: best quality, lowest cost, shortest lead time, best safety, highest morale. Two pillars hold it up, and a foundation supports both.

  • Pillar one, just-in-time: produce only what is needed, only when it is needed, only in the amount needed.
  • Pillar two, jidoka: build in quality by stopping when something goes wrong, so a defect is never passed forward.
  • Foundation: heijunka (level scheduling), standardized work, and kaizen (continuous improvement), plus stability in equipment and people.

A house with one pillar falls down. Just-in-time without jidoka is a system with no inventory and no defect detection, which stops constantly and produces scrap in volume. Jidoka without just-in-time catches defects but leaves the inventory that hides the problems causing them. The two are a matched pair, and most failed lean programs implement half of one.

Jidoka: the older half

Jidoka predates the car company. In 1896 Sakichi Toyoda invented an automatic loom that stopped itself the instant a thread broke. Before that, a loom could weave defective cloth for hours unattended. After it, a single operator could tend many machines, because each machine would signal rather than fail silently. Toyota describes the principle as automation with a human touch, which is a translation of a pun in the Japanese term.

Carried onto the assembly line, jidoka became the andon cord: any worker who sees a problem pulls a cord, which lights a signal and calls the team leader. If the problem is not resolved within the takt cycle, the line stops. The remarkable part is the culture around it. At Toyota, pulling the cord is expected and its use is tracked as a positive indicator, because a line that never stops is a line that is passing defects downstream. New workers at Toyota plants are explicitly told that failing to pull the cord is the error, not pulling it.

Related is poka-yoke, or mistake-proofing: designing the process so the error is physically impossible or immediately obvious. A fixture that only accepts a part in the correct orientation. A connector that will not mate the wrong way. A parts tray with exactly the right number of screws, so a leftover screw signals a missed step. Poka-yoke is more reliable than attention, and this is one of the very few places in operations where a cheap fixture genuinely eliminates a class of error rather than reducing it.

Key idea: Jidoka means stopping at the moment a defect appears rather than passing it forward, supported by mistake-proofing devices and a culture where stopping the line is rewarded.

Just-in-time, pull, and the kanban arithmetic

The insight that made just-in-time practical came from a supermarket. Taiichi Ohno, visiting the United States in the 1950s, was struck by how American supermarkets worked: shelves held a limited quantity, customers took what they wanted when they wanted it, and staff restocked only what had been removed. Nobody forecast each shelf and pushed inventory onto it. Consumption itself triggered replenishment.

That is pull, and it inverts the logic of a scheduled push system. In a push system, a central schedule tells each station what to make, and the station makes it whether or not the next station is ready. In a pull system, a downstream station's consumption authorizes upstream production, and nothing is made without that authorization.

The signal is the kanban, literally a sign or card. A container of parts carries a card; when the container is opened at the consuming station, the card goes back to the producing station and authorizes exactly one more container. The number of cards in the loop is a hard ceiling on work in process, which is what makes this so powerful: you control inventory by controlling cards, not by exhorting people.

How many cards? The standard calculation is:

Number of kanbans = (daily demand x replenishment lead time x (1 + safety factor)) / container size

Work it. A station consumes 500 units a day. The replenishment loop, meaning the time from card release to full container returning, takes half a day. Containers hold 25 units. Management wants a 20 percent buffer.

  • Numerator: 500 x 0.5 x 1.2 = 300 units.
  • Number of kanbans = 300 / 25 = 12 cards.
  • Maximum work in process in this loop = 12 x 25 = 300 units, and not one unit more, structurally.

Now cut the replenishment lead time in half, to a quarter day, by moving the supplying station closer and delivering more often:

  • Numerator: 500 x 0.25 x 1.2 = 150, so 6 cards and 150 units of work in process.

Notice that this is Little's Law again, wearing different clothes. Inventory equals rate times time, so halving the time halves the inventory. The kanban formula is just that equation with a safety factor and a container size.

Here is the part that distinguishes Toyota from a firm that merely bought some cards. Toyota deliberately removes kanbans to expose problems. The metaphor used inside the company is a boat on a lake: the water is inventory, the rocks are problems (long changeovers, unreliable machines, quality defects, unreliable suppliers). Lower the water and you hit a rock. The point is not to run aground; it is to find the rock and remove it, then lower the water again. Remove a card, watch where the line stalls, fix that cause, remove another card.

This only works if you actually fix the rock. Removing cards without fixing anything is not lean, it is just running short of parts. That distinction is the single most common failure in lean implementation and it is worth stating plainly: the inventory reduction is the measurement, not the intervention.

Key idea: Kanban caps work in process by physical card count, computed as demand times lead time divided by container size. Removing cards is a deliberate diagnostic that only works if the exposed problem is then fixed.

Heijunka: why level scheduling comes first

Pull systems break under lumpy demand. If Monday requires 300 units and Tuesday requires 20, the kanban loop sized for the average will starve on Monday and overflow on Tuesday, and the supplier upstream sees exactly the bullwhip pattern from the last lesson. So Toyota levels the schedule before it pulls, a practice called heijunka.

Take a plant building three models. Monthly demand is 1,000 of model A, 500 of B, and 500 of C, over 20 working days.

  • Daily requirement: 50 A, 25 B, 25 C, or 100 units a day.
  • Takt time with 480 available minutes: 480 / 100 = 4.8 minutes per unit.

A conventional plant runs ten days of A, then five of B, then five of C, because changeovers are expensive. Consider what that does upstream: the supplier of A-specific parts sees ten days of ferocious demand and ten days of nothing, and must hold inventory or capacity for the peak. Multiply across hundreds of parts and the whole supply base is carrying the cost of your batching decision.

The level alternative runs a repeating mixed sequence: A, B, A, C, A, B, A, C, and so on. Each cycle of four units delivers two A, one B, and one C, exactly matching the 50:25:25 ratio. Every supplier now sees smooth, predictable daily demand. The catch is that this requires changeovers between models measured in seconds, not hours, which is precisely why Toyota invested so heavily in setup reduction. Setup reduction is not a cost-saving project; it is the enabling condition for level scheduling, which is the enabling condition for pull. The chain of dependencies runs in that order, and programs that start with kanban cards are starting three steps too late.

The wastes, including the two that get dropped

Ohno's seven wastes, muda, are the standard list:

WasteWhat it looks like
OverproductionMaking more or sooner than needed. Ohno called it the worst waste because it creates all the others.
WaitingIdle time for people or parts, which Module 1 measured as flow time efficiency.
TransportationMoving material that the customer does not value being moved.
Over-processingTighter tolerances, extra approvals, or features nobody asked for.
InventoryStock beyond what the flow requires, which hides the problems above.
MotionPeople reaching, walking, searching. Distinct from transportation, which moves material.
DefectsScrap, rework, and the inspection built to catch them.

Many practitioners add an eighth, unused human talent, and it is not a decoration. It follows directly from the second pillar of the Toyota Way discussed below.

Two other Japanese terms travel with muda and are usually dropped in Western versions, which is a loss. Muri is overburden: pushing people or equipment beyond reasonable limits, which produces injuries, breakdowns, and quality failures. Mura is unevenness, the lumpiness that heijunka attacks. The three are causally linked: mura creates muri, and muri creates muda. A program that hunts for waste without addressing unevenness and overburden is treating symptoms, and it typically shows up as a workforce doing the same job faster under more strain. That is the diagnostic signature of a lean program that has gone wrong.

Standard work and kaizen

Standardized work specifies the current best-known method: the sequence, the timing, and the standard work in process at each station. It sounds like the opposite of empowerment and is in fact its precondition. Without a documented standard there is no baseline, so an improvement cannot be demonstrated, only asserted. Taiichi Ohno's formulation was that where there is no standard there can be no kaizen.

Kaizen is continuous improvement made by the people doing the work, in small increments, constantly. Its supporting tools are simple and specific: the five whys, which chases a cause chain past the first plausible answer; genchi genbutsu, going to the actual place to see the actual thing rather than discussing a report; and the A3, a single sheet of paper that forces a problem statement, current condition, analysis, countermeasures, and follow-up into one readable page.

The five whys in practice, on a stopped machine:

  1. Why did the machine stop? An overload tripped the fuse.
  2. Why was it overloaded? The bearing was not sufficiently lubricated.
  3. Why not? The lubrication pump was not pumping enough.
  4. Why not? The pump shaft was worn and rattling.
  5. Why was it worn? There was no filter, so metal shavings got in.

Stop at the first why and you replace a fuse, and the machine stops again next week. Reach the fifth and you install a filter. The number five is not magic; the discipline of not stopping at the first plausible cause is the whole technique.

Key idea: Standardized work is the baseline that makes improvement measurable, and kaizen is small continuous improvement by the people doing the work, disciplined by five whys, going to see, and the A3.

The half that gets dropped: respect for people

Toyota's own 2001 internal document, The Toyota Way, rests on two pillars, not one. The first is continuous improvement. The second is respect for people. Western lean programs adopted the first almost universally and the second almost not at all, and that omission explains a large share of failed implementations.

The clearest natural experiment in the history of manufacturing makes the case. General Motors' Fremont, California plant was, by GM's own account, the worst plant in its system by the late 1970s: absenteeism running around 20 percent, alcohol and drug use on the job, thousands of unresolved grievances, and quality among the poorest in the company. GM closed it in 1982.

In 1984 it reopened as NUMMI, a joint venture between GM and Toyota, building cars with the Toyota Production System. The decisive detail is the workforce: under the agreement with the United Auto Workers, most of the same workers were rehired, including the union leadership that GM had considered the problem. Within about a year, absenteeism had fallen to the low single digits and the plant's quality was among the best in the GM system.

Same people. Same building. Different system. The workers had not been the problem, and telling them they were had been part of the problem. This case is documented in detail, including in a widely heard 2010 public radio account, and it is the strongest available evidence for the claim that management systems rather than worker character drive most performance differences.

What did respect for people actually mean at NUMMI? Concretely: workers were trained in problem solving and expected to improve their own jobs; team leaders supported rather than policed; the andon cord gave every worker authority to stop a multi-million-dollar line; managers ate in the same cafeteria and parked in the same lot; and, critically, there was a commitment not to lay workers off as a result of productivity improvements. That last one is load-bearing. You cannot ask people to eliminate their own jobs. If improvement produces layoffs, improvement stops within one cycle, and no amount of training restarts it.

Do not romanticize this. Toyota's employment security applied to regular employees and was cushioned by a large tier of temporary workers who did not have the same protection, and by suppliers who absorbed pressure. The 2009 NUMMI closure, after GM's bankruptcy, put thousands out of work regardless. The honest claim is narrower and still important: the improvement half of lean does not function without the respect half, and firms that take the tools and skip the commitment get a short burst of savings and then nothing.

Key idea: The Toyota Way has two pillars, continuous improvement and respect for people. NUMMI showed the same workforce transformed by a different system, and employment security is a precondition for improvement rather than a reward for it.

Where lean fails when misapplied

Four failure modes recur, and all four are avoidable.

1. Lean as a layoff program. Practitioners have a sardonic acronym for it: L.A.M.E., lean as misguidedly executed. A program that starts by identifying surplus headcount destroys the improvement engine at the moment it needs it most. The workable version redeploys the freed capacity into growth, insourcing, or improvement work, and says so publicly before the first event.

2. Just-in-time imported into high-variability environments without buffers. Toyota levels its schedule before it removes inventory. An emergency department cannot level its arrivals, and a defense supplier facing lumpy orders cannot either. Applied there, the correct lean move is to reduce variability where possible and to size buffers deliberately where it is not, which is Module 2's arithmetic, not to drive inventory toward zero. Zero inventory in a high-variability system is not lean; it is a stockout schedule.

3. Cost-cutting wearing lean vocabulary. Single sourcing to get a better price, eliminating safety stock to hit a working capital target, and stretching supplier payment terms are all sometimes labeled lean. They are not. Toyota's own behavior is the counterexample: after the 2011 Tohoku earthquake disrupted its supply base, Toyota built a detailed multi-tier supply chain database precisely so it could see beyond its direct suppliers, and it asked suppliers of critical semiconductors to hold substantial buffer stock, on the order of several months. A company famous for eliminating inventory deliberately created inventory where the risk analysis called for it. That is lean thinking; a blanket inventory target is not.

4. Tool worship. The most common form is 5S theater: workplace organization audits scored monthly, tape outlines on benches, and a scoreboard, with no change in flow, no reduction in changeover time, and no problem-solving capability built. Kanban boards in offices have joined it. The tools are fine; detached from flow and problem solving they are decoration.

Steven Spear and H. Kent Bowen's 1999 study cut to what the tools sit on. They argued that Toyota's real system is four implicit rules: all work is highly specified as to content, sequence, timing, and outcome; every customer-supplier connection is direct and unambiguous; every product and service pathway is simple and direct; and any improvement is made using the scientific method, under a teacher, at the lowest possible level. Read that last rule carefully. Toyota's durable advantage is a company-wide habit of running experiments, coached, at the front line. Cards, cords, and boards are the visible residue of that habit, and copying the residue does not produce the habit.

Common misconceptions

  • "Lean means fewer people." Lean frees capacity. What a company does with freed capacity is a separate decision, and choosing layoffs ends the program.
  • "Lean means zero inventory." It means the right inventory in the right place, sized to the variability you have not yet removed. Toyota deliberately buffers critical components.
  • "Kanban is a software board with columns." A kanban is a physical authorization signal whose count caps work in process. A board without a work-in-process limit is a to-do list.
  • "You can start with the tools and add the culture later." The sequencing evidence runs the other way: setup reduction enables leveling, leveling enables pull, and problem-solving capability plus job security enable all of it.

Recap

  • The Toyota Production System came from postwar constraints: no capital, small fragmented demand, and no tolerance for scrap.
  • Its two pillars are just-in-time and jidoka, on a foundation of heijunka, standardized work, and kaizen; implementing one pillar alone fails.
  • Kanban count equals demand times lead time times a safety factor divided by container size: 12 cards and 300 units here, falling to 6 cards when lead time halved.
  • Removing kanbans is a deliberate diagnostic, useful only when the exposed problem is actually fixed.
  • Heijunka levels the mix before pull is attempted, which is why setup reduction is the enabling investment.
  • The seven wastes travel with muri and mura, and unevenness causes overburden which causes waste.
  • The Toyota Way's second pillar is respect for people; NUMMI transformed the same workforce with a different system, and employment security is a precondition for improvement.
  • Lean fails as a layoff program, as just-in-time without buffers in variable environments, as cost-cutting in lean vocabulary, and as tool worship.

Sources

  1. Wikipedia. (2025). Toyota Production System. en.wikipedia.org
  2. Toyota Motor Corporation. (2025). Toyota production system: vision and philosophy. global.toyota
  3. Lean Enterprise Institute. (2025). Lean thinking and practice. LEI. lean.org
  4. Spear, S., & Bowen, H. K. (1999). Decoding the DNA of the Toyota Production System. Harvard Business Review, 77(5), 96-106. hbr.org
  5. Wikipedia. (2025). NUMMI. en.wikipedia.org
Key terms
Jidoka
Building quality in by stopping the process the moment a defect appears, rather than passing it forward; expressed by the andon cord.
Poka-yoke
Mistake-proofing: designing a process or fixture so an error is physically impossible or immediately obvious.
Pull system
Production authorized by downstream consumption rather than by a central schedule pushing work forward.
Kanban
A card or signal authorizing production of exactly one container; the number of cards is a hard cap on work in process.
Heijunka
Level scheduling that spreads model mix and volume evenly across the period so pull systems and suppliers see steady demand.
Muri and mura
Overburden and unevenness; unevenness creates overburden, which creates waste, so both must be addressed before hunting muda.
Standardized work
The documented current best method covering sequence, timing, and standard work in process; the baseline that makes improvement measurable.
Respect for people
The second pillar of the Toyota Way, including problem-solving training, front-line authority, and employment security as a precondition for improvement.

Theory of Constraints and Improvement in Services

  • Apply Goldratt's five focusing steps to a multi-stage process and quantify the gain from exploitation alone.
  • Explain drum-buffer-rope and its relationship to Little's Law and work-in-process caps.
  • Choose a product mix using throughput per constraint minute rather than unit margin.
  • Adapt flow improvement to services and healthcare, including emergency department capacity computed with Little's Law.

The big picture

In 1984 a physicist turned consultant named Eliyahu Goldratt published a management book in the form of a novel. The Goal follows a plant manager with ninety days to save his factory, and its most memorable scene is not in a factory at all. He is chaperoning a scout troop hike, and the line of boys keeps stretching out. The fast walkers pull ahead and stop to wait; the gaps open again the moment they start. At the back is a heavy, slow boy named Herbie. No matter how the manager rearranges the front of the line, the troop as a whole moves at Herbie's pace, and the distance the troop covers is determined entirely by him.

The solution is what makes it a good story. He does not tell Herbie to walk faster. He puts Herbie at the front, so nobody can run ahead and create gaps, and then he redistributes the contents of Herbie's backpack among the other boys. The troop speeds up, and it speeds up because everyone else gave something to the constraint.

That is the theory of constraints in one image, and it is a genuine addition to the bottleneck arithmetic in Module 1. Module 1 taught you to find the bottleneck. This lesson is about what to do once you have found it, which turns out to be surprisingly counterintuitive and surprisingly cheap.

The five focusing steps

  1. Identify the constraint. The resource with the lowest capacity relative to demand. There is essentially always exactly one that matters at a time.
  2. Exploit the constraint. Get every possible unit out of it without spending capital. This step is skipped constantly and is where most of the free money is.
  3. Subordinate everything else to that decision. Every other resource runs at the pace the constraint can absorb, even though that means visible idleness elsewhere. This is the step organizations hate.
  4. Elevate the constraint. Now spend money: add a machine, add a shift, outsource some volume.
  5. If the constraint has moved, go back to step one, and do not let inertia become the new constraint. Policies written for the old bottleneck outlive it, and managing a resource that stopped being the constraint two years ago is a common and expensive habit.

Working the steps on a machine shop

A small shop runs three sequential operations on one machine each, with 480 available minutes a day.

OperationTime per unitDaily capacity
Cut12 minutes480 / 12 = 40 units
Weld20 minutes480 / 20 = 24 units
Finish15 minutes480 / 15 = 32 units

Step 1. Weld is the constraint at 24 units a day, and the shop ships 24 units a day.

Step 2, exploit. Walk the weld station for one day and write down every minute it is not welding.

  • The welder takes lunch and two breaks, during which the machine sits: 45 minutes. Stagger the breaks so a cross-trained operator covers the station, and available time rises to 525 minutes. Capacity: 525 / 20 = 26.25 units a day.
  • Changeovers consume 30 minutes a day. Apply the setup-reduction methods from Module 5's first lesson to get that to 10. Available welding time rises to 545 minutes. Capacity: 545 / 20 = 27.25 units a day.

That is a 13.5 percent throughput increase for zero capital, purchased with a break schedule and a changeover project. Now the subtler exploitation move. Suppose 8 percent of parts are found defective and scrapped at final inspection, after welding.

  • Good output today: 27.25 x 0.92 = 25.07 units a day. The constraint spent 8 percent of its time welding parts that were thrown away.
  • Move inspection before the weld station, so defective parts never consume constraint time. Good output becomes 27.25 units a day, an 8.7 percent gain, again for no capital.

The general rule falls out cleanly: never let the constraint process something that will not be sold. Inspect before the constraint, never after. Never let it run out of material. Never let it make something a customer has not ordered.

Step 3, subordinate. Cut can produce 40 a day and the constraint can only absorb 27. If Cut runs flat out, the shop accumulates 13 units of work in process a day in front of Weld, which costs money, lengthens flow time by Little's Law, and hides quality problems. So Cut is instructed to produce 27 and then stop, and its operator will be idle for part of every day. This is correct, and it is the hardest sentence in this lesson for a manager to accept. Local efficiency measures fight it directly: a supervisor rated on machine utilization will keep Cut running, and the shop will be worse off.

Step 4, elevate. Rent a second welder. Weld capacity becomes 54.5 a day, and the constraint moves to Finish at 32 a day. Throughput rises from 27.25 to 32, a real gain, but notice it is far smaller than the doubling of weld capacity, because the constraint moved. Anyone who justified the second welder by promising to double output has misled the room.

Step 5. Start again at Finish, and go delete the policies written for Weld.

Key idea: Exploiting a constraint, meaning removing idle time, changeovers, and scrapped work from it, typically buys double-digit throughput gains for no capital, and should always be completed before anyone proposes buying capacity.

Drum, buffer, rope

Goldratt's scheduling mechanism has three parts named for the hike. The drum is the constraint's schedule, which sets the beat for the whole plant. The buffer is a time buffer of work placed in front of the constraint, so that an upstream hiccup never starves it; note it is deliberately a buffer of time, sized by how long upstream disruptions typically last. The rope is the signal that releases new material into the plant only at the rate the constraint consumes it, exactly like a rope tying the front of the scout line to Herbie so nobody runs ahead.

Compare this to kanban and the family resemblance is obvious: both cap work in process, both release work based on consumption rather than a forecast, both shorten flow time. The difference is emphasis. Kanban caps inventory at every loop, which suits repetitive production with level demand. Drum-buffer-rope caps release at one point and protects one resource, which suits shops with variable routings and a clear bottleneck. Both are Little's Law used as a control policy: hold inventory down and flow time falls with it.

Throughput accounting and the product mix decision

Theory of constraints comes with its own measurement scheme, deliberately hostile to traditional cost allocation. It uses three numbers: throughput (sales revenue minus truly variable cost, essentially material), investment (money tied up, including inventory), and operating expense (everything spent to turn investment into throughput). The argument is that allocating fixed overhead per unit encourages producing inventory to absorb overhead, which looks profitable on paper and consumes cash in fact.

The most useful practical output is the product mix rule. A shop makes two products, and the weld station from earlier is the constraint with 2,400 available minutes a week.

Product PProduct Q
Selling price110130
Material cost4560
Throughput per unit6570
Constraint minutes per unit1530
Throughput per constraint minute4.332.33
Weekly demand10060

Q looks better on every conventional measure: higher price, higher unit throughput, probably a higher gross margin percentage. A sales force paid on revenue will push Q. Now run both orderings against the 2,400-minute constraint.

Prioritize P. 100 units x 15 minutes = 1,500 minutes, generating 100 x 65 = 6,500 of throughput. Remaining 900 minutes produce 900 / 30 = 30 units of Q, generating 30 x 70 = 2,100. Total: 8,600 per week.

Prioritize Q. 60 units x 30 minutes = 1,800 minutes, generating 60 x 70 = 4,200. Remaining 600 minutes produce 600 / 15 = 40 units of P, generating 40 x 65 = 2,600. Total: 6,800 per week.

The correct ordering is worth 1,800 dollars a week, about 26 percent more throughput from the same plant, the same people, and the same week. Nothing was bought. The only change was ranking products by throughput per constraint minute rather than by unit margin. Whenever a resource is genuinely scarce, the correct denominator is that scarce resource, and this generalizes well beyond factories: an operating room minute, a radiologist hour, a senior engineer's week.

Key idea: When a resource is the constraint, rank work by throughput per constraint minute, not by unit margin. Here the correct ranking produced 26 percent more throughput at zero cost.

An honest word on the evidence

The theory of constraints has a logic that is hard to argue with and an evidence base that is thinner than its enthusiasts suggest. Most published results are case studies, often written by consultants or by the firms themselves, with no comparison group and no accounting for the effect of simply paying close attention to a process. That is not a reason to dismiss it, because the underlying arithmetic is straightforwardly correct: throughput cannot exceed the constraint, so capacity recovered at the constraint is real. It is a reason to be skeptical of dramatic percentage claims and of the framing that treats it as a rival to lean and Six Sigma rather than a complement. In practice the three combine sensibly: constraints tell you where to work, lean tells you how to remove waste and shorten flow, and Six Sigma tells you how to reduce variation once you have a stable, measured process.

Improvement in services

Everything above transfers to service and office work, with three adaptations that matter.

First, the customer is inside the process, so you cannot buffer against their variability with inventory. You buffer with capacity or with scheduling instead, which is Module 2's argument.

Second, the waste is invisible. Nobody trips over a pile of half-finished insurance claims the way they trip over a pallet. The standard countermeasure is value stream mapping: walk the process end to end, record the actual clock time of each step and each wait, and compute the flow time efficiency from Module 1. A purchase requisition with 45 minutes of work inside six days of elapsed time made the ratio visible at 1.6 percent, and that single number reliably reorients a room that was about to buy faster software.

Third, handoffs are the constraint more often than people are. In office processes the queue is nearly always at a transition between departments, in an approval, or in a batching rule such as running payroll or purchase orders once a week.

Healthcare: where this matters most

Healthcare has adopted operations methods seriously since the early 2000s, and the results are genuinely instructive.

Little's Law in an emergency department. This is the single most useful calculation in hospital operations. An emergency department sees 60 patients per 12-hour shift, so arrivals run at 5 per hour. Average length of stay is 4 hours.

  • Patients present at any moment = 5 x 4 = 20 patients.
  • If the department has 18 treatment spaces, it is structurally over capacity. Patients will wait in the lobby or be boarded in hallways, and no amount of urging staff to hurry changes the arithmetic.
  • Two levers, and only two. Reduce length of stay to 18 / 5 = 3.6 hours, or reduce arrivals to 18 / 4 = 4.5 per hour.

Now the twist that makes this a constraints problem. Length of stay in an emergency department is often dominated not by emergency care but by boarding: admitted patients waiting for an inpatient bed. In that case the true constraint is not in the emergency department at all; it is the timing of inpatient discharges upstairs. Hospitals that pushed discharge orders and transport earlier in the day have reported meaningful reductions in emergency department crowding, while hospitals that added emergency department beds without touching discharge timing filled the new beds with boarders. Improving a non-constraint achieves nothing, exactly as step three predicts.

Smoothing elective schedules. A widely reported and counterintuitive finding, associated with the work of Eugene Litvak and colleagues and disseminated through the Institute for Healthcare Improvement, is that much hospital crowding comes not from unpredictable emergency arrivals but from the scheduled elective surgery calendar, which surgeons naturally clump into preferred days. Emergency arrivals are actually fairly stable week to week; the artificial peaks are self-inflicted. Hospitals that leveled elective schedules across the week, which is heijunka applied to operating rooms, reduced peak census, cancellations, and overtime without adding beds.

Virginia Mason. Virginia Mason Medical Center in Seattle adopted the Toyota Production System explicitly from 2002, naming its version the Virginia Mason Production System, and sent leaders to Japan to study it. It reported substantial reductions in inventory, staff walking distance, and lead times, and it built a patient safety alert system directly modeled on the andon cord, in which any staff member can halt a process they believe is unsafe. Treat the reported figures with the caution due any self-reported improvement program, and treat the andon transplant as the genuinely transferable idea: giving front-line staff the authority to stop is a structural change, not a slogan.

Where it fails. Healthcare lean goes wrong in a specific and predictable way: metrics that measure speed get applied to work whose value is not speed. Pushing length of stay down without regard to readmission, treating a patient encounter as a widget, and running staffing at the utilization levels a factory can tolerate all produce worse care and rapid burnout. Module 2's arithmetic says a high-variability system needs a large capacity cushion, and a hospital scheduled to 95 percent occupancy is not efficient, it is one bad night from diversion. The methods are sound; the target variable has to be chosen with more care than in a factory.

Key idea: In services and healthcare the constraint is usually a handoff, a batching rule, or a downstream resource, and Little's Law applied to an emergency department gives an exact capacity answer that exhortation cannot change.

Common misconceptions

  • "Improve everywhere at once." Improvement anywhere but the constraint produces no additional output. Find Herbie first.
  • "An idle worker is waste." Idleness at a non-constraint is free and often correct. Idleness at the constraint is the only idleness that costs anything.
  • "Sell the product with the highest margin." When a resource is scarce, rank by throughput per unit of that resource. The higher-margin product lost by 1,800 dollars a week here.
  • "Healthcare is too complex for operations methods." Arrival rates, length of stay, and bed counts obey Little's Law exactly. The complexity is in choosing which variable to improve, not in whether the arithmetic applies.

Recap

  • The five focusing steps are identify, exploit, subordinate, elevate, and repeat while guarding against inertia.
  • Exploiting the weld constraint through staggered breaks and setup reduction raised capacity from 24 to 27.25 units a day at no capital cost, and moving inspection upstream added another 8.7 percent of good output.
  • Subordination requires deliberate idleness at non-constraints, which local efficiency metrics actively fight.
  • Elevating moved the constraint from Weld to Finish, so throughput rose from 27.25 to 32 rather than doubling.
  • Drum-buffer-rope schedules the constraint, buffers time in front of it, and releases material at its pace; it is Little's Law used as a control policy.
  • Ranking by throughput per constraint minute produced 8,600 per week against 6,800 for the higher-margin ordering.
  • In an emergency department with 5 arrivals per hour and a 4-hour stay, 20 patients are present at any moment, and the real constraint is often inpatient discharge timing.

Sources

  1. Wikipedia. (2025). Theory of constraints. en.wikipedia.org
  2. Wikipedia. (2025). Eliyahu M. Goldratt. en.wikipedia.org
  3. Institute for Healthcare Improvement. (2025). Flow and variability in hospital systems. IHI. ihi.org
  4. Agency for Healthcare Research and Quality. (2025). Emergency department crowding and patient flow. AHRQ. ahrq.gov
  5. Lean Enterprise Institute. (2025). Value stream mapping. LEI. lean.org
Key terms
Theory of constraints
A management method holding that system output is governed by one constraint at a time, improved through five focusing steps.
Exploit the constraint
Recovering every available minute at the constraint without spending capital, by removing idle time, changeovers, and work that will be scrapped.
Subordination
Running non-constraint resources at the constraint's pace, accepting visible idleness rather than building work in process.
Drum-buffer-rope
Scheduling to the constraint's beat, protecting it with a time buffer, and releasing material only at its consumption rate.
Throughput accounting
Measuring throughput as sales minus truly variable cost, alongside investment and operating expense, instead of allocating overhead per unit.
Throughput per constraint minute
The ranking rule for product mix when a resource is scarce: unit throughput divided by the constraint time the unit consumes.
Value stream mapping
Walking a process end to end and recording work time and wait time at each step to expose flow time efficiency.
Boarding
Admitted patients held in an emergency department awaiting an inpatient bed, which makes inpatient discharge timing the real constraint.

Module 6: The Supply Chain

Where to put facilities, whether to make or buy, what it really costs to land a container, and how to build a chain that survives the shock nobody scheduled.

Network Design, Sourcing, and Make versus Buy

  • Apply weighted factor rating, locational break-even, and center of gravity to a facility location decision.
  • Evaluate a make-versus-buy decision on total cost, including the costs commonly omitted from a quote.
  • Classify purchases with the Kraljic matrix and match a sourcing approach to each quadrant.
  • Assess the strategic risks of outsourcing using documented company cases.

The big picture

Every process decision in the first five modules assumed the facilities already existed and the parts already came from somewhere. This module goes back and makes those decisions.

A supply chain is the network of organizations, people, activities, information, and resources that moves a product from raw material to end customer, and back again when it is returned. Its structure sets a floor under everything downstream. Choose a supplier eleven thousand miles away and you have chosen a six-week lead time, which by Module 3's arithmetic has chosen your safety stock, and by Module 4's arithmetic has chosen your forecast horizon and your bullwhip amplification. Network decisions are slow, expensive, and hard to reverse, which is why they deserve arithmetic rather than instinct.

A useful shared vocabulary comes from the SCOR model maintained by the Association for Supply Chain Management, which organizes supply chain activity into six processes: plan, source, make, deliver, return, and enable. This lesson covers source and the network structure that make and deliver sit on.

Location, method one: weighted factor rating

Location decisions involve factors that cannot be reduced to a single unit: labor cost, labor availability and skill, proximity to customers, proximity to suppliers, transport infrastructure, taxes and incentives, energy cost and reliability, political and legal risk, exchange rate exposure, and quality of life for the people you need to relocate.

Weighted factor rating makes the judgment explicit. Assign weights summing to 1, score each site from 0 to 100 on each factor, multiply and add. Three candidate sites for a distribution facility:

FactorWeightSite ASite BSite C
Labor cost0.30806090
Transportation access0.25609550
Labor availability0.20708060
Taxes and incentives0.15906070
Quality of life0.10708060
Weighted total1.0073.5074.7568.00

Site A computes as 0.30(80) + 0.25(60) + 0.20(70) + 0.15(90) + 0.10(70) = 24 + 15 + 14 + 13.5 + 7 = 73.5. Site B wins by 1.25 points.

Now do the thing most people skip, which is the only thing that makes this method honest: test whether the answer survives a change in the weights. Suppose labor cost really deserves 0.40 and transportation 0.15.

  • Site A: 32 + 9 + 14 + 13.5 + 7 = 75.5
  • Site B: 24 + 14.25 + 16 + 9 + 8 = 71.25
  • Site C: 36 + 7.5 + 12 + 10.5 + 6 = 72.0

The winner changed. A 1.25-point margin was never a real margin; it was noise inside judgments about weights and scores that nobody can defend to two decimal places. The correct use of this method is to structure a discussion and to reveal which factor is actually driving the decision, not to produce a number that ends the argument. If the ranking flips under plausible weights, say so.

Location, method two: locational break-even

When the differences really are monetary, plot cost against volume. Three sites with different capital intensity:

SiteAnnual fixed costVariable cost per unitTotal cost at 100,000 units
A (manual, leased)600,00018.002,400,000
B (semi-automated)900,00012.002,100,000
C (highly automated)1,400,0008.002,200,000
  • A versus B: 600,000 + 18Q = 900,000 + 12Q, so 6Q = 300,000 and Q = 50,000 units.
  • B versus C: 900,000 + 12Q = 1,400,000 + 8Q, so 4Q = 500,000 and Q = 125,000 units.

So A is cheapest below 50,000 units a year, B between 50,000 and 125,000, and C above 125,000. The decision therefore is not really about sites; it is about how confident you are in your volume forecast, and Module 4 told you that confidence declines with horizon. A firm forecasting 130,000 units with a 30 percent error band is choosing between B and C on a coin flip, and the right response is often to choose the more flexible option and pay a little more per unit for the right to be wrong.

Location, method three: center of gravity

For a distribution center serving known demand points, the center of gravity finds the location minimizing weighted straight-line distance. Weight each coordinate by volume:

Demand pointxyVolumex times volumey times volume
Store 1301202,00060,000240,000
Store 2901101,00090,000110,000
Store 360403,000180,000120,000
Total6,000330,000470,000
  • x coordinate = 330,000 / 6,000 = 55
  • y coordinate = 470,000 / 6,000 = 78.3

The center of gravity is at roughly (55, 78), pulled downward from the two northern stores by the heavier volume at Store 3. Its limitations are worth stating because people over-trust it: it assumes straight-line travel, cost proportional to volume times distance, no inbound freight considerations, and no differences in land, labor, or tax between candidate points. It is a starting point for a search, not an answer.

Key idea: Location methods structure a decision rather than settle it. Always test whether the ranking survives plausible changes in weights, volume forecasts, and unmodeled costs.

How many facilities?

The count matters more than any individual site. Consolidating cuts safety stock by the square root law from Module 3: nine warehouses holding 74 units of safety stock each need only about 223 in one location instead of 666. Consolidating also concentrates fixed costs and management attention.

Everything else pushes the other way. Fewer facilities mean longer outbound shipping distances, higher outbound freight per unit, longer delivery times, and a single point of failure. The classic pattern in United States distribution over the last two decades ran in both directions for exactly this reason: firms centralized in the 2000s to capture the inventory savings, then rebuilt dense regional and metropolitan networks once fast delivery became the competitive priority. Neither move was a mistake; the competitive priority changed, and Module 1 said the network should follow the priority.

Make versus buy, costed honestly

This is the decision where sloppy arithmetic costs the most, because the purchase quote is a single visible number and the costs of buying are scattered across four departments.

A company needs 40,000 units a year of a subassembly. A supplier quotes 14.60 dollars per unit delivered. Internal manufacturing would cost 15.95 per unit. It looks like an easy 9 percent saving. Build the real comparison.

Making it.

Cost elementAnnual amount
Equipment: 450,000 over 5 years90,000
Dedicated supervision and floor space (avoidable if not made)60,000
Materials at 6.20 per unit248,000
Direct labor at 4.10 per unit164,000
Variable overhead at 1.90 per unit76,000
Total to make638,000, or 15.95 per unit

Note that only avoidable fixed costs belong here. Allocated corporate overhead that continues whether or not you make this part is not a cost of making it, and including it is the single most common error in these analyses, biasing every decision toward outsourcing.

Buying it.

Cost elementAnnual amount
Purchase price, 40,000 at 14.60584,000
Incoming inspection and quality assurance18,000
Extra inventory: supplier is 6 weeks away, so about 6,000 more units on hand and in transit, carried at 25 percent21,900
Supplier management, audits, travel, expediting15,000
One-time qualification and tooling of 40,000, amortized over 4 years10,000
Total to buy648,900, or 16.22 per unit

Buying costs 10,900 dollars a year more, despite a quoted price 9 percent below the internal cost. The quote was never the cost. Inspection, inventory, supplier management, and qualification are all real and all invisible on a purchase order.

Now find the volume at which the answer flips, because it always turns on volume. Making has 150,000 of avoidable fixed cost and 12.20 of variable cost per unit. Buying has about 43,000 of fixed cost (inspection, supplier management, amortized qualification) and roughly 15.15 per unit once the price and the volume-scaled inventory carrying are included.

  • 150,000 + 12.20Q = 43,000 + 15.15Q
  • 107,000 = 2.95Q, so Q = about 36,300 units.

Above roughly 36,300 units a year, make. Below it, buy. Check the endpoints: at 60,000 units making costs 882,000 against about 952,000 to buy, so make wins comfortably. At 20,000 units making costs 394,000 against about 346,000 to buy, so buy wins. Fixed cost spread over volume is the whole story, which is why the same part can correctly be made at one plant and bought at another.

Key idea: Compare total cost, not quoted price, and include only avoidable fixed costs. Make-versus-buy usually turns on volume, because it is a fixed-cost-spreading problem.

The strategic half of make versus buy

Cost is necessary and not sufficient. Four strategic questions belong in every outsourcing decision.

Is this a core capability? Outsource what you are not distinctive at, keep what you are. The difficulty is that this is a judgment about the future, and capabilities you outsource are capabilities you stop having.

Are you creating a competitor? The canonical case is the IBM personal computer in 1981. Under intense time pressure IBM outsourced the microprocessor to Intel and the operating system to Microsoft, and negotiated neither exclusivity. The machine succeeded enormously, the architecture was cloned, and most of the value in personal computing accrued for decades to the two suppliers rather than to IBM. The decision was defensible on speed and correct on cost; it gave away the profitable positions.

Are you hollowing out your ability to innovate? Gary Pisano and Willy Shih argued that manufacturing capability and process knowledge form an industrial commons: once the plants, the engineers, and the suppliers leave a region, the design capability that depended on being near them erodes too, and cannot be repurchased quickly.

Can you actually manage the interface? Boeing's 787 program pushed an unusually large share of design and subassembly to partners, expecting lower cost and faster development. The program ran roughly three years late with major cost overruns, and Boeing ultimately bought back at least one struggling supplier and put its own engineers into others. The lesson is not that outsourcing fails; it is that outsourcing design work transfers coordination problems to an interface you control much less well, and the coordination cost is rarely in the business case.

Segmenting what you buy: the Kraljic matrix

Not every purchase deserves the same treatment. Peter Kraljic's 1983 framework sorts purchases on two axes, profit impact and supply risk:

QuadrantProfit impact / supply riskExamplesSourcing approach
Non-criticalLow / lowOffice supplies, fasteners, cleaning servicesMinimize transaction effort: catalogs, purchasing cards, consolidation, two-bin replenishment
LeverageHigh / lowStandard commodities with many suppliers, packaging, basic metalsCompete the business hard, use volume, negotiate price
BottleneckLow / highSmall custom parts with one qualified source, proprietary consumablesSecure supply: contracts, buffer stock, qualify alternatives, redesign the part out
StrategicHigh / highEngines, semiconductors, key subassembliesLong-term partnership, joint development, shared forecasts, supplier development

The matrix earns its place by preventing two symmetrical mistakes: running an intense competitive tender for a bottleneck item, which wins a low price on something you will later be unable to get, and running a cozy partnership on a leverage item, which leaves money on the table for no reason.

Supplier relationships, and the evidence

Two models compete. The transactional approach treats suppliers as interchangeable, competes every purchase, and moves the business to whoever is cheapest. The partnership approach concentrates volume with fewer suppliers, shares forecasts and designs, and invests in improving them.

The evidence in automotive is unusually clear because the industry is surveyed annually. Jeffrey Liker and Thomas Choi documented in 2004 how Toyota and Honda built supplier capability deliberately: understanding suppliers' operations in detail, sending engineers to improve them, sharing the gains from those improvements rather than simply repricing, and maintaining relationships across decades. The contrasting approach is well documented too, most notoriously the aggressive demands for retroactive price cuts at General Motors in the early 1990s, which extracted short-term savings while damaging supplier trust for years afterward. Annual North American supplier relations surveys, run for decades by an independent consultancy, have consistently ranked Toyota and Honda at or near the top and Detroit automakers lower, and have found that suppliers give their best technology and their most flexible capacity to the customers they rate highest. That is a direct competitive consequence of a relationship choice.

The practical process wrapped around all of this is unglamorous: spend analysis to see what you actually buy and from whom, requests for information and quotation, total cost of ownership evaluation rather than price comparison, contracting, and supplier scorecards tracking quality, delivery, cost, and responsiveness on a regular review cycle.

Key idea: Segment purchases before choosing a sourcing approach, and recognize that supplier relationships are an operational asset: firms suppliers rank highly get earlier technology and more flexible capacity.

Common misconceptions

  • "The cheaper quote wins." Inspection, inventory, supplier management, and qualification are costs of buying, and they moved a 9 percent apparent saving into a 10,900 dollar annual loss.
  • "Include full overhead in the make cost." Only avoidable costs belong. Allocated overhead that continues regardless systematically biases decisions toward outsourcing.
  • "Location is a real estate decision." It sets lead time, which sets safety stock, forecast horizon, and responsiveness for as long as the facility exists.
  • "Squeezing suppliers is good procurement." On leverage items competition is appropriate. On strategic and bottleneck items it buys a low price on something you may not be able to obtain when it matters.

Recap

  • Weighted factor rating gave Site B a 1.25-point win that reversed to Site A under plausible reweighting, which is the method's real lesson.
  • Locational break-even put site A below 50,000 units, B between 50,000 and 125,000, and C above 125,000, making the decision a bet on the volume forecast.
  • Center of gravity located a distribution center at (55, 78) but ignores roads, land cost, and inbound freight.
  • Facility count trades the square root inventory saving of consolidation against outbound freight, delivery speed, and concentration risk.
  • Make cost 638,000 against a buy total of 648,900 once inspection, inventory, supplier management, and qualification were counted, with the break-even at about 36,300 units.
  • Strategic risks of outsourcing include creating a competitor, as with IBM in 1981, eroding the industrial commons, and taking on coordination costs, as on the Boeing 787.
  • The Kraljic matrix segments purchases into non-critical, leverage, bottleneck, and strategic, each with a different sourcing approach.

Sources

  1. Association for Supply Chain Management. (2025). SCOR model and supply chain standards. ASCM. ascm.org
  2. Wikipedia. (2025). Kraljic matrix. en.wikipedia.org
  3. Liker, J. K., & Choi, T. Y. (2004). Building deep supplier relationships. Harvard Business Review, 82(12), 104-113. hbr.org
  4. Wikipedia. (2025). Boeing 787 Dreamliner. en.wikipedia.org
  5. OpenStax. (2018). Supply chain management and purchasing. In Introduction to Business. Rice University. openstax.org
Key terms
Supply chain
The network of organizations, activities, information, and resources that moves a product from raw material to customer and back on return.
SCOR model
A reference framework organizing supply chain activity into plan, source, make, deliver, return, and enable.
Weighted factor rating
A location method that assigns weights to factors and scores each site, useful for structuring a decision but sensitive to the weights chosen.
Locational break-even
Comparing sites by fixed and variable cost to find the volumes at which each becomes cheapest.
Center of gravity
A location method placing a facility at the volume-weighted average of demand point coordinates, assuming straight-line distance.
Avoidable fixed cost
Fixed cost that disappears if the activity stops; only these belong in a make-versus-buy comparison.
Total cost of ownership
The full cost of a purchased item including inspection, inventory, supplier management, qualification, and risk, not just the quoted price.
Kraljic matrix
A segmentation of purchases by profit impact and supply risk into non-critical, leverage, bottleneck, and strategic categories.

Logistics, Transportation, and Landed Cost

  • Explain how containerization reshaped global trade and what logistics costs in a modern economy.
  • Compare transportation modes on cost, speed, reliability, and suitability.
  • Compute the total landed cost of an imported shipment and compare it with a near-shore alternative.
  • Evaluate global versus near-shore sourcing including pipeline inventory and lead-time effects.

The big picture

On 26 April 1956 a converted tanker called the Ideal-X left Newark for Houston carrying fifty-eight metal boxes. The man behind it, Malcom McLean, was a trucking operator rather than a shipping man, and his insight was not about ships at all. He had noticed that the expensive part of moving freight was not the ocean voyage; it was the loading. Cargo arrived at a pier in barrels, crates, and sacks, and gangs of longshoremen moved each item by hand into the hold. Ships spent more time in port than at sea.

Marc Levinson's history of the container reports the scale of the change: loading loose cargo cost on the order of 5.86 dollars per ton, while loading the Ideal-X cost about 16 cents per ton. That is not an improvement, it is a different world. Once the container was standardized internationally in the late 1960s so that the same box could move between ship, rail, and truck without being unpacked, the cost of distance collapsed, and with it the assumption that you should manufacture near your customers. Nearly every global supply chain in this course exists because of that box.

This lesson covers what moving things actually costs, how to compute it properly, and how to compare a distant supplier with a near one on numbers rather than instinct.

What logistics is and what it costs

Logistics is the movement and storage of goods and the information about them: transportation, warehousing, materials handling, packaging, order fulfillment, and inventory in transit. It is a subset of supply chain management, which also covers sourcing, production planning, and supplier relationships.

In an advanced economy, business logistics costs run in the neighborhood of 8 percent of gross domestic product, which in the United States means a figure in the trillions of dollars. For an individual product the logistics share varies enormously: a few percent of the selling price of a laptop, and a large fraction of the delivered cost of bottled water or gravel, where the product is cheap and heavy. A useful rule for judging a supply chain decision is the ratio of value to weight and volume. High-value, low-weight goods can afford almost any transport mode. Low-value, high-bulk goods are ruled by freight cost and are made close to where they are used, which is why there is a concrete plant near you and no concrete plant that serves a continent.

The five modes

ModeRelative cost per ton-mileSpeedTypical useNotes
Water (ocean and barge)Lowest, on the order of a few cents or lessSlowestContainers, bulk grain, coal, ore, petroleumUnbeatable on cost per ton-mile; transit measured in weeks, and port congestion can add more
RailLow, roughly a few centsModerateIntermodal containers, bulk commodities, autosEfficient over long distances; requires drayage at both ends
TruckModerate, several times railFast over short and medium distancesAlmost everything, and every last mileDoor-to-door access is its decisive advantage
AirHighest, often an order of magnitude above truckFastestElectronics, pharmaceuticals, spare parts, perishables, fashionRational whenever the value or urgency is high relative to weight
PipelineVery low for what it carriesSlow but continuousOil, gas, refined products, some slurriesEnormous fixed cost, negligible variable cost, no flexibility

Two structural points matter more than the exact numbers, which move with fuel prices and capacity cycles.

First, intermodal shipping combines modes to get most of rail's cost with most of truck's reach: a container travels by rail between terminals and by truck for the short drayage moves at each end. Second, the last mile is disproportionately expensive. Delivering a parcel from a local facility to a house involves a vehicle stopping repeatedly, a driver walking to doors, and failed deliveries, and it commonly accounts for a large share, often cited around a third to a half, of total delivery cost for e-commerce. That is why so much operational innovation, from lockers to pickup points to route optimization, is aimed at those final few miles.

Shipment size matters too. Truckload shipping moves a full trailer point to point. Less-than-truckload consolidates shipments from many customers through terminals, costing more per unit and adding handling and transit time. Parcel handles small packages through a hub network. Choosing the wrong one is a common and expensive error: shipping eight pallets by less-than-truckload when a full truckload would have cost the same is routine in firms that do not measure it.

Key idea: Mode choice follows the ratio of value to weight and the urgency of the need, not a general preference for cheap freight. Air freight is often correct for high-value goods, and the last mile dominates parcel economics.

Total landed cost, worked

Here is the calculation that separates people who source internationally from people who think they do. A company imports 5,000 units of a small appliance. The supplier's quoted price is 22.00 dollars per unit, free on board at the origin port.

Cost elementBasisAmount
Goods, 5,000 at 22.00Supplier invoice110,000.00
Ocean freight, one 40-foot containerCarrier rate3,800.00
Drayage and terminal handlingPort to inland point650.00
Customs duty at 4.2 percentPercent of goods value4,620.00
Additional tariff at 7.5 percentPercent of goods value8,250.00
Harbor maintenance fee at 0.125 percentPercent of goods value137.50
Merchandise processing fee at 0.3464 percentPercent of goods value, subject to a cap381.00
Customs brokerPer entry175.00
Marine insurance at 0.4 percentPercent of goods plus freight455.00
Inland freight, port to distribution centerPer shipment1,900.00
Total landed cost130,368.50

Per unit that is 130,368.50 / 5,000 = 26.07 dollars, which is 18.5 percent above the quoted 22.00. Nothing here was exotic. Every line is ordinary and every line is invisible on the supplier's invoice.

Two cautions before we go on. Duty rates depend on the tariff classification of the specific product and on trade measures that change with policy, so the rates above are illustrative and must be verified against the current tariff schedule for the actual classification. And the merchandise processing fee has minimum and maximum values per entry, which matters for very small and very large shipments.

Now the cost the table still does not show. Ocean transit is about 32 days, and customs plus inland movement adds roughly 10 more, so the lead time is about 42 days. If annual demand is 43,000 units, that is about 118 units a day, and Little's Law gives the pipeline:

  • Pipeline inventory = 118 x 42 = about 4,956 units permanently in transit.
  • Carrying cost at 25 percent of the 26.07 landed value = 4,956 x 26.07 x 0.25 = about 32,300 dollars a year, or 0.75 dollars per unit.

So the true all-in cost is about 26.82 dollars per unit, and that still excludes the safety stock required by a long and variable lead time, which Module 3 showed is driven mostly by lead-time variance.

The near-shore comparison

A supplier in Mexico quotes 24.50 dollars per unit, 11 percent above the Asian quote. Truck freight is 2,100 dollars per shipment of 5,000 units, the goods qualify under the North American trade agreement so duty is zero, a broker charges 120 dollars, and transit is about 5 days with 2 more for customs and delivery.

Cost elementAmount
Goods, 5,000 at 24.50122,500.00
Truck freight2,100.00
Duty0.00
Broker120.00
Total landed cost124,720.00, or 24.94 per unit
  • Pipeline inventory = 118 x 7 = 826 units.
  • Carrying cost = 826 x 24.94 x 0.25 = about 5,150 dollars a year, or 0.12 dollars per unit.
  • All-in cost = 25.06 dollars per unit.

The supplier whose price was 11 percent higher is 1.76 dollars per unit cheaper delivered, which on 43,000 units a year is about 75,700 dollars. And that comparison still understates the near-shore advantage, because it has not counted the smaller safety stock a short reliable lead time permits, the shorter forecast horizon from Module 4, the ability to reorder in weeks rather than months when demand surprises you, and the reduced exposure to port congestion and freight rate spikes.

Be equally honest about the other direction. A different Asian supplier might quote 18 dollars rather than 22, which reverses the arithmetic entirely. The tooling and qualified capacity may exist only in Asia. Near-shore suppliers may lack the scale, the component ecosystem, or the labor. Duty-free treatment requires meeting rules of origin, which can require documentation the supplier cannot produce. The point is not that near-shoring wins; it is that the comparison must be made on total landed and carried cost, and that when firms began doing that arithmetic seriously, some of the offshoring decisions of the previous two decades did not survive it. Trade data reflects the shift: Mexico overtook China as the largest source of United States goods imports in 2023.

Key idea: Compare suppliers on total landed cost plus pipeline carrying cost, not on quoted price. In the worked case an 11 percent higher unit price delivered 1.76 dollars per unit cheaper.

Incoterms: who pays and where risk transfers

International shipment terms are standardized by the International Chamber of Commerce as Incoterms, and misunderstanding them is a reliable source of disputes and unbudgeted cost. The core idea is that each term sets two things: which party pays for each leg, and at what point the risk of loss passes from seller to buyer. Four you should recognize:

  • EXW (ex works): the buyer collects at the seller's door and bears everything from there. Cheapest quoted price, most work and risk for the buyer.
  • FOB (free on board): the seller delivers and clears for export, and risk passes when the goods are loaded on the vessel. Everything after that is the buyer's, which is exactly why the worked example above had ten more lines below the quoted price.
  • CIF (cost, insurance, and freight): the seller pays freight and insurance to the destination port, though risk still passes earlier than most buyers assume.
  • DDP (delivered duty paid): the seller delivers to the buyer's door with duties paid. Highest quoted price and the fewest surprises, which for a small importer is often the right trade.

One practical warning. In domestic commerce, especially in the United States, people use the phrase FOB loosely to mean who pays freight, without the precise Incoterms meaning. In an international contract that looseness produces expensive arguments about who owned the container when it was damaged.

Warehousing, and what it really does

A distribution center performs a defined sequence: receive, put away, store, pick, pack, and ship. Of these, picking dominates the cost, commonly cited at more than half of warehouse labor, because it is the step involving travel. That single fact drives most warehouse design decisions: slotting places fast-moving items near the packing area to cut travel, zone and batch picking reduce trips, and goods-to-person automation inverts the problem by bringing shelves to a stationary worker rather than sending workers to shelves.

Cross-docking skips storage altogether: inbound trucks are unloaded and their contents moved directly to outbound trucks, sometimes within hours. It cuts inventory and handling, and it demands precise timing and reliable inbound suppliers, which is why it became a signature capability of retailers with tight supplier integration rather than a general practice.

An honest note on the human side. Warehousing and transportation employ millions of people, and the sector has recorded injury and illness rates above the private-industry average in federal statistics. Productivity systems that pace work tightly interact badly with the ergonomics of repetitive lifting and reaching. The operations question is not whether to measure, but whether the target is set from what a process can sustainably deliver or from what its fastest hour looked like, which is Module 5's warning about muri, overburden, in a different setting.

The global footprint

Beyond cost, four factors shape where production sits. Tariffs and trade agreements change the arithmetic directly and change with politics. Currency moves the cost of a foreign supplier without anyone deciding anything. Intellectual property risk matters when the process itself is the asset. And lead-time risk, the theme of this whole module, ties up cash and forecast accuracy.

Firms have responded with recognizable patterns. China plus one keeps existing Chinese capacity while qualifying a second country to reduce concentration. Near-shoring moves production closer to the market, trading unit cost for responsiveness. Friend-shoring prefers politically aligned countries. Industrial policy has pushed in the same direction: the CHIPS and Science Act of 2022 provided roughly 52.7 billion dollars for semiconductor research and domestic manufacturing, explicitly to reduce dependence on concentrated overseas capacity for a component that the next lesson shows brought whole industries to a halt.

Common misconceptions

  • "Air freight is always too expensive." For high-value, low-weight, or urgent goods it frequently wins on total cost, because it removes weeks of pipeline inventory and lets you order later on a better forecast.
  • "The quoted unit price is the cost." Freight, duties, fees, insurance, inland movement, and pipeline inventory added 22 percent in the worked example.
  • "Logistics is a commodity you buy on price." Reliability determines safety stock, and Module 3 showed lead-time variance usually dominates the safety stock calculation.
  • "Near-shoring is always cheaper now." Sometimes it is and sometimes it is not. It won by 1.76 dollars a unit here and would have lost against an 18 dollar quote. Do the arithmetic each time.

Recap

  • Containerization collapsed the cost of loading freight from dollars to cents per ton and made distant manufacturing viable.
  • Logistics runs near 8 percent of gross domestic product; mode choice follows value-to-weight and urgency, and the last mile dominates parcel cost.
  • Total landed cost for the imported shipment was 26.07 dollars per unit against a 22.00 quote, 18.5 percent higher.
  • Pipeline inventory of about 4,956 units added roughly 0.75 dollars per unit, bringing the all-in cost to 26.82.
  • The near-shore supplier quoting 11 percent more delivered at 25.06 all-in, about 1.76 dollars per unit cheaper, before counting safety stock and forecast benefits.
  • Incoterms define who pays each leg and where risk transfers; EXW, FOB, CIF, and DDP shift very different amounts of work to the buyer.
  • Picking dominates warehouse cost, which drives slotting, batch picking, and goods-to-person automation.

Sources

  1. Bureau of Transportation Statistics. (2025). Freight transportation and modal statistics. U.S. Department of Transportation. bts.gov
  2. Wikipedia. (2025). Containerization. en.wikipedia.org
  3. U.S. Census Bureau. (2025). Foreign trade: top trading partners. census.gov
  4. International Chamber of Commerce. (2020). Incoterms 2020 rules. ICC. iccwbo.org
  5. U.S. Bureau of Labor Statistics. (2025). Injuries, illnesses, and fatalities in transportation and warehousing. bls.gov
Key terms
Containerization
The standardized intermodal shipping container system that collapsed loading costs and made long-distance manufacturing economical.
Total landed cost
The full delivered cost of goods including price, freight, duties, fees, insurance, broker charges, and inland transport.
Intermodal
Moving the same container across two or more modes, typically rail for the long haul and truck for drayage at each end.
Drayage
Short truck moves between a port or rail terminal and a nearby facility.
Last mile
The final delivery leg to the customer, disproportionately expensive because of stops, walking, and failed deliveries.
Incoterms
Standardized international trade terms specifying which party pays each transport leg and where risk of loss transfers.
Cross-docking
Transferring goods directly from inbound to outbound vehicles with little or no storage, requiring precise timing and reliable suppliers.
Near-shoring
Moving production closer to the market, trading a higher unit price for shorter lead time, lower pipeline inventory, and better responsiveness.

Resilience, Sustainability, and Careers

  • Explain factually what the disruptions of 2020 to 2023 revealed about supply chain fragility.
  • Apply time-to-recover and time-to-survive analysis and price a strategic buffer as insurance.
  • Describe supply chain sustainability, reverse logistics, and labor issues with their honest measurement limits.
  • Map operations and supply chain careers, including what certifications do and do not provide.

The big picture

Between 2020 and 2023 the global supply chain was subjected to a stress test nobody designed and everybody watched. Shelves emptied, factories idled for want of chips that cost a few dollars, and container rates rose severalfold and then collapsed. It was the most instructive period in the history of the field, and it is worth going through carefully, because the popular account of what happened is mostly wrong.

The popular account is that just-in-time failed. The accurate account is more specific and more useful: firms had removed buffers without removing the variability the buffers were absorbing, had concentrated sourcing to capture price, and could not see past their direct suppliers. Those are three separate management choices, all of them defensible in isolation, and all of them priced as though the probability of disruption were zero.

What actually happened

Toilet paper, 2020. The most-photographed shortage was mostly not a production failure. Consumption did not rise much; it moved. People stopped using bathrooms at offices, schools, restaurants, and airports and used bathrooms at home. Commercial tissue and consumer tissue are different products made on different machines with different packaging and sold through entirely different channels. A sudden shift of demand from one channel to the other cannot be met by a system tuned for the old split, and pantry loading by households on top of it emptied shelves. It is a nearly perfect teaching case: the constraint was the channel structure, not the number of paper machines.

Personal protective equipment, 2020. Here the shortage was real. Production of respirators was concentrated in a small number of countries, several of which restricted exports at the moment global demand spiked. The United States Strategic National Stockpile had been drawn down during the 2009 influenza pandemic and not fully replenished, a gap documented in subsequent Government Accountability Office reporting. Hospitals then engaged in exactly the shortage gaming from Module 4: multiple orders across multiple brokers, inflating apparent demand and making allocation harder.

Semiconductors, 2020 to 2022. This is the most instructive sequence. When vehicle sales collapsed in spring 2020, automakers cancelled chip orders. Semiconductor foundries reallocated that capacity to consumer electronics, where demand was surging because everyone was suddenly working and schooling from home. When vehicle demand rebounded far faster than expected, automakers went back to the queue and found it full, with lead times for some components stretching beyond a year. Then physical events compounded it: a fire at a Renesas fab in Japan in March 2021, the February 2021 Texas winter storm that shut down semiconductor plants in Austin, and drought in Taiwan affecting a water-intensive process. Automakers lost millions of units of planned production in 2021.

Notice the causal structure. The initial cause was not a natural disaster. It was a cancellation decision that made sense on a demand forecast, taken by an industry that had no visibility into the queue dynamics of an upstream industry whose capacity takes years and billions of dollars to add. The chips in question often cost a few dollars each and were holding up vehicles worth tens of thousands.

Ports and ocean freight, 2021 to 2022. Goods consumption rose sharply while service consumption fell, and the physical system that handles goods could not expand to match. Container ships waiting off Los Angeles and Long Beach peaked at over one hundred in January 2022. Spot container rates rose to several times their pre-2020 level, then fell back sharply through 2022 and 2023 as demand normalized, which is itself a lesson: the freight market is cyclical, and decisions made at the peak of a cycle often look foolish at the trough.

The Ever Given, March 2021. A single ultra-large container ship wedged across the Suez Canal for six days, blocking a waterway through which roughly a tenth of global trade passes. It is the cleanest illustration in modern memory that a network can have a physical single point of failure that no supplier scorecard would ever surface.

Why the system was fragile

  • Buffers removed without variability removed. Module 5 was explicit: Toyota levels demand before it removes inventory. Firms that cut inventory while facing unlevel, uncertain demand were not doing lean, they were removing the shock absorber and keeping the potholes.
  • Single sourcing for price. Concentrating volume with one supplier earns a discount and creates a single point of failure. The discount appears in this year's numbers; the failure appears in a year nobody budgeted.
  • No visibility past tier one. Most firms know their direct suppliers well and their suppliers' suppliers not at all. In 2011 and again in 2021, many companies discovered that their two carefully qualified alternate suppliers both depended on the same sub-tier plant.
  • Incentives. Inventory reduction is measurable, immediate, and rewarded. Avoided disruption is invisible and unrewarded. That asymmetry, more than any analytical error, is why buffers get cut.

Key idea: The disruptions did not disprove lean. They exposed inventory reduction without variability reduction, sourcing concentration priced as though disruption were free, and an inability to see past direct suppliers.

Resilience as an engineering problem

The productive response is not to hold more of everything. It is to find where exposure actually lies and price the fix.

Map the network. After the 2011 Tohoku earthquake, Toyota built a detailed multi-tier supplier database precisely so it could answer the question of which parts were at risk when a region went down, and it is generally credited with responding faster than peers in later disruptions as a result. Mapping is unglamorous and expensive and it is the prerequisite for everything else.

Time to recover and time to survive. David Simchi-Levi and colleagues developed a method, applied at Ford among others, that avoids the impossible task of estimating disruption probabilities. For each node, ask two answerable questions:

  • Time to recover (TTR): if this site went down, how long until it is restored or replaced?
  • Time to survive (TTS): how long could we keep operating without it, given current buffers and alternatives?

Where TTS is greater than or equal to TTR, you are covered. Where TTR exceeds TTS, you have exposure equal to the gap. A node with a 14-week time to recover and a 6-week time to survive is exposed by 8 weeks, and now you have three concrete options with computable costs: hold 8 more weeks of buffer, qualify a second source, or redesign the part out. The elegance of this framework is that it asks only questions your engineers can actually answer, and it surfaces the surprising nodes, which are usually cheap parts nobody thought about.

Price the buffer as insurance. Work an example. A component costs 3 dollars, the plant uses 400 a week, and a shortage stops a line producing 180,000 dollars a week of contribution margin.

  • Holding 10 extra weeks of this part means 4,000 units. Annual carrying cost at 25 percent = 4,000 x 3 x 0.25 = 3,000 dollars a year.
  • Suppose a 5 percent annual chance of a 6-week outage. Expected annual loss = 0.05 x 6 x 180,000 = 54,000 dollars a year.
  • The buffer costs 3,000 to avoid an expected 54,000. It is not a close call.

Now change one number. Make the component cost 900 dollars instead of 3.

  • Annual carrying cost = 4,000 x 900 x 0.25 = 900,000 dollars a year against the same 54,000 expected loss.
  • Buffering is now absurd, and the right answers are a second qualified source, a capacity reservation contract, or a design change that removes the dependency.

That contrast is the whole method. Resilience is an insurance-pricing problem, not a slogan. Cheap, small, catastrophic-to-lack parts should be buffered generously, and it is remarkable how often they are not, because a blanket inventory reduction target treats a 3 dollar clip and a 900 dollar module identically.

Flexibility beats dedicated backup. One genuinely elegant result deserves its own paragraph. William Jordan and Stephen Graves showed in the 1990s that a plant network does not need full flexibility to get most of full flexibility's benefit. If each plant can build two products and the assignments are arranged so that all plants and products form a single connected chain, the network captures nearly all the benefit of every plant building everything, at a small fraction of the cost. The principle generalizes: partial, well-connected flexibility is dramatically cheaper than universal flexibility and almost as good. Cross-training staff in a chain rather than training everyone on everything is the same idea in a service setting.

Key idea: Compare time to recover with time to survive to find real exposure, then choose among buffer, second source, and redesign by computing which is cheapest for that specific part.

Sustainability and reverse logistics

Emissions accounting splits into three scopes: Scope 1 is direct emissions from what you own, Scope 2 is purchased energy, and Scope 3 is everything else in the value chain, upstream and downstream. For most consumer-facing companies, Scope 3 is the overwhelming majority of the footprint, frequently cited in the range of 70 to 90 percent. Which means that for most firms, the environmental question is a supply chain question, and it is also the hardest to measure, because it depends on data from companies you do not control. Treat precise Scope 3 numbers with the same skepticism you would apply to any self-reported figure computed from estimated inputs.

Transport is the part operations controls most directly. Per ton-kilometer, ocean shipping is the lowest-emitting mode by a wide margin, rail is next, trucking is substantially higher, and air freight is higher again, commonly by more than an order of magnitude over ocean. That ordering has an immediate implication: expediting by air is not only expensive but carries a large emissions penalty, so the operational discipline that reduces expediting, better planning, shorter lead times, fewer stockouts, is also an emissions program. Packaging and right-sizing matter for the same reason, since shipping air inside boxes costs both money and fuel.

Reverse logistics is the flow backward: returns, repairs, recalls, recycling, and disposal. It is a large operation in its own right. Retail returns in the United States have run in the mid-teens as a percentage of sales in recent years, with online rates markedly higher than in-store, and each returned item has to be received, inspected, graded, and routed to restock, refurbishment, liquidation, recycling, or landfill. A significant share never returns to primary sale. Designing the returns process well matters both financially and environmentally, and it is chronically underinvested because it generates no revenue line.

The circular economy pushes further: design for disassembly, remanufacturing programs that rebuild used cores to original specification, and extended producer responsibility rules that make manufacturers accountable for end-of-life. Remanufacturing is the most operationally interesting because it is genuinely profitable in the right settings, notably heavy equipment and imaging hardware, where a rebuilt core sells at a substantial discount with comparable warranty and much lower material input.

Labor conditions in supply chains are an operations issue and not only a compliance one. The 2013 Rana Plaza building collapse in Bangladesh killed 1,134 garment workers and prompted binding safety agreements among brands. Social audits, the dominant tool for a generation, have well-documented limitations: they are announced, they are periodic, and they measure documentation as much as conditions. More recent regulatory approaches shift the burden, including the United States Uyghur Forced Labor Prevention Act of 2022, which presumes goods with inputs from a specific region are made with forced labor unless the importer proves otherwise. The practical consequence for an operations manager is concrete: you now need traceability deep into your supply base, which is the same capability resilience mapping requires.

Careers in operations and supply chain

This field hires steadily and hires people who can compute. A partial map of roles:

RoleWhat you actually do
Production or operations supervisorRun a shift: staffing, output, quality, safety, problem solving on the floor
Planner or schedulerConvert demand into a master schedule and material plan; live inside the MRP system
Demand plannerOwn the forecast, the error measurement, and the sales and operations planning input
Buyer, sourcing analyst, category managerSupplier selection, total cost analysis, negotiation, supplier performance
Inventory or supply analystSafety stock, reorder points, service levels, ABC policy, excess and obsolete
Logistics or transportation analystMode and carrier selection, freight cost, network and route analysis
Distribution center operations managerReceiving through shipping, labor planning, slotting, safety
Continuous improvement or industrial engineerProcess analysis, capacity, layout, kaizen events, standard work

On pay and outlook, the Bureau of Labor Statistics Occupational Outlook Handbook is the authoritative free source, and you should check the current figures rather than trusting any textbook. As of its recent editions, logisticians had a median annual wage of roughly 79,000 dollars with projected growth much faster than the average for all occupations; industrial production managers were near 117,000 dollars; purchasing managers near 136,000; and transportation, storage, and distribution managers near 99,000. Entry-level analyst roles sit well below those medians, which are for established practitioners.

Certifications, named honestly. The main credentials in this field come from the Association for Supply Chain Management, formerly APICS:

  • CPIM (Certified in Planning and Inventory Management): the deepest treatment of the material in Modules 3 and 4. Two exams, several hundred dollars each plus study materials, typically three to six months of preparation. It is well recognized in manufacturing planning roles.
  • CSCP (Certified Supply Chain Professional): broader and end-to-end, oriented toward Module 6 topics. One exam, similar cost and effort.
  • CLTD (Certified in Logistics, Transportation and Distribution): focused on the material in the previous lesson.
  • CPSM from the Institute for Supply Management, for sourcing and procurement roles.
  • Six Sigma belts: note carefully that the belt market is unregulated. An American Society for Quality Certified Six Sigma Black Belt is an examined credential with an experience requirement; many other belts are a weekend course and a certificate file. Employers vary widely in whether they know the difference.

The honest assessment: these certifications teach a real shared vocabulary and are genuinely useful for getting past a screening filter, especially if your degree is in something else. In some sectors, particularly manufacturing planning, hiring managers actively look for CPIM. In others they are indifferent. None of them substitutes for having actually run a schedule, walked a floor, or owned a number, and no certification will make up for being unable to build the models in this course in a spreadsheet.

What actually gets people hired into this field: fluency in spreadsheets beyond the basics, some SQL, the ability to build and explain the calculations in this course, and one concrete improvement project you can describe with before-and-after numbers. If you have access to any operation at all, a campus dining hall, a small business, a volunteer organization, run a real analysis on it and write it up. That artifact is worth more in an interview than any credential.

What this course did and did not do

Look back at what you can now compute. You can find a bottleneck and size a capacity change. You can apply Little's Law in three directions. You can explain why a queue forms at 80 percent utilization and quantify what a five percent capacity increase buys. You can build a control chart, judge capability, and price the cost of quality. You can compute an order quantity, size safety stock from two sources of variance, and solve a newsvendor problem. You can forecast, measure your error, detect bias, plan a workforce, and run an MRP explosion. You can explain the Toyota Production System accurately, including the part that usually gets dropped, find and exploit a constraint, and compute total landed cost.

What a text course cannot give you is everything that happens at the actual place. Standing on a floor at shift change. Watching where people walk and what they work around. Noticing the pallet of parts that has been in the corner for a year because moving it is someone else's job. Sitting in the meeting where the forecast everyone knows is wrong gets approved anyway, and understanding why. Toyota's insistence on genchi genbutsu, go and see, exists because processes on paper and processes in reality differ in ways that only physical presence reveals. Take the arithmetic in this course to a real operation as soon as you can, and expect the operation to teach you something the arithmetic did not.

Common misconceptions

  • "Resilience means holding more inventory." Inventory is one of several tools, and often the wrong one. For expensive components, a second source or a design change costs far less than a buffer.
  • "The 2020 to 2023 disruptions proved just-in-time was a mistake." They exposed inventory cuts made without variability reduction, sourcing concentrated for price, and no visibility past tier one. Toyota, the source of just-in-time, weathered the chip shortage comparatively well precisely because it had mapped its network and buffered critical parts deliberately.
  • "Sustainability always costs money." Reducing expediting, right-sizing packaging, and cutting returns all reduce cost and emissions together. Other measures genuinely do cost money, and conflating the two categories is how sustainability programs lose credibility.
  • "You need a certification to work in supply chain." You need to be able to do the arithmetic and explain it. Certifications help you get read; they do not substitute for the skill.

Recap

  • The 2020 to 2023 disruptions had specific documented causes: channel shifts, concentrated production, a cancellation-then-rebound sequence in semiconductors, port capacity limits, and one ship in a canal.
  • Fragility came from removing buffers without removing variability, single sourcing for price, no visibility past tier one, and incentives that reward inventory cuts and ignore avoided disruption.
  • Time to recover versus time to survive identifies exposure without requiring disruption probabilities; a 14-week recovery against a 6-week survival is 8 weeks of exposure.
  • A 3 dollar part justified a 3,000 dollar annual buffer against a 54,000 dollar expected loss; the same buffer on a 900 dollar part would cost 900,000 and calls for a second source or a redesign instead.
  • Partial, well-connected flexibility captures most of the benefit of full flexibility at a fraction of the cost.
  • Scope 3 emissions dominate most consumer-facing footprints, air freight carries a large emissions penalty, and reverse logistics handles returns running in the mid-teens as a share of retail sales.
  • Careers span planning, sourcing, inventory, logistics, and improvement, with ASCM credentials useful as a shared vocabulary and a screening aid rather than a substitute for skill.

Sources

  1. U.S. Government Accountability Office. (2022). COVID-19: Federal efforts and supply chain challenges. GAO. gao.gov
  2. U.S. Department of Commerce. (2022). Semiconductor supply chain assessment. commerce.gov
  3. Wikipedia. (2025). 2021 Suez Canal obstruction. en.wikipedia.org
  4. U.S. Bureau of Labor Statistics. (2025). Logisticians. In Occupational Outlook Handbook. bls.gov
  5. Association for Supply Chain Management. (2025). Certifications: CPIM, CSCP, and CLTD. ASCM. ascm.org
  6. U.S. Environmental Protection Agency. (2025). Scope 3 inventory guidance and supply chain emissions. EPA. epa.gov
Key terms
Time to recover
How long a disrupted node would take to be restored or replaced, one half of the exposure comparison.
Time to survive
How long an operation could continue without a given node, given current buffers and alternatives.
Strategic buffer
Inventory deliberately held against disruption, justified by comparing its carrying cost with the expected cost of the outage it prevents.
Process flexibility chaining
Arranging partial flexibility so plants and products form a connected chain, capturing most of full flexibility's benefit at a fraction of the cost.
Scope 3 emissions
Value chain emissions outside a firm's own operations and purchased energy, typically the large majority of a consumer company's footprint.
Reverse logistics
The backward flow of returns, repairs, recalls, recycling, and disposal, including receiving, grading, and routing returned goods.
Remanufacturing
Rebuilding used cores to original specification for resale, profitable in heavy equipment and imaging hardware.
CPIM
The ASCM Certified in Planning and Inventory Management credential, covering planning and inventory material through two examinations.

Open the interactive version with quizzes and progress →