➗ Mathematics · Undergraduate · MATH 151

Calculus I: Limits & Derivatives

A complete first course in single-variable calculus built around one big idea: the derivative, the instantaneous rate of change. You will start from functions and limits, build the derivative from its definition, master every core differentiation rule, and finish by applying derivatives to related rates and optimization before meeting the integral through antiderivatives. Every concept is taught…

Start the interactive course (quizzes, progress, videos) →

Free forever. No sign-up, no ads. 17 lessons. The full lesson text is below so you can read it right here.

Module 1: Functions and the Idea of a Limit

Review the functions calculus is built on, from their domains to composition, then meet the central concept of a limit that every later idea depends on. This module lays the vocabulary and intuition you will use for the rest of the course.

Functions and Precalculus Review

  • Describe a function using the vertical line test and interval notation for domain and range.
  • Identify the main function families and their basic shapes.
  • Evaluate and combine functions, including composition.

Press B4 on a vending machine and one specific snack drops out. Press B4 again tomorrow and the same snack drops out. That reliability is the whole idea we are about to make precise, and you have been using it for years without calling it anything. Calculus studies how quantities change, so before we touch anything new we will get comfortable with the machines that produce those quantities, one small step at a time. No question here is too small.

One input, one output

Feed 3 into the rule f(x) = 2x + 1 and 7 comes out. Feed 3 in again and 7 comes out again, today and next year. A function is any rule with that property: take an input, give back exactly one output. It matters because almost everything calculus studies, a speed, a cost, a height, is one number going in and one answer coming out.

Key idea: one input, one and only one output. That word "only" is what makes it a function.

Reading the notation out loud

We give a function a short name, usually the letter f, and we write f(x). Read that out loud as "f of x." It does not mean f times x. It means "the output of machine f when you feed in x." The little (x) is the input slot, not a multiplication. Read that once more if you like, because this one habit prevents a lot of confusion later.

Let us evaluate one together, slowly

Find f(3) when f(x) = 2x + 1.

  1. Read the rule: 2x + 1 means "two times the input, then add one." (Just translating.)
  2. The input is 3, so wherever we see x we gently put a 3: 2(3) + 1. (Filling the slot.)
  3. Do the multiply first: 2 × 3 = 6. Now we have 6 + 1. (Multiplying comes before adding.)
  4. Do the add: 6 + 1 = 7. (Last step.)

So f(3) = 7. That is the whole skill: take a rule, put in a number, read off one answer. Nice work, that is exactly what you will do with harder functions all course long.

Try it: with the same rule, what is f(5)?

Answer: Put 5 in the slot: 2(5) + 1 = 10 + 1 = 11. So f(5) = 11. See, you just did it again.

The vertical line test (a picture rule)

How can you look at a graph and tell if it is a function? Use the vertical line test: if any straight up-and-down line crosses the graph more than once, the graph is not a function. Here is why, in plain words: a vertical line stands over one input. If the graph is there twice, that one input has two outputs, which breaks the one-output rule. A U-shaped parabola passes the test. A full circle fails it, because a vertical line through the middle hits both the top and the bottom.

The families of functions you will keep meeting

Almost every problem this course throws at you is built from a small cast of characters. You do not have to master them today, just recognize their faces:

  • Polynomials like 3x2 − 5x + 1: smooth curves with no breaks or holes, defined for every number.
  • Rational functions like 1/x: a polynomial over a polynomial, undefined wherever the bottom is zero.
  • Root functions like √x (read "the square root of x"): only defined where the inside is not negative.
  • Trig functions like sin(x) and cos(x): wavy curves that repeat over and over.
  • Exponential and log functions like ex and ln(x): the language of growth and its reverse.

Interval notation: a shorthand for "which numbers are allowed"

We often need to say "all the numbers from here to there." Instead of a sentence, we use brackets and parentheses. A square bracket includes the endpoint; a parenthesis leaves it out. So [0, ∞) means "zero is included, and go up forever." Infinity always gets a parenthesis, because it is a direction, not a number you can land on. The allowed inputs of √x are [0, ∞). Do not worry about memorizing this; you will pick it up by seeing it a few times.

Finding the domain: which inputs are allowed

The domain is the set of legal inputs. When a function is just a formula, only two things can make an input illegal:

  • You cannot divide by zero. Any input that makes the bottom of a fraction zero is banned.
  • You cannot take the square root of a negative number (and stay in real numbers). Whatever sits under a square root must be zero or bigger.

Worked example. Find the domain of f(x) = 1 / (x − 4).

  1. Look only at the bottom: x − 4. (Fractions blow up when the bottom is zero.)
  2. Ask when it equals zero: x − 4 = 0 at x = 4. (Solve the little equation.)
  3. So 4 is the one banned input; everything else is fine.

Domain: all real numbers except 4. Read that as "any number you like, just not 4." The set of outputs a function actually produces is called its range, the up-and-down story to the domain's side-to-side story.

Try it: find the domain of g(x) = √(x − 2).

Answer: The inside must be zero or larger, so x − 2 ≥ 0 (the sign ≥ is read "greater than or equal to"). Add 2 to both sides: x ≥ 2. Domain: all numbers 2 or bigger.

Combining functions, and the one that matters most

You can add, subtract, multiply, and divide functions. The combination calculus cares about most is composition: feeding one function into another, written f(g(x)) and read "f of g of x." Think of two machines in a row, the output of the first becoming the input of the second. You always work from the inside out.

Worked example. Let f(x) = x2 + 1 and g(x) = 3x − 2. Find f(g(4)).

  1. Do the inside machine first: g(4) = 3(4) − 2. (Inside out.)
  2. Simplify: 3(4) = 12, then 12 − 2 = 10. So g(4) = 10.
  3. Feed that 10 into the outside machine: f(10) = 102 + 1.
  4. Simplify: 102 = 100, then 100 + 1 = 101.

So f(g(4)) = 101. Order matters here. If you reverse it, g(f(4)) = g(17) = 3(17) − 2 = 49, a completely different number. Spotting a function as "an outer wrapped around an inner" is exactly the skill you will use for the chain rule later, so this small practice pays off big.

Even and odd: a kind of mirror symmetry

A function is even if f(−x) = f(x), meaning its graph is a mirror image across the y-axis, like x2. It is odd if f(−x) = −f(x), meaning a spin-around symmetry through the center, like x3. Most functions are neither, and that is completely fine. These labels just name a nice symmetry when it happens.

Piecewise functions: different rules on different stretches

Some functions use one rule for part of the number line and another rule elsewhere. These are piecewise functions, and calculus leans on them constantly, so let us practice reading one. Define

f(x) = x + 1 when x < 2, and f(x) = x2 when x ≥ 2.

To evaluate, first ask which stretch the input lives on, then use only that rule.

  1. f(0): since 0 < 2, use the first rule: 0 + 1 = 1. (Check the condition before touching the formula.)
  2. f(2): since 2 ≥ 2, use the second rule: 22 = 4. (The boundary input belongs to the rule whose condition includes it.)
  3. f(3): since 3 ≥ 2, again the second rule: 32 = 9.

The most famous piecewise function is the absolute value, |x|, which equals x when x ≥ 0 and −x when x < 0. Its V-shaped graph has a sharp corner at 0, and that corner will star in a later lesson about where derivatives fail to exist.

A domain with two hazards at once

Find the domain of f(x) = √(x − 2) / (x − 5).

  1. The square root demands x − 2 ≥ 0, so x ≥ 2. (Hazard one: no negative insides.)
  2. The fraction demands x − 5 ≠ 0, so x ≠ 5. (Hazard two: no zero bottoms.)
  3. Both conditions must hold at once, so the domain is [2, 5) together with (5, ∞). (Start at 2, skip over 5, continue forever.)

Reading a formula for its hazards, then intersecting the conditions, is the complete domain method. It never gets harder than this; there are just more hazards to list.

Shifting and stretching a known graph

Once you know a parent shape like x2 or √x, you can read transformed versions at a glance: f(x) + 3 slides the graph up 3, f(x − 2) slides it right 2 (the sign is the famous surprise: subtracting inside moves right), 2f(x) stretches it vertically, and −f(x) flips it upside down. So g(x) = (x − 2)2 + 3 is the ordinary parabola moved right 2 and up 3, with its bottom now at the point (2, 3). Being able to picture a function without plotting points will pay off every time this course asks you to sanity-check a computed answer against a mental graph.

Two habits that prevent most errors

The most common stumble is reading f(x) as "f times x" and trying to multiply. It is not multiplication. Whenever you see f(3), say to yourself "put 3 into machine f," and you will not mix it up. The second common stumble is declaring a domain "all real numbers" without checking. Before you answer, always scan for a fraction bottom that could be zero and a square root that could go negative. Those are the only two troublemakers at this stage.

Common misconceptions

  • "f(x) means f times x." No. It is the output of f at input x.
  • "Any graph you can draw is a function." A sideways parabola or a full circle can be drawn but fails the vertical line test, because one input has two outputs.
  • "Composition is just multiplication, so f(g(x)) = g(f(x))." Composition is feeding one output into the next machine, and it is almost never reversible. Above, f(g(4)) = 101 but g(f(4)) = 49.
  • "Every formula allows all real numbers." Fractions ban a zero bottom; square roots ban a negative inside.

What to carry forward

A function gives each input exactly one output. Read f(x) as "f of x," and to evaluate it, drop the number into the slot and simplify one small step at a time. The domain is the legal inputs (never divide by zero, never square-root a negative), and the range is the outputs that actually come out. The vertical line test spots function graphs, interval notation like [0, ∞) writes sets of numbers compactly, and composition f(g(x)) runs two machines in a row from the inside out. You now hold the vocabulary the rest of calculus is built on. That was the hard part of getting started.

Sources

  1. OpenStax. (2016). 1.1 Review of functions. In Calculus volume 1. openstax.org
  2. OpenStax. (2016). 1.2 Basic classes of functions. In Calculus volume 1. openstax.org
  3. Dawkins, P. (n.d.). Calculus I course notes. Paul's Online Math Notes. tutorial.math.lamar.edu
  4. Math is Fun. (n.d.). What is a function? mathsisfun.com
  5. Khan Academy. (n.d.). Calculus 1 [Online course]. khanacademy.org
  6. MIT OpenCourseWare. (2010). 18.01SC Single variable calculus, Fall 2010. Massachusetts Institute of Technology. ocw.mit.edu
  7. Sanderson, G. (3Blue1Brown). (n.d.). The essence of calculus. 3blue1brown.com
Key terms
Function
A rule that assigns exactly one output to each input.
Domain
The set of all allowed inputs of a function.
Range
The set of all outputs the function actually produces.
Vertical line test
A graph is a function if no vertical line crosses it more than once.
Composition
Applying one function to the result of another, written f(g(x)).
Interval notation
A way to write sets of numbers using brackets and parentheses.
Even function
A function with f(-x) = f(x), symmetric about the y-axis.
Odd function
A function with f(-x) = -f(x), symmetric about the origin.

The Idea of a Limit

  • Explain intuitively what it means for a function to approach a limit.
  • Estimate a limit from a table or graph.
  • Distinguish the limit of a function from its value at a point.

Enter (x2 − 9)/(x − 3) into a calculator at x = 3. Both the numerator and denominator become zero, so the result is 0/0, which is undefined.

Now use an input close to 3. At 2.999 the output is 5.999; at 3.001 it is 6.001. Those outputs suggest a value near 6. A limit lets us describe this nearby behavior even though the function has no value at 3.

Near a point, not at it

Think of x as the input to f(x) and a as the number it approaches. A limit is a number L that the outputs can be made as close to as we wish by keeping x sufficiently close to a, without requiring x to equal a.

We write lim (x → a) f(x) = L. Read this as "the limit, as x approaches a, of f of x, equals L." For the two-sided limits in this lesson, we check inputs on both sides of a.

The distinction is between near a and at a. A limit concerns nearby outputs. The function's value at a may be missing, may equal L, or may be different.

The point: a limit describes nearby outputs. It does not require the function to have a value at the target input.

A picture: think of zooming in

Look at the graph near one input and follow the curve from each side. Does it approach the same height from the left and the right? That height suggests the limit. An open circle can mark a missing point without changing the surrounding curve.

A graph helps you see the idea. Its drawing scale can also hide details, so we will use algebra to check our conclusions.

Why the value at the point does not matter

For f(x) = (x2 − 9)/(x − 3), substituting x = 3 gives 0/0. The fraction is undefined there. That answers the question about the value at 3, but it does not answer the question about the limit.

To investigate the limit, compare outputs at inputs just below and just above 3:

x2.92.992.9993.0013.013.1
f(x)5.95.995.9996.0016.016.1

The outputs in the table suggest 6, so our estimate is lim (x → 3) f(x) = 6. This is possible even though f(3) does not exist.

We can check the estimate by factoring: (x2 − 9)/(x − 3) = (x − 3)(x + 3)/(x − 3) = x + 3 for every x except 3.

Near 3, the original fraction therefore has exactly the same outputs as x + 3. Those outputs approach 6. The algebra confirms the limit without assigning a value to the original fraction at 3.

Graph of (x squared minus 9) over (x minus 3) behaving like the line y = x + 3 with an open hole at (3, 6) hole at (3, 6) x y

Let us estimate one together, slowly

Estimate lim (x → 2) (x2 − 4)/(x − 2) from a table.

  1. Try a value just below 2, say x = 1.99: the output comes out about 3.99. (Sneaking up from the left.)
  2. Try a value just above 2, say x = 2.01: the output is about 4.01. (Sneaking up from the right.)
  3. The sampled outputs from both sides suggest the same number, 4. The algebra below will check whether that estimate is exact.

Now check the estimate algebraically. We have (x2 − 4)/(x − 2) = x + 2 whenever x is not 2. As x approaches 2, the simpler expression approaches 2 + 2 = 4. The limit is therefore exactly 4.

One side at a time

A left-hand limit uses inputs less than a. Its notation is x → a−, read "x approaches a from the left." A right-hand limit uses inputs greater than a, written x → a+.

For a finite two-sided limit to exist, both one-sided limits must exist and equal the same number. Check them separately before combining the results.

For example, a graph may approach height 2 on the left of a step and height 5 on the right. Neither side is ambiguous, but the two sides disagree. There is no single two-sided limit at that input.

Three ways a limit can fail to exist

  • A jump: the left and right limits disagree, common in piecewise functions (the staircase step).
  • A blow-up: the outputs grow without bound, as 1/x2 does near 0, so no finite number is approached.
  • Wild wobble: sin(1/x) near 0 swings between −1 and 1 faster and faster, never settling on one value.

These are three common ways a finite limit can fail. When outputs grow without bound, we may describe an infinite limit; that still does not give a finite real-number limit. A table or graph can suggest one of these patterns, while an argument about the function establishes it.

Why limits exist at all: the speed problem

Suppose a model gives the distance a rock has fallen after t seconds as s(t) = 16t2 feet. We want its speed at t = 1.

Average speed uses distance traveled during an interval divided by the interval's duration. Setting that duration to zero would produce 0/0, so it cannot define speed at an instant directly.

Instead, calculate average speeds over shorter and shorter intervals beginning at t = 1:

window1 to 1.11 to 1.011 to 1.001
average speed (ft/s)33.632.1632.016

Check the first interval. The distance fallen during it is 16(1.1)2 − 16(1)2 = 19.36 − 16 = 3.36 feet. Its duration is 0.1 seconds. Dividing gives an average speed of 3.36/0.1 = 33.6 feet per second.

The shorter intervals give averages closer to 32 ft/s. In this model, their limit defines the speed at t = 1. This is the same kind of limiting calculation that will define a derivative.

The honest definition: epsilon and delta

What does "close to L" mean precisely? We need to specify an allowed error, called a tolerance.

Use ε, pronounced "epsilon," for the allowed error in the output. Use δ, pronounced "delta," for how close the input must stay to a. The two symbols track two different distances.

The statement lim (x → a) f(x) = L has a precise requirement. For every output tolerance ε > 0, there must be an input distance δ > 0 with this guarantee:

Whenever an allowed input x is less than δ away from a, but is not equal to a, the output f(x) is less than ε away from L. For the two-sided setting here, assume the function is defined on both sides sufficiently near a, except possibly at a itself.

Suppose you require the outputs to be within 0.1 of L. I must choose an input window around a that guarantees this for every allowed input in that window, apart from the center.

One successful window is not enough to prove a limit. I must be able to choose a suitable window for any positive tolerance, however small. The condition excludes x = a, so the value at the center is irrelevant to this guarantee.

One full proof, for a linear function

We will prove lim (x → 3) (2x + 1) = 7. First work out how input error affects output error. Then use that relationship to choose delta.

  1. Let a challenge ε > 0 be given. We must produce a δ. (The skeptic moves first.)
  2. Look at the output error: |(2x + 1) − 7| = |2x − 6| = 2|x − 3|. (Simplify the distance between output and target.)
  3. So the output error is exactly twice the input error. To force 2|x − 3| < ε, it is enough to force |x − 3| < ε/2. (Work backward from the goal.)
  4. Choose δ = ε/2. Then whenever 0 < |x − 3| < δ, we get |(2x + 1) − 7| = 2|x − 3| < 2(ε/2) = ε. (The guarantee, verified.)

For example, the tolerance ε = 0.1 leads us to choose δ = 0.05. An input strictly less than 0.05 away from 3 produces an output strictly less than 0.1 away from 7.

At the boundary x = 3.05, the output is 7.1. That is exactly 0.1 away from 7, so this boundary is excluded by the strict inequalities in the proof.

For a line with nonzero slope m, the same argument gives δ = ε/|m|. A steeper line needs a narrower input window for the same output tolerance. If the slope is zero, the output is constant, so any positive input window works.

More complicated functions may take more algebra, but the task remains the same: choose an input window and prove that every allowed input in it meets the requested output tolerance.

A limit from two formulas

A piecewise function uses different formulas on different parts of its domain. Consider f(x) = x + 2 for x < 1, and f(x) = 5 − 2x for x > 1, with f(1) left undefined. Does lim (x → 1) f(x) exist?

Start from the left. Use the first formula because those inputs are below 1. Its outputs approach 1 + 2 = 3.

Then approach from the right using the second formula. Its outputs approach 5 − 2 = 3. Both results are 3, so the two-sided limit exists and equals 3.

Now imagine changing the second formula to 4 − 2x. Its right-hand limit would be 2, while the left-hand limit would still be 3. The disagreement would make the two-sided limit fail. Always choose the formula for the side you are approaching from.

Plugging in too soon, and tables that stop too far out

Substitution works when the function is continuous at the target: its nearby outputs approach its value there. If substitution produces 0/0, the original expression is undefined at that input, and more work is needed to find its limit.

When using a table, try inputs on both sides and progressively closer to the target, such as 2.9, 2.99, and 2.999. Treat the pattern as evidence for an estimate. A finite table cannot prove that the function continues behaving that way between or beyond the sampled inputs.

Common misconceptions

  • "The limit is just f(a)." This equality holds at a point of continuity. A limit describes nearby behavior even when f(a) is missing or different, as with the hole at (3, 6).
  • "If f is undefined at a, the limit cannot exist." A hole does not kill a limit. (x2 − 9)/(x − 3) is undefined at 3 yet has limit 6.
  • "A limit is the biggest value the function reaches." No. It is the single value the outputs head toward near a point.
  • "Left-hand and right-hand limits are always equal." At a jump they differ, and then the two-sided limit does not exist.

The short version

A limit describes the outputs near a target input. Read lim (x → a) f(x) = L as "as x approaches a, f of x approaches L." The value at a does not decide the limit. This is why a function can have a limit at a hole in its graph.

For a finite two-sided limit, check that both one-sided limits exist and agree. A jump, unbounded growth, or persistent oscillation can prevent that.

Use a table or graph to form an estimate. Use algebra or the epsilon-delta definition to justify the result. Keep those two jobs distinct as you work through the examples.

Sources

  1. OpenStax. (2016). 2.2 The limit of a function. In Calculus volume 1. openstax.org
  2. OpenStax. (2016). 2.5 The precise definition of a limit. In Calculus volume 1. openstax.org
  3. Dawkins, P. (n.d.). The limit. Paul's Online Math Notes. tutorial.math.lamar.edu
  4. Dawkins, P. (n.d.). The definition of the limit. Paul's Online Math Notes. tutorial.math.lamar.edu
  5. Sanderson, G. (3Blue1Brown). (n.d.). Limits, L'Hopital's rule, and epsilon delta definitions. 3blue1brown.com
  6. Math is Fun. (n.d.). Limits (an introduction). mathsisfun.com
  7. Tall, D., & Vinner, S. (1981). Concept image and concept definition in mathematics with particular reference to limits and continuity. Educational Studies in Mathematics, 12(2), 151-169. doi.org/10.1007/BF00305619
  8. Grabiner, J. V. (1983). Who gave you the epsilon? Cauchy and the origins of rigorous calculus. The American Mathematical Monthly, 90(3), 185-194. doi.org/10.2307/2975545
Key terms
Limit
The value a function approaches as the input approaches a given number.
Left-hand limit
The value approached using inputs slightly less than the target, written x to a^-.
Right-hand limit
The value approached using inputs slightly greater than the target, written x to a^+.
Removable discontinuity
A hole in a graph where the limit exists but the function value does not match.
Does not exist (DNE)
Failure to approach one finite value, for example because one-sided limits disagree, outputs grow without bound, or oscillation does not settle.
Oscillation
Repeated variation in outputs. It prevents a limit when the outputs do not settle toward one value, as with sin(1/x) near 0.

Module 2: Computing Limits and Continuity

Turn the intuition of a limit into reliable algebraic techniques, from direct substitution through factoring, rationalizing, and the Squeeze Theorem. Then define continuity precisely and meet the Intermediate Value Theorem it powers.

Computing Limits Algebraically

  • Apply the limit laws and direct substitution.
  • Resolve 0/0 indeterminate forms by factoring or rationalizing.
  • Use the Squeeze Theorem and the special trig limit sin(x)/x.

The last lesson's table gave outputs of 5.999 and 6.001 near the input 3. Those numbers suggest a limit of 6, but they do not prove that it is exactly 6.

Algebra can settle the question. We will factor the numerator, cancel a shared nonzero factor, and then evaluate a simpler limit. The key is to explain why each new expression has the same nearby values as the original.

What 0/0 is telling you

Start by substituting the target input. If the function is continuous there, this gives the limit. If substitution gives 0/0, both parts of the fraction have become zero.

The fraction is undefined at that input. Its limit is still an open question. Factoring or working with a square-root expression may reveal the nearby behavior.

In short: a 0/0 form does not determine the limit. Find an equivalent expression for nearby inputs, then evaluate that limit.

Move 0: direct substitution (try the easy thing first)

A function is continuous at a point when its limit there equals its value. If f is continuous at the target, then lim (x → a) f(x) = f(a): you can substitute directly.

Polynomials are continuous everywhere. Rational functions are continuous where their denominators are nonzero. The familiar root, trigonometric, exponential, and logarithmic functions are continuous on their domains. For a two-sided limit, also check that the domain supplies nearby inputs on both sides of the target.

  1. Identify the function in lim (x → 2) (3x2 − x). It is a polynomial, so it is continuous at 2 and direct substitution is valid.
  2. Replace every x with 2: 3(22) − 2. Square the 2 before multiplying by 3.
  3. Simplify: 3(4) = 12, then 12 − 2 = 10.

The limit is 10. Substitution is a useful first check, but whether it gives an answer depends on the function and the target input.

The limit laws (why substitution is allowed)

The limit laws tell us how to combine limits. Assume the individual limits exist as finite numbers at the same target input.

  • Addition and subtraction: combine the individual limits using the same operation.
  • Multiplication: multiply the limits. A constant multiplier can be kept outside the limit.
  • Division: divide the limits only if the denominator's limit is nonzero.
  • Powers and roots: take a positive integer power of the limit, or take its root when the root is defined. For an even root, the function values used must be nonnegative, and so must their limit.

A polynomial is built from powers, constant multiples, and sums. Applying these laws to its pieces explains why direct substitution works.

The tricky case: 0/0, and Move 1, factor and cancel

A quotient has the indeterminate form 0/0 when its numerator and denominator both approach zero. "Indeterminate" means that this information alone does not determine the limit.

When the numerator and denominator are polynomials, try factoring them. Look for a common factor that you can cancel at nearby inputs where it is nonzero.

Worked example. Find lim (x → 3) (x2 − 9)/(x − 3).

  1. Try substitution: you get 0/0. (So we switch to algebra.)
  2. Factor the top: x2 − 9 = (x − 3)(x + 3). (A difference of squares.)
  3. Cancel the matching (x − 3) on top and bottom, leaving x + 3. (Legal, because near 3 that factor is not zero.)
  4. Now substitute into the simpler form: 3 + 3 = 6.

The limit is 6. Notice the restriction behind the cancellation: we are using inputs near 3, but not x = 3 itself. At those inputs, x − 3 is nonzero.

The original fraction and the simplified expression have the same values throughout that nearby region. They therefore have the same limit, even though only the simplified expression has a value at 3.

Try it: find lim (x → 1) (x2 + 2x − 3)/(x − 1).

Answer: factor the numerator as (x − 1)(x + 3). At inputs other than 1, cancel the common factor (x − 1).

The remaining expression is continuous at 1, so substitute to obtain 1 + 3 = 4. Check that you canceled a whole factor, rather than a term joined to other terms by addition.

Move 2: rationalize (when a square root is in the way)

If substitution gives 0/0 and a square-root difference is involved, try rationalizing: rewrite the expression so that this difference no longer contains a root.

The conjugate keeps the same two terms but reverses the sign between them. Multiply both the numerator, the top of the fraction, and the denominator, the bottom, by that conjugate. Their product can then use the difference-of-squares identity.

Worked example. Find lim (x → 0) (√(x + 9) − 3)/x.

  1. Substitution gives 0/0. (Move to algebra.)
  2. Multiply the numerator and denominator by the conjugate √(x + 9) + 3. This multiplies the fraction by 1 wherever the conjugate is nonzero, including inputs near zero.
  3. Multiply the numerator using the difference of squares. It becomes (x + 9) − 9 = x, giving x / [x(√(x + 9) + 3)]. Keep the denominator factored so you can see the common x.
  4. Cancel the x, leaving 1/(√(x + 9) + 3), then substitute: 1/(3 + 3) = 1/6.

The limit is 1/6. Multiplying by the conjugate gave a factor that could be canceled for nonzero x. We then evaluated the simpler expression at zero.

The Squeeze Theorem and one famous limit

The Squeeze Theorem compares three functions near the same target input. Suppose g(x) ≤ f(x) ≤ h(x) throughout a nearby interval, except possibly at the target itself.

If the lower function g and the upper function h both approach the same finite number L, the middle function f must approach L too. There is no room for a different limit.

This idea establishes the following trigonometric limit. The angle must be measured in radians.

lim (x → 0) sin(x)/x = 1

Direct substitution gives 0/0, so it cannot settle this limit. A geometric argument places sin(x)/x between cos(x) and 1 near zero, excluding zero itself. Both bounds approach 1, so the Squeeze Theorem gives the limit 1.

A related limit is lim (x → 0) (1 − cos(x))/x = 0. Use radians for this one too. These two results will help us derive the trigonometric differentiation rules.

Reshaping to reveal a known limit

Find lim (x → 0) sin(3x)/x. The known sine limit has the same expression inside the sine and in the denominator. Here they do not match yet, so rewrite the fraction to make them match.

  1. Multiply and divide by 3: sin(3x)/x = 3 × sin(3x)/(3x). (Same value, now the bottom matches the inside.)
  2. As x → 0, the inside 3x → 0 too, so sin(3x)/(3x) → 1. (The famous limit, in disguise.)
  3. What is left is 3 × 1 = 3.

The limit is 3. We introduced a factor of 3 in the denominator and compensated with a factor of 3 outside the fraction. That kept the value unchanged while revealing the known limit.

Naming the laws as they work

Evaluate lim (x → 2) (x2 + 3x)/(x + 1) and identify the law used at each step. This makes the reasoning behind direct substitution visible.

  1. Quotient law: the limit of a quotient is the quotient of the limits, provided the bottom limit is not zero. Bottom first: lim (x + 1) = 3 by the sum law (limit of a sum is the sum of the limits), and 3 is not zero, so the law applies.
  2. Sum law on the top: lim (x2 + 3x) = lim x2 + lim 3x.
  3. Product and constant multiple laws: lim x2 = (lim x)(lim x) = 2 × 2 = 4, and lim 3x = 3 × 2 = 6. So the top's limit is 10.
  4. Assemble: 10/3.

Direct substitution also gives 10/3. The worked steps show which limit laws justify that result.

The same laws work when you know limits but do not know the formulas. Suppose lim f(x) = 5 and lim g(x) = −2 at the same target input. Apply the constant multiple and difference laws to get lim [2f(x) − 3g(x)] = 2(5) − 3(−2) = 16. Apply the product law to get lim [f(x)g(x)] = −10.

A limit that is secretly a slope

Evaluate lim (h → 0) [(2 + h)3 − 8]/h. Here h is the changing input, and its target is zero. This quotient will also give us a first example of a curve's slope.

  1. Substitution gives (8 − 8)/0 = 0/0. (Indeterminate, so simplify.)
  2. Expand the cube: (2 + h)3 = 8 + 12h + 6h2 + h3. (Binomial expansion, done carefully.)
  3. Subtract 8: the numerator is 12h + 6h2 + h3. Each remaining term contains a factor of h, so the entire numerator can be divided by h at nonzero inputs.
  4. Divide by h: 12 + 6h + h2. (The 0/0 is gone.)
  5. Let h → 0: the limit is 12.

The answer 12 is the slope of the cubic curve at the input 2. In the derivative definition, we will repeat this calculation with x in place of 2. That process is called differentiating x3. Evaluating the result at x = 2 will recover the slope 12.

Making a limit exist: a parameter puzzle

The letter c represents a constant we can choose. Which choice makes lim (x → 2) f(x) exist?

The function uses f(x) = cx + 3 when x < 2, and f(x) = x2 + c when x ≥ 2. Work out the limit from each side using the formula assigned to that side. Then require the two answers to agree.

  1. Left-hand limit: approach 2 through the first rule: c(2) + 3 = 2c + 3.
  2. Right-hand limit: approach 2 through the second rule: 22 + c = 4 + c.
  3. A two-sided limit exists only when the sides agree: 2c + 3 = 4 + c, so c = 1.
  4. Check: with c = 1 both sides give 5. The limit exists and equals 5.

Choosing the constant this way makes the two formulas meet at the same height. In this example, that also makes the function continuous at the join. Matching heights alone does not guarantee matching slopes.

Two worries about 0/0, answered

Keep the expression and its limit separate. The expression 0/0 is undefined. When it appears after substitution in a limit problem, it tells you that substitution has not determined the limit.

Canceling (x − 3) is permitted only at inputs where that factor is nonzero. That is enough here because a limit depends on nearby inputs, not the value at the exact target.

Common misconceptions

  • "Seeing 0/0 tells me the limit." The expression 0/0 is undefined. A quotient whose numerator and denominator both approach zero has an indeterminate limit form: further work may reveal a finite limit, unbounded behavior, or no limit.
  • "You can never divide by something that becomes zero." For a limit you may cancel a factor that is only zero at the target point, because the limit ignores that exact point.
  • "sin(x)/x is undefined at 0, so its limit does not exist." The value at 0 is undefined, but the limit is a clean 1.
  • "Rationalizing changes the problem." Multiplying by a conjugate over itself multiplies by 1, so the value is unchanged; only the appearance is.

Putting it together

First check direct substitution. If the function is continuous at the target, it gives the limit. If substitution gives 0/0, inspect the expression: factoring or multiplying by a conjugate may produce an equivalent, simpler expression for nearby inputs.

The Squeeze Theorem gives another route by placing a function between two bounds with a shared limit. With angles in radians, the key trigonometric results are sin(x)/x → 1 and (1 − cos x)/x → 0.

Before accepting an answer, state why your method applies. In particular, check which denominators must stay nonzero and whether both sides approach the same value.

Sources

  1. OpenStax. (2016). 2.3 The limit laws. In Calculus volume 1. openstax.org
  2. Dawkins, P. (n.d.). Computing limits. Paul's Online Math Notes. tutorial.math.lamar.edu
  3. Math is Fun. (n.d.). Evaluating limits. mathsisfun.com
  4. Khan Academy. (n.d.). Calculus 1 [Online course]. khanacademy.org
  5. Sanderson, G. (3Blue1Brown). (n.d.). Limits, L'Hopital's rule, and epsilon delta definitions. 3blue1brown.com
  6. MIT OpenCourseWare. (2010). 18.01SC Single variable calculus, Fall 2010. Massachusetts Institute of Technology. ocw.mit.edu
  7. Grabiner, J. V. (1983). Who gave you the epsilon? Cauchy and the origins of rigorous calculus. The American Mathematical Monthly, 90(3), 185-194. doi.org/10.2307/2975545
Key terms
Direct substitution
Evaluating a limit by replacing the input with its target value when the function is continuous there.
Limit laws
Rules that let the limit of a sum, product, or quotient be built from simpler limits.
Indeterminate form
A limit form such as 0/0 that does not determine the limit from the individual numerator and denominator limits alone.
Conjugate
The expression formed by switching the sign between two terms, used to rationalize.
Squeeze Theorem
If a function is trapped between two others that share a limit, it shares that limit too.
Radian
The angle measure that makes sin(x)/x approach 1, required for clean trig calculus.

One-Sided Limits and Limits at Infinity

  • Evaluate one-sided limits and infinite limits at vertical asymptotes.
  • Find limits at infinity by comparing leading terms.
  • Identify horizontal asymptotes from limits at infinity.

Evaluate 1/(x − 2) at 2.1, then 2.01, then 2.001: the outputs are 10, then 100, then 1000. Come at 2 from the other side, at 1.999, and the output is −1000. Same formula, same point, two runaways in opposite directions. Now push the same function the other way, out to x = 1,000,000, and it flattens to almost exactly zero. Those two behaviors, the wall up close and the cruising height far away, are what this lesson reads off.

Walls nearby, cruising heights far away

Two new questions. First: near a spot where the bottom of a fraction hits zero, the graph can race off toward a wall. That wall is a vertical asymptote. Second: as x runs off to the far right or left, the graph often levels out toward a cruising height. That height is a horizontal asymptote. Both are read off with limits.

Remember: vertical asymptote = a wall the graph races up or down near a point; horizontal asymptote = the cruising height the graph settles toward far away.

Infinite limits and vertical asymptotes

Look at f(x) = 1/x near zero. As x approaches 0 from the right (tiny positive numbers), 1/x grows huge, so lim (x → 0+) 1/x = +∞. From the left (tiny negative numbers) it plunges, so lim (x → 0−) 1/x = −∞. The two sides disagree, so the two-sided limit does not exist, and the line x = 0 is a vertical asymptote. Writing ∞ here is shorthand for "grows without bound," not a number the function reaches.

To get the sign right, do a quick check. For lim (x → 2+) 1/(x − 2):

  1. Inputs just above 2 make x − 2 a tiny positive number. (Just bigger than 2.)
  2. One divided by a tiny positive is a huge positive. (Small bottom, big result.)
  3. So the limit is +∞.

Just below 2, x − 2 is a tiny negative, so the same fraction goes to −∞. Tracking that one sign is the whole skill.

Limits at infinity

To ask what a function does for enormous x, we compute lim (x → ∞) f(x). For a fraction of polynomials, the answer is decided by the leading terms, the highest powers on top and bottom, because for giant x the smaller-power terms barely matter. A clean method is to divide every term by the highest power of x in the bottom.

Worked example. Find lim (x → ∞) (3x2 − 5x + 2)/(6x2 + 4x − 1).

  1. Divide every term by x2: it becomes (3 − 5/x + 2/x2)/(6 + 4/x − 1/x2). (Same value, tidier.)
  2. As x → ∞, every piece with an x on the bottom fades to 0. (A number over something huge is tiny.)
  3. What remains is 3/6, which is 1/2.

So the limit is 1/2, and the line y = 1/2 is a horizontal asymptote, the cruising height the graph flattens toward far out.

The three cases, at a glance

DegreesLimit at infinity
top degree less than bottom degree0
top degree equal to bottom degreeratio of the leading numbers
top degree greater than bottom degreeplus or minus infinity (no horizontal asymptote)

So lim (x → ∞) (2x + 1)/(x2 + 3) = 0 because the bottom grows faster, while lim (x → ∞) (5x3 + 2)/(2x2 + 7) = ∞ because the top wins. Matching leading terms lets you predict the far-off behavior at a glance, no table needed.

Try it: find lim (x → ∞) (4x2 + 1)/(2x2 − x).

Answer: The degrees match (both 2), so take the ratio of the leading numbers: 4/2 = 2. Nice, you read the end behavior straight off the leading terms.

Beyond fractions of polynomials

The same question fits other families, and there is a simple pecking order: exponentials beat powers beat logarithms. So ex outruns every polynomial, giving lim (x → ∞) x2/ex = 0. A shrinking exponential fades away: lim (x → ∞) e−x = 0, which is why y = 0 is the cruising line of a cooling cup of coffee. The logarithm climbs forever but painfully slowly, so lim (x → ∞) ln(x) = ∞ even though it never levels off. Knowing this order resolves many limits by inspection.

Why this matters

Far-off behavior is the backbone of curve sketching and of real models. A population that levels off has a horizontal asymptote at its carrying capacity; a saturating reaction approaches a ceiling; a dose-response curve flattens at high dose. Each is a limit at infinity, and reading it off the leading terms turns a messy formula into a one-line prediction about the long run.

A complete sign check, narrated

Evaluate lim (x → 3−) (x + 1)/(x − 3), the approach to 3 from the left.

  1. Check the top near 3: x + 1 → 4, a solid positive number. (The top is not the problem.)
  2. Check the bottom from the left: inputs like 2.9 and 2.99 make x − 3 equal −0.1 and −0.01, tiny and negative. (This is the side-dependent part.)
  3. A positive top over a tiny negative bottom is a huge negative number: 4/(−0.01) = −400, and it only gets more extreme. (Sign logic, not memorization.)
  4. Conclusion: the limit is −∞, and from the right it is +∞ by the same reasoning with a positive bottom.

Write the three checks out every time: sign of the top, sign of the bottom on the given side, then divide. This thirty-second ritual eliminates the most common asymptote errors.

Limits at negative infinity, and a square-root trap

Evaluate lim (x → ∞) (3x + 2)/√(x2 + 1) and then the same limit as x → −∞.

  1. Divide top and bottom by x. For the bottom, note that x = √(x2) when x is positive, so √(x2 + 1)/x = √(1 + 1/x2).
  2. As x → ∞: top → 3, bottom → √1 = 1, so the limit is 3.
  3. Now x → −∞. Here is the trap: for negative x, √(x2) = |x| = −x, not x. Dividing the root by a negative x flips its sign: √(x2 + 1)/x = −√(1 + 1/x2).
  4. So the limit is 3/(−1) = −3.

One function, two different horizontal asymptotes, y = 3 to the right and y = −3 to the left. This is why "a function has at most one horizontal asymptote" is false: it can have one in each direction, and square roots are the classic source.

Slant asymptotes and polynomial end behavior

When the top degree is exactly one more than the bottom degree, the graph does not level off, but it does snuggle up to a slanted line. Divide (x2 + 1)/x = x + 1/x. Far out, the 1/x piece fades to 0, so the graph hugs the line y = x: a slant asymptote. For plain polynomials, end behavior is even simpler: only the leading term matters. For giant x, x3 − 100x2 behaves like x3, because at x = 1000 the first term is a billion while the second is a hundred million, a tenth its size and shrinking in proportion as x grows. So odd-degree polynomials run from −∞ up to +∞ (or the reverse if the leading coefficient is negative), and even-degree polynomials point both ends the same way. Reading end behavior off the leading term will be a standard final check when we sketch curves in Module 6.

A full asymptote inventory

Find every asymptote of f(x) = (2x2 − 8)/(x2 − 1).

  1. Factor: f(x) = 2(x − 2)(x + 2) / [(x − 1)(x + 1)]. (Everything about asymptotes lives in the factors.)
  2. Vertical candidates where the bottom is zero: x = 1 and x = −1. Check the top there: at x = 1 the top is 2(1) − 8 = −6, not zero, so the fraction blows up: a genuine vertical asymptote. Same at x = −1. (Nonzero over zero means a wall.)
  3. Sign near x = 1: just above 1, the bottom (x − 1)(x + 1) is a tiny positive times 2, and the top is near −6, so f → −∞; just below 1 the bottom flips sign and f → +∞.
  4. Horizontal: degrees are equal, so y = 2/1 = 2, the ratio of leading coefficients, in both directions.

Contrast the lookalike (x2 − 4)/(x − 2): at x = 2 the top is also zero, the factor cancels, and the graph has a hole at height 4, not a wall. The diagnostic is always the same: bottom zero with top nonzero gives an asymptote; bottom and top both zero means simplify first and look again.

Infinity is not a value, and not every fraction levels off

The first snag is treating ∞ like an ordinary number and saying "the limit equals infinity, so it exists." Infinity is not a value the function reaches; we write it to describe how the limit fails, by growing without bound. The second snag is assuming every fraction of polynomials levels off. It only does when the top degree is at most the bottom degree. When the top wins, the graph climbs forever and there is no horizontal asymptote.

Common misconceptions

  • "Infinity is a number the function hits." No. It records unbounded growth; strictly, the limit does not exist, and ∞ says how.
  • "Every rational function has a horizontal asymptote." Only when the top degree is less than or equal to the bottom degree.
  • "A graph can never cross a horizontal asymptote." It can, in the middle. An asymptote only governs the far-off behavior, not the interior.
  • "The plus or minus in a one-sided infinite limit does not matter." The sign tells you whether the graph races up or down that wall, which is the whole picture.

What you now know

Near a point where the bottom is zero but the top is not, a fraction blows up to +∞ or −∞, giving a vertical asymptote; a quick sign check on the bottom picks the direction. Far out, a limit at infinity gives the cruising height: compare leading terms, or divide every term by the highest power on the bottom. Top degree smaller gives 0, equal gives the ratio of leading numbers, larger gives infinity. And exponentials beat powers beat logarithms. You can now describe both the near-singular and the long-range shape of a graph.

Sources

  1. OpenStax. (2016). 4.6 Limits at infinity and asymptotes. In Calculus volume 1. openstax.org
  2. Dawkins, P. (n.d.). One-sided limits. Paul's Online Math Notes. tutorial.math.lamar.edu
  3. Dawkins, P. (n.d.). Infinite limits. Paul's Online Math Notes. tutorial.math.lamar.edu
  4. Dawkins, P. (n.d.). Limits at infinity, part I. Paul's Online Math Notes. tutorial.math.lamar.edu
  5. Math is Fun. (n.d.). Limits to infinity. mathsisfun.com
  6. Khan Academy. (n.d.). Calculus 1 [Online course]. khanacademy.org
  7. MIT OpenCourseWare. (2010). 18.01SC Single variable calculus, Fall 2010. Massachusetts Institute of Technology. ocw.mit.edu
Key terms
Infinite limit
A limit where the function grows or falls without bound, written as plus or minus infinity.
Vertical asymptote
A vertical line the graph approaches as the function blows up near a point.
Limit at infinity
The value a function approaches as x grows arbitrarily large or small.
Horizontal asymptote
A horizontal line the graph approaches as x goes to plus or minus infinity.
Leading term
The term with the highest power, which controls end behavior.
Slant asymptote
A diagonal line the graph approaches when the top degree is one more than the bottom.

Continuity

  • State the three-part definition of continuity at a point.
  • Classify removable, jump, and infinite discontinuities.
  • Apply the Intermediate Value Theorem.

The polynomial x3 − x − 1 equals −1 at x = 1 and +5 at x = 2. It must therefore be zero somewhere in between, and a computer can hunt that root down to any number of decimals without ever solving the equation. The whole guarantee rests on one property of the graph: no gaps between those two inputs. This lesson turns "no gaps" into a three-item checklist you can run on any function, and then collects the payoff.

No gaps, stated as a checklist

Continuity is the property of being connected, of the graph and the point agreeing with no surprises. It matters because nearly every big theorem ahead, and the very idea of a derivative, quietly assumes it. When a modeler says a function is "nice," continuity is usually what they mean.

The core of it: continuous at a point means the pen never lifts there. The limit heads to a value, the point exists, and the two match.

The three-part checklist

Formally, f is continuous at x = a when all three of these hold:

  1. f(a) is defined. (The point actually exists.)
  2. lim (x → a) f(x) exists. (The two sides head to the same value.)
  3. lim (x → a) f(x) = f(a). (Where the graph was heading is where the point actually sits.)

If any one fails, f has a discontinuity at a. So checking continuity is just running the checklist: does the point exist, does the limit exist, do they match. Polynomials pass all three everywhere. Rational, root, trig, exponential, and log functions pass everywhere in their domains, which is exactly why "just plug in" worked for their limits last module. Continuity is the property that makes plugging in legal.

Three kinds of break

  • Removable (a hole): the limit exists but does not match f(a), or f(a) is missing. You could patch it by filling in one point, which is why it is called removable.
  • Jump: the left and right limits both exist but differ, so the graph leaps. Think postage rates or tax brackets that step up suddenly.
  • Infinite: the function blows up, as 1/x does at 0. There is a vertical asymptote, and no single point can patch it.

Let us classify one together, slowly

Where is f(x) = (x2 − 1)/(x − 1) discontinuous, and what type is it?

  1. Find where the bottom is zero: x − 1 = 0 at x = 1, so the function is undefined there. (Checklist item 1 fails.)
  2. Check the limit anyway: factor the top, x2 − 1 = (x − 1)(x + 1), cancel, get x + 1. (Same move as the limit lessons.)
  3. So lim (x → 1) f(x) = 1 + 1 = 2. The limit exists, but the value is missing.
  4. Limit exists, value missing means a removable discontinuity, a hole at the point (1, 2).

That is the whole classification method: run the checklist and see which item fails.

Try it: is f(x) = (x2 − 4)/(x − 2) continuous at x = 2?

Answer: No. It is undefined at 2, but the limit is x + 2 → 4. Limit exists, value missing, so it is a removable discontinuity, a hole at (2, 4). Well spotted.

A piecewise seam

Let f(x) = x + 1 for x < 2, and f(x) = x2 for x ≥ 2. Is it continuous at x = 2? Coming from the left, the values head to 2 + 1 = 3. Coming from the right, they head to 22 = 4. Since 3 and 4 disagree, the two-sided limit does not exist, and there is a jump of size 1 at x = 2. A piecewise function is continuous at a seam only when the pieces actually meet, when the two one-sided limits and the value all line up.

The payoff: the Intermediate Value Theorem

Continuity has a lovely consequence. The Intermediate Value Theorem (IVT) says that if f is continuous on a closed interval [a, b], then it takes every value between f(a) and f(b) at least once. In plain terms, a connected graph cannot skip a height. The practical use is finding roots: if f is continuous with f(1) = −2 and f(2) = 3, then it changes sign, so it must equal 0 somewhere between 1 and 2. This is why a temperature that goes from below freezing to above freezing must pass through exactly 0 degrees at some moment.

Why continuity matters downstream

Continuity is a prerequisite for the big results ahead. A function must be continuous even to have a chance at a derivative. The theorem that guarantees an optimization problem on a closed interval actually has a highest and lowest value requires continuity too. It is the quiet assumption behind most of calculus.

Choosing a constant to force continuity

Engineers often need two formulas to hand off smoothly. Find the value of k that makes f continuous everywhere, where f(x) = kx2 for x ≤ 2 and f(x) = 3x + 2 for x > 2.

  1. Away from the seam, each piece is a polynomial, continuous on its own stretch. Only x = 2 is in question. (Isolate the risky point.)
  2. Value at the seam: the first rule owns x = 2, so f(2) = 4k. (Checklist item 1.)
  3. Left-hand limit: through the first rule, k(2)2 = 4k. Right-hand limit: through the second rule, 3(2) + 2 = 8. (Checklist item 2 needs these to agree.)
  4. Set them equal: 4k = 8, so k = 2. With that choice all three checklist items pass, and the two formulas weld into one continuous function.

Check: with k = 2, both sides meet at the point (2, 8). Done.

Using the IVT for real: trapping a root

The Intermediate Value Theorem is not just a promise; it powers a genuine numerical method called bisection. Show that f(x) = x3 − x − 1 has a root, and trap it.

  1. f(1) = 1 − 1 − 1 = −1, negative. f(2) = 8 − 2 − 1 = 5, positive. f is a polynomial, hence continuous, so by the IVT a root lies in (1, 2). (Sign change plus continuity equals a guaranteed crossing.)
  2. Test the midpoint: f(1.5) = 3.375 − 1.5 − 1 = 0.875, positive. The sign change is now between 1 and 1.5, so the root lives in (1, 1.5). (Half the suspects eliminated.)
  3. Again: f(1.25) = 1.953125 − 1.25 − 1 = −0.296875, negative. Root in (1.25, 1.5).
  4. Again: f(1.375) = 2.599609 − 1.375 − 1 = 0.224609, positive. Root in (1.25, 1.375).

Each step halves the interval; ten more steps would pin the root (about 1.3247) to three decimal places. This is exactly how calculators and computers find roots of equations no algebra can solve, and the IVT is the reason the method cannot miss.

Continuity survives composition

Sums, differences, products, and quotients (away from zero bottoms) of continuous functions are continuous, and so is a composition: if g is continuous at a and f is continuous at g(a), then f(g(x)) is continuous at a. That is why a beast like h(x) = cos(√(x2 + 1)) needs no checklist: x2 + 1 is a polynomial (continuous, and always positive so the root is safe), the square root is continuous on its domain, and cosine is continuous everywhere, so the whole stack is continuous everywhere. Spotting "built from continuous parts, legally assembled" saves enormous time.

One more refinement matters for the theorems ahead: on a closed interval [a, b], continuity at the endpoints means one-sided continuity, the limit from inside matching the value. A function continuous on a closed interval is the exact hypothesis both the IVT here and the Extreme Value Theorem in Module 6 demand, and neither theorem survives without it.

One function, both kinds of break

Classify every discontinuity of f(x) = (x + 3)/(x2 − 9).

  1. The bottom factors as (x − 3)(x + 3), so f is undefined at x = 3 and x = −3: two suspects. (Domain scan first.)
  2. Simplify: for x away from −3, f(x) = 1/(x − 3). (The common factor cancels.)
  3. At x = −3: the limit is 1/(−3 − 3) = −1/6, a finite number. Limit exists, value missing: a removable discontinuity, a hole at (−3, −1/6).
  4. At x = 3: the simplified form 1/(x − 3) blows up, positive from the right, negative from the left: an infinite discontinuity with vertical asymptote x = 3.

Same formula, two completely different failures, distinguished by whether the troublesome factor canceled. This is the standard exam problem for this lesson, and the method is fully mechanical: find the undefined points, simplify, then take the limit at each suspect and let the checklist name the break.

The limit is only one of the three items

The most common mix-up is thinking "the limit exists, so the function is continuous." The limit is only one of the three checklist items. The point can still be missing (a hole) or sitting at a different height, and either way the function is not continuous there. Always check all three: point exists, limit exists, and they match.

Common misconceptions

  • "If the limit exists, f is continuous." Not necessarily. The point may be missing or at a different height, leaving a removable discontinuity.
  • "A hole and a jump are the same kind of break." A hole (removable) has a limit; a jump has two different one-sided limits.
  • "The IVT tells you where the root is." It only promises a root exists in the interval. Finding it takes further work.
  • "Every function is continuous." Piecewise steps, asymptotes, and holes all break continuity at specific points.

Pulling it together

Continuous at a point means you can draw through it without lifting your pen, captured by a three-item checklist: f(a) exists, the limit exists, and they are equal. Breaks come in three flavors: removable holes, jumps, and infinite blow-ups. Continuity is what makes direct substitution valid, and it powers the Intermediate Value Theorem, which guarantees a continuous function that changes sign must cross zero. It is also the ground floor for derivatives and optimization ahead.

Sources

  1. OpenStax. (2016). 2.4 Continuity. In Calculus volume 1. openstax.org
  2. Dawkins, P. (n.d.). Continuity. Paul's Online Math Notes. tutorial.math.lamar.edu
  3. Math is Fun. (n.d.). Continuous functions. mathsisfun.com
  4. Khan Academy. (n.d.). Calculus 1 [Online course]. khanacademy.org
  5. MIT OpenCourseWare. (2010). 18.01SC Single variable calculus, Fall 2010. Massachusetts Institute of Technology. ocw.mit.edu
  6. Tall, D., & Vinner, S. (1981). Concept image and concept definition in mathematics with particular reference to limits and continuity. Educational Studies in Mathematics, 12(2), 151-169. doi.org/10.1007/BF00305619
  7. OpenStax. (2016). 2.5 The precise definition of a limit. In Calculus volume 1. openstax.org
Key terms
Continuous at a point
f(a) exists, the limit exists, and the limit equals f(a).
Discontinuity
A point where a function fails to be continuous.
Removable discontinuity
A hole where the limit exists but the value is missing or different.
Jump discontinuity
A break where the left and right limits exist but disagree.
Infinite discontinuity
A break where the function grows without bound, giving a vertical asymptote.
Intermediate Value Theorem
A continuous function on [a, b] hits every value between f(a) and f(b).

Module 3: The Derivative

Build the derivative from the limit of a difference quotient and interpret what it measures, connecting the slope of a tangent line to the instantaneous rate of change. This is the conceptual heart of the whole course.

The Definition of the Derivative

  • Write the difference quotient and take its limit to define the derivative.
  • Compute a derivative from the definition for polynomials and simple functions.
  • Connect the derivative to the slope of a tangent line.

Your speedometer reads 60 mph. That is a speed at the moment you look, not an average for the whole trip. How can we define a rate at one instant?

Average speed divides distance traveled by elapsed time over an interval. If we set the interval's duration to zero, that calculation gives 0/0, which is undefined. The derivative takes a different route: calculate average rates over nonzero intervals, then find their limit as the intervals shrink.

You have already worked with limits whose direct substitution gives 0/0. We will use that same reasoning to define an instantaneous rate.

Shrinking the gap without dividing by zero

Start with an average rate over an interval. Then calculate rates over shorter intervals around the same input. If these rates approach one finite number as the interval shrinks, that number is the derivative at the input.

Each quotient still uses a nonzero interval. We use a limit to describe what happens as its size approaches zero.

Why this matters: a derivative turns a limit of average rates into a rate at one input. The limit must exist as a finite number.

Starting from average rate

For a function f, start at the input x and move to x + h. The change in input is h, which may be positive or negative but cannot be zero in the quotient. Divide the change in output by this change in input to get the average rate of change.

This fraction is called the difference quotient:

[f(x + h) − f(x)] / h

Read the fraction as "change in output divided by change in input." The letter h records the input change. Hold x fixed while you let h approach zero from either side.

If the quotient approaches a finite number, that limit is the derivative. We write it f'(x), pronounced "f prime of x":

f'(x) = lim (h → 0) [f(x + h) − f(x)] / h

Setting h = 0 in the original quotient produces 0/0, which is undefined. Keep h nonzero while you simplify or otherwise analyze the quotient. Then take its limit.

The picture: secant lines becoming a tangent

A secant line passes through two points on the graph: one at x and another at x + h. Its slope is the difference quotient.

Keep the first point fixed and move the second closer. If the secant slopes approach one finite number, that number is the slope of the tangent line through the fixed point.

This describes the curve's local rate of change. It does not require the tangent to touch the graph at only one point, and a tangent may cross the curve.

Let us build one together, slowly

Find f'(x) for f(x) = x2 from the definition.

  1. Find the shifted output f(x + h). Replace each x with x + h, then expand: (x + h)2 = x2 + 2xh + h2. The input x stays fixed while h changes.
  2. Subtract the original output f(x) = x2. The x2 terms cancel, leaving 2xh + h2. This is the change in output, the numerator of the difference quotient.
  3. Divide by the input change h: (2xh + h2)/h = 2x + h. This cancellation is valid for every nonzero h.
  4. Take the limit as h → 0. The simplified expression gives 2x + 0 = 2x.

We have found f'(x) = 2x, a formula for the slope at each input. For example, at x = 3 the slope is 2(3) = 6.

Step 3 explains why the limit became easier. For nonzero h, the original difference quotient equals the simpler expression. We take the limit of that expression instead of trying to evaluate 0/0.

Try it: use the definition on f(x) = x2 + 3x.

Answer: expand the shifted function and subtract the whole original function. This gives f(x + h) − f(x) = 2xh + h2 + 3h.

Next divide each term by nonzero h to obtain 2x + h + 3. Finally, let h → 0. The result is f'(x) = 2x + 3.

A second example with more terms

Find f'(x) for f(x) = 3x2 − 5x + 1.

  1. Expand f(x + h) = 3(x + h)2 − 5(x + h) + 1 = 3x2 + 6xh + 3h2 − 5x − 5h + 1.
  2. Subtract the entire expression f(x), using parentheses to keep its signs together. The 3x2, −5x, and +1 terms cancel, leaving 6xh + 3h2 − 5h.
  3. Divide by h: 6x + 3h − 5.
  4. Let h → 0: 6x − 5.

The derivative is f'(x) = 6x − 5. The added constant +1 canceled when we subtracted the original function from the shifted function. It contributes no change in output when the input changes.

It works beyond polynomials

For f(x) = 1/x, use a nonzero input x and keep x + h nonzero too. The difference quotient is [1/(x + h) − 1/x]/h.

First combine the two fractions in the numerator using a common denominator. Their difference is −h / [x(x + h)], and the whole result is still divided by h.

For nonzero h, cancel that common factor to obtain −1/[x(x + h)]. Now let h → 0. The derivative is −1/x2, valid for x other than zero.

Notation, and why we do the hard way first

You will see several notations for the derivative: f'(x), pronounced "f prime of x"; dy/dx, pronounced "d y by d x"; and D f(x).

The notation dy/dx reminds us that we are measuring change in y with respect to change in x. At this stage, treat it as a symbol for the derivative defined by a limit, not as ordinary division by a zero input change.

The definition tells you what the derivative measures. The rules in the next module will make the calculation quicker.

A square root, tamed by the conjugate

Find f'(x) for f(x) = √x. Here we work at a positive input x and keep h small enough that x + h stays positive. To simplify the quotient, use the conjugate method from the limits lesson.

  1. Write the difference quotient: [√(x + h) − √x]/h. Substituting h = 0 gives 0/0, as always. (The starting point.)
  2. Multiply top and bottom by the conjugate √(x + h) + √x. The top becomes (x + h) − x = h. (The root difference collapses to a plain h.)
  3. The quotient is now h / [h(√(x + h) + √x)]; cancel the h to get 1/(√(x + h) + √x). (The 0/0 is cleared.)
  4. Let h → 0: the answer is 1/(2√x).

For positive x, the result is f'(x) = 1/(2√x). At x = 4 the slope is 1/4, while at x = 100 it is 1/20. Both slopes are positive, but the second is smaller: the curve is rising more gradually there.

At x = 0, this derivative formula is not defined. The right-hand slopes grow without bound, so there is no finite derivative at the endpoint.

Where derivatives fail: corners, cusps, jumps

A function is differentiable at an input when its derivative exists there as a finite number. This requires more than continuity. A graph can be continuous at a point and still lack a derivative there.

  • A corner: for f(x) = |x| at 0, the difference quotient is |h|/h, which equals +1 for positive h and −1 for negative h. The one-sided limits disagree, so f'(0) does not exist. The graph is perfectly connected; it just bends too sharply for a single slope.
  • A vertical tangent: the cube root x1/3 at 0 has difference quotient h1/3/h = 1/h2/3 → ∞. The tangent exists as a picture, but its slope is infinite, so the derivative does not exist as a number.
  • Any discontinuity: a jump or hole always kills the derivative, because of the fact below.

Differentiable implies continuous. Suppose f'(a) exists. As x → a, write the output change as the difference quotient multiplied by the input change:

f(x) − f(a) = [the difference quotient] × (x − a) → f'(a) × 0 = 0

The quotient approaches a finite derivative, and the input change approaches zero. Their product therefore approaches zero. This means f(x) → f(a), which is continuity at a.

The reverse implication is false. The absolute-value function is continuous at zero, but its left and right slopes disagree. Continuity by itself does not guarantee a derivative.

The derivative as a function, tabulated

Keep the notation's two roles separate. f' names the derivative function, while f'(3) is its value at one input. For f(x) = x2, the derivative rule is f'(x) = 2x:

x−2−10123
slope f'(x)−4−20246

Negative entries describe a graph that falls as you move to the right. The zero entry gives a horizontal tangent. Positive entries describe a graph that rises.

For this quadratic, the negative slopes become less steep as x approaches zero, and the positive slopes become steeper after zero. These are slopes of the original graph, not its heights.

Estimating a derivative numerically

You can also estimate a derivative by evaluating a difference quotient at a small nonzero input change. Use f(x) = x2 at x = 3.

A forward difference uses the target input and an input just to its right. With h = 0.01, the estimate is [f(3.01) − f(3)]/0.01 = (9.0601 − 9)/0.01 = 6.01. The exact derivative is 6, so this estimate is high by 0.01.

A symmetric difference uses inputs equally far to the left and right. Their separation is 0.02 here, so the estimate is [f(3.01) − f(2.99)]/0.02 = (9.0601 − 8.9401)/0.02 = 0.1200/0.02 = 6.0000.

The symmetric estimate is exact for this example. That does not make every numerical estimate exact. Keep track of the chosen inputs, the size of the interval, and whether the function values themselves are rounded or measured.

Setting h to zero too early, and what a tangent really is

Setting h = 0 in the original quotient gives 0/0. In the algebraic examples here, we first simplify so that the denominator factor h cancels at nonzero inputs. Only then do we let h approach zero. Other functions may need other ways to evaluate the limit.

A second common error is defining a tangent by the number of times it touches the graph. The tangent's slope comes from the limit of secant slopes at the chosen point. It can also meet the curve elsewhere.

Common misconceptions

  • "Just plug in h = 0." That gives 0/0. Simplify first to cancel the h, then take the limit.
  • "The derivative is a single fixed number." The notation f'(x) describes a function that assigns a slope to each input where the derivative exists. Its values can vary, although some functions have the same slope at every input.
  • "A tangent line touches the curve at exactly one point." Not required. It is the limit of secant slopes; it may cross the curve elsewhere.
  • "Average rate and instantaneous rate are the same." The instantaneous rate is the limit of average rates as the gap shrinks to zero.

Where this leaves us

The difference quotient [f(x + h) − f(x)]/h gives an average rate over a nonzero input change. The derivative is its finite limit: f'(x) = lim (h → 0) [f(x + h) − f(x)]/h.

For the polynomial examples, the steps are: expand the shifted function, subtract the whole original function, divide by nonzero h, and take the limit. Fractions and roots may need a common denominator or a conjugate before cancellation is possible.

On the graph, the same limit gives the tangent's slope. A corner, a discontinuity, or unbounded secant slopes can prevent a finite derivative. Check that the limit exists before interpreting it as a slope or rate.

Sources

  1. OpenStax. (2016). 3.1 Defining the derivative. In Calculus volume 1. openstax.org
  2. OpenStax. (2016). 3.2 The derivative as a function. In Calculus volume 1. openstax.org
  3. Dawkins, P. (n.d.). The definition of the derivative. Paul's Online Math Notes. tutorial.math.lamar.edu
  4. Math is Fun. (n.d.). Introduction to derivatives. mathsisfun.com
  5. Sanderson, G. (3Blue1Brown). (n.d.). The paradox of the derivative. 3blue1brown.com
  6. MIT OpenCourseWare. (2010). 1. Differentiation. 18.01SC Single variable calculus. Massachusetts Institute of Technology. ocw.mit.edu
  7. Grabiner, J. V. (1983). The changing concept of change: The derivative from Fermat to Weierstrass. Mathematics Magazine, 56(4), 195-206. doi.org/10.2307/2689807
Key terms
Difference quotient
The average rate of change [f(x+h) - f(x)]/h for a nonzero input change h.
Derivative
The finite limit of the difference quotient as h approaches 0, when it exists; the instantaneous rate of change.
Secant line
A line through two points on a curve.
Tangent line
At a differentiable point, the line through the point whose slope is the limit of nearby secant slopes.
Instantaneous rate of change
The rate at a single instant, given by the derivative.
Leibniz notation
The derivative written dy/dx, echoing rise over run with the gap taken to zero.

Interpreting the Derivative: Slope and Rate

  • Write the equation of a tangent line using the derivative.
  • Interpret the derivative as velocity in motion problems.
  • Read where a function increases or decreases from the sign of its derivative.

A factory's total cost for q items is C(q) = 500 + 4q + 0.01q2 dollars. Building the 101st item actually costs 6.01 dollars. The derivative at q = 100 says 6 dollars per item, in one line, without building anything. Same function, a second reading: C'(q) is not the steepness of a picture here, it is money per item, a quantity the plant manager can act on. Every derivative you compute carries both readings at once, and this lesson is about learning to hear them.

One number, two readings

The derivative is a two-in-one tool. Geometrically it is the steepness of the graph. Physically it is how fast the quantity is changing. Same number, two stories. If you can read a speedometer and read a hill, you can read a derivative.

The upshot: the value of f'(x) is a slope; the sign of f'(x) tells you uphill (rising) or downhill (falling).

The equation of the tangent line

The tangent line at x = a passes through the point (a, f(a)) and has slope f'(a). Using point-slope form, its equation is:

y − f(a) = f'(a) (x − a)

Worked example. Find the tangent line to f(x) = x2 at x = 3.

  1. Find the point: f(3) = 32 = 9, so the curve passes through (3, 9). (Where we are.)
  2. Find the slope: f'(x) = 2x, so f'(3) = 6. (How steep, from last lesson.)
  3. Fill in point-slope: y − 9 = 6(x − 3). (Plug the point and slope in.)
  4. Tidy up: y = 6x − 18 + 9 = 6x − 9.

So the tangent line is y = 6x − 9. Near x = 3, this straight line and the curve x2 give almost the same output. That is the idea of linear approximation: a curve, up close, looks like its tangent line, which is how calculators and error estimates lean on calculus.

Try it: find the tangent line to f(x) = x2 − 2x at x = 1.

Answer: f(1) = −1 and f'(x) = 2x − 2, so f'(1) = 0. A zero slope means a flat line: y = −1, the bottom of the parabola. You found a horizontal tangent.

The derivative as velocity

If s(t) gives the position of an object at time t, then s'(t) is the velocity, the instantaneous rate at which position changes. Positive velocity means moving forward, negative means backward, and zero means momentarily at rest, like the top of a tossed ball's arc. Differentiate once more and you get acceleration s''(t), the rate at which velocity changes.

This was the original problem calculus was invented to solve, and the pattern generalizes: the derivative of a quantity with respect to time is its rate of flow. Marginal cost is the derivative of total cost; electric current is the derivative of charge; a reaction rate is the derivative of concentration.

Worked example: a thrown ball

A ball's height is s(t) = −16t2 + 32t feet after t seconds.

  1. Velocity is the derivative: s'(t) = −32t + 32. (Rate of the height.)
  2. At the start, t = 0: s'(0) = 32 ft/s upward. (Thrown up.)
  3. It is momentarily at rest when s'(t) = 0: solve −32t + 32 = 0 to get t = 1 second. (The peak.)
  4. Acceleration is the next derivative: s''(t) = −32 ft/s per second, constant gravity pulling down.

After the peak the velocity is negative and the ball falls. Notice how zero velocity marks the very top, the turning point.

Sign of the derivative: uphill or downhill

The sign of f'(x) reveals the shape of the graph:

  • Where f'(x) > 0, the tangent slopes up, so f is increasing (going uphill).
  • Where f'(x) < 0, the tangent slopes down, so f is decreasing (going downhill).
  • Where f'(x) = 0, the tangent is flat, a possible peak or valley.

For f(x) = x2, we have f'(x) = 2x, which is negative for x < 0 (the left side falls) and positive for x > 0 (the right side rises), with a flat spot at x = 0. That matches the U shape of a parabola exactly. This link between the sign of the derivative and the direction of the graph is the engine behind curve sketching and optimization ahead.

Units: the fastest way to read a derivative

Whenever a derivative models something real, its units are the output units divided by the input units, and saying the units out loud usually explains the number. Suppose a factory's total cost is C(q) = 500 + 4q + 0.01q2 dollars for producing q items. Then C'(q) = 4 + 0.02q dollars per item: the marginal cost, the approximate cost of the next single item.

  1. At q = 100: C'(100) = 4 + 2 = 6 dollars per item. (The 101st item will cost about 6 dollars to make.)
  2. Check against reality: C(101) − C(100) = [500 + 404 + 102.01] − [500 + 400 + 100] = 6.01 dollars. The derivative said 6; the exact answer is 6.01. (The derivative is the clean approximation to "one more step.")

The same reading works everywhere: if P(t) is population in people and t is in years, P'(t) is people per year; if V(t) is liters and t is seconds, V'(t) is liters per second, a flow rate.

A full motion story with a sign chart

A particle moves along a line with position s(t) = t3 − 6t2 + 9t meters at time t seconds, for t in [0, 5]. Where is it moving forward, and where backward?

  1. Velocity: v(t) = s'(t) = 3t2 − 12t + 9 = 3(t2 − 4t + 3) = 3(t − 1)(t − 3). (Differentiate, then factor to expose the zeros.)
  2. Zeros at t = 1 and t = 3: the only moments the particle can turn around.
  3. Sign of v between the zeros: at t = 0, 3(−1)(−3) = 9 > 0, forward. At t = 2, 3(1)(−1) = −3 < 0, backward. At t = 4, 3(3)(1) = 9 > 0, forward again.
  4. Story: it advances until t = 1 (reaching s(1) = 1 − 6 + 9 = 4 m), backs up until t = 3 (down to s(3) = 27 − 54 + 27 = 0 m), then advances for good. Acceleration a(t) = 6t − 12 is negative before t = 2 and positive after, so the velocity itself bottoms out at t = 2.

This forward-backward reading from the sign of the derivative is precisely the "sign chart" discipline Module 6 will formalize for arbitrary graphs.

Linear approximation: the tangent line as a calculator

Because a curve hugs its tangent line near the touch point, the tangent's outputs approximate the curve's. The recipe: near x = a, f(x) ≈ L(x) = f(a) + f'(a)(x − a). Estimate √4.1 by hand.

  1. Pick a friendly nearby point: a = 4, where √4 = 2 exactly. (Anchor at easy arithmetic.)
  2. Slope there: f'(x) = 1/(2√x), so f'(4) = 1/4. (From last lesson's hard-won formula.)
  3. Tangent line: L(x) = 2 + (1/4)(x − 4).
  4. Evaluate at 4.1: L(4.1) = 2 + (1/4)(0.1) = 2.025.

The true value is 2.02485..., so the estimate is off by about 0.00015, from ten seconds of arithmetic. Why it works: zoom in on any differentiable curve and it straightens into its tangent; the smaller the step from a, the better the line stands in for the curve. Why it can fail: step too far (using the same line to estimate √9 gives 3.25, badly off from 3), because far from the touch point the curve and line part ways. Linear approximation, done at scale, is how computers evaluate roots, logs, and trig functions internally, and it returns in Calculus II as the first term of Taylor series.

Average versus instantaneous, side by side

The thrown ball with s(t) = −16t2 + 32t makes the distinction sharp. Over the first second, the average velocity is total change over total time: [s(1) − s(0)]/(1 − 0) = (16 − 0)/1 = 16 ft/s. But the instantaneous velocity v(t) = −32t + 32 tells a richer story: 32 ft/s at launch, 16 ft/s at t = 0.5, 0 at the peak. The average is one number for the whole interval, the slope of the secant line from (0, 0) to (1, 16); the derivative is a number for each instant, the slope of the tangent at that instant. Notice the average value 16 is actually achieved, at t = 0.5, exactly halfway: no coincidence, and Module 6's Mean Value Theorem will turn that observation into a guarantee. When a homework problem says "velocity," check which one it wants; mixing them is the single most common motion-problem error.

Height is not direction

A common mix-up is thinking a big value of f(x) means the function is increasing there. Not so. Increase depends on the derivative, not on how high the graph is. A function can be way up high yet falling (just past a peak) or low down yet rising. Always look at the sign of f'(x), not the size of f(x). A second mix-up is reading zero velocity as "stopped for good." It is only a momentary pause, like the top of the ball's flight, where it is about to reverse.

Common misconceptions

  • "A large f(x) means f is increasing there." Direction comes from the sign of f'(x), not the height of f(x).
  • "Zero velocity means the object stopped permanently." It is instantaneous rest, as at a ball's peak, with nonzero acceleration turning it around.
  • "The tangent line only matters in geometry." It is also the best linear approximation to the function near the point.
  • "Acceleration is just speed." Acceleration is the rate of change of velocity, the second derivative of position.

Summing up

The derivative reads as both slope and rate. The tangent line at x = a is y − f(a) = f'(a)(x − a), and near a point the curve looks like that line (linear approximation). For motion, s'(t) is velocity and s''(t) is acceleration. And the sign of f'(x) tells direction: positive means rising, negative means falling, zero means a flat spot. To understand any graph, find where its derivative is positive, negative, or zero.

Sources

  1. OpenStax. (2016). 3.4 Derivatives as rates of change. In Calculus volume 1. openstax.org
  2. OpenStax. (2016). 3.2 The derivative as a function. In Calculus volume 1. openstax.org
  3. OpenStax. (2016). 4.2 Linear approximations and differentials. In Calculus volume 1. openstax.org
  4. Dawkins, P. (n.d.). Rates of change. Paul's Online Math Notes. tutorial.math.lamar.edu
  5. Dawkins, P. (n.d.). Linear approximations. Paul's Online Math Notes. tutorial.math.lamar.edu
  6. Sanderson, G. (3Blue1Brown). (n.d.). The paradox of the derivative. 3blue1brown.com
  7. Khan Academy. (n.d.). Calculus 1 [Online course]. khanacademy.org
Key terms
Tangent line equation
y - f(a) = f'(a)(x - a), the line touching the curve at x = a.
Linear approximation
Using the tangent line as the best straight-line estimate of a curve near a point.
Velocity
The derivative of position with respect to time.
Acceleration
The derivative of velocity, the second derivative of position.
Increasing
A function is increasing where its derivative is positive.
Decreasing
A function is decreasing where its derivative is negative.

Module 4: Differentiation Rules

Replace the slow limit definition with fast, reliable rules for differentiating any elementary function, from the power rule through the product, quotient, and chain rules. By the end you can differentiate essentially anything a first calculus course throws at you.

The Power, Constant, and Sum Rules

  • Apply the power rule to any power of x, including negative and fractional exponents.
  • Use the constant multiple and sum or difference rules.
  • Differentiate any polynomial quickly.

Differentiating x2 from the definition took four steps and most of a page. Doing 7x4 − 2x3 + x − 9 that way would take four separate limits and a long afternoon. With the rule on this page it takes about eight seconds, and the answer is 28x3 − 6x2 + 1. Nothing here is a trick: every shortcut is provable from the definition you just used, and we will prove the main one before the page is out.

The rule that replaces the limit

The power rule is a two-word recipe for differentiating any power of x. Once you have it, polynomials, roots, and reciprocals all fall in seconds instead of minutes. This is the workhorse of the whole module.

Bottom line: for x to a power, bring the exponent down in front, then knock the exponent down by one.

The power rule

d/dx (xn) = n x(n − 1)

Read the left side "the derivative of x to the n." In words: the exponent hops down to the front as a multiplier, and the new exponent is one less. It works for every real exponent, positive, negative, or fractional. For example, d/dx (x5) = 5x4. It even agrees with the hard way: last module we found d/dx (x2) = 2x, and the power rule gives 2x1 = 2x instantly. Same answer, far less work.

Three helper rules

  • Constant rule: the derivative of a plain number is 0, because a flat line has no slope. So d/dx (7) = 0.
  • Constant multiple rule: a number multiplying the function just comes along for the ride. d/dx (5x3) = 5 × 3x2 = 15x2.
  • Sum and difference rule: differentiate one term at a time. d/dx (f ± g) = f' ± g'.

Together these let you differentiate any polynomial in a single pass, term by term, with no limits in sight. Each one follows from the matching limit law in Module 2, so they are trustworthy, not magic.

Let us differentiate a polynomial together, slowly

Find the derivative of f(x) = 7x4 − 2x3 + x − 9, one term at a time.

  1. d/dx (7x4) = 7 × 4x3 = 28x3. (Exponent down, minus one.)
  2. d/dx (−2x3) = −2 × 3x2 = −6x2. (Same move, keep the sign.)
  3. d/dx (x) = 1, because x = x1 gives 1 × x0 = 1. (Anything to the 0 is 1.)
  4. d/dx (−9) = 0. (A constant has no slope.)

Put the pieces together: f'(x) = 28x3 − 6x2 + 1. Handy check: each term dropped one degree, so a degree-4 polynomial gives a degree-3 derivative. Nicely done, that is the whole skill for polynomials.

Try it: differentiate f(x) = 5x3 − x2 + 4x − 8.

Answer: Term by term: 15x2 − 2x + 4 (the −8 becomes 0). You just did four derivatives at once.

Negative and fractional exponents (the step people skip)

Roots and reciprocals are secretly powers. Rewrite them as powers first, then use the same rule. This one habit is where most errors are avoided.

  • √x = x(1/2), so d/dx √x = (1/2) x(−1/2) = 1/(2√x).
  • 1/x = x(−1), so d/dx (1/x) = −1 × x(−2) = −1/x2.
  • 1/x3 = x(−3), so its derivative is −3 x(−4) = −3/x4.
  • A cube root x(1/3) has derivative (1/3) x(−2/3).

These match the definition-based answers from Module 3 exactly, but now they take seconds. The routine is always: rewrite every term as a power, apply the rule, then translate back to root or fraction form for a clean answer.

A mixed example

Differentiate f(x) = 4√x + 2/x − 3.

  1. Rewrite as powers: 4x(1/2) + 2x(−1) − 3. (The key setup step.)
  2. First term: 4 × (1/2) x(−1/2) = 2x(−1/2).
  3. Second term: 2 × (−1) x(−2) = −2x(−2).
  4. Constant: −3 gives 0.

Translate back: f'(x) = 2/√x − 2/x2. Rewrite, apply, translate back, and roots and reciprocals become routine.

Why the power rule is true

The power rule is not handed down from on high; it falls out of the definition you already own. Run the definition on f(x) = x3:

  1. Expand: (x + h)3 = x3 + 3x2h + 3xh2 + h3. (Binomial expansion.)
  2. Subtract x3: the top is 3x2h + 3xh2 + h3.
  3. Divide by h: 3x2 + 3xh + h2.
  4. Let h → 0: everything with an h dies, leaving 3x2.

Now look at the pattern instead of the particular. For any positive whole number n, the binomial theorem says (x + h)n = xn + n x(n−1) h + (terms each carrying h2 or higher). Subtract xn, divide by h, and every term still carrying an h vanishes in the limit. The lone survivor is n x(n−1). That is the power rule, and the reason the exponent "hops down": it counts how many ways the single h can pair with the remaining (n − 1) copies of x. Extending the rule to negative and fractional exponents takes the quotient and chain rules (coming soon), but the answers stay the same shape, which is why we state it for all real exponents now.

A full mixed derivative, start to finish

Differentiate f(x) = x3 + 5√x − 2/x2 + 7.

  1. Rewrite everything as powers: f(x) = x3 + 5x(1/2) − 2x(−2) + 7. (The non-negotiable setup step.)
  2. Term 1: 3x2. Term 2: 5 × (1/2)x(−1/2) = (5/2)x(−1/2). (Exponent down, minus one.)
  3. Term 3: −2 × (−2)x(−3) = +4x(−3). (Two negatives make the sign flip; this is the step to slow down for.)
  4. Term 4: the constant 7 gives 0.
  5. Translate back: f'(x) = 3x2 + 5/(2√x) + 4/x3.

Sanity check at x = 1: f'(1) = 3 + 2.5 + 4 = 9.5, a definite number with no algebra left over, which is what a finished derivative should produce.

Using the derivative immediately: a tangent line

Find the tangent line to f(x) = x3 − 4x at x = 2.

  1. Point: f(2) = 8 − 8 = 0, so the line passes through (2, 0).
  2. Slope: f'(x) = 3x2 − 4, so f'(2) = 12 − 4 = 8.
  3. Assemble: y − 0 = 8(x − 2), that is, y = 8x − 16.

Ten seconds of rules replaced the limit computation that took a whole page last module. That speed is the entire point of this module.

Derivatives of derivatives

Nothing stops you from differentiating again. For f(x) = x4: f'(x) = 4x3, f''(x) = 12x2, f'''(x) = 24x, f''''(x) = 24, and the fifth derivative is 0. Each pass drops the degree by one, so a degree-n polynomial dies after n + 1 differentiations. The second derivative f'' already has a physical name, acceleration, and it will carry the geometry of "bending" in Module 6. Higher derivatives power Taylor series in Calculus II, where they encode a function's entire local personality.

Why the helper rules are legal, and a first application

The helper rules inherit their truth from the limit laws of Module 2 in one line each. For the constant multiple rule: [c f(x + h) − c f(x)]/h = c × [f(x + h) − f(x)]/h, and the constant-multiple limit law lets the c ride outside the limit. For the sum rule, the sum limit law splits the quotient in two. Nothing about differentiation is new machinery; it is limit machinery wearing work clothes.

Application: finding flat spots. Where does f(x) = x3 − 3x have a horizontal tangent?

  1. Differentiate: f'(x) = 3x2 − 3. (Power and constant multiple rules, term by term.)
  2. Horizontal means slope zero: solve 3x2 − 3 = 0, so x2 = 1, giving x = 1 and x = −1.
  3. Locate the points: f(1) = 1 − 3 = −2 and f(−1) = −1 + 3 = 2. Flat tangents at (1, −2) and (−1, 2).

Those two flat spots are the crest and trough of the cubic's S-curve, and finding them by solving f'(x) = 0 is a two-line preview of the entire optimization machinery in Module 6.

Two places the power rule does not go

The most common error is using the power rule on 2x. It does not apply there. The power rule is only for a variable base raised to a constant power, like xn. When the variable is up in the exponent, as in 2x, that is an exponential function with its own rule, coming later this module. The second stumble is trying to differentiate √x without rewriting it. Always convert to x(1/2) first; there is no shortcut that skips that step.

Common misconceptions

  • "The power rule works on 2 to the x." No. That is an exponential; the power rule is for xn, a constant exponent.
  • "The derivative of a root is the root of the derivative." Rewrite √x as x(1/2) and apply the rule; the answer is 1/(2√x).
  • "Constants disappear only sometimes." Every plain constant differentiates to 0, always.
  • "You can differentiate a fraction like 1/x directly." Rewrite it as x(−1) first, then use the power rule.

The takeaway

The power rule, d/dx (xn) = n x(n − 1), brings the exponent down and drops it by one, for any real exponent. Constants differentiate to 0, constant factors ride along, and sums differentiate term by term, so any polynomial falls in one pass. For roots and reciprocals, rewrite as powers first, apply the rule, then translate back. This is the fast machinery that replaces the limit definition for everyday derivatives.

Sources

  1. OpenStax. (2016). 3.3 Differentiation rules. In Calculus volume 1. openstax.org
  2. Dawkins, P. (n.d.). Differentiation formulas. Paul's Online Math Notes. tutorial.math.lamar.edu
  3. Dawkins, P. (n.d.). Proof of various derivative properties. Paul's Online Math Notes. tutorial.math.lamar.edu
  4. Math is Fun. (n.d.). Derivative rules. mathsisfun.com
  5. Khan Academy. (n.d.). Calculus 1 [Online course]. khanacademy.org
  6. MIT OpenCourseWare. (2010). 1. Differentiation. 18.01SC Single variable calculus. Massachusetts Institute of Technology. ocw.mit.edu
  7. Sanderson, G. (3Blue1Brown). (n.d.). Higher order derivatives. 3blue1brown.com
Key terms
Power rule
d/dx (x^n) = n x^(n-1); drop the exponent in front and reduce it by one.
Constant rule
The derivative of any constant is 0.
Constant multiple rule
A constant factor passes through the derivative unchanged.
Sum rule
The derivative of a sum is the sum of the derivatives.
Exponent form
Rewriting roots and reciprocals as powers of x so the power rule applies.
Term-by-term differentiation
Differentiating a polynomial one term at a time using the sum rule.

The Product and Quotient Rules

  • Differentiate products of functions with the product rule.
  • Differentiate quotients with the quotient rule.
  • Recognize which rule a given expression requires.

Here is a guess almost everyone makes. The derivative of a sum is the sum of the derivatives, so surely the derivative of a product is the product of the derivatives. Test it on the simplest product there is: x × x. That is x2, whose derivative is 2x. The guess predicts 1 × 1 = 1. One counterexample, and the guess is finished. This page works out what is true instead, and why the correct rule has to carry two terms.

Why two factors need two terms

When two functions are multiplied or divided, they interact: nudge the input and both pieces move at once. The derivative has to account for each of them changing while the other holds still for a moment, which is why each of these rules comes in two parts. Learn the shape once and they become automatic.

What matters here: for a product, differentiate one factor at a time and add; for a quotient, do the same with a subtraction and divide by the bottom squared.

Why the naive guess fails

Suppose the derivative of a product were just the product of the derivatives. Test it on x × x = x2. The real derivative is d/dx (x2) = 2x. But the product of the separate derivatives is 1 × 1 = 1, which is wrong. One counterexample settles it: products interact, and the real rule captures that.

The product rule

For a product of two functions u and v:

d/dx (u v) = u' v + u v'

In words: "derivative of the first times the second, plus the first times derivative of the second." A helpful chant is exactly that sentence, repeated until it sticks. The two terms say that as a product grows, each factor contributes its own change while the other is held still for a moment.

Worked example. Differentiate f(x) = x2 sin(x).

  1. Label the parts: u = x2 so u' = 2x; and v = sin(x) so v' = cos(x). (Name everything before assembling.)
  2. Slot into the rule: u' v + u v' = 2x sin(x) + x2 cos(x).

So f'(x) = 2x sin(x) + x2 cos(x). Labeling u, u', v, v' first keeps the pieces from getting jumbled.

Try it: differentiate f(x) = x2 ex.

Answer: u = x2, u' = 2x, v = ex, v' = ex, so f'(x) = 2x ex + x2 ex. You handled the product rule.

The quotient rule

For a quotient u/v:

d/dx (u/v) = (u' v − u v') / v2

Order matters here because of the subtraction: "derivative of the top times the bottom, minus the top times derivative of the bottom, all over the bottom squared." A classic memory aid is "low d-high minus high d-low, over low-low," where "d" means "derivative of." Swap the order and you flip the sign of the whole answer, so the mnemonic is worth getting exactly right.

Worked example. Differentiate f(x) = x / (x2 + 1).

  1. Label: u = x so u' = 1; and v = x2 + 1 so v' = 2x.
  2. Slot in: (u' v − u v')/v2 = [1(x2 + 1) − x(2x)] / (x2 + 1)2.
  3. Simplify the top: x2 + 1 − 2x2 = 1 − x2.

So f'(x) = (1 − x2) / (x2 + 1)2. Keep the bottom squared, and do not cancel too early.

A second quotient example

Differentiate f(x) = (2x + 3)/(x − 1). With u = 2x + 3, u' = 2, v = x − 1, v' = 1, the top is 2(x − 1) − (2x + 3)(1) = 2x − 2 − 2x − 3 = −5. So f'(x) = −5/(x − 1)2. The constant top is a good reminder to expand and combine carefully before declaring the answer.

Which rule, and when to simplify first

Spotting the structure is half the job. Two things multiplied means the product rule; one thing divided by another means the quotient rule; only added means the plain sum rule is enough. And sometimes a little algebra dodges the harder rule entirely: (x2 + x)/x = x + 1 simplifies before you differentiate, and 3/x2 = 3x(−2) is easier by the power rule than by the quotient rule. Always pause and ask whether the expression simplifies before you commit to a rule.

Where the product rule comes from

The rule is provable in six lines from the definition, using one clever move: adding and subtracting the same quantity. Start with the difference quotient for u(x)v(x):

[u(x + h)v(x + h) − u(x)v(x)] / h

  1. Insert − u(x + h)v(x) + u(x + h)v(x) in the middle of the top. Adding zero changes nothing, but it splits the expression into two workable pieces. (The trick of the proof.)
  2. Group: u(x + h)[v(x + h) − v(x)]/h + v(x)[u(x + h) − u(x)]/h. (Each bracket is now a difference quotient we recognize.)
  3. Let h → 0. The first piece: u(x + h) → u(x) (differentiable functions are continuous) and the bracket over h heads to v'(x), giving u(x)v'(x).
  4. The second piece heads to v(x)u'(x). Sum: u'v + uv'. Done.

Read the two terms as a story: when both factors are nudged, the total change is (change in u) times v, plus u times (change in v); the tiny change-times-change piece dies in the limit. The quotient rule then comes free: write Q = u/v, so Qv = u. Differentiate both sides with the product rule: Q'v + Qv' = u'. Solve for Q': Q' = (u' − Qv')/v = (u' − (u/v)v')/v = (u'v − uv')/v2. The bottom-squared and the minus sign both fall out of the algebra; nothing is arbitrary.

One derivative, two roads (a built-in answer check)

Differentiate f(x) = (x2 + 1)(x3 − 3x) by the product rule, then again by expanding first.

  1. Product rule: u = x2 + 1, u' = 2x, v = x3 − 3x, v' = 3x2 − 3.
  2. Assemble: 2x(x3 − 3x) + (x2 + 1)(3x2 − 3).
  3. Expand each piece: 2x4 − 6x2 and 3x4 − 3x2 + 3x2 − 3 = 3x4 − 3. Sum: 5x4 − 6x2 − 3.
  4. Road two: expand first. f(x) = x5 − 3x3 + x3 − 3x = x5 − 2x3 − 3x, so f'(x) = 5x4 − 6x2 − 3.

Identical answers, as they must be. When a product of polynomials is small, expanding first is often less error-prone; when a factor is sin x or ex, expanding is impossible and the product rule is the only road. Having both roads means you can audit yourself on the polynomial cases until the rule feels trustworthy.

The reciprocal shortcut, and a sign-error autopsy

A quotient with a constant top is faster by pattern: d/dx [1/v] = −v'/v2 (that is the quotient rule with u = 1, u' = 0). So d/dx [1/(x2 + 1)] = −2x/(x2 + 1)2 in one step.

Now watch the classic sign error happen to f(x) = x/(x2 + 1) so you can recognize it later. The wrong order computes (uv' − u'v)/v2 = [x(2x) − 1(x2 + 1)]/(x2 + 1)2 = (x2 − 1)/(x2 + 1)2, exactly the negative of the true answer (1 − x2)/(x2 + 1)2. A negated derivative is poison downstream: it says the function is falling where it is rising, flips every increasing/decreasing conclusion, and reverses every max/min call in Module 6. Quick audit: at x = 0 the true derivative gives 1 > 0, and indeed x/(x2 + 1) rises through the origin. Test a point whenever you doubt your order.

Simplify first when algebra allows

Differentiate f(x) = (3x2 − 5)/x. The quotient rule works, but splitting the fraction is faster: f(x) = 3x − 5/x = 3x − 5x(−1), so f'(x) = 3 + 5x(−2) = 3 + 5/x2. Same answer, half the writing, and far fewer chances to drop a sign. The order of operations for a seasoned differentiator is: simplify, then choose the lightest rule that fits.

Reading a derivative after computing it

The answer f'(x) = (1 − x2)/(x2 + 1)2 for f(x) = x/(x2 + 1) is not just a formula to box; it is information. The bottom is always positive, so the sign of f' is the sign of 1 − x2: positive between −1 and 1, negative outside. So the curve rises through the middle and falls on both wings, with horizontal tangents at x = ±1, where f(1) = 1/2 and f(−1) = −1/2: a small crest and dip on an otherwise flattening curve. Ten seconds of sign-reading sketched the graph. Building the habit of interrogating each derivative you compute, where is it zero, where positive, is what turns the rules of this module into the applications of Module 6.

The wrong product, and the flipped quotient

The most common error is the naive one we disproved: writing d/dx (u v) = u' v'. It is wrong; the correct rule has two terms, u' v + u v'. The second common error is the quotient-rule order. It is top-derivative first: u' v − u v'. Getting it backward negates your answer. When in doubt, test your memory on a tiny example like x/1 or x2/x, where you already know the answer.

Common misconceptions

  • "The derivative of a product is the product of the derivatives." False. It is u' v + u v'. Counterexample: d/dx (x × x) = 2x, not 1.
  • "Order does not matter in the quotient rule." The top derivative comes first; swapping negates the whole answer.
  • "You must use the quotient rule on every fraction." Often you can simplify first, or rewrite as a negative power and use the power rule.
  • "The v squared is optional." The bottom squared is part of the quotient rule; leaving it off changes the answer.

Looking back

Products and quotients interact, so their rules have two parts. Product rule: d/dx (u v) = u' v + u v'. Quotient rule: d/dx (u/v) = (u' v − u v')/v2, top-derivative first, all over the bottom squared. Label u, u', v, v' before you assemble, watch the quotient order, and simplify with algebra when you can to avoid extra work. Those habits make these rules dependable.

Sources

  1. OpenStax. (2016). 3.3 Differentiation rules. In Calculus volume 1. openstax.org
  2. Dawkins, P. (n.d.). Product and quotient rule. Paul's Online Math Notes. tutorial.math.lamar.edu
  3. Dawkins, P. (n.d.). Proof of various derivative properties. Paul's Online Math Notes. tutorial.math.lamar.edu
  4. Math is Fun. (n.d.). Derivative rules. mathsisfun.com
  5. Sanderson, G. (3Blue1Brown). (n.d.). Visualizing the chain rule and product rule. 3blue1brown.com
  6. Khan Academy. (n.d.). Calculus 1 [Online course]. khanacademy.org
  7. MIT OpenCourseWare. (2010). 1. Differentiation. 18.01SC Single variable calculus. Massachusetts Institute of Technology. ocw.mit.edu
Key terms
Product rule
d/dx (u v) = u'v + u v'.
Quotient rule
d/dx (u/v) = (u'v - u v')/v^2.
Factor
One of the functions being multiplied in a product.
Numerator and denominator
The top (u) and bottom (v) of a quotient being differentiated.
Rule selection
Diagnosing whether an expression is a product, quotient, or sum to pick the right rule.

The Chain Rule

  • Recognize a composite function as an outer function wrapped around an inner one.
  • Apply the chain rule to differentiate composites.
  • Combine the chain rule with the power, product, and quotient rules.

Differentiate (3x2 + 1)4. The power rule seems to answer 4(3x2 + 1)3, and that answer is wrong. The true derivative is 24x(3x2 + 1)3, six times x larger. That missing 6x is the derivative of what sits inside the brackets, and leaving it out is the single most common slip in all of differential calculus. This page is about that factor: where it comes from, why it cannot be dropped, and how to spot every place it is owed.

Layers, and the factor people drop

A composite function is one machine feeding another: an inner function does its job first, and an outer function acts on the result. The chain rule differentiates the outer and the inner separately, then multiplies. That extra inner factor is the whole point, and forgetting it is the most common slip in all of differential calculus.

Worth holding on to: derivative of the outside (leaving the inside alone), times the derivative of the inside.

The rule

If y = f(g(x)), then:

d/dx f(g(x)) = f'(g(x)) × g'(x)

Read the pieces: f'(g(x)) is "differentiate the outer function but keep the inner one inside it," and g'(x) is "the derivative of the inner function." People often say "derivative of the outside, times the derivative of the inside." That short phrase is the entire rule.

Why the extra factor makes sense

Think in rates, like gears. Suppose y changes 3 times as fast as u, and u changes 2 times as fast as x. Then y changes 3 × 2 = 6 times as fast as x. The rates simply multiply. In Leibniz notation this reads dy/dx = (dy/du)(du/dx), which is exactly why the inner derivative has to appear.

Let us do one together, slowly: a power of a bundle

Differentiate f(x) = (3x2 + 1)4.

  1. Spot the pieces: the outer is ( )4 and the inner is 3x2 + 1. (Name inside and outside.)
  2. Differentiate the outer, keeping the inside: 4(3x2 + 1)3. (Power rule on the bundle.)
  3. Differentiate the inner: d/dx (3x2 + 1) = 6x. (The extra factor.)
  4. Multiply: 4(3x2 + 1)3 × 6x = 24x (3x2 + 1)3.

So f'(x) = 24x (3x2 + 1)3. If you forgot the 6x, the answer would be wrong; that inner factor is the reason the chain rule exists.

Try it: differentiate f(x) = (2x + 1)10.

Answer: Outer derivative 10(2x + 1)9 times inner derivative 2, giving 20(2x + 1)9. Nicely spotted.

A trig composite

Differentiate f(x) = sin(x2).

  1. Outer is sin( ), inner is x2. (Identify the nest.)
  2. Derivative of the outer, keeping the inside: cos(x2).
  3. Times the inner derivative 2x: cos(x2) × 2x = 2x cos(x2).

Compare this with (sin x)2, which is a different nest (a power of sine) and gives 2 sin(x) cos(x). Reading which part is inside and which is outside changes the answer completely, so slow down and look.

A square root

Differentiate f(x) = √(x2 + 1).

  1. Rewrite as a power: (x2 + 1)(1/2). (Roots are powers.)
  2. Outer derivative: (1/2)(x2 + 1)(−1/2).
  3. Times inner derivative 2x: (1/2)(x2 + 1)(−1/2) × 2x = x / √(x2 + 1).

Combining with the other rules

The chain rule teams up with the product and quotient rules, and nested composites use it more than once. For f(x) = sin(cos(x)), the outer is sine and the inner is cos(x): cos(cos(x)) × (−sin(x)) = −sin(x) cos(cos(x)).

For f(x) = (2x + 1)3 ex, use the product rule with one factor needing the chain rule: u = (2x + 1)3 gives u' = 3(2x + 1)2 × 2 = 6(2x + 1)2, and v = ex, so f'(x) = 6(2x + 1)2 ex + (2x + 1)3 ex. Whenever something sits inside something else, a power of a bundle, a trig of a bundle, a root of a bundle, the chain rule applies.

Why composition multiplies rates

Here is the chain rule seen honestly, not as a formula but as a fact about compounding sensitivities. Suppose y depends on u, and u depends on x. Nudge x by a tiny amount dx. The middle variable moves by roughly du = (du/dx) × dx. That movement in u then moves y by roughly dy = (dy/du) × du. Substitute the first into the second: dy = (dy/du)(du/dx) × dx. The sensitivity of y to x is the product of the two sensitivities, because the second machine amplifies whatever the first machine hands it. Gears mesh the same way: if the middle gear turns twice per turn of the crank, and the last gear turns three times per turn of the middle, the last gear turns six times per crank. One honest caution: treating dy/du × du/dx as fractions that cancel is a memory aid, not a proof, because each is a limit; the real proof patches the case where du is momentarily zero. But the intuition it captures is exactly right.

A concrete rate chain with numbers

A spherical balloon is inflated so its radius grows as r(t) = 2t cm after t seconds, and volume is V = (4/3)πr3. How fast is the volume growing at t = 3?

  1. Outer rate: dV/dr = 4πr2. (How volume responds to radius.)
  2. Inner rate: dr/dt = 2 cm/s. (How radius responds to time.)
  3. Chain them: dV/dt = (dV/dr)(dr/dt) = 4πr2 × 2 = 8πr2.
  4. At t = 3, r = 6, so dV/dt = 8π(36) = 288π, about 905 cm3/s.

Notice the same steady 2 cm/s of radius produces ever more volume per second as the balloon grows; the outer rate depends on where you are. This computation is the entire engine of Module 5's related rates.

Three layers deep

Differentiate f(x) = sin3(2x), which means [sin(2x)]3. Peel from the outside in.

  1. Outermost layer: a cube. Derivative: 3[sin(2x)]2, insides untouched. (Layer one.)
  2. Middle layer: sine. Derivative: cos(2x), its own inside untouched. (Layer two.)
  3. Innermost layer: 2x. Derivative: 2. (Layer three.)
  4. Multiply all three: f'(x) = 3 sin2(2x) × cos(2x) × 2 = 6 sin2(2x) cos(2x).

The pattern generalizes: one factor per layer, outermost first, and you are finished only when the innermost derivative (here the plain 2) has been collected. Writing the layers as a vertical list before multiplying prevents nearly all omissions.

The chain rule explains the exponential rule

The next lesson will hand you the rule d/dx (ax) = ax ln a. Where does that ln a come from? From the chain rule, and you can watch it appear right now. Any base can be rewritten on base e: ax = ex ln a, because eln a = a. Now differentiate the right side as a composite: outer e( ) keeps itself, times the inner derivative d/dx (x ln a) = ln a. Result: ex ln a × ln a = ax ln a. The rule arrives already explained, which is the usual situation in calculus: the shortcut rules are compressed proofs.

Chain plus quotient in one line

Differentiate f(x) = [x/(x + 1)]2. Outer square: 2[x/(x + 1)]. Inner quotient: [(1)(x + 1) − x(1)]/(x + 1)2 = 1/(x + 1)2. Multiply: f'(x) = 2x/(x + 1)3. Two rules, one factor each, assembled without drama, which is what fluency looks like.

Three fast reps to make it automatic

Fluency comes from reps where you name the layers out loud. Cover the answers and try these.

  1. d/dx (5x2 − 3x)7: outer power, inner polynomial. Answer: 7(5x2 − 3x)6 × (10x − 3).
  2. d/dx cos3(x), meaning (cos x)3: outer cube, inner cosine. Answer: 3 cos2(x) × (−sin x) = −3 cos2(x) sin(x).
  3. d/dx √(9 − x2): outer half-power, inner 9 − x2. Answer: (1/(2√(9 − x2))) × (−2x) = −x/√(9 − x2).

That last one is worth a second look: the graph of √(9 − x2) is the top half of a circle of radius 3, and the formula says the slope is negative x over y, exactly what implicit differentiation will re-derive for circles next module from a completely different direction. When two methods you trust give one answer, that is calculus quietly auditing itself.

Stopping too soon, and reading the nest backwards

The classic mistake is stopping too soon, writing d/dx (3x2 + 1)4 = 4(3x2 + 1)3 and forgetting the inner derivative 6x. Build a tiny habit: after you differentiate the outside, immediately ask "and what is the derivative of the inside?" and multiply by it. The second mistake is confusing sin(x2) with (sin x)2. Ask which operation happens last: in sin(x2) you square first then take sine; in (sin x)2 you take sine first then square. The outer function is whichever happens last.

Common misconceptions

  • "Just apply the power rule to the outside." That omits the inner derivative. d/dx (3x2 + 1)4 = 24x(3x2 + 1)3, not 4(3x2 + 1)3.
  • "sin(x^2) and (sin x)^2 have the same derivative." They are different nests: 2x cos(x2) versus 2 sin(x) cos(x).
  • "The chain rule is a rare special case." It is the most used rule; most realistic functions are composites.
  • "You cancel the du like a fraction to prove it." That is a memory aid, not a proof; dy/du and du/dx are limits, not literal fractions.

What to remember

The chain rule differentiates composites: d/dx f(g(x)) = f'(g(x)) g'(x), the derivative of the outside (inside left alone) times the derivative of the inside. It comes from rates multiplying, like gears. Spot the inner and outer, differentiate the outer keeping the inside, then multiply by the inner derivative, and never skip that last factor. Combined with the power, product, and quotient rules, this lets you differentiate essentially any function in a first calculus course.

Sources

  1. OpenStax. (2016). 3.6 The chain rule. In Calculus volume 1. openstax.org
  2. Dawkins, P. (n.d.). Chain rule. Paul's Online Math Notes. tutorial.math.lamar.edu
  3. Sanderson, G. (3Blue1Brown). (n.d.). Visualizing the chain rule and product rule. 3blue1brown.com
  4. Math is Fun. (n.d.). Derivative rules. mathsisfun.com
  5. OpenStax. (2016). 3.9 Derivatives of exponential and logarithmic functions. In Calculus volume 1. openstax.org
  6. Khan Academy. (n.d.). Calculus 1 [Online course]. khanacademy.org
  7. MIT OpenCourseWare. (2010). 1. Differentiation. 18.01SC Single variable calculus. Massachusetts Institute of Technology. ocw.mit.edu
Key terms
Chain rule
d/dx f(g(x)) = f'(g(x)) g'(x); outer derivative times inner derivative.
Composite function
A function formed by nesting one function inside another, f(g(x)).
Outer function
The function applied last, differentiated first in the chain rule.
Inner function
The function applied first, whose derivative is the extra chain-rule factor.
Nested composite
A composition inside a composition, requiring the chain rule more than once.

Derivatives of Trigonometric, Exponential, and Logarithmic Functions

  • State and apply the derivatives of the six basic trig functions.
  • Differentiate exponential functions, including e^x and a^x.
  • Differentiate logarithmic functions, including ln x and log base a.

At x = 2 the curve ex stands 7.389 units high, and its slope there is also 7.389. At x = 5 the height is 148.41, and so is the slope. Height equals steepness at every single point, which no polynomial and no root ever manages, and it is why e turns up in every model where growth is proportional to current size. This lesson collects that fact and a handful of others into a short list you will use for the rest of the course.

A short list, then the chain rule does the rest

You do not rederive these every time. You memorize a short list, then let the chain rule handle anything nested inside them. Think of this lesson as stocking your toolbox with a few reliable, named tools.

Key idea: learn a small set of base derivatives, then reach for the chain rule for the inside of any composite.

Trig derivatives

Built from the special limit sin(x)/x → 1, the core results are:

FunctionDerivative
sin(x)cos(x)
cos(x)−sin(x)
tan(x)sec2(x)
cot(x)−csc2(x)
sec(x)sec(x) tan(x)
csc(x)−csc(x) cot(x)

Here is a friendly pattern to lean on: the three "co" functions (cos, cot, csc) all carry a minus sign. Only sine and cosine are truly fundamental; the other four are ratios of them and follow by the quotient rule. With the chain rule, d/dx cos(5x) = −sin(5x) × 5 = −5 sin(5x), and d/dx tan(x2) = sec2(x2) × 2x = 2x sec2(x2).

Exponential derivatives

The function ex has a remarkable property: it is its own derivative.

d/dx (ex) = ex

Read that once more, because it is genuinely special. At every point, the slope of ex equals its own height. That is the defining feature of the number e (about 2.718). For a general base a, an extra factor of ln a shows up: d/dx (ax) = ax ln a. So d/dx (2x) = 2x ln 2. With the chain rule, d/dx (e3x) = e3x × 3 = 3 e3x, and d/dx e(x2) = e(x2) × 2x = 2x e(x2). This self-copying behavior is why ex models anything whose growth rate matches its current size.

Logarithmic derivatives

The natural logarithm has a beautifully simple derivative:

d/dx (ln x) = 1/x

For a general base, d/dx (loga x) = 1/(x ln a). With the chain rule, the log of a bundle gives the inner derivative over the inside: d/dx ln(g(x)) = g'(x)/g(x). For example, d/dx ln(x2 + 1) = 2x/(x2 + 1). That "inner derivative over the inside" pattern shows up so often it is worth recognizing instantly.

Try it: differentiate f(x) = 3 sin(x) + ex − ln(x).

Answer: Straight from the list: f'(x) = 3 cos(x) + ex − 1/x. You just used three of the new rules in one line.

Combining rules

Differentiate f(x) = x2 ex with the product rule.

  1. Label: u = x2 so u' = 2x; v = ex so v' = ex.
  2. Product rule: 2x ex + x2 ex, which factors to x ex (2 + x).

And f(x) = ex sin(x): with u = ex, u' = ex, v = sin(x), v' = cos(x), the product rule gives ex sin(x) + ex cos(x) = ex(sin x + cos x). With this short list plus the four rules from the module, you can differentiate essentially any function a first course will ask.

Where sin' = cos comes from

The sine rule is not folklore; it falls out of the limit definition plus the two special limits from Module 2. Run the definition on f(x) = sin x:

  1. Difference quotient: [sin(x + h) − sin x]/h.
  2. Apply the angle-addition identity sin(x + h) = sin x cos h + cos x sin h. (Precalculus, earning its keep.)
  3. Regroup: sin x (cos h − 1)/h + cos x (sin h)/h. (Everything depending on h is isolated in two famous quotients.)
  4. Let h → 0: from Module 2, (cos h − 1)/h → 0 and (sin h)/h → 1. The survivors: sin x × 0 + cos x × 1 = cos x.

So d/dx sin x = cos x, and the same computation with cosine's addition formula gives d/dx cos x = −sin x. This derivation is also exactly why radians are mandatory: the limit (sin h)/h → 1 is a radian fact. In degrees that limit is π/180, and the ugly constant would infect every formula.

The other four, by quotient rule

Tangent is sin x / cos x, so the quotient rule gives:

  1. d/dx tan x = [cos x × cos x − sin x × (−sin x)] / cos2 x. (Top-derivative times bottom, minus top times bottom-derivative.)
  2. Top: cos2 x + sin2 x = 1. (The Pythagorean identity collapses it.)
  3. So d/dx tan x = 1/cos2 x = sec2 x.

The secant, cosecant, and cotangent rules all fall the same way; none needs separate memorization so much as five minutes of quotient-rule practice. Knowing the table and where it comes from means a forgotten entry is recoverable rather than fatal.

Why e is the natural base

Run the definition on f(x) = ex: the quotient is [ex + h − ex]/h = ex × (eh − 1)/h, using the exponent law ex + h = exeh. Everything hinges on the number (eh − 1)/h as h shrinks. Watch it numerically:

h0.10.010.001
(eh − 1)/h1.05171.005021.0005

The quotient heads to exactly 1, and that is not a coincidence: e is defined as the one base that makes this limit equal 1. For base 2 the analogous limit is ln 2 ≈ 0.693; for base 10 it is ln 10 ≈ 2.303. So every exponential is its own derivative times a base-dependent constant, and e is the base whose constant is 1. That is the entire reason e is "natural," and why scientists write growth laws in base e.

The logarithm inherits its rule from the exponential

Since y = ln x means ey = x, the two graphs are mirror images across the line y = x, and their slopes are reciprocals at mirrored points. Differentiating ey = x (treating y as a function of x, a move next lesson makes routine) gives ey × dy/dx = 1, so dy/dx = 1/ey = 1/x. The general base follows from the change-of-base identity loga x = ln x / ln a: divide the derivative by the constant ln a to get 1/(x ln a).

Worked combo. Differentiate f(x) = x2 ln x. Product rule: u = x2, u' = 2x, v = ln x, v' = 1/x, so f'(x) = 2x ln x + x2(1/x) = 2x ln x + x. At x = 1: f'(1) = 0 + 1 = 1, a positive slope, matching the graph rising through the point (1, 0).

Two consistency checks worth internalizing

Check one: a famous tangent line. The tangent to y = ex at x = 0 passes through (0, 1) with slope e0 = 1, giving the line y = 1 + x. That is the workhorse approximation ex ≈ 1 + x for small x, used constantly in physics and finance: 3 percent growth compounds to roughly e0.03 ≈ 1.03. The full curve always sits above this line (it is concave up everywhere), so the approximation slightly undershoots, and knowing the direction of the error is half the value.

Check two: differentiating an identity. The Pythagorean identity says sin2x + cos2x = 1 for every x. A constant has derivative zero, so differentiating the left side must give zero. Try it with the chain rule: 2 sin x cos x + 2 cos x (−sin x) = 0. It does, identically. Running your derivative rules against known identities like this is a free error detector: if the check had failed, one of the memorized formulas would have to be wrong. Mathematicians audit themselves this way all the time, and you should too.

Not a power, and not in degrees

The most common error is using the power rule on ex, writing something like x e(x−1). That is wrong. The power rule is for a variable base with a constant exponent, like x2; here the variable is up in the exponent, so it is exponential, and d/dx (ex) = ex. The second trap is using trig derivatives in degrees. They hold only in radians, because the whole thing rests on sin(x)/x → 1, which is a radian fact. In degrees, an ugly factor of π/180 would clutter every formula.

Common misconceptions

  • "d/dx (e^x) = x e^(x-1)." No. That is the power rule, which does not apply. ex is its own derivative.
  • "Trig derivative formulas work in degrees." Only in radians. Radian measure is what makes d/dx sin(x) = cos(x) clean.
  • "The derivative of ln x is complicated." It is just 1/x, one of the simplest results in calculus.
  • "a^x and e^x differentiate the same way." ex gives ex; a general base gives an extra ln a factor.

Recap

Memorize a short list: sine gives cosine, cosine gives negative sine, tangent gives secant squared (the "co" functions carry minus signs). The exponential ex is its own derivative, a general base gives ax ln a, and the natural log gives 1/x, with ln(g(x)) giving g'(x)/g(x). Use the chain rule for anything nested and the product or quotient rule for combinations. Keep everything in radians. You now have derivatives for every elementary function.

Sources

  1. OpenStax. (2016). 3.5 Derivatives of trigonometric functions. In Calculus volume 1. openstax.org
  2. OpenStax. (2016). 3.9 Derivatives of exponential and logarithmic functions. In Calculus volume 1. openstax.org
  3. OpenStax. (2016). 3.7 Derivatives of inverse functions. In Calculus volume 1. openstax.org
  4. Dawkins, P. (n.d.). Derivatives of trig functions. Paul's Online Math Notes. tutorial.math.lamar.edu
  5. Dawkins, P. (n.d.). Derivatives of exponential and logarithm functions. Paul's Online Math Notes. tutorial.math.lamar.edu
  6. Math is Fun. (n.d.). Derivative rules. mathsisfun.com
  7. Khan Academy. (n.d.). Calculus 1 [Online course]. khanacademy.org
Key terms
Derivative of sine
d/dx sin(x) = cos(x).
Derivative of cosine
d/dx cos(x) = -sin(x); the co-functions carry minus signs.
Derivative of e^x
e^x is its own derivative.
General exponential
d/dx (a^x) = a^x ln a.
Derivative of ln x
d/dx (ln x) = 1/x.
Logarithmic differentiation
Taking ln of both sides to differentiate messy products or variable exponents.

Module 5: Implicit Differentiation and Related Rates

Differentiate equations that are not solved for y, then use derivatives to link the rates of change of connected quantities as they evolve in time. These two techniques turn the chain rule into a tool for real applied problems, from tangents on a circle to a ladder sliding down a wall.

Implicit Differentiation

  • Explain when implicit differentiation is needed.
  • Differentiate both sides of an equation, treating y as a function of x.
  • Solve for dy/dx and evaluate a slope on a curve.

The circle x2 + y2 = 25 passes through the point (3, 4). What is its slope there? Every method so far wants a function y = f(x), and this equation refuses to be one: above x = 3 the circle has two points, at height 4 and at height −4, so no single rule assigns one output to that input. You could split it into two half-circles full of square roots and differentiate each. Implicit differentiation gets the slope in three lines without splitting anything, and it is really the chain rule wearing a different hat.

Treating y as a hidden function of x

When an equation mixes x and y and you cannot isolate y, you do not have to. You treat y as a hidden function of x, differentiate the whole equation, and solve for the slope dy/dx at the end.

The point: pretend y is secretly a function of x. Every time you differentiate a y-term, the chain rule tacks on a dy/dx.

The one new habit

Treat y as a function of x, even though you never write that function down. Differentiate both sides of the equation with respect to x. Whenever you differentiate a term with y in it, the chain rule attaches a dy/dx factor, because y is the inner function. So d/dx (y2) = 2y × dy/dx, not just 2y. Then solve the resulting equation for dy/dx. That extra dy/dx factor is the whole method in a nutshell.

Let us do the circle together, slowly

Find dy/dx for x2 + y2 = 25.

  1. Differentiate both sides with respect to x. (Same operation on both sides, keeping balance.)
  2. The x2 term: d/dx (x2) = 2x. (Ordinary power rule.)
  3. The y2 term: d/dx (y2) = 2y × dy/dx. (Chain rule, because y hides a function of x.)
  4. The right side: d/dx (25) = 0. (A constant.)
  5. Put it together: 2x + 2y (dy/dx) = 0. Solve: 2y (dy/dx) = −2x, so dy/dx = −x/y.

That one formula gives the slope anywhere on the circle. At (3, 4) the slope is −3/4; at (3, −4) it is 3/4. It makes sense that the slope depends on both coordinates, since a circle has an upper point and a lower point for most x values, with opposite slopes.

Try it: using dy/dx = −x/y, what is the slope of the circle at (3, 4)?

Answer: −x/y = −3/4. You read a slope off a curve you could not even solve for. That is the power of the method.

A mixed term (where care pays off)

Find dy/dx for x2 + xy + y2 = 7.

  1. d/dx (x2) = 2x. (Power rule.)
  2. The middle term needs the product rule (it is x times y): d/dx (xy) = 1 × y + x × dy/dx = y + x (dy/dx).
  3. d/dx (y2) = 2y (dy/dx). (Chain rule.)
  4. Combine: 2x + y + x (dy/dx) + 2y (dy/dx) = 0.
  5. Gather the dy/dx terms: (x + 2y)(dy/dx) = −(2x + y), so dy/dx = −(2x + y)/(x + 2y).

The xy term is the spot that rewards attention: it is a product of two things that both depend on x, so the product rule applies, and its second piece carries the dy/dx.

The reliable recipe

  • Differentiate every term on both sides with respect to x.
  • Attach a dy/dx whenever you differentiate a y-term (chain rule), and use the product rule for any xy mix.
  • Collect all dy/dx terms on one side, everything else on the other.
  • Factor out dy/dx and divide to solve.

A famous curve, start to finish

The folium of Descartes, x3 + y3 = 6xy, is a looping curve no one can solve for y. Find the tangent line at the point (3, 3). First confirm the point is on the curve: 27 + 27 = 54 and 6(3)(3) = 54. Good.

  1. Differentiate every term with respect to x: 3x2 + 3y2(dy/dx) = 6y + 6x(dy/dx). (Left side: power rule, then chain rule on the y-term. Right side: product rule on 6xy, giving 6y + 6x(dy/dx).)
  2. Collect the dy/dx terms on one side: 3y2(dy/dx) − 6x(dy/dx) = 6y − 3x2.
  3. Factor and divide: dy/dx = (6y − 3x2)/(3y2 − 6x) = (2y − x2)/(y2 − 2x). (Canceling the common 3.)
  4. Evaluate at (3, 3): dy/dx = (6 − 9)/(9 − 6) = −1.
  5. Tangent line: y − 3 = −1(x − 3), that is, y = −x + 6.

A slope of exactly −1 at the symmetric point is no accident: the folium is mirror-symmetric across the line y = x, and the tangent at the point on that mirror must be perpendicular to it. When geometry and algebra agree like this, you can trust the computation.

The second derivative, implicitly

Implicit curves have concavity too. For the circle x2 + y2 = 25 we found dy/dx = −x/y. Differentiate that quotient again with respect to x, remembering y is still a function of x:

  1. Quotient rule on −x/y: d2y/dx2 = −[1 × y − x × (dy/dx)]/y2.
  2. Substitute the known dy/dx = −x/y: the bracket becomes y − x(−x/y) = y + x2/y = (y2 + x2)/y. (Common denominator.)
  3. So d2y/dx2 = −(x2 + y2)/y3 = −25/y3, using the original equation to replace x2 + y2 with 25.

Read the answer: on the upper semicircle (y > 0) the second derivative is negative, concave down, and on the lower semicircle it is positive, concave up. That is exactly how a circle bends. Two habits appeared here that mark clean implicit work: substitute the first derivative back in, and simplify with the original equation whenever it shows up.

Logarithmic differentiation: for functions the rules refuse

What is the derivative of y = xx? The power rule is illegal (the exponent is not constant) and the exponential rule is illegal (the base is not constant). The trick: take the natural log of both sides first, then differentiate implicitly.

  1. ln y = ln(xx) = x ln x. (The log pulls the exponent down, its superpower.)
  2. Differentiate both sides: left side (1/y)(dy/dx) by the chain rule; right side, product rule: 1 × ln x + x × (1/x) = ln x + 1.
  3. Solve: dy/dx = y(ln x + 1) = xx(ln x + 1).

A bonus insight falls out free: the derivative is zero where ln x = −1, at x = 1/e ≈ 0.368, which is where xx bottoms out (its minimum value is about 0.692). Logarithmic differentiation also tames big products: to differentiate y = x2√(x + 1)/(x + 2)3, take logs, turn the mess into a sum 2 ln x + (1/2)ln(x + 1) − 3 ln(x + 2), differentiate term by term, and multiply back by y at the end. It converts multiplication problems into addition problems, which is what logarithms were invented for.

An ellipse, for contrast

The method is identical for any conic. Find the tangent to the ellipse x2 + 4y2 = 8 at the point (2, 1). First confirm the point: 4 + 4 = 8. Good.

  1. Differentiate both sides: 2x + 8y(dy/dx) = 0. (Chain rule supplies the dy/dx on the y-term.)
  2. Solve: dy/dx = −2x/(8y) = −x/(4y).
  3. Evaluate at (2, 1): dy/dx = −2/4 = −1/2.
  4. Tangent line: y − 1 = −(1/2)(x − 2), that is, y = −x/2 + 2.

Compare the circle's slope formula −x/y with the ellipse's −x/(4y): the squashed geometry shows up as the 4, flattening the slopes. Implicit differentiation handles every curve in the conic family, circles, ellipses, parabolas, hyperbolas, with the same four moves, which is why it is the standard tool for tangents to orbits and lenses, where these curves actually live.

The missing dy/dx

The defining mistake is differentiating y2 as just 2y, forgetting the dy/dx. Remember: y is a hidden function of x, so it is an inner function and the chain rule always attaches its derivative. A small ritual helps: every time you differentiate a y-term, immediately write "times dy/dx" before moving on. The second stumble is thinking you must solve for y first. You never do; that is the entire point of the technique.

Common misconceptions

  • "d/dx of y^2 is 2y." No. Since y depends on x, it is 2y × dy/dx.
  • "You must solve the equation for y before differentiating." The whole point is that you do not; you differentiate as is and solve for dy/dx.
  • "The xy term just uses the power rule." It is a product of x and y, so it needs the product rule, giving y + x(dy/dx).
  • "A slope formula that divides by zero is a mistake." Often it is honest: at a circle's far left and right edges, y = 0, the tangent is vertical, and the formula rightly blows up.

What to carry forward

Implicit differentiation finds dy/dx for curves you cannot solve for y. Treat y as a hidden function of x, differentiate every term with respect to x, attach a dy/dx to each y-term via the chain rule, use the product rule for xy mixes, then collect and solve. For the circle x2 + y2 = 25, the slope is dy/dx = −x/y. This same idea powers related rates next lesson, where the hidden variable is time.

Sources

  1. OpenStax. (2016). 3.8 Implicit differentiation. In Calculus volume 1. openstax.org
  2. Dawkins, P. (n.d.). Implicit differentiation. Paul's Online Math Notes. tutorial.math.lamar.edu
  3. Dawkins, P. (n.d.). Logarithmic differentiation. Paul's Online Math Notes. tutorial.math.lamar.edu
  4. Math is Fun. (n.d.). Implicit differentiation. mathsisfun.com
  5. Sanderson, G. (3Blue1Brown). (n.d.). Implicit differentiation, what's going on here? 3blue1brown.com
  6. Khan Academy. (n.d.). Calculus 1 [Online course]. khanacademy.org
  7. MIT OpenCourseWare. (2010). 1. Differentiation. 18.01SC Single variable calculus. Massachusetts Institute of Technology. ocw.mit.edu
Key terms
Implicit equation
An equation relating x and y that is not solved for y.
Implicit differentiation
Differentiating both sides of an equation while treating y as a function of x.
dy/dx factor
The chain-rule term attached whenever a y-containing expression is differentiated.
Slope on a curve
The value of dy/dx at a specific point, which may depend on both x and y.
Vertical tangent
A point where dy/dx is undefined because the denominator is zero, as at a circle's left and right edges.

Module 6: Applications and an Introduction to the Integral

Use derivatives to find extrema, analyze the shape of curves, and solve real optimization problems, then reverse differentiation with antiderivatives to open the door to integration. This capstone module turns every earlier skill into a problem-solving toolkit and previews Calculus II.

Extrema and the First Derivative Test

  • Find critical points where the derivative is zero or undefined.
  • Classify critical points as local maxima or minima with the first derivative test.
  • Distinguish local from absolute extrema on a closed interval.
  • State the Mean Value Theorem and use its corollary to justify the sign-of-f' reasoning.

The curve y = x2/3 has an obvious lowest point. It sits at 0 when x = 0 and climbs on both sides. Now hunt for that minimum the usual way, by solving f'(x) = 0, and you will find nothing at all: the derivative 2/(3x1/3) is never zero anywhere. The minimum is real; the search was incomplete. This lesson lays out the full method for finding peaks and valleys, including the ones no horizontal tangent will ever point to.

Peaks, valleys, and the only places they can hide

Picture a hilly landscape. A hilltop is a local maximum, higher than the ground right around it; a valley bottom is a local minimum. The very highest and lowest points over a whole stretch are the absolute extrema. At a smooth peak or valley, the ground is momentarily flat, which is exactly what the derivative detects.

In short: at a smooth peak or valley the tangent is flat, so hunt where f'(x) = 0 (or is undefined), then test which kind it is.

Critical points: the only candidates

Because the tangent is flat at a smooth peak or valley, extrema can only happen where f'(x) = 0 or where f'(x) is undefined. Those inputs are called critical points. They are the suspects. Not every suspect is guilty, though, which is the content of Fermat's Theorem: if f has a local extremum at a point and the derivative exists there, then the derivative is 0. The reverse is not guaranteed, so we screen every critical point with a test.

The first derivative test

To classify a critical point, look at the sign of f' just to its left and just to its right. The sign tells you whether the function is rising or falling on each side:

  • f' goes from positive to negative: the curve rises then falls, a local maximum (a hilltop).
  • f' goes from negative to positive: the curve falls then rises, a local minimum (a valley).
  • f' does not change sign: neither, just a flat pause on a curve that keeps going the same way, like x3 at 0.

Let us classify one together, slowly

Find and classify the critical points of f(x) = x3 − 3x2 + 1.

  1. Differentiate: f'(x) = 3x2 − 6x = 3x(x − 2). (Factor to find the zeros.)
  2. Set it to zero: 3x(x − 2) = 0, so the critical points are x = 0 and x = 2. (The suspects.)
  3. Test the sign on each side. Left of 0 (try x = −1): 3(−1)(−3) = 9, positive, rising. Between 0 and 2 (try x = 1): 3(1)(−1) = −3, negative, falling. Right of 2 (try x = 3): 3(3)(1) = 9, positive, rising.
  4. Read the changes: at x = 0, plus to minus, a local maximum with f(0) = 1. At x = 2, minus to plus, a local minimum with f(2) = −3.

A sign chart, a number line marked with the critical points and the sign of f' in each gap, keeps all this organized. Draw one every time; it prevents mistakes.

Try it: find and classify the critical points of f(x) = x2 − 4x + 1.

Answer: f'(x) = 2x − 4 = 0 gives x = 2. The sign of f' goes minus to plus, so it is a local minimum, with f(2) = −3. Nicely done.

Absolute extrema on a closed interval

On a closed interval [a, b], find the absolute highest and lowest with the closed interval method: evaluate f at every critical point inside the interval and at both endpoints, then pick the largest and smallest outputs. Endpoints matter, because the top or bottom value might sit at an edge rather than a peak. The Extreme Value Theorem promises this works: a function continuous on a closed interval always attains both an absolute maximum and an absolute minimum somewhere on it.

Worked example. Find the absolute extrema of f(x) = x3 − 3x2 + 1 on [−1, 3]. The interior critical points are x = 0 and x = 2. Check all four candidates: f(−1) = −3, f(0) = 1, f(2) = −3, f(3) = 1. So the absolute maximum is 1 (at both x = 0 and x = 3) and the absolute minimum is −3 (at both x = −1 and x = 2). Notice the endpoints tied the interior points, which is exactly why you must check them.

Why the Extreme Value Theorem needs its fine print

"Continuous on a closed interval" sounds like lawyer talk until you watch each clause fail.

  • Drop the closed interval: f(x) = x on the open interval (0, 1) has no maximum and no minimum. Outputs get arbitrarily close to 1 and to 0, but no input achieves either, because the endpoints that would do the job are excluded.
  • Drop boundedness: f(x) = 1/x on (0, 1] has no maximum at all; it blows up without bound as x nears 0.
  • Drop continuity: define f(x) = x on [0, 1] except f(1) = 0. The single removed point at the top destroys the maximum: outputs approach 1 but never reach it.

With both hypotheses intact, none of these escapes is possible: the function must actually attain a highest and lowest value. That is why every optimization argument in this course either works on a closed interval or supplies a replacement justification (a sign chart showing one interior peak, or end behavior running to +∞ on both sides for a minimum). "Because the theorem applies" is a real reason; get in the habit of saying which theorem and checking its fine print.

A critical point where f' does not exist

Critical points come in two species, and the second is easy to forget. Take f(x) = x2/3, the curve of cube-root-squared.

  1. Differentiate: f'(x) = (2/3)x−1/3 = 2/(3x1/3). (Power rule with a fractional exponent.)
  2. Note f' is never zero (the top is the constant 2), but it is undefined at x = 0. So x = 0 is a critical point of the second species.
  3. Sign test: for x < 0, the cube root of a negative is negative, so f' is negative, falling. For x > 0, positive, rising.
  4. Minus to plus: a genuine local minimum at x = 0, with f(0) = 0, even though no tangent is horizontal there. The graph comes to a sharp cusp, infinitely steep on both sides.

Hunting only for f'(x) = 0 would have missed this minimum entirely. The full definition, zero or undefined, is not pedantry; it is where cusps and corners hide.

The Mean Value Theorem: the law behind the tests

Every sign chart above rests on a claim we have not yet justified: that a positive derivative really does force the function to climb. One theorem signs that guarantee. Drive 150 miles in 3 hours and your average speed is 50 mph; had the needle stayed under 50 the whole way, you could not have covered 150 miles. At some instant it read exactly 50. That everyday certainty, made precise, is the Mean Value Theorem (MVT): if

  1. f is continuous on the closed interval [a, b], and
  2. f is differentiable on the open interval (a, b),

then there is at least one point c with a < c < b where

f'(c) = [f(b) − f(a)]/(b − a)

The right side is the slope of the secant line through the two endpoints: two points, no calculus. The left side is the slope of the tangent at c. So the theorem promises an interior point where the tangent runs parallel to that secant. Notice the deliberate asymmetry in the fine print: continuity is demanded on the closed interval, endpoints included, but differentiability only on the open interval, which is why √x on [0, 1] still qualifies despite its vertical tangent at 0. When the endpoint values happen to match, f(a) = f(b), the secant is horizontal and the promised tangent is flat, f'(c) = 0. That special case has its own name, Rolle's Theorem, and the general statement is simply Rolle's picture tilted.

Worked check. For f(x) = x2 on [0, 2] the average rate is (4 − 0)/(2 − 0) = 2; setting f'(c) = 2c = 2 gives c = 1, safely inside. Now one whose answer surprises people. For f(x) = √x on [1, 4] the secant slope is (2 − 1)/(4 − 1) = 1/3, and f'(x) = 1/(2√x), so solve 1/(2√c) = 1/3: cross-multiplying gives 2√c = 3, so √c = 3/2 and c = 9/4 = 2.25. The midpoint of [1, 4] is 2.5, and c is not there. Landing on the midpoint is a private habit of parabolas, not a rule; solve f'(c) = secant slope every time, and keep only the roots that lie strictly between a and b.

Both hypotheses are load-bearing

Delete either clause and the promise genuinely dies, which is worth watching once so the fine print stops looking like decoration.

  • Without differentiability: f(x) = |x| on [−1, 1] has f(−1) = f(1) = 1, so Rolle's Theorem would owe us a flat tangent. But f' is −1 to the left of 0 and +1 to the right, and never 0. The corner at one interior point voids the guarantee, with no contradiction anywhere.
  • Without continuity: let f(x) = x on [0, 1) but set f(1) = 0. The average rate over [0, 1] is (0 − 0)/1 = 0, yet f'(x) = 1 at every interior point, never 0. One broken endpoint is enough to sink it.

From the theorem to the sign chart

Here is the payoff for this lesson. Suppose f'(x) > 0 at every point of an interval, and take any two inputs u < v inside it. The MVT applies on [u, v] (a differentiable function is automatically continuous), so for some c between them f(v) − f(u) = f'(c)(v − u): positive times positive, hence positive. So f(v) > f(u) for every such pair, which is exactly what "f is increasing" means, now proved rather than assumed. The same two lines deliver the rest of the family:

  • f' > 0 throughout an interval: f is increasing there.
  • f' < 0 throughout: f is decreasing there, since the product turns negative.
  • f' = 0 throughout: f is constant there, since the product is always zero.
  • If f' = g' throughout, then f − g has zero derivative and so is constant: f and g differ by a constant. Hold on to that one. It is the entire reason for the + C you will meet with antiderivatives at the end of this module.

The theorem settles arguments outside the classroom too. Average-speed cameras photograph a car at two points 5 miles apart, and a car covers the gap in 4 minutes, which is 1/15 of an hour: an average of 5 ÷ (1/15) = 75 mph. Position is a continuous, differentiable function of time, so at some instant between the cameras the speedometer read exactly 75. Not "probably over the limit": exactly 75, provable with no radar at that instant. Several countries issue speeding tickets on precisely this reasoning.

A closed-interval rep with a fractional term

Find the absolute extrema of f(x) = x + 1/x on [1/2, 3].

  1. Differentiate: f'(x) = 1 − 1/x2. (Rewrite 1/x as a power first.)
  2. Critical points: f'(x) = 0 gives x2 = 1, so x = ±1; also f' is undefined at x = 0. Only x = 1 lies inside [1/2, 3]. (Discard candidates outside the interval, including the undefined point.)
  3. Evaluate all candidates: f(1/2) = 1/2 + 2 = 2.5; f(1) = 1 + 1 = 2; f(3) = 3 + 1/3 ≈ 3.33.
  4. Verdict: absolute minimum 2 at x = 1, absolute maximum 10/3 at the endpoint x = 3.

This example rehearses every discipline at once: rewrite before differentiating, collect both kinds of critical point, discard the ones outside the interval, and let the endpoint win when it wins. The function itself is a classic: x + 1/x models total cost when one part grows linearly and another shrinks inversely, and its minimum at x = 1 is the balance point between them.

A flat tangent is a candidate, not a verdict

The biggest trap is assuming every place where f'(x) = 0 is a max or a min. It is only a candidate. At x = 0, the function x3 has f'(0) = 0, yet it is neither a peak nor a valley, because f' stays positive on both sides; it is a flat pause. Always run the sign test. The second trap is forgetting the endpoints when hunting an absolute extremum on a closed interval; the winner is often sitting quietly at an edge.

Common misconceptions

  • "Every point with f'(x) = 0 is an extremum." Only a candidate. Test the sign change; x3 at 0 is neither.
  • "For the absolute max, just find the local maxima." The absolute extreme can live at an endpoint, where there is no flat tangent.
  • "Critical points are only where f' = 0." Also where f' is undefined, like a corner.
  • "The Extreme Value Theorem always applies." It needs continuity and a closed, bounded interval; drop either and a max may not exist.
  • "The Mean Value Theorem tells you where c is." It promises only that some c exists strictly inside the interval. Finding it is your algebra, more than one c can qualify, and it is rarely the midpoint.

Putting it together

Extrema are the peaks and valleys of a function. They can only occur at critical points, where f'(x) = 0 or is undefined, and the cusp of x2/3 is the reminder that the second kind is not optional. The first derivative test classifies each by how the sign of f' changes: plus to minus is a local max, minus to plus a local min, no change is neither. For absolute extrema on a closed interval, compare f at all interior critical points and both endpoints, backed by the Extreme Value Theorem. Underneath all of it sits the Mean Value Theorem, which is what licenses the move from "the derivative is positive here" to "the function actually climbs here." Draw a sign chart, and never skip the endpoints.

Sources

  1. OpenStax. (2016). 4.3 Maxima and minima. In Calculus volume 1. openstax.org
  2. OpenStax. (2016). 4.4 The mean value theorem. In Calculus volume 1. openstax.org
  3. Dawkins, P. (n.d.). Minimum and maximum values. Paul's Online Math Notes. tutorial.math.lamar.edu
  4. Dawkins, P. (n.d.). The shape of a graph, part I. Paul's Online Math Notes. tutorial.math.lamar.edu
  5. Dawkins, P. (n.d.). The mean value theorem. Paul's Online Math Notes. tutorial.math.lamar.edu
  6. Math is Fun. (n.d.). Finding maxima and minima using derivatives. mathsisfun.com
  7. Khan Academy. (n.d.). Calculus 1 [Online course]. khanacademy.org
Key terms
Extremum
A maximum or minimum value of a function.
Local maximum
A point higher than all nearby points.
Critical point
An input where the derivative is zero or undefined, the only candidate for a local extremum.
First derivative test
Classifying a critical point by how the sign of f' changes around it.
Closed interval method
Finding absolute extrema by checking critical points and endpoints.
Extreme Value Theorem
A continuous function on a closed interval attains an absolute max and min.
Mean Value Theorem
If f is continuous on [a, b] and differentiable on (a, b), some c in (a, b) has f'(c) equal to the average rate of change over [a, b].
Rolle's Theorem
The special case of the Mean Value Theorem with f(a) = f(b), where the guaranteed tangent is horizontal.

Concavity and the Second Derivative Test

  • Interpret the second derivative as concavity.
  • Locate points of inflection where concavity changes.
  • Classify critical points with the second derivative test.

One company's revenue rises by 3, then 5, then 7 million dollars over three straight quarters. A second company's rises by 7, then 5, then 3. Both are growing every quarter, so both have a positive first derivative throughout, and no analyst would treat them as the same story. What separates them is not direction but bending: the first is accelerating, the second is running out of steam. That property has a name, concavity, and one more derivative is all it takes to read it.

What bending adds to direction

The second derivative f''(x) (read "f double prime of x") is just the derivative of the derivative. Since f' is the slope, f'' tells you whether the slope is growing or shrinking, which is the same as how the curve bends. It fills in a picture the first derivative alone cannot.

The upshot: concave up bends like a cup (a smile); concave down bends like a frown. The sign of f'' tells you which.

Concave up and concave down

  • Where f''(x) > 0, the curve is concave up, shaped like a cup that could hold water; the slope is increasing.
  • Where f''(x) < 0, the curve is concave down, shaped like a frown; the slope is decreasing.

A quick memory aid: concave up looks like a cup or a smile; concave down looks like a frown. A point where the bending switches from one to the other is a point of inflection. It usually happens where f''(x) = 0 and the sign of f'' actually flips across it. A zero of f'' that does not flip sign is not an inflection point.

What concavity means in real life

If f is position, then f'' is acceleration, so concave up means speeding up. In business, a concave-down profit curve means diminishing returns: each extra unit adds a little less than the one before. When the news says inflation is "still rising but slowing," that is a positive first derivative (still rising) with a negative second derivative (slowing), concave down. Reading concavity turns a graph into a story about acceleration and change.

The second derivative test

Concavity gives a fast way to classify a critical point where f'(c) = 0:

  • If f''(c) > 0, the curve is concave up there, a valley, so c is a local minimum.
  • If f''(c) < 0, the curve is concave down there, a peak, so c is a local maximum.
  • If f''(c) = 0, the test is inconclusive, and you fall back on the first derivative test.

Let us do one together, slowly

Use the second derivative test on f(x) = x3 − 3x2 + 1, whose critical points are x = 0 and x = 2.

  1. First derivative: f'(x) = 3x2 − 6x. Second derivative: f''(x) = 6x − 6. (Differentiate twice.)
  2. At x = 0: f''(0) = −6, which is negative, so concave down, a local maximum.
  3. At x = 2: f''(2) = 6, which is positive, so concave up, a local minimum.
  4. Inflection point: set f''(x) = 6x − 6 = 0, giving x = 1, where the bending flips from frown to cup.

This matches the first derivative test from last lesson, but with less sign-checking. When the second derivative is easy to compute and not zero at the critical point, this is usually the quicker route.

Try it: for f(x) = x2, find f''(x) and state the concavity.

Answer: f'(x) = 2x, so f''(x) = 2, which is positive everywhere. The parabola is concave up everywhere, a cup shape. Makes sense.

When the test cannot decide

The f''(c) = 0 case genuinely needs care, because all three outcomes are possible. Look at x4, −x4, and x3 at the origin: each has f'(0) = 0 and f''(0) = 0, yet x4 has a minimum, −x4 a maximum, and x3 neither. The second derivative test simply cannot tell them apart, so return to the first derivative test's sign-change analysis. Knowing the limits of a tool is as valuable as knowing how to use it.

A complete curve sketch, start to finish

Everything from the last two lessons assembles into a portrait of f(x) = x4 − 4x3, drawn without plotting a single random point.

  1. Intercepts and end behavior. Factor: f(x) = x3(x − 4), so the graph crosses zero at x = 0 and x = 4. The leading term x4 sends both ends up to +∞.
  2. First derivative and its chart. f'(x) = 4x3 − 12x2 = 4x2(x − 3). Critical points: x = 0 and x = 3. Signs: at x = −1: 4(1)(−4) < 0. At x = 1: 4(1)(−2) < 0. At x = 4: 4(16)(1) > 0. Chart: negative, negative, positive.
  3. Read the chart. Falling on both sides of 0 (no sign change: x = 0 is a flat pause, not an extremum), still falling to x = 3, then rising forever. So the one local (and absolute) minimum is at x = 3, f(3) = 81 − 108 = −27.
  4. Second derivative and its chart. f''(x) = 12x2 − 24x = 12x(x − 2). Zeros at x = 0 and x = 2. Signs: at x = −1: 12(−1)(−3) > 0, concave up. At x = 1: 12(1)(−1) < 0, concave down. At x = 3: 12(3)(1) > 0, concave up.
  5. Inflection points. The sign flips at both zeros, so both are genuine inflections: (0, 0) and (2, f(2)) = (2, 16 − 32) = (2, −16).
  6. Confirm with the second derivative test. At the critical point x = 3: f''(3) = 36 > 0, concave up, a minimum, agreeing with the first-derivative chart. At x = 0: f''(0) = 0, inconclusive, and the first-derivative chart already settled it (no sign change, so neither).

Now narrate the shape left to right: the curve sweeps down from +∞ concave up, flattens momentarily at the origin while switching to concave down (a flat inflection, an S-bend with a horizontal tangent), keeps falling, switches back to concave up at (2, −16), bottoms out at (3, −27), then climbs through (4, 0) and away to +∞. Every feature came from two factored derivatives and two sign charts. That is the entire craft of curve sketching: f gives points, f' gives direction, f'' gives bending, and the three stories must fit together.

A reading habit: the four shape types

Every smooth stretch of any graph is one of exactly four shapes, and naming them speeds up both sketching and reading: rising concave up (speeding up uphill, like compound interest), rising concave down (uphill but tiring, like a saturating population), falling concave down (accelerating downhill), and falling concave up (a leveling descent, like a cooling coffee approaching room temperature). When you sketch, label each interval between critical points and inflections with its type and the picture nearly draws itself. When you read a news graph, the type is the story: "cases still rising but concave down" is the phrase "the surge is slowing" in mathematics.

Concavity from raw data: second differences

Concavity is visible even without a formula. Take equally spaced outputs of f(x) = x2: at x = 1, 2, 3, 4, 5 the values are 1, 4, 9, 16, 25. First differences (each value minus the last): 3, 5, 7, 9, climbing, so the slope is increasing: concave up. Difference the differences: 2, 2, 2, constant, echoing f''(x) = 2. A data table whose first differences shrink instead signals concave down. Analysts eyeball exactly this when they scan quarterly numbers: revenue up 3, then 5, then 7 is an accelerating (concave up) story even before anyone fits a curve. Second differences are the discrete shadow of the second derivative.

One more test rep: a W-shaped curve

Classify the critical points of f(x) = x4 − 2x2.

  1. f'(x) = 4x3 − 4x = 4x(x − 1)(x + 1): critical points at x = −1, 0, 1.
  2. f''(x) = 12x2 − 4. Evaluate: f''(−1) = 8 > 0 (min), f''(0) = −4 < 0 (max), f''(1) = 8 > 0 (min).
  3. Heights: f(±1) = 1 − 2 = −1 and f(0) = 0: a W shape with twin valley bottoms at (±1, −1) and a local crest at the origin.

Three critical points, three clean verdicts, no sign charts needed: this is the second derivative test at its best, on a function whose f'' is easy and nonzero at every critical point. The inflections sit where f''(x) = 0, at x = ±1/√3 ≈ ±0.577, where the W's walls change their bend.

Zeros of f'' are not automatically inflection points

The most common error is thinking every zero of f'' is an inflection point. It is only an inflection point if f'' actually changes sign there. For x4, f''(0) = 0 but the curve stays concave up on both sides, so there is no inflection. The second stumble is reading "inconclusive" as "not an extremum." Inconclusive just means this particular test cannot decide; x4 has a clear minimum at 0 even though f''(0) = 0. Switch tools and use the first derivative test.

Common misconceptions

  • "Wherever f'' = 0 there is an inflection point." Only if f'' changes sign. x4 at 0 has f''(0) = 0 but no inflection.
  • "f''(c) = 0 means c is not an extremum." It means the test is inconclusive; x4 still has a minimum there.
  • "Concave up means increasing." Concavity is about bending, not direction. A concave-up curve can still be going downhill.
  • "The second derivative test always works." It stalls when f''(c) = 0; then use the first derivative test.

What you now know

The second derivative measures bending. Where f'' > 0 the curve is concave up (a cup), where f'' < 0 it is concave down (a frown), and where it flips sign there is a point of inflection. The second derivative test classifies a critical point fast: concave up gives a local min, concave down a local max, and f''(c) = 0 is inconclusive, so fall back on the first derivative test. Concavity also reads as acceleration and as diminishing returns in the real world.

Sources

  1. OpenStax. (2016). 4.5 Derivatives and the shape of a graph. In Calculus volume 1. openstax.org
  2. Dawkins, P. (n.d.). The shape of a graph, part I. Paul's Online Math Notes. tutorial.math.lamar.edu
  3. Dawkins, P. (n.d.). The shape of a graph, part II. Paul's Online Math Notes. tutorial.math.lamar.edu
  4. Math is Fun. (n.d.). Concave upward and downward. mathsisfun.com
  5. Math is Fun. (n.d.). Second derivative. mathsisfun.com
  6. Sanderson, G. (3Blue1Brown). (n.d.). Higher order derivatives. 3blue1brown.com
  7. Khan Academy. (n.d.). Calculus 1 [Online course]. khanacademy.org
Key terms
Second derivative
The derivative of the derivative, written f''(x).
Concave up
A cup-shaped bend where f'' > 0 and the slope increases.
Concave down
A frown-shaped bend where f'' < 0 and the slope decreases.
Point of inflection
A point where concavity changes, typically where f'' = 0 and switches sign.
Second derivative test
Classifying a critical point by the sign of f'' there.
Inconclusive test
When f''(c) = 0 the second derivative test cannot decide, so the first derivative test is used.

Optimization

  • Translate a word problem into an objective function and constraint.
  • Reduce the objective to one variable and differentiate.
  • Find and justify the optimal value.

A farmer has 100 ft of fence and a straight river that needs no fencing. A pen 10 ft deep by 80 ft wide encloses 800 square feet. Try 20 by 60 and you get 1200. Try 30 by 40 and you get 1200 again, exactly the same. So the best pen lies somewhere between those two, and no amount of guessing will land on it exactly. Two lines of calculus will: 25 ft deep, 50 ft wide, 1250 square feet, and a proof that nothing beats it. The calculus is the easy part here; the work is the setup, and we will do that slowly.

Objective and constraint, every time

Every optimization problem has two pieces: something you want to make as big or small as possible (the objective), and a rule that ties the variables together (the constraint). You use the constraint to boil the objective down to one variable, then let the derivative find the best value, exactly the extrema skill from the last two lessons.

What matters here: write the objective, write the constraint, use the constraint to get one variable, then set the derivative to zero.

A step-by-step strategy

  1. Name the variables, and draw a picture if it helps. A clear sketch often turns a muddle into a solvable problem.
  2. Write the objective function, the quantity to optimize.
  3. Write the constraint, the equation that limits the variables.
  4. Use the constraint to rewrite the objective in one variable.
  5. Differentiate, set the derivative to zero, and solve for the critical points.
  6. Confirm it is the max or min you want (a derivative test or the endpoints), then answer in context, with units.

Spotting which equation is the objective and which is the constraint is the key skill. Every optimization problem has both.

Let us do the biggest pen together, slowly

A farmer has 100 ft of fencing to enclose a rectangular pen against a straight river, so no fence is needed on the river side. What dimensions give the largest area?

  1. Name things: let x be each of the two sides perpendicular to the river, and w the side parallel to it. (Draw it.)
  2. Constraint (only three sides are fenced): 2x + w = 100. (What limits us.)
  3. Objective (what to maximize): area A = x w. (What we want big.)
  4. Reduce to one variable: solve the constraint for w = 100 − 2x, then substitute, giving A(x) = x(100 − 2x) = 100x − 2x2. (One variable now.)
  5. Differentiate and set to zero: A'(x) = 100 − 4x = 0, so x = 25. (Find the critical point.)
  6. Confirm a maximum: A''(x) = −4, which is negative, concave down, so x = 25 is the max. Then w = 100 − 50 = 50, and the area is 25 × 50 = 1250 square feet.

The best pen is 25 ft deep and 50 ft wide, enclosing 1250 square feet. A quick sanity check: x = 20, w = 60 gives 1200, less than 1250, so our answer holds up. You just solved a real optimization problem.

A quick second example

Find two nonnegative numbers whose sum is 20 and whose product is as large as possible. Let them be x and 20 − x, so the product is P(x) = x(20 − x) = 20x − x2. Then P'(x) = 20 − 2x = 0 gives x = 10, and the other number is also 10, for a maximum product of 100. The equal split wins, a pattern that also explains why a square encloses more area than any other rectangle of the same perimeter.

Try it: find two nonnegative numbers with sum 12 whose product is largest.

Answer: Let them be x and 12 − x; product P = 12x − x2, so P' = 12 − 2x = 0 gives x = 6. Both are 6, product 36. Equal split again.

The classic can problem

A staple of every calculus course: use the least metal for a cylindrical can holding a fixed volume. The objective is surface area S = 2π r2 + 2π r h, and the constraint is the fixed volume V = π r2 h. Solve the constraint for h, substitute into S to get a function of r alone, differentiate, and set to zero. The tidy result: the cheapest can has height equal to its diameter, h = 2r. Real cans differ because of lids, labels, and stacking, but the calculus gives the ideal baseline, which is exactly how optimization guides engineering.

A caution about the domain

Always respect the physical domain. A length cannot be negative, so x in the pen problem lives in [0, 50]. Here the best value fell safely inside, but sometimes the true optimum sits at an endpoint of the allowed range, so checking the ends is part of a complete solution, just like the closed interval method.

A geometry problem on an unbounded domain

Closed intervals justify themselves with the Extreme Value Theorem. When the domain runs to infinity, you owe a different justification, so let us do one honestly. Which point on the parabola y = x2 is closest to the point (0, 3)?

  1. Objective. Distance from (x, x2) to (0, 3) is √(x2 + (x2 − 3)2). Minimizing a square root is the same as minimizing what is under it (the square root is an increasing function), so minimize the squared distance D(x) = x2 + (x2 − 3)2 instead and spare yourself the root. (A standard, legal simplification; say it out loud when you use it.)
  2. Differentiate with the chain rule: D'(x) = 2x + 2(x2 − 3)(2x) = 2x[1 + 2(x2 − 3)] = 2x(2x2 − 5).
  3. Critical points: x = 0 and 2x2 = 5, so x = ±√(5/2) ≈ ±1.58.
  4. Compare candidates. D(0) = 0 + 9 = 9. At x2 = 5/2: D = 5/2 + (5/2 − 3)2 = 5/2 + 1/4 = 2.75.
  5. Justify the winner without EVT. The domain is all of x, so check the ends: as x → ±∞, D(x) → ∞. A continuous function that runs to infinity in both directions attains its minimum at one of its critical points, and the smallest critical value is 2.75. (The point x = 0 is in fact a local maximum of D, the top of the dip between the two winners; the sign of D' confirms it: positive just left of 0, negative just right.)

Answer: the two symmetric points (±√(5/2), 5/2), at distance √2.75 ≈ 1.66. The symmetry was predictable: the target sits on the parabola's axis, so ties come in mirror pairs.

A bonus tool from derivatives: L'Hopital's rule

Optimization endpoint checks often produce limits of tug-of-war form, so this is the natural place to meet the last derivative technique of the course. L'Hopital's rule: if lim f(x)/g(x) lands on exactly 0/0 or ∞/∞, and the functions are differentiable near the point, then

lim f(x)/g(x) = lim f'(x)/g'(x)

provided the right-hand limit exists. Differentiate top and bottom separately (this is not the quotient rule), then take the limit again.

Worked, twice over. lim (x → 0) (1 − cos x)/x2: substituting gives 0/0, so pass to (sin x)/(2x), which is 0/0 again, so pass once more to (cos x)/2, which is 1/2 by substitution. Each pass requires re-checking that the form is still indeterminate before differentiating again.

Growth contests. lim (x → ∞) (ln x)/x is ∞/∞; one pass gives (1/x)/1 → 0. This proves the pecking order claimed back in Module 2: logs lose to powers. A product like lim (x → 0+) x ln x is the form 0 × (−∞), not directly eligible; rewrite it as a quotient (ln x)/(1/x), now −∞/∞, and one pass gives (1/x)/(−1/x2) = −x → 0. That little limit is exactly the endpoint check that certifies xx → 1 as x → 0+ in the last lesson's logarithmic-differentiation example.

The warning that saves you. L'Hopital applies only to 0/0 and ∞/∞. Apply it to a determinate form and it manufactures a wrong answer with total confidence: lim (x → 1) (x2 + 1)/(x + 1) is simply 2/2 = 1 by substitution, but blind L'Hopital would say 2x/1 → 2. Always substitute first, name the form you see, and reach for the rule only when the form is indeterminate.

Where the compound-interest number comes from

L'Hopital settles one more famous indeterminate form, 1∞, and it is worth seeing because the answer is a celebrity. What is lim (x → 0+) (1 + x)1/x? (Equivalently, compound interest at rate 1 split into ever more periods.) The base heads to 1, the exponent to infinity: a tug of war. Take logs to bring the exponent down: ln L = lim ln(1 + x)/x, which is honest 0/0, so L'Hopital applies: pass to [1/(1 + x)]/1 → 1. So ln L = 1, meaning L = e ≈ 2.71828. The banker's number is a limit, and the log-first maneuver you just watched, convert a power form to a quotient, apply L'Hopital, exponentiate back, is the standard recipe for every 1∞, 00, and ∞0 form on any exam.

Which equation is the objective?

The first snag is mixing up the objective and the constraint. The objective is what you optimize (area, cost); the constraint is the separate equation that ties the variables (a fixed perimeter or volume). You always use the constraint to shrink the objective to one variable. The second snag is stopping the moment you solve f'(x) = 0. A critical point still needs to be confirmed as the right kind of extremum, and the physical domain still needs checking, or you may report a minimum when you wanted a maximum.

Common misconceptions

  • "The objective and the constraint are the same thing." Opposite roles: the objective is optimized; the constraint limits the variables.
  • "Once f'(x) = 0 is solved, you are done." Justify it is the desired max or min, and check the domain, including endpoints.
  • "Optimization needs a fancy new method." It is the extrema toolkit applied to a one-variable objective built from the constraint.
  • "The domain does not matter." Physical limits (nonnegative lengths, fixed volumes) shape the answer and can put the optimum at an edge.

Pulling it together

Optimization finds the best outcome. Name the variables, write the objective and the constraint, use the constraint to reduce the objective to one variable, then set the derivative to zero and confirm the extremum with a derivative test or the endpoints. Respect the physical domain throughout. The recurring equal-split and symmetric answers, like the pen or the square, are the calculus quietly finding balance.

Sources

  1. OpenStax. (2016). 4.7 Applied optimization problems. In Calculus volume 1. openstax.org
  2. OpenStax. (2016). 4.8 L'Hopital's rule. In Calculus volume 1. openstax.org
  3. Dawkins, P. (n.d.). Optimization. Paul's Online Math Notes. tutorial.math.lamar.edu
  4. Dawkins, P. (n.d.). L'Hospital's rule and indeterminate forms. Paul's Online Math Notes. tutorial.math.lamar.edu
  5. Sanderson, G. (3Blue1Brown). (n.d.). Limits, L'Hopital's rule, and epsilon delta definitions. 3blue1brown.com
  6. Math is Fun. (n.d.). Finding maxima and minima using derivatives. mathsisfun.com
  7. Khan Academy. (n.d.). Calculus 1 [Online course]. khanacademy.org
Key terms
Optimization
Finding the maximum or minimum value of a quantity.
Objective function
The quantity being maximized or minimized.
Constraint
An equation that limits the variables in the problem.
Single-variable reduction
Using the constraint to express the objective in one variable.
Optimal value
The maximum or minimum output, justified by a derivative test.
Feasible domain
The range of variable values allowed by the physical context, whose endpoints must also be checked.

Antiderivatives: An Introduction to the Integral

  • Define an antiderivative and the role of the constant of integration.
  • Reverse the power rule and other basic derivatives to integrate.
  • Use an initial condition to find a specific antiderivative.

A falling object accelerates at a steady 32 ft/s per second, and that is the only fact you are handed. Where is it? The acceleration alone cannot say: an object dropped from a roof and one dropped from an airplane obey the identical rate. Recovering a function from its rate of change is the reverse of everything this course has done so far, and it works, right up to one stubborn unknown that no rate can ever pin down: where the motion started. That reverse process is antidifferentiation, or integration, and it is the doorway to the second half of calculus.

Running differentiation backwards

An antiderivative of f is a function F whose derivative is f, written F'(x) = f(x). Since d/dx (x2) = 2x, an antiderivative of 2x is x2. Where the derivative asked "what is the rate of change," the antiderivative asks the reverse: "what function produced this rate."

Why this matters: integration is differentiation in reverse. To integrate, ask "what would I differentiate to get this?"

The constant of integration

Here is a gentle twist. Both x2 + 5 and x2 − 7 also have derivative 2x, because the derivative of any constant is 0. So a function has infinitely many antiderivatives, all differing by a constant. We capture them all with a constant of integration C, writing the general antiderivative as x2 + C. In integral notation, ∫ 2x dx = x2 + C, where the symbol ∫ is read "the integral of." Picture a whole family of parallel curves, each a vertical shift of the others, all with the same slope everywhere. Forgetting the + C is the most common integration slip, so make it a habit from day one.

Reversing the power rule

The power rule brought the exponent down and dropped it by one. To undo it, do the opposite: add one to the exponent and divide by the new exponent.

∫ xn dx = x(n+1)/(n+1) + C (for n not equal to −1)

Why exclude n = −1? Because it would force division by zero. That one missing case is filled by the logarithm: ∫ (1/x) dx = ln|x| + C, which ties back to d/dx ln x = 1/x.

Let us integrate one together, slowly

Find ∫ (4x3 − 6x + 5) dx, term by term.

  1. ∫ 4x3 dx = 4 × x4/4 = x4. (Add one to the exponent, divide by 4.)
  2. ∫ −6x dx = −6 × x2/2 = −3x2. (Same move.)
  3. ∫ 5 dx = 5x. (The antiderivative of a constant is that constant times x.)
  4. Add the constant of integration: x4 − 3x2 + 5x + C.

Here is the nicest part: you can always check by differentiating. d/dx (x4 − 3x2 + 5x + C) = 4x3 − 6x + 5, the original. Unlike most of calculus, integration lets you instantly verify your own answer.

Try it: find ∫ (3x2 + 2x) dx.

Answer: x3 + x2 + C. Check by differentiating: 3x2 + 2x. It matches, so you did it.

A short table of basic antiderivatives

FunctionAntiderivative
xn (n not −1)x(n+1)/(n+1) + C
1/xln|x| + C
exex + C
cos(x)sin(x) + C
sin(x)−cos(x) + C

Every row is just a derivative rule read backward. Because sine differentiates to cosine, cosine integrates to sine; because cosine differentiates to negative sine, sine integrates to negative cosine. Know your derivatives well and the antiderivatives come almost for free.

Pinning down C with a starting value

Often a little extra information selects one antiderivative from the family. Suppose f'(x) = 6x and f(0) = 2.

  1. General antiderivative: f(x) = 3x2 + C. (Reverse the power rule.)
  2. Use the starting value: f(0) = 3(0)2 + C = 2, so C = 2. (Solve for the constant.)
  3. The specific function: f(x) = 3x2 + 2.

This is exactly how physics recovers position from velocity: integrate the velocity, then use the known starting position to fix C.

The road ahead

Antiderivatives are the bridge to the definite integral and the Fundamental Theorem of Calculus, which opens Calculus II. That theorem reveals something wonderful: this same reverse-of-slopes idea also computes exact areas under curves, by adding up infinitely many thin strips. The area under f from a to b turns out to equal F(b) − F(a), where F is any antiderivative. You have reached the doorway to integration, and you got here by mastering derivatives. Every rule you learned is about to pay off again, in reverse. That is a real accomplishment.

The area problem: adding infinitely many slivers

Integration has a second birthplace, apparently unrelated to undoing derivatives: measuring curved area. What is the area under f(x) = x2 from 0 to 1? No triangle or rectangle formula applies, but we can approximate with rectangles and then take a limit, the same shrink-to-zero move that built the derivative.

  1. Slice [0, 1] into 4 strips of width Δx = 1/4. On each strip, erect a rectangle whose height is the function's value at the strip's right edge: heights (1/4)2, (1/2)2, (3/4)2, 12.
  2. Add the areas: (1/4)[1/16 + 4/16 + 9/16 + 16/16] = (1/4)(30/16) = 0.46875. (An overestimate: on a rising curve, right edges poke above.)
  3. Left edges instead give (1/4)[0 + 1/16 + 4/16 + 9/16] = 0.21875, an underestimate. The true area is trapped between.
  4. Refine: with n = 100 right-edge strips the sum is 0.33835; with the algebraic identity for 12 + 22 + ... + n2 one can show the sums close in on exactly 1/3 as n → ∞.

These rectangle totals are Riemann sums, and their limit is the definite integral, written ∫01 x2 dx = 1/3. Read the notation as its own picture: the elongated S is a sum, f(x) is a rectangle's height, dx is its vanishing width, and the numbers 0 and 1 are the endpoints of the ground being covered.

The Fundamental Theorem: why area and slope are inverses

Here is the miracle that fuses the two halves of calculus. Define the accumulation function A(x) = the area under f from 0 to x, a function that grows as x sweeps right. How fast does it grow? Nudge x by a sliver h: the new area gained is a skinny strip of width h and height essentially f(x), so A(x + h) − A(x) ≈ f(x) × h. Divide by h and shrink: A'(x) = f(x). The rate at which area accumulates is the height of the curve. That statement is Part 1 of the Fundamental Theorem of Calculus: every continuous function has an antiderivative, namely its own accumulation function.

Part 2 is the payoff for computation. If F is any antiderivative of f, then

∫ab f(x) dx = F(b) − F(a)

because A and F differ only by a constant, which cancels in the subtraction. Watch it demolish the rectangle problem: an antiderivative of x2 is x3/3, so ∫01 x2 dx = 1/3 − 0 = 1/3, agreeing with the Riemann sums but taking five seconds. (For definite integrals the + C may be dropped: any one antiderivative works, since the constant subtracts away.) Slope-finding and area-finding, invented centuries apart, are one subject read in two directions, and this theorem is the hinge.

Substitution: the chain rule in reverse

Each derivative rule reverses into an integration technique, and the reversed chain rule, called u-substitution, is the workhorse. Find ∫ 2x(x2 + 1)5 dx.

  1. Spot an inside function whose derivative is also a factor: inside u = x2 + 1 has derivative 2x, present. (The tell-tale pattern.)
  2. Trade variables: du = 2x dx, so the integral becomes ∫ u5 du.
  3. Integrate in u: u6/6 + C.
  4. Translate back: (x2 + 1)6/6 + C. Check by differentiating: chain rule gives 6(x2 + 1)5/6 × 2x = 2x(x2 + 1)5. Perfect.

For a definite integral, convert the endpoints too and never return to x: ∫01 2x(x2 + 1)5 dx has u running from u(0) = 1 to u(1) = 2, so it equals ∫12 u5 du = (26 − 16)/6 = 63/6 = 21/2. A simpler cousin handles linear insides by inspection: ∫ cos(3x) dx = sin(3x)/3 + C, dividing by the inner derivative 3. When the required factor is missing entirely, substitution does not apply, and Calculus II supplies further tools (integration by parts, the reversed product rule, among others). You now stand exactly at that doorway, with the complete differential calculus behind you and the integral calculus opening ahead.

The + C, and the one power the rule cannot undo

The most common slip is dropping the + C. It is not decoration; it stands for the whole family of antiderivatives, and in starting-value problems it is precisely the number you solve for. The second slip is using the reverse power rule on 1/x. The rule x(n+1)/(n+1) divides by zero when n = −1, so 1/x is the special case whose antiderivative is ln|x| + C, not a power at all.

Common misconceptions

  • "The + C is optional." It represents infinitely many antiderivatives and is exactly what you solve for in starting-value problems.
  • "The reverse power rule works for every power." It fails at n = −1; the antiderivative of 1/x is ln|x| + C.
  • "A function has one antiderivative." It has infinitely many, all differing by a constant.
  • "You cannot check an integral." You can, and should: differentiate your answer and see if you get the original function back.

Looking back

An antiderivative F satisfies F'(x) = f(x), so integration is differentiation in reverse. Because constants vanish under differentiation, always include + C. Reverse the power rule by adding one to the exponent and dividing (except n = −1, which gives ln|x|), and read the other derivative rules backward for a table of antiderivatives. A starting value pins down C. This opens the door to the definite integral and areas in Calculus II. You made it to the end, and you built the whole thing yourself.

Sources

  1. OpenStax. (2016). 4.10 Antiderivatives. In Calculus volume 1. openstax.org
  2. OpenStax. (2016). 5.1 Approximating areas. In Calculus volume 1. openstax.org
  3. OpenStax. (2016). 5.3 The fundamental theorem of calculus. In Calculus volume 1. openstax.org
  4. OpenStax. (2016). 5.5 Substitution. In Calculus volume 1. openstax.org
  5. Dawkins, P. (n.d.). Indefinite integrals. Paul's Online Math Notes. tutorial.math.lamar.edu
  6. Dawkins, P. (n.d.). Definition of the definite integral. Paul's Online Math Notes. tutorial.math.lamar.edu
  7. Sanderson, G. (3Blue1Brown). (n.d.). Integration and the fundamental theorem of calculus. 3blue1brown.com
  8. Math is Fun. (n.d.). Introduction to integration. mathsisfun.com
Key terms
Antiderivative
A function F whose derivative is f, so F'(x) = f(x).
Integration
The process of finding antiderivatives, the inverse of differentiation.
Constant of integration
The +C that accounts for all antiderivatives differing by a constant.
Reverse power rule
integral of x^n dx = x^(n+1)/(n+1) + C for n not equal to -1.
Initial condition
A known value like f(0) that determines the constant C.
Fundamental Theorem of Calculus
The result linking antiderivatives to areas, giving the integral from a to b as F(b) - F(a).

Open the interactive version with quizzes and progress →