🔭 Physics · Graduate · PHYS 430

Statistical Mechanics & Thermodynamics

A graduate course that derives thermodynamics instead of asserting it. You begin by counting. The multiplicity of a two-state paramagnet, the multiplicity of an Einstein solid worked out with real binomial coefficients, and the moment when the peak in a multiplicity function becomes so narrow that the second law stops being a tendency and becomes a certainty. Entropy arrives as k ln(Omega),…

Start the interactive course (quizzes, progress, videos) →

Free forever. No sign-up, no ads. 17 lessons. The full lesson text is below so you can read it right here.

Module 1: Counting Microstates

Build thermodynamics from arithmetic. Count the microstates of a paramagnet and an Einstein solid exactly, watch the multiplicity peak sharpen until deviations become unobservable, and derive entropy and temperature from the count rather than assuming them.

The Two-State Paramagnet, and a Peak Sharper Than Any Instrument

  • Distinguish microstate, macrostate and multiplicity, and compute the multiplicity of a two-state paramagnet exactly for small numbers of dipoles.
  • Apply Stirling's approximation to obtain the Gaussian form of the multiplicity function and show that its relative width scales as one over the square root of N.
  • State the statistical definition of entropy and evaluate it for a mole of free spins.

Take twenty coins and shake them. There are 220 = 1,048,576 ways they can land, and exactly one of those ways is all heads. There are 184,756 ways to land ten heads and ten tails. That ratio, 184,756 to 1, is the second law of thermodynamics in miniature. The rest of this course is largely about what happens to that ratio when twenty becomes 1022, which is roughly the number of electron spins in a grain of paramagnetic salt you could balance on a fingernail.

Nothing in the argument is physics in the usual sense. There is no force law, no equation of motion, no conservation principle beyond the ones you already have. There is counting, and then there is the observation that when you count things at the scale of Avogadro's number, probabilities stop behaving like probabilities and start behaving like certainties.

Three questions to hold while you read

  • If every arrangement of the spins is equally likely, why does the magnetization always come out near zero?
  • What exactly gets large when N gets large: the probability of the typical outcome, or something else?
  • Why should the logarithm of a count of arrangements have anything to do with heat?

Microstate, macrostate, multiplicity

The system is a two-state paramagnet: N magnetic dipoles fixed on lattice sites, in an external field B, each dipole pointing either along the field or against it. Nothing else. Call the magnetic moment of one dipole m, so an aligned dipole has energy -mB and an anti-aligned one has +mB. Real systems that behave this way include the nuclear spins in a lithium fluoride crystal and the electron spins in a dilute salt such as chrome alum, and you can read the general phenomenon under paramagnetism.

A microstate is a complete specification: which way dipole 1 points, which way dipole 2 points, and so on to dipole N. There are 2N of them. A macrostate is what you can actually measure, which here is just the number pointing up, call it Nup, since that fixes both the total energy and the magnetization. The multiplicity of a macrostate is the number of microstates it contains:

Omega(N, N_up) = N! / [N_up! (N - N_up)!], the binomial coefficient.

Write out the whole distribution for ten dipoles, because you should see the shape before you see the formula for it.

Nup012345678910
Omega1104512021025221012045101

The eleven entries sum to 1024 = 210, as they must. The peak holds 252 of those, which is 24.6 percent. The two extreme macrostates hold one microstate each, 0.1 percent each. Ten dipoles is not a large number, and already the distribution is lopsided.

The one physical assumption is the fundamental postulate: for an isolated system in equilibrium, every accessible microstate is equally likely. Not every macrostate. Every microstate. The all-up configuration is exactly as probable as the specific configuration up-down-up-up-down-down-up-down-up-up. What differs is how many microstates sit inside each macrostate.

Key idea: microstates are equiprobable; macrostates are not, and the entire asymmetry comes from the count.

Stirling, and what the peak actually looks like

You cannot evaluate 1022 factorial. You can evaluate its logarithm, using Stirling's approximation:

ln(N!) = N ln N - N + (1/2) ln(2 pi N) + ...

Test it at N = 100. The exact value of ln(100!) is 363.7394. Stirling gives 100 ln(100) - 100 + 0.5 ln(628.32) = 460.517 - 100 + 3.221 = 363.738. The error is in the fifth decimal place, and it shrinks as N grows. For 1022 the last term is utterly negligible and the working form is ln(N!) = N ln N - N.

Now put N_up = N/2 + x and N_down = N/2 - x, so x measures the excess of up spins over the balanced value. Taking logs and applying Stirling to all three factorials,

ln Omega = N ln N - (N/2 + x) ln(N/2 + x) - (N/2 - x) ln(N/2 - x).

Write ln(N/2 + x) = ln(N/2) + ln(1 + 2x/N) and expand the logarithms to second order in the small quantity 2x/N, using ln(1 + u) = u - u^2/2. The terms linear in x cancel against each other, which is why the peak sits at x = 0, and what survives is

ln Omega = N ln 2 - 2x^2/N, so Omega(x) = 2^N e^(-2x^2/N).

That is a Gaussian in x with standard deviation sqrt(N)/2. Including the subleading Stirling term sharpens the prefactor to Omega_max = 2^N sqrt(2/(pi N)). Check that against a number you can compute exactly: at N = 100, 2^100 sqrt(2/(100 pi)) = 1.2677 x 10^30 x 0.079788 = 1.0113 x 10^29, while the exact C(100, 50) = 1.00891 x 10^29. The approximation is high by two parts in a thousand.

Why 10 to the 22 is a different kind of number

Fractional magnetization is the useful variable: mag = 2x/N, running from -1 to +1. Substituting into the Gaussian,

Omega(mag) / Omega_max = e^(-N mag^2 / 2).

Everything follows from that exponent. The width of the peak in the variable mag is 1/sqrt(N). For ten coins that is 32 percent, which is why the table above is broad. For a mole it is 1/sqrt(6 x 10^23) = 1.3 x 10^-12.

Do the suppression explicitly. Take a crystal with N = 1022 spins and ask for the multiplicity of a macrostate whose magnetization is one part in a billion away from zero, mag = 10^-9. The exponent is N mag^2 / 2 = 10^22 x 10^-18 / 2 = 5000, so the macrostate is suppressed by e^-5000. Converting, 5000 / ln(10) = 2171, so that macrostate is less likely than the balanced one by a factor of 102171.

To feel that size: suppose the crystal reshuffled its spins once every Planck time, 5.4 x 10^-44 seconds, for the whole 13.8 billion year age of the universe. That is about 8 x 10^60 tries. Multiply 10^-2171 by 10^61 and you still have 10^-2110. A deviation of one part in a billion has never happened anywhere and never will.

The upshot: the second law is not a new law layered on top of mechanics. It is the statement that a Gaussian of relative width 10^-11 looks, to any instrument, like a delta function.

Common misconceptions

  • "An ordered microstate is less likely than a disordered one." No. Every microstate has probability 2^-N. All-heads and any particular scrambled sequence are equally likely. It is the macrostate labelled "half heads" that is likely, because 1029 different sequences carry that label and only one carries "all heads".
  • "The second law forbids fluctuations." It does not. It makes them small in relative terms, and the smallness is the 1/sqrt(N) above. For a colloidal particle a micron across, with only about 1010 molecules hitting it per second from each side, the imbalance is visible under a microscope as Brownian motion. Lesson 5 makes fluctuations quantitative and shows they are not noise but measurable physics.
  • "Entropy is disorder." The metaphor is a liability. A crystal of ice at 200 K has a perfectly ordered lattice and a substantial entropy, and there are ordered-looking states with high entropy and messy-looking states with low entropy. Entropy is the logarithm of a count of microstates compatible with what you know about the system, and nothing more.
  • "Bigger N just makes the typical outcome more probable." That undersells it. The peak probability of any single macrostate actually falls as N grows, since Omega_max/2^N = sqrt(2/(pi N)) shrinks. What grows is sharpness: the band of macrostates carrying essentially all the probability narrows to nothing.

Entropy, and the number on Boltzmann's monument

Multiplicities multiply. Put two independent systems side by side and the combined multiplicity is Omega_1 Omega_2. Experimentally, though, the quantity that adds when you put two systems together is entropy. There is exactly one function that turns products into sums, so define

S = k ln(Omega),

with k = 1.380649 x 10^-23 J/K, now an exactly defined constant since the 2019 revision of the SI. This is Boltzmann's entropy formula, and it is the hinge of the subject: it connects a count, which is combinatorics, to a thermodynamic variable that appears in engine efficiencies and chemical equilibrium constants.

A historical detail worth knowing. The inscription S = k log W on Boltzmann's grave in the Vienna Zentralfriedhof was carved after his death in 1906. Boltzmann never wrote the relation in that compact form; the notation, and the constant k as a named quantity, are due to Planck around 1900. Boltzmann's own 1877 paper works with the combinatorics and the proportionality without the modern shorthand.

Evaluate S for the paramagnet with the field off, so all 2N microstates are accessible. Then S = k ln(2^N) = N k ln 2. Per mole, S = R ln 2 = 8.314 x 0.6931 = 5.76 J/(mol K). That number is not decorative. It is the entropy that magnetic cooling removes: William Giauque's 1933 demonstration of adiabatic demagnetization worked by aligning spins in a strong field, which strips out this R ln 2 per mole of spins, then removing the field under thermal isolation so the spins must take the entropy back from the lattice, cooling the sample. He reached 0.25 K, and the technique later reached microkelvins. Giauque received the 1949 Nobel Prize in Chemistry for it.

Try it: eight dipoles, exactly and approximately

Exercise. For N = 8, compute the multiplicity of the macrostate with three dipoles up, exactly and by the Gaussian formula, and compare.

Exact. C(8, 3) = 8!/(3! 5!) = 40320/(6 x 120) = 56. The total is 2^8 = 256, so this macrostate carries 21.9 percent of the microstates.

Gaussian. Here x = 3 - 4 = -1. The peak value is Omega_max = 2^8 sqrt(2/(8 pi)) = 256 x 0.28209 = 72.22, and the exact peak C(8,4) = 70, so the prefactor is high by 3 percent at this tiny N. The Gaussian factor is e^(-2(1)^2/8) = e^-0.25 = 0.7788, giving 72.22 x 0.7788 = 56.2 against the exact 56.

Change one input. Now ask for seven up, x = 3. Exact: C(8,7) = 8. Gaussian: 72.22 x e^(-18/8) = 72.22 x 0.1054 = 7.6. Still close, but notice the trend: the Gaussian is a small-x expansion, and out in the tail it drifts. At N = 1022 the region where it is inaccurate lies so far from anything observable that the drift never matters.

Where this leaves us

A microstate is a complete configuration; a macrostate is what you can measure; multiplicity is how many of the first sit inside the second. For the two-state paramagnet the multiplicity is a binomial coefficient, and Stirling's approximation turns it into a Gaussian of relative width 1/sqrt(N). At laboratory N that width is around 10-11, which is why macroscopic systems appear to obey deterministic laws even though the microscopic dynamics is nothing of the kind. Entropy is k ln(Omega), chosen so that it adds when multiplicities multiply, and a mole of free spins carries R ln 2 = 5.76 J/(mol K) of it.

The paramagnet has one weakness as a teaching system: its energy is bounded above, which will produce something startling in Lesson 3. Before that, the next lesson counts a system with no such ceiling, puts two copies of it in thermal contact, and watches energy flow from one to the other for no reason except that more microstates lie in that direction.

Sources

  1. Wikipedia contributors. (n.d.). Multiplicity (statistical mechanics). Wikipedia. en.wikipedia.org
  2. Wikipedia contributors. (n.d.). Boltzmann's entropy formula. Wikipedia. en.wikipedia.org
  3. Greytak, T., and Kardar, M. (2013). 8.044 Statistical Physics I. MIT OpenCourseWare. ocw.mit.edu
  4. Schroeder, D. V. (2000). An introduction to thermal physics, Chapter 2. Addison-Wesley.
  5. Kittel, C., and Kroemer, H. (1980). Thermal physics (2nd ed.), Chapter 1. W. H. Freeman.
Key terms
Microstate
A complete specification of every degree of freedom of a system, such as the orientation of every dipole.
Macrostate
A specification of only the measurable bulk quantities, such as total energy or magnetization.
Multiplicity
The number of microstates belonging to a given macrostate, written Omega.
Fundamental postulate
For an isolated system in equilibrium, every accessible microstate is equally probable.
Stirling's approximation
ln(N!) = N ln N - N + (1/2)ln(2 pi N), accurate to five decimal places already at N = 100.
Boltzmann entropy
S = k ln(Omega), the definition that makes entropy additive when multiplicities multiply.
Two-state paramagnet
N fixed dipoles, each pointing along or against an applied field, with energy -mB or +mB.
Adiabatic demagnetization
Cooling by aligning spins in a field, isolating the sample, then removing the field so the spins absorb entropy from the lattice.

The Einstein Solid, Worked with Actual Multiplicities

  • Derive the multiplicity formula for an Einstein solid by the dots-and-dividers argument and evaluate it exactly for small systems.
  • Build the full multiplicity table for two solids in thermal contact and identify the equilibrium macrostate by inspection.
  • Show that in the large-N limit the multiplicity of the combined system is a Gaussian of relative width one over the square root of N, and convert that into a statement about the second law.

In November 1906 Einstein sent a four-page paper to the Annalen der Physik that treated a solid as a collection of identical quantum oscillators, every one vibrating at the same frequency. As a description of a real crystal the model is wrong, and Lesson 12 shows precisely how and repairs it. As a machine for counting microstates it is close to perfect, because its multiplicity has a closed form that you can evaluate by hand and check against a list you write out yourself.

Here is the model. Take N one-dimensional quantum harmonic oscillators, all with the same frequency f. Quantum mechanics allows each one the energies 0, hf, 2hf, 3hf and so on, so measure energy in units of hf and say that oscillator i holds ni units. The macrostate is the total, q = n_1 + n_2 + ... + n_N. The microstate is the full list of n values. The question is how many lists give a particular total.

Dots and dividers

Represent a microstate as a row of symbols: q dots for the quanta, and N - 1 vertical bars separating the oscillators. The row . . | | . means two quanta in oscillator 1, none in oscillator 2, one in oscillator 3. Every arrangement of the row is a distinct microstate, and every microstate corresponds to exactly one arrangement. There are q + N - 1 symbols in total and you are choosing which q of the positions hold dots, so

Omega(N, q) = (q + N - 1)! / [q! (N - 1)!].

Test it where you can check by hand. Three oscillators, three quanta: Omega = 5!/(3! 2!) = 10. Write out all ten and confirm: (3,0,0), (0,3,0), (0,0,3), (2,1,0), (2,0,1), (1,2,0), (0,2,1), (1,0,2), (0,1,2), (1,1,1). Ten, and no eleventh.

Notice what the formula does not contain. There is no upper limit on q. Unlike the paramagnet of Lesson 1, whose energy is capped once every dipole is anti-aligned, an Einstein solid can absorb energy without limit, and Omega grows monotonically with q. That difference will matter enormously in Lesson 3.

Two solids in contact, counted exactly

Now the experiment. Solid A has N_A = 3 oscillators, solid B has N_B = 3, and between them they hold q = 6 quanta. Wrap the pair in perfect insulation so the total is fixed, but let them exchange energy freely. Each value of qA is a macrostate of the combined system, and its multiplicity is Omega_A(q_A) x Omega_B(6 - q_A), because any configuration of A can pair with any configuration of B.

For three oscillators, Omega(3, q) = (q + 2)(q + 1)/2, giving 1, 3, 6, 10, 15, 21, 28 for q = 0 to 6.

qAOmegaAqBOmegaBOmegaA x OmegaBShare
01628286.1%
135216313.6%
264159019.5%
31031010021.6%
415269019.5%
521136313.6%
62801286.1%

The column of products sums to 462, and you can check that independently: the combined system is one Einstein solid with six oscillators and six quanta, so Omega = C(11, 6) = 462. The peak sits at equal sharing, holding 100 microstates out of 462.

Read the table as a physicist rather than a combinatorialist. Suppose you prepare the pair with all six quanta in A and then allow contact. The system does not move toward q_A = 3 because energy prefers to be shared. It moves there because 100 microstates sit at q_A = 3 and 28 sit at q_A = 6, and the system wanders among microstates without preference. Heat flow is a random walk on a state space whose regions have wildly different sizes.

The point: nothing in this argument mentions temperature, and nothing forbids energy flowing back into A. With six quanta it happens 6.1 percent of the time.

Change one input: make B a hundred times the size

Keep q = 6 but let N_A = 3 and N_B = 30. Now Omega_B(q) = C(q + 29, 29), which runs 1, 30, 465, 4960, 40920, 278256, 1623160.

qA0123456
OmegaA x OmegaB1,623,160834,768245,52049,6006,97563028

The peak has moved to q_A = 0, which carries 58.8 percent of the 2,760,681 microstates. The small solid gives up essentially everything to the large one, which is what you would say in ordinary language as: a hot pebble dropped into a lake ends at the lake's temperature, not halfway.

There is an exact result hiding in that table, and it is worth extracting because it is one of the few things in this subject that is true for every N and every q. All N_A + N_B oscillators are interchangeable in the counting, so the average number of quanta per oscillator is the same for every one of them, namely q/(N_A + N_B). Therefore

average q_A = q N_A / (N_A + N_B), exactly.

Check it on the table: 6 x 3/33 = 0.5455. Computing the mean directly from the seven products gives 1,505,826/2,760,681 = 0.5455. The two agree to every digit. Note that the average, 0.55, is not the most probable value, 0. At these tiny numbers the distribution is strongly skewed, and mean and mode part company. They converge as N grows, which is the next thing to establish.

The large-N limit, and where the sharpness comes from

For q much larger than N, and N itself large, write the multiplicity as (q + N)!/(q! N!) to leading order and apply Stirling to all three factorials:

ln Omega = (q + N) ln(q + N) - q ln q - N ln N.

Expand ln(q + N) = ln q + N/q and keep terms through order N. The q ln q pieces cancel, one factor of N survives from the cross term, and you are left with

ln Omega = N ln(q/N) + N, that is, Omega = (e q / N)^N.

Put two such solids together, each with N oscillators, sharing q quanta, and set q_A = q/2 + x:

Omega_total = (e/N)^(2N) (q_A q_B)^N = (e/N)^(2N) (q^2/4 - x^2)^N.

Factor out the peak and expand the logarithm: N ln(1 - 4x^2/q^2) = -4Nx^2/q^2. So

Omega_total(x) = Omega_max e^(-4 N x^2 / q^2),

a Gaussian whose standard deviation in x is q/(2 sqrt(2N)). In relative terms, the energy of solid A fluctuates by a fraction 1/(2 sqrt(2N)) about half the total.

Numbers make this concrete. One gram of copper contains about 9.5 x 10^21 atoms, hence roughly 2.8 x 10^22 oscillators. Take N = 10^22. The relative fluctuation in how the energy divides between two such grams is 1/(2 sqrt(2 x 10^22)) = 3.5 x 10^-12. To see a one-part-per-million imbalance you would need the exponent 4 N (x/q)^2 = 4 x 10^22 x 10^-12 = 4 x 10^10, so that macrostate is suppressed by e^(-4 x 10^10), a number with seventeen billion zeros after the decimal point.

Why this matters: the second law is not an extra postulate. It is e^(-4 x 10^10).

Common misconceptions

  • "Energy spreads out because systems seek disorder." Systems do not seek anything. The dynamics wanders among microstates with no preference, and the destination is simply where nearly all the microstates are. Replace the verb and the mystery goes away.
  • "Equilibrium means the energy split is exactly equal." It means the split sits within the width of the peak. For two identical solids that peak is centred on equal sharing, but for unequal solids it is not, as the second table shows. What equalises at equilibrium is temperature, which Lesson 3 defines.
  • "The reverse process is impossible." It is merely rare, and the rarity is quantified. At q = 6 with two three-oscillator solids, the state with all the energy back in A has probability 28/462, about one time in sixteen. At laboratory sizes it is one time in e^(4 x 10^10), which is a different situation but not a different principle.
  • "Omega grows without bound, so entropy is infinite." Omega does grow with q, but S = k ln(Omega) grows only logarithmically. For the high-temperature Einstein solid, S = Nk[ln(q/N) + 1], so doubling the energy adds only Nk ln 2.

Entropy, and what actually increased

Apply S = k ln(Omega) to the two-solid experiment. Before contact, all six quanta sit in A and the pair has 28 x 1 = 28 accessible microstates, so S_initial = k ln 28 = 3.33 k. After contact, all 462 microstates of the combined system are accessible, so S_final = k ln 462 = 6.14 k. The increase is k ln(462/28) = 2.80 k.

That is the whole content of the second law of thermodynamics in this model: removing a constraint makes more microstates accessible, and the logarithm of the count goes up. Note carefully what is not being claimed. Entropy did not increase because anything irreversible happened at the level of the equations of motion, which are perfectly reversible. It increased because the accessible region of state space got bigger when the insulating partition came out. Lesson 17 takes that observation and pushes on it until it hurts.

In the large-N limit the same statement reads S_total(x) = S_max - 4 N k x^2/q^2, a parabola whose curvature is set by N. Maximising the entropy over how the energy divides is exactly maximising the multiplicity, because the logarithm is monotonic.

Try it: four and four

Exercise. Two Einstein solids, four oscillators each, sharing eight quanta. Build the table and find the equilibrium share.

Solution. Omega(4, q) = C(q + 3, 3), which runs 1, 4, 10, 20, 35, 56, 84, 120, 165 for q = 0 to 8. The products Omega_A(q_A) Omega_B(8 - q_A) are 165, 480, 840, 1120, 1225, 1120, 840, 480, 165. They sum to 6435, and the independent check is C(15, 8) = 6435. The peak at q_A = 4 carries 1225/6435 = 19.0 percent.

Compare the three symmetric cases you now have: with three oscillators each the peak held 21.6 percent, with four it holds 19.0 percent. The peak share is falling, because the distribution is spreading over more macrostates. What is rising is the sharpness in the fractional variable, and only that. Never confuse the two.

Putting it together

The Einstein solid has multiplicity C(q + N - 1, q), derived by arranging q dots among N - 1 dividers. Two solids in thermal contact form a combined system whose macrostates are labelled by the energy split, with multiplicity the product of the two. Written out by hand for small numbers, the table shows heat flow as nothing more than the system finding the macrostate with the most microstates. In the limit of many oscillators and many quanta, Omega = (eq/N)^N per solid, and the combined multiplicity is a Gaussian of relative width 1/(2 sqrt(2N)), which is 3.5 x 10^-12 for a gram of copper. Entropy is the logarithm of the count, and it increases when a constraint is removed because more microstates become accessible.

Both systems so far, the paramagnet and the solid, have been described entirely by counting. Neither lesson has said what temperature is. That is the next question, and the answer is a derivative that looks arbitrary until you see what goes wrong with every other candidate.

Sources

  1. Wikipedia contributors. (n.d.). Einstein solid. Wikipedia. en.wikipedia.org
  2. Likharev, K. K. (n.d.). Essential graduate physics: Statistical mechanics. Physics LibreTexts. phys.libretexts.org
  3. Einstein, A. (1907). Die Plancksche Theorie der Strahlung und die Theorie der spezifischen Waerme. Annalen der Physik, 22, 180-190.
  4. Schroeder, D. V. (2000). An introduction to thermal physics, Sections 2.2 to 2.6. Addison-Wesley.
Key terms
Einstein solid
N identical quantum harmonic oscillators of the same frequency, sharing q quanta of energy.
Dots and dividers
The combinatorial argument giving Omega = C(q + N - 1, q) by arranging q dots among N - 1 bars.
Thermal contact
A connection allowing two systems to exchange energy while their total remains fixed.
Equilibrium macrostate
The energy division with the largest multiplicity, which nearly all microstates of the combined system belong to.
High-temperature limit
The regime q much greater than N, where Omega = (eq/N)^N and the entropy is Nk[ln(q/N) + 1].
Sharpness
The relative width of the multiplicity peak, which falls as one over the square root of the number of oscillators.
Exchangeability
The property that all oscillators are equivalent in the counting, which makes the mean energy share exactly proportional to N.

Why Temperature Has to Be 1/T = dS/dU

  • Show where the definition of temperature as average kinetic energy succeeds and identify three systems on which it fails outright.
  • Derive the equilibrium condition for two systems in thermal contact and define temperature by 1/T = the partial derivative of S with respect to U.
  • Apply the definition to the Einstein solid and the two-state paramagnet, obtaining Dulong-Petit and negative temperature respectively.
  • State the third law and explain the residual entropy measured in ice and carbon monoxide.

Kinetic theory gives a clean and correct result: for a monatomic ideal gas the mean translational kinetic energy per molecule is (3/2)kT, so at 300 K every argon atom carries 1.5 x 1.380649 x 10^-23 x 300 = 6.21 x 10^-21 joules. From there it is one short step to a definition: temperature is a measure of average kinetic energy. That step is a mistake, and the fastest way to see it is to point at a system that plainly has a temperature and has no kinetic energy whatsoever.

Where the kinetic-energy definition works, and where it stops

The definition survives a monatomic gas because for that system, and only through a coincidence you will derive in Lesson 7, energy really is proportional to T. Now take it elsewhere.

Failure one: the paramagnet. Lesson 1's dipoles are pinned to lattice sites. They do not move. Their entire energy is -mB or +mB per dipole, and none of it is kinetic. Yet a spin system has a temperature, you can measure it, and it can be brought into equilibrium with a gas. A definition that cannot assign a temperature to it is not a definition of temperature.

Failure two: the cold solid. Diamond at 100 K has a molar heat capacity near 0.1 J/(mol K), against the 3R = 24.9 a classical solid would show. Its energy per vibrational mode is nowhere near kT, and the energy-temperature relation is exponential rather than linear. Define temperature as energy per mode and diamond at 100 K gets assigned a few kelvin, whereupon it refuses to equilibrate with anything else at that temperature.

Failure three, the fatal one. Even where the definition gives a number, it does not explain what temperature is for. Why does energy flow from hot to cold and not the reverse? Nothing in the phrase "average kinetic energy" answers that. What you need is a quantity that is equal at equilibrium and whose imbalance drives energy in a definite direction. Find it and you have found temperature, whatever it looks like.

What actually equalises

Lesson 2 built the apparatus. Two systems, A and B, in thermal contact, total energy U fixed. The multiplicity of the combined system as a function of the split is Omega_A(U_A) Omega_B(U - U_A). Equilibrium is the maximum of that product, so set the derivative to zero. Taking logarithms first, since sums are easier than products:

d/dU_A [ln Omega_A(U_A) + ln Omega_B(U - U_A)] = 0

d(ln Omega_A)/dU_A = d(ln Omega_B)/dU_B,

where the minus sign from the chain rule has been used to flip the second derivative onto U_B. Multiply through by Boltzmann's constant and this reads

(dS_A/dU_A) = (dS_B/dU_B).

There it is. The quantity taking a common value when two systems stop exchanging energy is the derivative of entropy with respect to energy. It is defined for any system whose microstates you can count, and it emerges from the counting rather than being imposed on it.

Now check the direction of flow. If dS_A/dU_A is larger than dS_B/dU_B, then moving a small amount of energy from B into A raises the total entropy, so the combined system drifts that way. A large value of dS/dU therefore means "energy is drawn here", which is what we ordinarily call cold. Since we want the hot body to have the large number, define

1/T = (dS/dU) at fixed volume and particle number.

The reciprocal is not cosmetic. It turns "the system that gains entropy fastest per joule" into "the colder system", which is the convention every thermometer follows, and it reproduces the ideal-gas scale exactly, as the next section shows. The combination beta = 1/kT appears so often that it has its own name, and it is arguably the more natural variable.

What matters here: temperature is not a property a system has in isolation so much as the exchange rate between energy and entropy at that system's current state.

Testing the definition on systems you have already counted

Take the high-temperature Einstein solid from Lesson 2, where Omega = (eq/N)^N and U = q times the quantum hf. Then

S = Nk[ln(U/(N hf)) + 1], so 1/T = dS/dU = Nk/U, giving U = NkT.

Every oscillator carries kT on average. A real three-dimensional solid has three oscillators per atom, so U = 3NkT and the molar heat capacity is C = 3R = 24.94 J/(mol K). That is the Dulong-Petit law of 1819, and here it drops out of counting quanta. Measured room-temperature values: copper 24.44, aluminium 24.20, silver 25.35, lead 26.44 J/(mol K). The agreement is good to a few percent for these metals and catastrophic for diamond, which is the story of Lesson 12.

Do the same for a monatomic ideal gas, whose entropy you will derive properly in Lesson 6. Its energy-dependence is S = Nk[(3/2) ln U + terms independent of U], so 1/T = (3/2)Nk/U and U = (3/2)NkT. The kinetic-theory result comes back out, now as a theorem about one particular system rather than as a definition.

The same machinery gives the other two intensive variables for free. Allowing the volume to be exchanged as well and maximising the total entropy gives (dS/dV) at fixed U and N = P/T, and allowing particles to be exchanged gives (dS/dN) at fixed U and V = -mu/T, where chemical potential mu is the subject of Lesson 5. All three intensive quantities are slopes of one surface, S(U, V, N).

Common misconceptions

  • "Temperature is average kinetic energy." True for a monatomic ideal gas and false in general. Spin systems have no kinetic energy; solids below their Debye temperature have energy per mode far below kT; and a relativistic gas has a different constant of proportionality again.
  • "A hotter body contains more heat." Heat is energy in transit, not a stored property. A bathtub at 30 C holds far more internal energy than a spark at 1000 C. Temperature says which way energy moves, not how much is there.
  • "Absolute zero is where all motion stops." Quantum mechanics forbids it: an oscillator keeps its zero-point energy hf/2 at T = 0. What vanishes is entropy, not motion.
  • "Negative temperature means colder than absolute zero." The opposite. A negative-temperature system gives up energy to any positive-temperature system it touches, so it is hotter than all of them. The scale of hotness runs 0+, then up through positive values to plus infinity, which is the same physical state as minus infinity, then up through negative values to 0-.

Negative temperature, and a Harvard basement in 1951

Return to the paramagnet, whose energy is bounded above. Let x be the fraction of dipoles aligned minus the fraction anti-aligned, so U = -N m B x. The entropy from Lesson 1 is

S = -Nk[((1+x)/2) ln((1+x)/2) + ((1-x)/2) ln((1-x)/2)],

which is zero at x = 1, rises to Nk ln 2 at x = 0, and falls back to zero at x = -1. Differentiating and using dU = -N m B dx,

1/T = (k/(2mB)) ln((1 + x)/(1 - x)).

Follow the consequences. When most dipoles are aligned, x is near 1, the logarithm is large, and T is small and positive. As energy is added, x falls to zero and T runs off to plus infinity. Add one more joule so that more than half the dipoles are anti-aligned, and the logarithm changes sign: T is now negative. This is not a pathology of the algebra. It is what happens whenever a system has a maximum possible energy, because then S must eventually decrease with U, and dS/dU must eventually be negative.

The state is real and it has been made. In 1951 Edward Purcell and Robert Pound at Harvard polarised the nuclear spins of a lithium fluoride crystal in a strong field, then reversed the field faster than the spins could follow, leaving a population inversion. Because the spin-lattice relaxation time in LiF was around five minutes, far longer than the spin-spin equilibration time, the spins were internally in equilibrium and could be assigned a temperature. That temperature was negative, and the sign says the spin system will give up energy to anything at any positive temperature.

Remember: negative temperatures require an upper bound on the energy. A gas has none, because you can always make the molecules faster, which is why negative temperatures are a spin-system phenomenon rather than a general one.

The third law, and the entropy ice refuses to give up

What happens as T approaches zero? The system settles into its lowest-energy configuration. If that ground state is unique, Omega = 1 and S = k ln 1 = 0. That is the third law of thermodynamics in its Planck form, and it does real work: it fixes the zero of the entropy scale, so that absolute entropies can be tabulated rather than only differences.

The interesting cases are those where the ground state is not unique. Ice is the classic. Each oxygen sits in a tetrahedral cage of four hydrogen bonds, and the Bernal-Fowler rules require exactly two hydrogens close and two far, a constraint satisfiable in many global ways. Linus Pauling estimated in 1935 that the count is about (3/2)^N, predicting a residual molar entropy of R ln(3/2) = 3.37 J/(mol K). William Giauque and John Stout measured ice calorimetrically in 1936 and found about 3.4. Carbon monoxide behaves similarly: its dipole moment is tiny, head-to-tail ordering is nearly free, and the predicted R ln 2 = 5.76 is measured as roughly 4.6.

That is not a failure of the third law. These crystals never reach their true ground state on any laboratory timescale, so they freeze in a disordered configuration and keep its entropy. The measured number counts the frozen-in configurations, which is S = k ln(Omega) read directly off a calorimeter.

Try it: which way does the heat go?

Exercise. Solid A has three oscillators holding five quanta; solid B has three oscillators holding one. Which is hotter, and by how much?

Solution. Use a symmetric difference on the discrete entropy, 1/T = [S(q+1) - S(q-1)]/(2 hf). For A, Omega(3, 6) = 28 and Omega(3, 4) = 15, so 1/T_A = k ln(28/15)/(2hf) = 0.3121 k/hf and T_A = 3.20 hf/k. For B, Omega(3, 2) = 6 and Omega(3, 0) = 1, so 1/T_B = k ln(6)/(2hf) = 0.8959 k/hf and T_B = 1.12 hf/k.

A is nearly three times hotter, so energy flows from A to B. Check that against the table in Lesson 2: the multiplicity peak for these two solids sits at q_A = 3, and the system is starting at q_A = 5, so it does indeed drift downward in q_A. The temperature calculation and the multiplicity table are two views of the same arithmetic.

Change one input. Put four quanta in A and two in B. Then 1/T_A = k ln(21/10)/(2hf) = 0.3710 k/hf and 1/T_B = k ln(10/3)/(2hf) = 0.6020 k/hf. A is still hotter and energy still flows the same way, but now the imbalance in 1/T is a factor of 1.6 rather than 2.9, and the system has only one quantum left to move.

What to carry forward

Temperature is not average kinetic energy; that identification is a special result for the monatomic ideal gas. The quantity that equalises when two systems exchange energy is dS/dU, and thermodynamic temperature is defined by 1/T = (dS/dU) at fixed V and N. On the Einstein solid it gives U = NkT per oscillator and the Dulong-Petit value 3R; on an ideal gas, (3/2)NkT; on a paramagnet above half occupancy, a negative temperature, which Purcell and Pound made in 1951. The third law fixes S = 0 at T = 0 for a unique ground state, and ice's residual 3.4 J/(mol K) against Pauling's 3.37 is a calorimetric reading of a combinatorial count.

You now have entropy, temperature and the second law, all from counting isolated systems. But isolated systems are rare. Almost everything you care about sits in a room, trading energy with surroundings far larger than itself. The next lesson asks what the counting looks like from inside such a system, and the answer is one exponential that runs the rest of the subject.

Sources

  1. Wikipedia contributors. (n.d.). Negative temperature. Wikipedia. en.wikipedia.org
  2. Wikipedia contributors. (n.d.). Third law of thermodynamics. Wikipedia. en.wikipedia.org
  3. National Institute of Standards and Technology. (2019). Boltzmann constant. CODATA fundamental physical constants. physics.nist.gov
  4. Purcell, E. M., and Pound, R. V. (1951). A nuclear spin system at negative temperature. Physical Review, 81(2), 279-280.
  5. Kittel, C., and Kroemer, H. (1980). Thermal physics (2nd ed.), Chapters 2 and 3. W. H. Freeman.
Key terms
Thermodynamic temperature
Defined by 1/T = the partial derivative of entropy with respect to energy at fixed volume and particle number.
Thermal equilibrium
The condition dS_A/dU_A = dS_B/dU_B, reached when the combined multiplicity is maximal.
Beta
The inverse temperature 1/kT, which runs smoothly through zero at infinite temperature.
Dulong-Petit law
The classical molar heat capacity of a solid, 3R = 24.94 J/(mol K), recovered from U = 3NkT.
Negative temperature
A state with dS/dU less than zero, possible only when the energy has an upper bound, and hotter than any positive temperature.
Third law
Entropy tends to zero as temperature tends to zero when the ground state is unique, fixing the zero of the entropy scale.
Residual entropy
Entropy frozen into a crystal that cannot reach its true ground state, such as 3.4 J/(mol K) in ice.
Chemical potential
Defined by the partial derivative of S with respect to N at fixed U and V, equal to -mu/T.

Module 2: Ensembles and Partition Functions

Move from isolated systems to systems in contact with the world. Derive the Boltzmann factor from a reservoir, build the canonical and grand canonical ensembles, extract thermodynamics from partition functions, and recover the classical ideal gas with its entropy computed to four figures.

The Boltzmann Factor, and What a Reservoir Does to a System

  • Derive the Boltzmann factor by expanding the reservoir entropy to first order and state the condition under which the expansion is valid.
  • Define the canonical partition function and obtain average energy, entropy, free energy and heat capacity from it.
  • Work the two-level system and the quantum harmonic oscillator completely, including their limiting behaviour at high and low temperature.

Pulling a copper atom out of its lattice site and parking it at the surface costs about 1.28 electronvolts. At 1300 K, just below copper's melting point, kT is 0.112 eV, so the factor e^(-1.28/0.112) = e^-11.43 = 1.1 x 10^-5, and roughly one lattice site in ninety thousand stands empty. Cool the same crystal to 300 K and the factor becomes e^-49.5 = 3.2 x 10^-22. With 8.5 x 10^19 atoms in a cubic millimetre of copper, that is about one vacancy for every forty cubic millimetres of metal.

Two lines, one exponential, and a prediction about diffusion rates, creep and annealing that materials scientists use daily. The exponential is the Boltzmann factor, and this lesson derives it from the counting you already have, then shows that almost everything else in equilibrium statistical mechanics is bookkeeping on top of it.

The problem with isolated systems

Everything in Module 1 assumed a fixed total energy, which let you count microstates and be done. But the copper crystal is not isolated, and neither is a molecule in the air above your desk, which trades energy with its neighbours billions of times a second. Its energy is not fixed, its microstates cannot be enumerated without enumerating the room's, and Lesson 1's postulate says nothing directly about it.

The fix is to keep the postulate and change what you apply it to. Let the small system be S and everything else be a reservoir R, and treat the combination as isolated. Now every microstate of the pair is equally likely, and the probability that S is in one specific microstate s, with energy E_s, is proportional to the number of reservoir microstates compatible with it:

P(s) is proportional to Omega_R(U_tot - E_s).

Note carefully that there is no Omega_S in that expression. The state s is one microstate, so it has multiplicity one. All the counting has moved into the reservoir.

The expansion, and why it terminates

Take logarithms and expand about U_tot, since E_s is minuscule by comparison:

ln Omega_R(U_tot - E_s) = ln Omega_R(U_tot) - E_s (d ln Omega_R / dU) + (1/2) E_s^2 (d^2 ln Omega_R / dU^2) - ...

The first derivative is exactly what Lesson 3 named: d(ln Omega_R)/dU = (1/k)(dS_R/dU) = 1/(kT). So the linear term is -E_s/kT. The second derivative is d(1/kT)/dU = -1/(kT^2 C_R), where C_R is the reservoir's heat capacity. Since C_R grows without bound as the reservoir grows, that term goes to zero, and so does everything after it. This is the precise sense in which "reservoir" means "large": large enough that its temperature does not shift measurably when the system borrows energy from it. Exponentiating,

P(s) = e^(-E_s/kT) / Z, with Z = sum over all microstates s of e^(-E_s/kT).

Z is the canonical partition function, and the set of probabilities is the canonical ensemble, a name Gibbs gave it in his 1902 book Elementary Principles in Statistical Mechanics. Write beta = 1/kT throughout; it keeps the algebra clean.

The core of it: the reservoir has been eliminated entirely. Its only surviving trace is one number, T.

Everything from Z

Z looks like a normalising denominator and is in fact a generating function. Differentiate it with respect to beta:

dZ/dbeta = -sum E_s e^(-beta E_s), so average E = -(1/Z) dZ/dbeta = -d(ln Z)/dbeta.

Differentiate once more and you get the mean square spread:

average of (E - average E)^2 = d^2(ln Z)/dbeta^2 = k T^2 C_V.

Entropy comes from the Gibbs formula S = -k sum P_s ln P_s. Substituting ln P_s = -beta E_s - ln Z gives S = k beta (average E) + k ln Z = U/T + k ln Z. Rearranging, U - TS = -kT ln Z, and the left side is the Helmholtz free energy. So

F = -kT ln Z,

from which S = -(dF/dT) at fixed V, P = -(dF/dV) at fixed T, and mu = (dF/dN) at fixed T and V. Compute Z and you have the complete thermodynamics of the system. That is why the rest of this course is largely a sequence of partition functions.

Worked example: the two-level system

One particle with two states, energies 0 and eps. Then Z = 1 + e^(-beta eps), and

average E = eps e^(-beta eps)/(1 + e^(-beta eps)) = eps/(e^(beta eps) + 1).

Check the limits. At T = 0 the average energy is zero: the system sits in its ground state. At T infinite it is eps/2: both states equally occupied, and never more than half, because there is nowhere else for the energy to go. Differentiating with respect to T for N independent copies,

C = Nk (beta eps)^2 e^(beta eps)/(e^(beta eps) + 1)^2.

This function is zero at both ends and has a hump in between, called a Schottky anomaly. Maximising numerically, the peak sits at beta eps = 2.40, that is kT = 0.417 eps, with height C = 0.439 Nk. For a splitting of 1 meV, and since 1 meV corresponds to 11.60 K, the anomaly peaks near 4.8 K. Experimentalists use exactly this: measure the temperature of the hump in a low-temperature heat capacity and you have read off a level splitting without ever doing spectroscopy.

Worked example: the quantum harmonic oscillator

Levels E_n = (n + 1/2) h_bar omega for n = 0, 1, 2, .... The sum is geometric:

Z = e^(-beta h_bar omega/2) sum over n of e^(-n beta h_bar omega) = e^(-beta h_bar omega/2)/(1 - e^(-beta h_bar omega)).

Then ln Z = -beta h_bar omega/2 - ln(1 - e^(-beta h_bar omega)), and differentiating,

average E = h_bar omega/2 + h_bar omega/(e^(beta h_bar omega) - 1).

The first term is the zero-point energy, present at every temperature. The second is the thermal part, and the factor 1/(e^(beta h_bar omega) - 1) is the average number of quanta in the oscillator. Keep that expression in view: in Lesson 10 it will reappear as the Bose-Einstein distribution, and in Lesson 11 it is what makes Planck's law work.

Two limits. When kT is much larger than h_bar omega, expand the exponential as 1 + beta h_bar omega and the thermal part becomes kT, recovering the classical result of Lesson 3. When kT is much smaller, the thermal part is h_bar omega e^(-beta h_bar omega), exponentially small, and so is the heat capacity: C = k (beta h_bar omega)^2 e^(-beta h_bar omega). The oscillator is frozen. That exponential freeze-out is Einstein's 1907 explanation of why diamond's heat capacity collapses, and Lesson 12 shows both what it gets right and what it gets wrong.

Common misconceptions

  • "The Boltzmann factor gives the probability of an energy." It gives the probability of a microstate. To find the probability of an energy you must multiply by the number of microstates at that energy: P(E) = g(E) e^(-beta E)/Z. Forgetting the degeneracy g(E) is the single most common error in the subject.
  • "High energies are effectively impossible." They are exponentially suppressed, but degeneracy often wins over a range. The Maxwell speed distribution peaks at a nonzero speed precisely because the number of velocity states grows as v^2 while the Boltzmann factor falls as e^(-mv^2/2kT), and the product has an interior maximum.
  • "The canonical ensemble is an approximation to the microcanonical." Neither is an approximation to the other. They describe different physical situations, fixed energy against fixed temperature, and they agree for macroscopic systems because the energy fluctuations are of relative size 1/sqrt(N).
  • "Z is just a normalisation constant." Z carries the whole thermodynamics through F = -kT ln Z. Its logarithm is a potential whose derivatives are the entropy, the pressure and the chemical potential.

How big are the fluctuations?

The canonical ensemble lets the energy vary, which sounds alarming for a system you thought had a definite energy. Use average of (Delta E)^2 = kT^2 C_V on a monatomic ideal gas, where C_V = (3/2)Nk and U = (3/2)NkT:

Delta E / U = sqrt((3/2) N k^2 T^2)/((3/2) N k T) = sqrt(2/(3N)).

For a mole that is 1.1 x 10^-12. Fixing the temperature therefore fixes the energy to twelve significant figures, which is why you may use whichever ensemble is algebraically easier and expect identical answers. It also says something you will exploit in Lesson 5: a fluctuation is not noise around the physics, it is a heat capacity. Measure how much the energy of a small system wanders and you have measured C_V.

Try it: the atmosphere as a Boltzmann factor

Exercise. A nitrogen molecule at height z has extra potential energy mgz. Predict the density profile of the atmosphere at 288 K and check it against experience.

Solution. The Boltzmann factor gives n(z) = n(0) e^(-mgz/kT), so the density falls by 1/e over a scale height H = kT/mg. For N2, m = 28.013 x 1.6605 x 10^-27 = 4.652 x 10^-26 kg, so mg = 4.563 x 10^-25 N, and kT = 1.3806 x 10^-23 x 288 = 3.976 x 10^-21 J. Then H = 3.976 x 10^-21/4.563 x 10^-25 = 8710 m, about 8.7 km.

Test it: at the 5500 m altitude of a high Andean pass, the prediction is e^(-5500/8710) = 0.53, so a little over half of sea-level density, which matches the standard atmosphere tables and matches how it feels to walk there. The real atmosphere is not isothermal, so 8.7 km is an idealisation, but the exponential form and the order of magnitude come straight out of one factor.

Change one input. Replace nitrogen with helium, m = 4.0026 u, seven times lighter. The scale height becomes 61 km, so helium is far less bound to the planet than nitrogen. Push the same logic to escape velocity and you have the reason Earth retains nitrogen and oxygen but has lost nearly all of its primordial helium and hydrogen.

The short version

A system in contact with a large reservoir is described by P(s) = e^(-E_s/kT)/Z, derived by expanding the reservoir's entropy to first order in the system's energy and using dS/dU = 1/T. The canonical partition function Z generates the thermodynamics: average E = -d(ln Z)/dbeta, F = -kT ln Z, and the energy variance is kT^2 C_V. Worked out for a two-level system it produces the Schottky anomaly, peaking at kT = 0.417 eps with height 0.439 Nk; worked out for a harmonic oscillator it produces the zero-point energy plus a thermal term that becomes kT at high temperature and freezes exponentially at low temperature. Energy fluctuations in a mole are one part in 10^12.

What the canonical ensemble cannot do is let particles come and go. For an electron gas in a metal, or a photon gas in a cavity, or any question about adsorption or chemical equilibrium, the particle number is not fixed either, and the next lesson builds the ensemble that handles it, then compares all three side by side.

Sources

  1. Wikipedia contributors. (n.d.). Boltzmann distribution. Wikipedia. en.wikipedia.org
  2. Wikipedia contributors. (n.d.). Partition function (statistical mechanics). Wikipedia. en.wikipedia.org
  3. Kardar, M. (2013). 8.333 Statistical Mechanics I: Statistical Mechanics of Particles. MIT OpenCourseWare. ocw.mit.edu
  4. Gibbs, J. W. (1902). Elementary principles in statistical mechanics. Charles Scribner's Sons.
  5. Kardar, M. (2007). Statistical physics of particles, Chapter 4. Cambridge University Press.
Key terms
Reservoir
A system so large that its temperature is unchanged by the energy a small system borrows from it.
Boltzmann factor
The weight e^(-E/kT) attached to a single microstate of energy E.
Canonical ensemble
The probability distribution over microstates of a system held at fixed temperature, volume and particle number.
Partition function
Z, the sum of Boltzmann factors over all microstates, whose logarithm generates the thermodynamics.
Helmholtz free energy
F = U - TS = -kT ln Z, the potential minimised at fixed T, V and N.
Degeneracy
The number of microstates sharing a given energy, which must multiply the Boltzmann factor to give the probability of that energy.
Schottky anomaly
The hump in heat capacity from a two-level system, peaking at kT = 0.417 times the splitting with height 0.439Nk.
Zero-point energy
The residual h_bar omega/2 that a quantum oscillator retains at absolute zero.

Three Ensembles, and the Fluctuations That Connect Them

  • Set out the microcanonical, canonical and grand canonical ensembles side by side, naming what each holds fixed and which thermodynamic potential each generates.
  • Derive the grand canonical weight and the grand potential, and interpret the chemical potential numerically for a real gas.
  • Relate energy fluctuations to heat capacity and particle-number fluctuations to isothermal compressibility, and state the fluctuation-dissipation relation with two measured examples.

Seal a cavity, hold its walls at 3000 K, and count the photons inside: about 5 x 10^17 per cubic metre. Count again an instant later and the number has changed, because the walls create and destroy photons freely and nothing conserves their number. Lesson 4's partition function cannot describe that cavity, since Z sums over the microstates of a system with a fixed particle count. The same difficulty arises for conduction electrons in a metal in contact with a lead, for molecules adsorbing onto a catalyst, and for any chemical reaction at equilibrium.

The repair is to run the reservoir argument once more, letting the reservoir supply particles as well as energy. What comes out is the third of Gibbs's ensembles, and once it is in place you have a complete toolkit, so this lesson also lays the three of them out together and shows what each one is good for.

The three ensembles, side by side

EnsemblePhysical pictureHeld fixedWeight of one microstateSumPotential
MicrocanonicalSealed, insulated boxU, V, NEqual for every accessible stateOmegaS = k ln Omega
CanonicalClosed box in a heat bathT, V, Ne-beta EZF = -kT ln Z
Grand canonicalOpen box in a heat and particle bathT, V, muebeta(mu N - E)XiPhi = -kT ln Xi = -PV

Reading down the table, each step trades a fixed extensive quantity for a fixed intensive one that controls it: energy for temperature, then particle number for chemical potential. Each trade makes a hard constrained sum into an easier unconstrained one, which is why the grand canonical ensemble is the natural setting for quantum statistics in Lesson 10, even though nobody genuinely has a box with a fixed chemical potential and a wandering electron count.

Deriving the grand canonical weight

Repeat Lesson 4's argument with a reservoir that shares both energy and particles. The probability of finding the system in a microstate s with energy E_s and particle number N_s is proportional to Omega_R(U_tot - E_s, N_tot - N_s). Expand its logarithm to first order in both arguments:

ln Omega_R = const - E_s (dS_R/dU)/k - N_s (dS_R/dN)/k.

Lesson 3 gave both derivatives: dS/dU = 1/T and dS/dN = -mu/T. Substituting,

P(s) = e^(beta(mu N_s - E_s)) / Xi, with Xi = sum over all s of e^(beta(mu N_s - E_s)).

It is often useful to group the sum by particle number: Xi = sum over N of z^N Z_N, where z = e^(beta mu) is the fugacity and Z_N is the ordinary canonical partition function for exactly N particles. The grand canonical ensemble is a generating function in z for the canonical ensembles.

The corresponding potential is the grand potential Phi = -kT ln Xi, and there is a tidy identity behind it. Since Phi = U - TS - mu N = F - mu N, and Euler's relation for a homogeneous system gives G = U - TS + PV = mu N, it follows that F = mu N - PV and therefore

Phi = -PV.

So ln Xi = PV/kT. Computing a grand partition function hands you the equation of state directly, which is exactly what Lesson 10 will exploit for a gas of bosons and a gas of fermions.

What the chemical potential actually is

The usual definitions are mu = (dU/dN) at fixed S and V and, more usefully, mu = (dF/dN) at fixed T and V. In words: the free energy cost of adding one more particle to the system while holding the temperature and volume fixed. Particles flow from high chemical potential to low, exactly as energy flows from high temperature to low, and equilibrium against a particle exchange means equal mu.

The sign trips people, so put a number on it. For a classical ideal gas, Lesson 6 will derive

mu = kT ln(n / n_Q), where n_Q = (2 pi m kT / h^2)^(3/2) is the quantum concentration.

Take helium at 300 K and 1 bar. The thermal de Broglie wavelength is lambda = h/sqrt(2 pi m k T) = 5.04 x 10^-11 m, so n_Q = 1/lambda^3 = 7.82 x 10^30 per cubic metre. The actual density is n = P/kT = 2.41 x 10^25 per cubic metre. Their ratio is 3.09 x 10^-6, whose logarithm is -12.69, so

mu = -12.69 kT = -5.26 x 10^-20 J = -0.328 eV.

Negative, and by a lot. That is normal for a dilute gas: there are so many more single-particle states available than particles to fill them that adding a particle raises the entropy substantially, and the free energy F = U - TS falls. Chemical potential goes to zero only when n approaches n_Q, which is the boundary of quantum degeneracy and the subject of Lessons 13 and 14.

Why this matters: the ratio n/n_Q is the single number deciding whether a gas is classical. At 3 x 10^-6, room-temperature helium is emphatically classical. Cool it to 4 K and compress it and the ratio approaches 1, and the physics changes character.

Common misconceptions

  • "The three ensembles are three approximations, and one must be the right one." They describe three different physical arrangements, and for a macroscopic system they give identical answers for every measurable quantity, because the fluctuations that distinguish them are of relative size 1/sqrt(N). Choose whichever makes the algebra easiest. For small systems, and near a critical point, the equivalence genuinely fails, and then you must model the actual boundary conditions.
  • "Chemical potential is a kind of energy per particle, so it must be positive." It is a free energy per particle, and the entropy term makes it strongly negative for any dilute gas. Only when a system is degenerate, as with electrons in a metal, does mu become positive and large.
  • "A fluctuation is experimental noise." It is the signal. The variance of the energy is kT^2 C_V, and the variance of the particle number is proportional to the compressibility, so both are laboratory quantities you can independently measure.
  • "Photons have a chemical potential like anything else." Photon number is not conserved, so there is no constraint to attach a Lagrange multiplier to, and mu = 0. That single fact is what makes Planck's law come out right in Lesson 11.

Fluctuations are response functions

Lesson 4 showed average of (Delta E)^2 = kT^2 C_V in the canonical ensemble. The grand canonical ensemble gives the analogous statement for particle number. Differentiating average N = kT (d ln Xi/d mu) once more with respect to mu,

average of (Delta N)^2 = kT (d average N/d mu) at fixed T and V.

A short thermodynamic manipulation converts the right-hand side into the isothermal compressibility kappa_T = -(1/V)(dV/dP) at fixed T:

average of (Delta N)^2 = (average N)^2 k T kappa_T / V.

Check it on an ideal gas. There kappa_T = 1/P, so the right side is N^2 kT/(PV) = N. The particle number is Poisson distributed, and the relative fluctuation is 1/sqrt(N), one part in 10^12 for a mole, exactly as expected.

Now notice what happens when a fluid approaches its critical point. There kappa_T diverges, so density fluctuations become macroscopic, and you can see them: a fluid held at its critical density and temperature scatters light strongly and turns milky, an effect called critical opalescence. Lesson 16 returns to it. The general pattern is worth stating plainly: a large fluctuation and a large response are the same phenomenon seen from two sides.

The fluctuation-dissipation relation, and two experiments that used it

The general form of that pattern is the fluctuation-dissipation theorem: the way a system in equilibrium fluctuates spontaneously and the way it dissipates energy when driven are governed by the same function. Two cases made the point historically.

Brownian motion. In 1905 Einstein derived D = mu_mob k T, where D is the diffusion coefficient of a suspended particle and mu_mob is its mobility, the drift velocity per unit applied force. For a sphere of radius a in a fluid of viscosity eta, Stokes gives mu_mob = 1/(6 pi eta a), so D = kT/(6 pi eta a). Diffusion is the fluctuation, viscous drag is the dissipation, and one equation ties them. Jean Perrin then measured the mean square displacement of resin particles under a microscope between 1908 and 1911, extracted k, and since the gas constant R was already known, obtained Avogadro's number in the region of 6 to 7 x 10^23, against the modern defined value 6.02214076 x 10^23. That work ended the serious scientific dispute about whether atoms exist, and won Perrin the 1926 Nobel Prize in Physics.

Johnson-Nyquist noise. In 1928 John Johnson at Bell Labs found that any resistor generates a small random voltage across its terminals even with no current flowing, and Harry Nyquist explained it in the same year with a thermodynamic argument: average of V^2 = 4 k T R (bandwidth). Again the fluctuation, the noise voltage, is fixed by the dissipation, the resistance. Put numbers on it: a 1 megohm resistor at 300 K, measured over a 10 kHz bandwidth, delivers sqrt(4 x 1.381 x 10^-23 x 300 x 10^6 x 10^4) = 1.29 x 10^-5 V, about 13 microvolts RMS. Every low-noise amplifier design starts from that number, and it is a statistical mechanics calculation.

Try it: choose the ensemble

Exercise. For each system, say which ensemble is the natural choice and why: (a) a gas of 1020 argon atoms in a rigid, insulated flask; (b) the same gas in a thin-walled flask sitting in a laboratory; (c) the conduction electrons in a copper wire attached to a battery terminal; (d) five protein molecules bound to a strand of DNA in a cell.

Solution. (a) Microcanonical: energy, volume and particle number are all fixed. In practice you would still compute with the canonical ensemble, because Delta E/U = sqrt(2/3N) = 8 x 10^-11 makes the two indistinguishable, and Z is far easier than Omega. (b) Canonical: the room fixes T, the walls fix V and N. (c) Grand canonical: the wire exchanges both energy and electrons with the rest of the circuit, and the chemical potential of the electrons is exactly what a voltmeter measures differences of. (d) Grand canonical, and this time it matters: with only five molecules the relative number fluctuation is 1/sqrt(5) = 45 percent, so the ensembles do not agree and the fluctuations are the biology.

Pulling it together

The microcanonical, canonical and grand canonical ensembles correspond to a sealed box, a box in a heat bath, and a box open to both heat and particles, and they generate S = k ln Omega, F = -kT ln Z and Phi = -kT ln Xi = -PV respectively. The grand canonical weight e^(beta(mu N - E)) follows from expanding the reservoir entropy in both energy and particle number. Chemical potential is the free energy cost of one more particle, equal to kT ln(n/n_Q) for a classical gas, which is -0.33 eV for helium at room temperature and pressure. Energy fluctuations give the heat capacity, number fluctuations give the compressibility, and the general statement is the fluctuation-dissipation theorem, which Perrin used to weigh the atom and Nyquist used to predict the 13 microvolts of noise across a megohm resistor.

You now have three ways to compute anything. The next lesson spends all three on the oldest problem in the subject, a box of monatomic gas, and gets an entropy that agrees with a calorimeter to four significant figures. It also produces a paradox that took thirty years and quantum mechanics to settle.

Sources

  1. Wikipedia contributors. (n.d.). Grand canonical ensemble. Wikipedia. en.wikipedia.org
  2. Wikipedia contributors. (n.d.). Johnson-Nyquist noise. Wikipedia. en.wikipedia.org
  3. The Nobel Foundation. (1926). The Nobel Prize in Physics 1926: Jean Baptiste Perrin. NobelPrize.org. nobelprize.org
  4. Pathria, R. K., and Beale, P. D. (2011). Statistical mechanics (3rd ed.), Chapters 3 and 4. Academic Press.
  5. Kardar, M. (2007). Statistical physics of particles, Chapter 4. Cambridge University Press.
Key terms
Grand canonical ensemble
The distribution for a system exchanging both energy and particles with a reservoir, with weight exp(beta(muN - E)).
Grand partition function
Xi, the sum of grand canonical weights, satisfying ln Xi = PV/kT.
Fugacity
z = exp(beta mu), the variable in which Xi is a generating function for the canonical partition functions Z_N.
Grand potential
Phi = -kT ln Xi = U - TS - muN, which equals -PV for a homogeneous system.
Quantum concentration
n_Q = (2 pi m kT/h^2)^(3/2), the density at which the interparticle spacing equals the thermal de Broglie wavelength.
Isothermal compressibility
kappa_T = -(1/V)(dV/dP) at constant T, which fixes the size of density fluctuations.
Fluctuation-dissipation theorem
The statement that spontaneous equilibrium fluctuations and the response to a driving force are governed by the same function.
Johnson-Nyquist noise
The thermal voltage across a resistor, with mean square 4kTR per unit bandwidth.

The Ideal Gas, the Sackur-Tetrode Equation, and the Paradox Gibbs Left Behind

  • Compute the single-particle partition function for a particle in a box and identify the thermal de Broglie wavelength and the quantum concentration.
  • Derive the Sackur-Tetrode equation and evaluate it numerically against measured standard molar entropies for the noble gases.
  • State the Gibbs paradox precisely, show how the factor of N factorial removes it, and explain why only quantum mechanics justifies that factor.

The standard molar entropy of argon gas at 298.15 K and 1 bar, obtained with a calorimeter by integrating the heat capacity up from near absolute zero and adding the latent heats, is 154.846 joules per mole per kelvin. By the end of this lesson you will compute that number from Planck's constant, Boltzmann's constant and the mass of an argon atom, using nothing measured about argon except its mass, and you will get 154.84.

That agreement is among the strongest confirmations that entropy really is the logarithm of a count of quantum states, because the count is what supplies the h, and a classical theory has no h to supply.

Step 1: one particle in a box

Put one particle of mass m in a cubical box of side L with rigid walls. The energy levels are

E(n_x, n_y, n_z) = (h^2/(8 m L^2))(n_x^2 + n_y^2 + n_z^2), with each n a positive integer.

At room temperature in a laboratory-sized box these levels are absurdly close together. For an argon atom in a 1 cm box, h^2/(8mL^2) = 8.3 x 10^-40 J against kT = 4.1 x 10^-21 J, a ratio of 2 x 10^-19, so the sum over states becomes an integral with no measurable error:

Z_1 = (V/h^3) times the integral over all momenta of e^(-beta p^2/2m) d^3p.

That V d^3p/h^3 is the semiclassical counting rule: one quantum state per volume h^3 of phase space. It is not a convention, it follows from the level spacing above, and it is why h appears in a classical-looking answer. The Gaussian integral factorises into three standard ones, giving

Z_1 = V (2 pi m k T / h^2)^(3/2) = V / lambda^3, where lambda = h / sqrt(2 pi m k T).

The length lambda is the thermal de Broglie wavelength, roughly the de Broglie wavelength of a particle carrying thermal energy. For argon at 298.15 K it is 1.60 x 10^-11 m, about a sixth of an atomic diameter, which is a preview of why argon behaves classically. Its reciprocal cube n_Q = 1/lambda^3 is the quantum concentration from Lesson 5.

Step 2: N particles, and a mistake worth making first

For N non-interacting particles the energy is a sum, so the Boltzmann factor factorises and the obvious guess is Z_N = (Z_1)^N. Follow it through. Then F = -NkT ln(V/lambda^3) and, differentiating with respect to temperature,

S_wrong = Nk[ln(V/lambda^3) + 3/2].

Now test whether that is extensive, meaning that doubling the size of the system doubles it. Replace N by 2N and V by 2V:

S_wrong(2N, 2V) = 2Nk[ln(2V/lambda^3) + 3/2] = 2 S_wrong(N, V) + 2Nk ln 2.

It is not extensive. There is a spurious extra 2Nk ln 2. Something is wrong, and it is not arithmetic.

The fault is in (Z_1)^N. That expression counts the configuration "particle 1 in state a, particle 2 in state b" as different from "particle 1 in state b, particle 2 in state a". For identical atoms there is no such distinction: nothing whatsoever distinguishes the two situations. Since almost every term in the sum has all N particles in different states, the overcount is close to a uniform factor of N!, so write

Z_N = (Z_1)^N / N!.

Notice how uncomfortable that is inside classical physics. Classical particles have trajectories, so in principle you could follow atom 1 and label it. The N! has no classical justification; it is smuggled in because the answer is wrong without it. Only quantum mechanics, where the many-body wave function is symmetric or antisymmetric under exchange, supplies a reason, and Lesson 10 does that properly.

Step 3: the Sackur-Tetrode equation

With the N! in place, and using Stirling's ln N! = N ln N - N,

F = -kT ln Z_N = -NkT[ln(V/(N lambda^3)) + 1].

Now differentiate. The only temperature dependence beyond the explicit factor sits in lambda^3, which is proportional to T^(-3/2), so ln(V/(N lambda^3)) contributes a term (3/2) ln T. Then S = -(dF/dT) at fixed V and N gives

S = Nk[ln(V/(N lambda^3)) + 5/2],

the Sackur-Tetrode equation, found independently in 1912 by Otto Sackur in Breslau and Hugo Tetrode in Amsterdam, who was still a teenager. This is now manifestly extensive, since V/N is intensive. And it contains Planck's constant inside lambda, which was the startling part in 1912: h had shown up in blackbody radiation and in specific heats, and here it was again, fixing the absolute entropy of an ordinary gas.

Now the arithmetic for argon. Take T = 298.15 K and P = 100000 Pa, the standard state.

  • m = 39.948 x 1.66054 x 10^-27 = 6.6335 x 10^-26 kg.
  • kT = 1.380649 x 10^-23 x 298.15 = 4.1164 x 10^-21 J.
  • 2 pi m k T = 6.2832 x 6.6335 x 10^-26 x 4.1164 x 10^-21 = 1.7157 x 10^-45, and its square root is 4.1421 x 10^-23.
  • lambda = 6.62607 x 10^-34 / 4.1421 x 10^-23 = 1.59969 x 10^-11 m, so lambda^3 = 4.0936 x 10^-33 m3.
  • V/N = kT/P = 4.1164 x 10^-21 / 100000 = 4.1164 x 10^-26 m3.
  • V/(N lambda^3) = 4.1164 x 10^-26 / 4.0936 x 10^-33 = 1.00560 x 10^7, whose logarithm is 16.1237.
  • S = R(16.1237 + 2.5) = 8.31446 x 18.6237 = 154.84 J/(mol K).

The calorimetric value is 154.846. The two agree to five significant figures, and the only input specific to argon was its atomic mass.

Five gases, one formula

GasM (u)lambda at 298.15 K (pm)S predicted, J/(mol K)S measured, J/(mol K)
Helium4.00350.54126.14126.153
Neon20.18022.51146.33146.328
Argon39.94816.00154.84154.846
Krypton83.79811.05164.07164.085
Xenon131.2938.82169.69169.685

The spread down that column is entirely the mass dependence, entering as (3/2) ln m, and each measured value comes from an independent calorimetric integration. The noble gases are the clean test because each has a non-degenerate electronic ground state; a gas with electronic degeneracy needs an extra Nk ln g, and a molecular gas needs the rotational and vibrational terms of Lesson 7.

Key idea: five numbers computed from h, k and a mass, five numbers measured with a calorimeter, agreeing to four or five digits. Entropy is a count of quantum states.

Gibbs's paradox

Return to the missing extensivity, because Gibbs wrote about it in the 1870s and his argument is sharper than the algebra above. Take a box divided by a partition, with N atoms of the same gas at the same temperature and pressure on each side. Pull out the partition. Nothing observable happens, and you can slide it back in and be exactly where you started, so the entropy change must be zero.

Computed with S_wrong, though, removing the partition doubles both N and V and the entropy rises by 2Nk ln 2. Worse, that holds however slight the difference between the two gases, so the entropy change would be discontinuous: 2Nk ln 2 for gases differing by anything at all, zero for gases that are identical. That is the Gibbs paradox.

With the N! in place, S = Nk[ln(V/(N lambda^3)) + 5/2] depends on the density V/N, which is unchanged, so removing the partition between identical gases changes nothing. Now do the other experiment: argon on the left, krypton on the right, same T and P. Each gas expands into twice the volume, and the entropy of mixing is

Delta S = 2 N k ln 2 = 11.53 J per kelvin per mole of mixture,

which is real, measurable, and is exactly the work you would have to do to unmix them. The N! did not delete the mixing entropy; it deleted it only for the case where mixing is not a physical process at all.

Common misconceptions

  • "The N! is a classical bookkeeping convention." It cannot be justified classically, because classical particles have trajectories and are therefore labellable in principle. The correct statement is that classical statistical mechanics is internally inconsistent about entropy, and the resolution requires quantum indistinguishability.
  • "Entropy of mixing arises because the molecules get jumbled." It arises because each gas expands into a larger volume. Mix argon with argon and there is just as much jumbling and no entropy change at all.
  • "Sackur-Tetrode is exact for a real gas." It is exact for a non-interacting monatomic gas with a non-degenerate ground state. Real argon at 1 bar is close to that, but at 100 bar the interactions matter, and Lesson 15's van der Waals equation is the first correction.
  • "Absolute entropy is meaningless; only differences matter." That is true within classical thermodynamics, which is why the third law and Planck's constant are both needed to fix the constant. The agreement in the table above is the evidence that it is fixed correctly.

What else falls out of F

Since F = -NkT[ln(V/(N lambda^3)) + 1] is the full thermodynamics, three more results come for free.

P = -(dF/dV) at fixed T and N = NkT/V, which is the ideal gas law, now derived rather than assumed.

U = F + TS = (3/2)NkT, recovering the kinetic-theory energy of Lesson 3.

mu = (dF/dN) at fixed T and V = kT ln(N lambda^3/V) = kT ln(n/n_Q), which is the expression used in Lesson 5 and evaluated there as -0.328 eV for helium.

Where the formula stops working

Look at the bracket. When ln(V/(N lambda^3)) drops below -5/2, the predicted entropy goes negative, which is impossible. That happens when n/n_Q exceeds e^(5/2) = 12.2. The formula is announcing its own failure: once the density approaches the quantum concentration, so that atoms are packed closer than their thermal wavelength, the assumption that almost every particle sits in a different state collapses, the crude division by N! is no longer right, and you need the exact quantum counting.

For helium at 1 bar that boundary sits near 2 K; for copper's conduction electrons, far lighter and far denser, it sits above 80,000 K, which is why a metal's electrons are never classical. Lesson 10 builds the exact counting and Lessons 13 and 14 spend it.

Try it: neon at a different state point

Exercise. Compute the molar entropy of neon at 500 K and 10 bar.

Solution. Scaling is easier than starting over. Relative to 298.15 K, lambda falls as T^(-1/2), so lambda^3 falls by (500/298.15)^(3/2) = 2.172, raising V/(N lambda^3) by that factor. But V/N = kT/P changes by (500/298.15) x (1/10) = 0.1677. The net factor is 2.172 x 0.1677 = 0.3643, whose logarithm is -1.0098. So the bracket goes from 15.100 to 14.090, and S = 8.31446 x (14.090 + 2.5) = 137.9 J/(mol K).

Change one input. Drop the pressure back to 1 bar at the same 500 K. Now the net factor is 2.172 x 1.677 = 3.643, logarithm +1.293, so the bracket is 16.393 and S = 157.2 J/(mol K). Expanding a gas tenfold at fixed temperature adds R ln 10 = 19.1 J/(mol K), and indeed 157.2 minus 137.9 is 19.3, the small discrepancy being rounding.

What you now know

A particle in a box has Z_1 = V/lambda^3 with lambda = h/sqrt(2 pi m kT), and the semiclassical rule of one state per h^3 of phase space is what puts Planck's constant into a classical answer. For N identical particles the correct partition function is (Z_1)^N/N!, and dropping the N! makes the entropy non-extensive, which is the Gibbs paradox. With it, the Sackur-Tetrode equation S = Nk[ln(V/(N lambda^3)) + 5/2] reproduces the measured standard molar entropies of five noble gases to four or five significant figures. The same free energy delivers the ideal gas law, U = (3/2)NkT, and mu = kT ln(n/n_Q), and it announces its own breakdown when the density approaches the quantum concentration.

One thing has been assumed rather than derived: that each translational degree of freedom carries kT/2. The next lesson proves that in general, then finds four systems where it fails so badly that the failures were, historically, among the first clues that classical physics was finished.

Sources

  1. Wikipedia contributors. (n.d.). Sackur-Tetrode equation. Wikipedia. en.wikipedia.org
  2. Wikipedia contributors. (n.d.). Gibbs paradox. Wikipedia. en.wikipedia.org
  3. National Institute of Standards and Technology. (n.d.). Argon: gas phase thermochemistry data. NIST Chemistry WebBook. webbook.nist.gov
  4. Sackur, O. (1912). Die universelle Bedeutung des sogenannten elementaren Wirkungsquantums. Annalen der Physik, 40, 67-86.
  5. Schroeder, D. V. (2000). An introduction to thermal physics, Section 6.7. Addison-Wesley.
Key terms
Thermal de Broglie wavelength
lambda = h/sqrt(2 pi m kT), 16.0 pm for argon at room temperature.
Semiclassical counting
The rule of one quantum state per h^3 of phase space, which follows from the particle-in-a-box level spacing.
Sackur-Tetrode equation
S = Nk[ln(V/(N lambda^3)) + 5/2], the absolute entropy of a monatomic ideal gas.
Extensivity
The requirement that doubling N and V doubles S, which fails without the factor of N factorial.
Gibbs paradox
The spurious entropy increase on removing a partition between two samples of the same gas, resolved by indistinguishability.
Entropy of mixing
The real 2Nk ln 2 that arises when two different gases each expand into the whole volume.
Quantum concentration
n_Q = 1/lambda^3; the Sackur-Tetrode entropy goes negative once n exceeds 12.2 n_Q, signalling its own failure.

Module 3: The Architecture of Thermodynamics

Recover the classical laws from the statistics and see exactly where they break. Prove equipartition and find the four systems that defeat it, build the thermodynamic potentials as Legendre transforms, and derive the Maxwell relations and use one to predict a measurable inversion temperature.

Equipartition, and the Four Places It Fails

  • Prove the equipartition theorem for a quadratic degree of freedom in the classical canonical ensemble and state its generalisation to other powers.
  • Predict molar heat capacities for monatomic, diatomic and polyatomic gases and for solids, and compare with measured values.
  • Identify the four regimes where equipartition fails and give the quantitative condition that decides whether a mode is active.

Measure the constant-volume molar heat capacity of hydrogen gas and you get roughly 12.5 J/(mol K) near 60 K, 20.5 at 300 K, and about 28.8 at 3000 K. Those are 3R/2, 5R/2 and 7R/2 to within a couple of percent. A gas that changes its heat capacity in steps, as though degrees of freedom were being switched on one at a time as it warms, is not something classical mechanics can produce. James Clerk Maxwell said so publicly in 1875, calling the specific heat discrepancy the greatest difficulty the molecular theory had yet met, and he was right: it was not resolved until quantum mechanics.

This lesson proves the classical theorem that predicts a single fixed number, works out what it predicts for real substances, and then examines the four distinct ways the prediction fails. The failures are more useful than the theorem.

The theorem, proved in four lines

Suppose the energy of a system contains a term a x^2, where x is any coordinate or momentum that ranges over the whole real line and appears nowhere else in the energy. In the classical canonical ensemble the average of that term is

average of a x^2 = [integral of a x^2 e^(-beta a x^2) dx] / [integral of e^(-beta a x^2) dx].

Rather than evaluate both integrals, notice that the numerator is minus the beta-derivative of the denominator. So the whole ratio is -d/dbeta of ln[integral of e^(-beta a x^2) dx]. The Gaussian integral is sqrt(pi/(beta a)), whose logarithm is -(1/2) ln beta + constant. Differentiating,

average of a x^2 = 1/(2 beta) = kT/2.

The constant a has vanished. That is the equipartition theorem: every quadratic term in the energy carries kT/2 on average, regardless of its stiffness, its mass, or what it physically represents. A stiff spring and a floppy one hold the same thermal energy.

The same trick handles any power. If the energy contains a |x|^n, the substitution u = beta^(1/n) x gives a denominator proportional to beta^(-1/n) and hence

average of a |x|^n = kT/n.

That matters more often than it looks. An ultrarelativistic gas has E = pc, linear in momentum, so each particle carries 3kT rather than (3/2)kT, which is why the radiation-dominated early universe has a different equation of state from a cold gas.

What it predicts

SystemQuadratic terms per particlePredicted CVMeasured near 298 K
Monatomic gas (Ar)3 translational1.5R = 12.4712.47
Diatomic, rigid (N2)3 translational + 2 rotational2.5R = 20.7920.8
Diatomic, vibrating+ 2 vibrational3.5R = 29.10N2 reaches this only above 3000 K
Solid (Cu)3 kinetic + 3 potential3R = 24.9423.7 (CV)
Solid (diamond)3 kinetic + 3 potential3R = 24.946.11

Two rows agree beautifully and one is off by a factor of four. That is the pattern to explain.

Failure one: rotations that are switched off

A rigid rotor has energy levels E_J = J(J+1) h_bar^2/(2I), so the natural temperature scale is Theta_rot = h_bar^2/(2 I k). Compute it for hydrogen. The reduced mass is half a proton mass, 8.368 x 10^-28 kg, and the bond length is 74.14 pm, giving I = 4.60 x 10^-47 kg m2 and

Theta_rot = (1.0546 x 10^-34)^2 / (2 x 4.60 x 10^-47 x 1.3806 x 10^-23) = 87.6 K.

Below that temperature there is not enough thermal energy to reach the first excited rotational level, and the rotations contribute nothing: the molecule cannot rotate a little bit. Hydrogen is the only common gas with this problem, because it is uniquely light and tightly bound. Do nitrogen the same way, with reduced mass 7.00 u and bond length 109.8 pm, and you get Theta_rot = 2.88 K, far below any temperature at which nitrogen is a gas. That is why nitrogen sits at 5R/2 and hydrogen does not.

Arnold Eucken measured hydrogen's heat capacity down to 60 K in 1912 and saw exactly the predicted collapse toward 3R/2. Einstein and Otto Stern used the result the following year in an early argument about zero-point energy.

Failure two: vibrations that are switched off

The vibrational scale is Theta_vib = h c (wavenumber)/k, and the conversion is 1 cm-1 = 1.4388 K. For hydrogen, whose stretch is at 4401 cm-1, that is 6332 K; for nitrogen, 2359 cm-1, it is 3394 K. Both are enormously above room temperature, which is why neither gas shows its vibrational pair of degrees of freedom until it is nearly hot enough to dissociate.

Chlorine is the instructive counterexample. Its stretch is at 560 cm-1, so Theta_vib = 806 K, and at 298 K the mode is partly excited. Use the Einstein heat capacity function C/R = x^2 e^x/(e^x - 1)^2 with x = 806/298 = 2.705: e^x = 14.96, and C/R = 7.32 x 14.96/(13.96)^2 = 0.562. So the prediction is C_V = 2.5R + 0.562R = 3.06R = 25.5 J/(mol K), against a measured 25.6. Partial excitation, computed to three digits.

The point: a mode is not on or off but continuously activated, and the Einstein function is the dial.

Failure three: diamond

Every solid should sit at 3R by the Dulong-Petit argument. Copper does, at 23.7 J/(mol K) once you convert the measured constant-pressure value to constant volume. Diamond, at 6.11 J/(mol K), is nowhere near. The reason is the same as for hydrogen: carbon is light and its bonds are extraordinarily stiff, so its vibrational quanta are large and room temperature is cold for it. Lesson 12 puts a number on that and shows what Einstein's 1907 model gets right and what Debye had to fix.

Failure four: the electrons that were not there

This one nearly killed the theory of metals. In Drude's 1900 model a metal contains roughly one free electron per atom, forming a classical gas. Equipartition then demands an extra (3/2)R = 12.47 J/(mol K) of heat capacity per mole of copper, on top of the lattice's 24.9, so copper should measure about 37 J/(mol K). It measures 24.4. The electrons carry the current, so they are certainly there, and they carry essentially no heat.

Measured carefully at low temperature, copper's electronic heat capacity is gamma T with gamma = 0.695 mJ/(mol K2), which at 300 K is 0.21 J/(mol K), sixty times smaller than the classical prediction. Sommerfeld resolved it in 1927 by replacing Boltzmann statistics with Fermi-Dirac: only electrons within about kT of the Fermi energy can absorb heat at all, and for copper kT/E_F is 4 x 10^-3 at room temperature. Lesson 13 does the calculation.

Common misconceptions

  • "Every degree of freedom gets kT/2." Only quadratic ones, and only when the level spacing is small compared with kT. A term linear in the variable gets kT, and a frozen mode gets essentially nothing.
  • "A diatomic molecule has three rotational degrees of freedom, so it should be 3R." Rotation about the internuclear axis has a moment of inertia set by the electron cloud, some five orders of magnitude smaller than the other two, putting its Theta_rot in the millions of kelvin. It is frozen at every temperature at which the molecule exists.
  • "Heat capacity is a fixed property of a substance." It is a strong function of temperature for anything but a monatomic gas, and the temperature dependence is precisely the record of which modes have woken up.
  • "Stiffer bonds store more thermal energy." Classically they store exactly the same, kT/2 per quadratic term, since the theorem is independent of the coefficient. Quantum mechanically stiffer means a larger quantum, hence less thermal energy, which is the opposite of the intuition.

Try it: carbon dioxide, mode by mode

Problem. Predict C_V for CO2 at 298 K. It is linear, so 3 translational and 2 rotational degrees of freedom give 5R/2. It has four vibrational modes: a doubly degenerate bend at 667 cm-1, a symmetric stretch at 1333, and an asymmetric stretch at 2349.

Solution. Convert each to a temperature with the factor 1.4388 K per cm-1: 960 K, 1918 K and 3380 K. Then evaluate x = Theta/T and C/R = x^2 e^x/(e^x - 1)^2 for each.

Modecm-1Theta (K)x at 298 KC/R per modeDegeneracy
Bend6679603.220.4542
Symmetric stretch133319186.440.0691
Asymmetric stretch2349338011.340.0021

The vibrational total is 2(0.454) + 0.069 + 0.002 = 0.979 R, so C_V = 2.5R + 0.979R = 3.48R = 28.9 J/(mol K). The measured value, from C_p = 37.13 minus R, is 28.8. The agreement is better than one percent, and the bend alone supplies 93 percent of the vibrational contribution because it is the softest mode.

Change one input. Take CO2 to 1000 K. Now x becomes 0.96, 1.92 and 3.38, giving C/R of 0.925, 0.727 and 0.443. The vibrational total is 2(0.925) + 0.727 + 0.443 = 3.02R, close to the classical maximum of 4R, so C_V = 5.52R = 45.9 J/(mol K) against a measured 45.2. The molecule has warmed into its classical regime and equipartition is nearly right again.

The takeaway

Equipartition assigns kT/2 to every quadratic term in the energy, independent of its coefficient, and kT/n to a term of power n. It predicts 3R/2 for a monatomic gas, 5R/2 or 7R/2 for a diatomic, and 3R for a solid, and it is right for argon, for nitrogen at room temperature and for copper. It fails whenever the spacing between quantum levels is comparable with or larger than kT: for hydrogen's rotations below 87.6 K, for almost every molecular vibration at room temperature, for diamond, and for the conduction electrons in a metal, where it overpredicts by a factor of sixty. The condition to remember is the one comparison, kT against the quantum, with the Einstein function x^2 e^x/(e^x-1)^2 interpolating between the two limits.

Having taken the statistics as far as heat capacities, it is worth stepping back to the structure of thermodynamics itself. The next lesson shows that the various energies you have been using, U, F, H and G, are four faces of one object connected by a single mathematical operation.

Sources

  1. Wikipedia contributors. (n.d.). Equipartition theorem. Wikipedia. en.wikipedia.org
  2. National Institute of Standards and Technology. (n.d.). Carbon dioxide: gas phase thermochemistry data. NIST Chemistry WebBook. webbook.nist.gov
  3. Maxwell, J. C. (1875). On the dynamical evidence of the molecular constitution of bodies. Nature, 11, 357-359.
  4. Kittel, C., and Kroemer, H. (1980). Thermal physics (2nd ed.), Chapter 3. W. H. Freeman.
Key terms
Equipartition theorem
Each quadratic term in the classical energy carries kT/2 on average, independent of its coefficient.
Generalised equipartition
A term proportional to |x|^n carries kT/n, so an ultrarelativistic particle with E = pc carries 3kT.
Rotational temperature
Theta_rot = h_bar^2/(2Ik); 87.6 K for hydrogen and 2.88 K for nitrogen.
Vibrational temperature
Theta_vib = hc(wavenumber)/k, using 1.4388 K per inverse centimetre.
Einstein heat capacity function
C/R = x^2 e^x/(e^x - 1)^2 with x = Theta/T, which interpolates between a frozen and a fully active mode.
Freeze-out
The exponential suppression of a mode's heat capacity once its quantum exceeds kT.
Electronic heat capacity
The linear term gamma T in a metal, 0.695 mJ/(mol K^2) for copper, sixty times below the classical prediction.

Legendre Transforms, and the Four Faces of Energy

  • Define the Legendre transform and explain why it exchanges a variable for its conjugate slope without losing information.
  • Generate the enthalpy, Helmholtz free energy, Gibbs free energy and grand potential from the internal energy, and give the natural variables and minimum principle for each.
  • Compute the maximum electrical work and the open-circuit voltage of a hydrogen fuel cell from tabulated thermodynamic data.

A hydrogen fuel cell at 25 degrees Celsius and 1 bar puts 1.229 volts across its terminals. Not 1.481 volts, which is what you would get if all the energy released by combining hydrogen and oxygen were available as electrical work. The missing 17 percent is not lost to resistance or leakage or poor engineering. It is forbidden, and the number 1.229 comes out of a mathematical manoeuvre that Legendre invented for classical mechanics and Gibbs imported into thermodynamics.

The reason there are four common energy functions rather than one is that different experiments hold different things fixed. A bomb calorimeter fixes volume; a beaker on a bench fixes pressure; a reaction in a thermostat fixes temperature. Each constraint calls for its own potential, and the operation that produces them is always the same.

The transform itself

Take a function f(x) and let p = df/dx be its slope. Define

g = f - p x,

and regard g as a function of p rather than of x. Check what its differential is: dg = df - p dx - x dp = (p dx) - p dx - x dp = -x dp. The independent variable really has changed from x to p, and the new slope is -x.

Nothing has been thrown away. Geometrically, a convex curve can be specified either by listing its points or by listing the family of tangent lines that envelope it, and g(p) is the intercept of the tangent line of slope p. The two descriptions carry the same information, and the Legendre transform is invertible: apply it twice and you are back where you started. That is the whole reason the technique is trusted, and it is also why it fails for non-convex functions, a fact that returns with a vengeance in Lesson 15.

The four potentials

Start from the fundamental relation for a simple compressible system, which is just the first law with the definitions of Lesson 3 substituted in:

dU = T dS - P dV + mu dN.

So U is naturally a function of S, V and N, with T = (dU/dS), -P = (dU/dV) and mu = (dU/dN). Transforming on any subset of those pairs gives the standard potentials.

PotentialDefinitionDifferentialNatural variablesExtremal principle
Internal energy U-T dS - P dV + mu dNS, V, NMinimum at fixed S, V, N
Enthalpy HU + PVT dS + V dP + mu dNS, P, NMinimum at fixed S, P, N
Helmholtz free energy FU - TS-S dT - P dV + mu dNT, V, NMinimum at fixed T, V, N
Gibbs free energy GU - TS + PV-S dT + V dP + mu dNT, P, NMinimum at fixed T, P, N
Grand potential PhiF - mu N-S dT - P dV - N d(mu)T, V, muMinimum at fixed T, V, mu

Read a row by looking at which differentials survive. In dF = -S dT - P dV + mu dN, the coefficient of dT is -S, so S = -(dF/dT) at fixed V and N, which is exactly the step used in Lesson 6 to get the Sackur-Tetrode equation from F = -kT ln Z. Every such coefficient is a usable formula.

Where the minimum principles come from

Only one law is at work, the maximisation of total entropy, and the potentials are it in disguise. Put the system in contact with a reservoir at temperature T and let it exchange energy but not volume. When the system absorbs dU, the reservoir loses it, so dS_R = -dU/T and

dS_total = dS - dU/T = -(1/T)(dU - T dS) = -dF/T at fixed T.

Total entropy increases exactly when the system's Helmholtz free energy decreases. Repeat with a reservoir that also fixes the pressure, so the system does work P dV against it, and the same manipulation gives dS_total = -dG/T. There is no separate principle that free energy seeks a minimum. There is one principle, and F and G are the convenient ways to apply it while only looking at the system.

The core of it: the -TS in F = U - TS is not a correction to the energy. It is the reservoir's entropy budget, converted into the system's units.

What each one is good for

The enthalpy is what a constant-pressure calorimeter measures. At fixed P with no other work, dH = T dS = dQ, so the heat released by a reaction in an open beaker is -Delta H. Tables of standard enthalpies of formation exist because of that identity.

The Helmholtz free energy bounds the work available at fixed temperature. Since dF = dU - T dS and dU = dQ + dW with dQ at most T dS, the work you can extract is at most -Delta F.

The Gibbs free energy does the same at fixed temperature and pressure, but excludes the expansion work the system must do against the atmosphere. What is left, -Delta G, is the maximum useful work: electrical, mechanical, chemical. That is the quantity a fuel cell delivers.

Two structural identities are worth having. Because U is extensive and its natural variables are all extensive, scaling the system by a factor gives Euler's relation, U = TS - PV + mu N, hence G = mu N. Differentiating that and comparing with dU gives the Gibbs-Duhem relation, S dT - V dP + N d(mu) = 0, which says the three intensive variables cannot be varied independently: fix two and the third follows.

Common misconceptions

  • "Free energy is the part of the energy that is free to do work." Not in general. For a reaction with Delta S positive, |Delta G| exceeds |Delta H|, and the system absorbs heat from the surroundings to make up the difference and does more work than the energy change alone would allow. The TS term is an entropy ledger, not a stock of unusable energy.
  • "Systems minimise their energy." Only at fixed entropy, which is an unusual constraint. At fixed T and V they minimise F, at fixed T and P they minimise G, and in every case the underlying statement is that the entropy of system plus surroundings increases.
  • "The Legendre transform loses information." It is invertible for a convex function, which is what makes it legitimate. Where convexity fails, as inside the coexistence region of a real fluid, the transform genuinely does lose information, and repairing that loss is exactly the Maxwell construction of Lesson 15.
  • "Enthalpy is heat content." H equals the heat only at constant pressure and only when no work other than expansion is done. A fuel cell operates at constant pressure and does electrical work, and its heat output is T Delta S, not Delta H.

Worked example: where 1.229 volts comes from

The cell reaction is H_2 + (1/2) O_2 to H_2O(liquid) at 298.15 K and 1 bar. Take three tabulated standard entropies, in J/(mol K): S(H_2) = 130.68, S(O_2) = 205.15, S(H_2O, liquid) = 69.95. And take the standard enthalpy of formation of liquid water, Delta H = -285.83 kJ/mol.

Step 1, the entropy change. Delta S = 69.95 - 130.68 - 0.5(205.15) = -163.31 J/(mol K). It is strongly negative because three moles of gas become one mole of liquid.

Step 2, the heat that must be dumped. T Delta S = 298.15 x (-163.31) = -48,690 J = -48.69 kJ/mol. The reaction destroys entropy internally, so it must export at least that much entropy as heat, and that heat cannot be converted to work.

Step 3, the free energy. Delta G = Delta H - T Delta S = -285.83 + 48.69 = -237.14 kJ/mol.

Step 4, the voltage. Two electrons pass through the external circuit per molecule of water formed, so the electrical work per mole is n F E with n = 2 and Faraday's constant F = 96,485 C/mol. Setting that equal to -Delta G,

E = 237,140/(2 x 96,485) = 1.229 V.

That is the measured open-circuit voltage of a hydrogen fuel cell, to three decimal places. The efficiency ceiling is Delta G/Delta H = 237.14/285.83 = 83.0 percent, and the 1.481 V quoted at the start is the so-called thermoneutral voltage Delta H/(nF), which you would reach only if the 48.69 kJ did not have to be shed.

Worth holding on to: a chemist's table, two subtractions and a Legendre transform predict a laboratory voltage to four significant figures.

Try it: which potential, and what does it bound?

Exercise. For each situation, name the potential that is minimised and the quantity it bounds. (a) A gas expanding in a rigid insulated container. (b) Ice melting in a glass of water on a table. (c) A protein folding in a cell at fixed temperature and pressure. (d) Electrons in a metal wire connected to a reservoir at fixed voltage.

Solution. (a) Nothing is exchanged, so the operative statement is the original one: entropy is maximised at fixed U, V and N. (b) Fixed T and P, so G is minimised, and the melting point is where G(ice) = G(water), which is the equality of chemical potentials. (c) Also G, and the bound is on the useful work the folding can deliver, which is why Delta G of folding, typically a few tens of kJ/mol, is the number quoted in the literature rather than Delta H, which is much larger and largely cancelled by T Delta S. (d) Fixed T, V and mu, so the grand potential is minimised, and Lesson 5 showed that it equals -PV.

Summing up

The Legendre transform swaps a variable for the slope conjugate to it, exchanging f(x) for g(p) = f - px with dg = -x dp, and it loses nothing when the function is convex. Applied to U(S, V, N) it generates the enthalpy, the Helmholtz free energy, the Gibbs free energy and the grand potential, each with its own natural variables and its own minimum principle, and each of those principles is the maximisation of total entropy rewritten so that only the system appears. Euler's relation gives G = mu N, and Gibbs-Duhem says the intensive variables are not independent. Applied to hydrogen and oxygen at 298.15 K, the machinery predicts a fuel cell voltage of 1.229 V and an efficiency ceiling of 83 percent.

All five thermodynamic potentials are smooth functions of two or three variables, and a smooth function has equal mixed second derivatives no matter which order you differentiate in. That trivial fact turns out to generate a set of relations connecting quantities that no experiment could otherwise link, and the next lesson derives them and uses one to predict something you can measure with a valve and a thermometer.

Sources

  1. Wikipedia contributors. (n.d.). Thermodynamic potential. Wikipedia. en.wikipedia.org
  2. Wikipedia contributors. (n.d.). Gibbs free energy. Wikipedia. en.wikipedia.org
  3. National Institute of Standards and Technology. (n.d.). Water: condensed phase thermochemistry data. NIST Chemistry WebBook. webbook.nist.gov
  4. Callen, H. B. (1985). Thermodynamics and an introduction to thermostatistics (2nd ed.), Chapters 5 and 6. John Wiley and Sons.
  5. Schroeder, D. V. (2000). An introduction to thermal physics, Chapter 5. Addison-Wesley.
Key terms
Legendre transform
The operation g = f - px with p = df/dx, which exchanges a variable for its conjugate slope and is invertible for convex functions.
Enthalpy
H = U + PV, whose change equals the heat exchanged at constant pressure with no other work.
Helmholtz free energy
F = U - TS, minimised at fixed T, V and N, and bounding the total work extractable at fixed temperature.
Gibbs free energy
G = U - TS + PV, minimised at fixed T, P and N, and bounding the useful non-expansion work.
Natural variables
The set in which a potential's differential contains only differentials of independent variables, so that all its derivatives are simple.
Euler relation
U = TS - PV + muN, which follows from extensivity and implies G = muN.
Gibbs-Duhem relation
S dT - V dP + N d(mu) = 0, so the three intensive variables cannot all be varied independently.
Thermoneutral voltage
Delta H/(nF) = 1.481 V for hydrogen, the voltage a cell would give if no heat had to be rejected.

The Maxwell Relations, and Why Dewar Needed Liquid Air

  • Derive the four Maxwell relations from the equality of mixed second derivatives of the thermodynamic potentials.
  • Use a Maxwell relation to compute how internal energy depends on volume for an ideal and a van der Waals gas.
  • Compute Joule-Thomson inversion temperatures from van der Waals constants and compare them with measured values for four gases.

In 1895 Carl von Linde forced compressed air through a throttle valve, let it expand, recycled the cooled gas to chill the incoming stream, and produced liquid air. Three years later James Dewar tried to run hydrogen through the same kind of apparatus starting from room temperature. Hydrogen throttled at room temperature comes out warmer, not colder. Dewar succeeded on 10 May 1898 only by pre-cooling his hydrogen with liquid air first, and a decade later Heike Kamerlingh Onnes liquefied helium by pre-cooling with liquid hydrogen.

The difference between those gases is a sign, and the sign is predictable. It comes from a relation between partial derivatives that Maxwell wrote down in 1871, whose entire derivation is the observation that mixed second derivatives commute.

Where the relations come from

Take the Helmholtz free energy, whose differential Lesson 8 gave as dF = -S dT - P dV at fixed N. Reading off the coefficients,

(dF/dT) at fixed V = -S and (dF/dV) at fixed T = -P.

Now differentiate the first with respect to V and the second with respect to T. Both are the mixed second derivative of F, and for any function with continuous second derivatives the order does not matter. So -(dS/dV) at fixed T = -(dP/dT) at fixed V, that is,

(dS/dV) at fixed T = (dP/dT) at fixed V.

Pause on how odd that is. On the left is a derivative of entropy, a quantity no instrument measures directly. On the right is a derivative of pressure with respect to temperature at fixed volume, which you can obtain with a sealed vessel, a heater and a pressure gauge in an afternoon. The Maxwell relations are a currency exchange between the unmeasurable and the measurable, and they have real content: they would fail if entropy were not a state function.

The four of them

FromDifferentialRelation
UT dS - P dV(dT/dV) at fixed S = -(dP/dS) at fixed V
HT dS + V dP(dT/dP) at fixed S = (dV/dS) at fixed P
F-S dT - P dV(dS/dV) at fixed T = (dP/dT) at fixed V
G-S dT + V dP(dS/dP) at fixed T = -(dV/dT) at fixed P

The two from F and G are the workhorses, because their left sides involve entropy at fixed temperature, which is what almost every practical question needs. Notice the minus sign in the G relation and where it comes from: +V dP rather than -P dV. Sign errors here are the commonest mistake in the subject, and the cure is to rederive the relation from the differential each time rather than memorise a diagram.

Use one: does a gas cool when it expands into a vacuum?

Ask how internal energy depends on volume at fixed temperature. From dU = T dS - P dV,

(dU/dV) at fixed T = T (dS/dV) at fixed T - P = T (dP/dT) at fixed V - P,

using the F relation. This is the thermodynamic energy equation, and it needs only an equation of state.

Ideal gas. P = NkT/V, so (dP/dT) at fixed V = Nk/V = P/T, and the whole expression is T(P/T) - P = 0. The internal energy of an ideal gas does not depend on volume at all. Joule looked for this effect in 1845 by letting compressed air expand into an evacuated vessel and measuring the water bath's temperature; he found no change, within his precision.

Van der Waals gas. With P = RT/(v - b) - a/v^2 per mole, (dP/dT) at fixed v = R/(v - b), so

(dU/dV) at fixed T = RT/(v - b) - RT/(v - b) + a/v^2 = a/v^2.

Positive, so expanding a real gas at fixed temperature costs energy: you are pulling molecules apart against their mutual attraction. Do it with no heat supplied and the temperature falls. The quantity a/v^2 is called the internal pressure, and for carbon dioxide at molar volume 1 L/mol it is 0.364/(10^-3)^2 = 3.6 x 10^5 Pa, about 3.6 bar of hidden inward pull.

Use two: the throttling valve

Linde's process is not free expansion. Gas is pushed steadily through a porous plug or a valve from pressure P_1 to pressure P_2, and the work done on it upstream, P_1 V_1, and by it downstream, P_2 V_2, exactly cancel the internal energy change. The process conserves enthalpy. So the question is the sign of the Joule-Thomson coefficient

mu_JT = (dT/dP) at fixed H.

Use the triple product rule to write it as -(dH/dP)_T / C_P, then attack the numerator with dH = T dS + V dP and the G relation:

(dH/dP) at fixed T = T (dS/dP) at fixed T + V = -T (dV/dT) at fixed P + V = V(1 - alpha T),

where alpha = (1/V)(dV/dT) at fixed P is the volume expansion coefficient. Therefore

mu_JT = (V/C_P)(alpha T - 1).

For an ideal gas alpha = 1/T exactly, so mu_JT = 0: throttling an ideal gas changes nothing. Cooling on throttling is entirely an effect of intermolecular forces, and its sign depends on which force wins, the attraction that costs energy to overcome or the repulsion that gives energy back.

Working the van der Waals equation to first order in the small parameters gives a clean criterion. The coefficient changes sign at the inversion temperature, and in the low-pressure limit

T_inv = 2a/(R b).

Gasa (Pa m6/mol2)b (m3/mol)Tinv predicted (K)Tinv measured (K)
Carbon dioxide0.36404.267 x 10-52050about 1500
Nitrogen0.13703.87 x 10-5852621
Hydrogen0.024762.661 x 10-5224about 205
Helium0.003462.38 x 10-535about 43

Now read Linde and Dewar off the table. Nitrogen at 293 K is far below its inversion temperature of 621 K, so throttling cools it, and Linde's regenerative cycle can bootstrap its way down to 77 K. Hydrogen at 293 K is above its inversion temperature of 205 K, so throttling warms it, and no amount of recycling helps: Dewar had to get below 205 K first, which is what the liquid air jacket was for. Helium, above 43 K, warms as well, so Kamerlingh Onnes needed liquid hydrogen. The whole history of low-temperature physics up to 1908, and the 1913 Nobel Prize that followed, is contained in one column of that table.

The van der Waals prediction is 20 to 40 percent high for the heavier gases and about 20 percent low for helium, which is honest performance for a two-parameter equation. It gets the ordering and the order of magnitude right, which is all that was needed to know which gas to try next.

Why this matters: a relation derived by swapping the order of two derivatives told nineteenth-century engineers which gases would liquefy in their apparatus and which would not.

Use three: squeezing water

The G relation, (dS/dP) at fixed T = -(dV/dT) at fixed P = -V alpha, predicts the entropy change on isothermal compression from measurable expansion data alone. For liquid water at 298 K, alpha = 2.57 x 10^-4 per kelvin and the molar volume is 1.807 x 10^-5 m3. Compressing from 1 bar to 1000 bar, and treating both quantities as constant over that range,

Delta S = -V alpha Delta P = -(1.807 x 10^-5)(2.57 x 10^-4)(9.99 x 10^7) = -0.464 J/(mol K),

so the water must give up T Delta S = 298 x 0.464 = 138 J/mol of heat to stay at the same temperature. Now do the same at 2 degrees Celsius, where water's expansion coefficient is negative because of its density maximum at 4 degrees. The sign of Delta S flips: compressing cold water raises its entropy. The relation does not care that the behaviour is peculiar; it simply reports what the expansion data imply.

Common misconceptions

  • "Maxwell relations are algebra with no physics in them." They encode the claim that entropy is a state function, so that dS is an exact differential. If it were not, mixed partials would not commute and the relations would be false. Their content is the second law.
  • "Throttling always cools a gas." Only below the inversion temperature. Above it, throttling warms the gas, which is why hydrogen and helium had to be pre-cooled and why a leaking hydrogen line can warm rather than frost.
  • "Joule-Thomson cooling is the gas doing work as it expands." The net external work in a steady throttling process is zero by construction, since enthalpy is conserved. The temperature change comes from the internal potential energy of the molecules, and vanishes exactly for an ideal gas.
  • "You have to memorise the thermodynamic square." The mnemonic works, but rederiving from the differential takes fifteen seconds and never produces a sign error, which the mnemonic frequently does.

Try it: a rubber band on your lip

Rubber gives you a Maxwell relation you can feel. For a one-dimensional elastic system the work term is f dL instead of -P dV, so the analogue of the F relation is

(dS/dL) at fixed T = -(df/dT) at fixed L.

The right-hand side is measurable: hang a weight on a rubber band and warm it, and the band contracts, which means the tension needed to hold a fixed length rises with temperature. So (df/dT) at fixed L is positive, and therefore (dS/dL) at fixed T is negative. Stretching rubber lowers its entropy, because the long polymer chains have fewer configurations when extended than when coiled.

Now predict the experiment. Stretch the band fast enough that no heat escapes, so S is constant. The entropy lost to stretching must be paid for by entropy gained from heating, so the band warms. Take a thick rubber band, hold it against your upper lip, and stretch it sharply: it is noticeably warm. Hold it stretched until it returns to room temperature, then let it snap back while still touching your lip, and it goes cold.

Change one input. Do the same with a steel spring. Nothing happens, because steel's elasticity is energetic rather than entropic: stretching a metal changes bond lengths, not the number of accessible configurations. That contrast is the whole difference between an energy-driven and an entropy-driven material, and you can detect it with your lip.

Looking back

Maxwell relations follow from the equality of mixed second derivatives of the thermodynamic potentials, and each one trades a derivative of entropy for a derivative of pressure, volume or temperature that an instrument can read. The F relation gives the energy equation (dU/dV)_T = T(dP/dT)_V - P, which is zero for an ideal gas and a/v^2 for a van der Waals gas. The G relation gives the Joule-Thomson coefficient mu_JT = (V/C_P)(alpha T - 1), zero for an ideal gas, and van der Waals then predicts inversion temperatures of 852 K for nitrogen, 224 K for hydrogen and 35 K for helium, against measured values of 621, 205 and 43. Those numbers explain why Linde could liquefy air directly, why Dewar needed liquid air to reach hydrogen, and why Kamerlingh Onnes needed liquid hydrogen to reach helium. The same relations, applied to a rubber band, predict a warming you can feel on your lip.

That completes the classical structure. Everything from here on is quantum: the next module derives the two distributions that govern identical particles and spends them on radiation, solids, metals and a cloud of rubidium atoms colder than anything in nature.

Sources

  1. Wikipedia contributors. (n.d.). Maxwell relations. Wikipedia. en.wikipedia.org
  2. Wikipedia contributors. (n.d.). Joule-Thomson effect. Wikipedia. en.wikipedia.org
  3. National Institute of Standards and Technology. (n.d.). Thermophysical properties of fluid systems. NIST Chemistry WebBook. webbook.nist.gov
  4. Callen, H. B. (1985). Thermodynamics and an introduction to thermostatistics (2nd ed.), Chapter 7. John Wiley and Sons.
  5. Zemansky, M. W., and Dittman, R. H. (1997). Heat and thermodynamics (7th ed.), Chapter 11. McGraw-Hill.
Key terms
Maxwell relation
An identity between partial derivatives obtained by equating the mixed second derivatives of a thermodynamic potential.
Thermodynamic energy equation
(dU/dV) at fixed T = T(dP/dT) at fixed V - P, which vanishes for an ideal gas and equals a/v^2 for a van der Waals gas.
Internal pressure
The quantity a/v^2 measuring the energy cost per unit volume of pulling attracting molecules apart.
Throttling
Steady flow through a restriction, during which enthalpy is conserved.
Joule-Thomson coefficient
mu_JT = (V/C_P)(alpha T - 1), the temperature change per unit pressure drop in throttling.
Inversion temperature
The temperature above which throttling warms a gas; 2a/(Rb) in the van der Waals model.
Volume expansion coefficient
alpha = (1/V)(dV/dT) at fixed P, equal to 1/T for an ideal gas and negative for water below 4 degrees Celsius.
Entropic elasticity
Restoring force arising from a reduction in the number of configurations, as in rubber, rather than from changed bond energies.

Module 4: Quantum Statistics

Count states rather than particles. Derive the Bose-Einstein and Fermi-Dirac distributions from the grand canonical ensemble, then spend them on the photon gas and the Stefan-Boltzmann constant, on the heat capacity of real solids, and on the electrons in copper and inside a white dwarf.

Bose-Einstein and Fermi-Dirac from the Grand Canonical Ensemble

  • Show by explicit enumeration that dividing by N factorial is not a valid correction once occupation numbers exceed one.
  • Derive the Fermi-Dirac and Bose-Einstein distributions by factorising the grand partition function over single-particle orbitals.
  • Recover the Maxwell-Boltzmann distribution as a limit and state the condition on density and temperature under which it holds.

Two identical particles, two available single-particle states. Count the microstates. If the particles carry labels there are four: both in a, both in b, particle 1 in a with particle 2 in b, and the reverse. Lesson 6's recipe divides by 2! = 2 and reports two. The correct answer for helium-4 atoms or photons is three; for electrons it is one. Neither equals two, and the gap is not a small correction that improves with more particles.

ConfigurationLabelled particlesBosonsFermions
Both in state a11forbidden
Both in state b11forbidden
One in each211
Total431

The division by N factorial repairs only the last row, where the two particles occupy different states. It has nothing to say about the first two rows, where there is nothing to overcount. That is why Lesson 6's Sackur-Tetrode entropy went negative at high density: doubly occupied states become common there, and the recipe was never valid for them.

Stop counting particles

The fix is to change what you enumerate. Instead of asking which particle is in which state, specify how many particles occupy each single-particle state, or orbital. A microstate is the list of occupation numbers {n_1, n_2, n_3, ...}, and since identical particles carry no labels, that list is a complete description. Then

E = sum over k of n_k eps_k and N = sum over k of n_k.

Quantum mechanics supplies exactly one further rule, from the symmetry of the many-body wave function under exchange. For bosons, integer spin, the wave function is symmetric and n_k may be any non-negative integer. For fermions, half-integer spin, the wave function is antisymmetric and vanishes if two particles share an orbital, so n_k is 0 or 1. That is the Pauli exclusion principle, and it is a statement about counting rather than about forces.

Why the canonical ensemble cannot do this

Try the obvious route. In the canonical ensemble with N fixed,

Z_N = sum over all {n_k} with sum n_k = N of e^(-beta sum n_k eps_k).

Without the constraint, that sum would factorise into an independent sum for each orbital, and each of those is a two-term or geometric series you can do in your head. The constraint destroys the factorisation, because choosing a large n_1 restricts what the other occupations may be. There is no closed form.

Lesson 5 built the tool that removes constraints: hand the particle number over to a reservoir and fix mu instead. In the grand canonical ensemble the sum runs over all occupation lists with no restriction at all:

Xi = sum over all {n_k} of e^(beta sum n_k (mu - eps_k)) = product over k of [sum over n_k of e^(beta n_k (mu - eps_k))].

The sum and product exchange places because the exponent is a sum over independent orbitals. The whole problem has collapsed into one single-orbital calculation done twice.

Key idea: the grand canonical ensemble is not a physical claim about wandering particle numbers. It is the device that makes the constraint disappear.

Fermions

For one orbital, n is 0 or 1, so

Xi_k = 1 + e^(beta(mu - eps_k)).

The average occupancy is (0 x 1 + 1 x e^(beta(mu - eps)))/Xi_k. Dividing numerator and denominator by the exponential,

average n(eps) = 1/(e^((eps - mu)/kT) + 1),

the Fermi-Dirac distribution, obtained independently by Enrico Fermi and Paul Dirac in 1926. Three features are worth fixing in memory. It never exceeds 1, as exclusion requires. At eps = mu it equals exactly 1/2, at every temperature, which is the cleanest operational definition of the chemical potential for fermions. And it falls from 0.881 to 0.119 as eps - mu runs from -2kT to +2kT, so the transition from full to empty occupies a window a few kT wide and nothing else.

Bosons

For one orbital, n runs over all non-negative integers, and the sum is geometric:

Xi_k = 1/(1 - e^(beta(mu - eps_k))),

which converges only if mu < eps_k. Since that must hold for every orbital, the chemical potential of a boson gas can never exceed the lowest single-particle energy. Differentiating ln Xi_k with respect to mu and multiplying by kT,

average n(eps) = 1/(e^((eps - mu)/kT) - 1),

the Bose-Einstein distribution, which Bose derived for photons in 1924 and Einstein extended to material particles the same year. One sign has changed and the consequences are enormous: as eps approaches mu the occupancy diverges rather than saturating. A single orbital can hold an unlimited number of bosons, and Lesson 14 shows what happens when it does.

For photons there is no conservation law on particle number, so there is no constraint to attach a multiplier to and mu = 0. The distribution reduces to 1/(e^(h_bar omega/kT) - 1), which is precisely the average quantum number of the harmonic oscillator computed in Lesson 4. Two calculations that looked unrelated are the same calculation.

The classical limit, located precisely

When e^((eps - mu)/kT) is much larger than 1, the plus or minus one in the denominator is irrelevant and both distributions become

average n(eps) = e^((mu - eps)/kT),

the Maxwell-Boltzmann result. The condition is that the exponential be large for every orbital including the lowest, which means e^(-mu/kT) must be large, which means mu must be large and negative. Lesson 6 gave mu = kT ln(n/n_Q), so the condition is simply

n much less than n_Q.

That is a joint statement about density and temperature, not about temperature alone. Room-temperature helium at 1 bar has n/n_Q = 3 x 10^-6 and is classical. Conduction electrons in copper have n/n_Q of order 100 even at 1000 K, because they are four orders of magnitude lighter and four orders denser, so they are never classical at any temperature a metal survives.

(eps - mu)/kTFermi-DiracBose-EinsteinMaxwell-Boltzmann
0.10.47509.50830.9048
10.26890.58200.3679
30.04740.05240.0498
50.006690.006780.00674

Read the last row: by five kT above the chemical potential the three distributions agree to three significant figures. Quantum statistics matters only in the crowded region near and below mu.

Common misconceptions

  • "Exclusion is a repulsive force between electrons." No force appears anywhere in the derivation. Exclusion is the restriction n_k in {0, 1}, a statement about which states exist. It nevertheless produces a very real pressure, as Lesson 13 computes for copper and for a white dwarf, which is a good demonstration that pressure does not require a force law.
  • "Bosons attract one another." The bosons here are non-interacting by assumption. The apparent clustering comes from the counting: configurations with multiple occupancy are single microstates rather than many, so they are relatively more likely than distinguishable-particle intuition suggests.
  • "The classical limit is the high-temperature limit." It is the limit n much less than n_Q, and since n_Q depends on both mass and temperature, density enters on equal footing. Heating a metal does not make its electrons classical.
  • "For fermions, mu is the Fermi energy." Only at absolute zero. As temperature rises mu falls, slowly at first and eventually becoming negative when the gas turns classical, which is how a semiconductor's chemical potential behaves.

Try it: how sharp is the Fermi surface?

Exercise. Copper at 300 K. What fraction of orbitals 0.1 eV above the chemical potential are occupied, and what fraction 0.1 eV below?

Solution. At 300 K, kT = 0.02585 eV, so 0.1 eV is 3.868 kT. Above: 1/(e^3.868 + 1) = 1/48.85 = 0.0205, about 2 percent occupied. Below: 1/(e^-3.868 + 1) = 1/1.0209 = 0.9795, about 98 percent occupied. The Fermi surface is smeared over a window of about 0.1 eV, against a Fermi energy of 7.0 eV, so 98.6 percent of the electron sea is untouched by room temperature.

Change one input. Look 0.5 eV above instead. Now eps - mu = 19.34 kT and the occupancy is 4 x 10^-9. Moving out five times as far in energy cut the occupancy by a factor of five million, because the tail is exponential. That steepness is exactly why only the electrons within a few kT of mu participate in heat capacity, conduction and everything else a metal does.

Recap

Identical quantum particles must be counted by occupation number, not by which particle sits where, and the division by N factorial fails as soon as an orbital holds more than one particle. Fixing N makes the canonical sum unfactorisable, so the natural setting is the grand canonical ensemble, where the sum over unconstrained occupation lists factorises into one independent calculation per orbital. Fermions give 1/(e^((eps-mu)/kT) + 1), never exceeding 1 and equal to 1/2 at eps = mu; bosons give 1/(e^((eps-mu)/kT) - 1), unbounded as eps approaches mu, which forces mu below the ground-state energy. Both reduce to e^((mu-eps)/kT) once n is far below n_Q, and agree with it to three digits by 5kT above mu.

The next three lessons are three uses of these two formulas. Photons, with mu = 0, give Planck's law and the Stefan-Boltzmann constant. Phonons give the heat capacity of solids down to a few kelvin. Electrons, packed a hundred times above their quantum concentration, give the pressure that holds up a copper crystal and a white dwarf.

Sources

  1. Wikipedia contributors. (n.d.). Fermi-Dirac statistics. Wikipedia. en.wikipedia.org
  2. Wikipedia contributors. (n.d.). Bose-Einstein statistics. Wikipedia. en.wikipedia.org
  3. Kardar, M. (2013). 8.333 Statistical Mechanics I, Lectures on quantum statistical mechanics. MIT OpenCourseWare. ocw.mit.edu
  4. Dirac, P. A. M. (1926). On the theory of quantum mechanics. Proceedings of the Royal Society A, 112(762), 661-677.
  5. Pathria, R. K., and Beale, P. D. (2011). Statistical mechanics (3rd ed.), Chapter 6. Academic Press.
Key terms
Orbital
A single-particle quantum state, the unit that occupation numbers count.
Occupation-number representation
Describing a microstate by the list of how many particles occupy each orbital, which requires no particle labels.
Boson
An integer-spin particle whose many-body wave function is symmetric, allowing any occupation number.
Fermion
A half-integer-spin particle whose antisymmetric wave function restricts each orbital to occupancy 0 or 1.
Fermi-Dirac distribution
Average occupancy 1/(exp((eps - mu)/kT) + 1), equal to 1/2 at eps = mu at any temperature.
Bose-Einstein distribution
Average occupancy 1/(exp((eps - mu)/kT) - 1), which diverges as eps approaches mu.
Classical limit
The regime n much less than n_Q, where both quantum distributions collapse to exp((mu - eps)/kT).
Fermi surface smearing
The few-kT window around mu over which occupancy falls from near one to near zero.

The Photon Gas: Planck's Law, and Sigma Computed from h, c and k

  • Count the electromagnetic modes in a cavity and combine the count with the Bose-Einstein occupancy at zero chemical potential to derive Planck's law.
  • Integrate Planck's law to obtain the Stefan-Boltzmann constant numerically from h, c and k, and derive Wien's displacement constant from a transcendental equation.
  • Compute the photon number density, energy density and radiation pressure of a photon gas and apply them to the cosmic microwave background and to stars.

On Sunday 7 October 1900 Heinrich Rubens and his wife came to lunch at the Plancks' house in Berlin. Rubens brought news from his laboratory: in the far infrared, out at wavelengths of 50 micrometres and beyond, the cavity radiation he and Ferdinand Kurlbaum were measuring did not follow Wien's law. The energy density was simply proportional to T, exactly as an old classical argument said it should be, and Wien's exponential said it should not.

Planck spent that evening constructing a formula that would reduce to Wien's law at short wavelengths and to the linear behaviour at long ones, and sent it to Rubens on a postcard. He read it to the German Physical Society on 19 October. It fitted every measurement. What it did not have was a derivation, and on 14 December he supplied one by assuming that the energy of an oscillator comes in units of h nu, a step he later called an act of desperation.

This lesson takes the modern route, which is shorter and shows where the physics is: count the modes, put the Bose-Einstein occupancy of Lesson 10 into each one, and integrate. The Stefan-Boltzmann constant then falls out as a number you can check against NIST to seven digits.

Counting the modes

Put the electromagnetic field in a cubical cavity of side L with conducting walls. The allowed standing waves have wave vectors k = (pi/L)(n_x, n_y, n_z) with positive integers n, and each wave vector carries two independent polarisations. Count the modes with wave number below k: they occupy an octant of a sphere of radius k in a lattice of cells of side pi/L, so

N(k) = 2 x (1/8)(4/3)pi k^3 x (L/pi)^3 = V k^3/(3 pi^2).

Differentiating and using omega = ck, the number of modes per unit angular frequency is

dN/d(omega) = V omega^2/(pi^2 c^3).

The omega^2 is the whole difficulty of the classical theory: modes proliferate without limit as the frequency rises, and giving each one kT by equipartition makes the total energy infinite. Ehrenfest named that the ultraviolet catastrophe in 1911, more than a decade after Planck had already avoided it.

Putting a photon gas in the modes

Each mode is a harmonic oscillator of frequency omega, and Lesson 4 gave its thermal occupancy as 1/(e^(beta h_bar omega) - 1). Lesson 10 gave the same expression from the Bose-Einstein distribution with mu = 0, and the two agree because they are the same calculation seen from two directions. Photon number is not conserved, so mu is zero, and that is the single physical input that makes everything work.

Multiply the mode density by the occupancy and by the energy per photon:

u(omega) d(omega) = [h_bar omega^3 / (pi^2 c^3 (e^(h_bar omega/kT) - 1))] d(omega),

the energy per unit volume per unit angular frequency. This is Planck's law. Check its two limits against Rubens's lunchtime news. When h_bar omega is much less than kT, expand the exponential as 1 + h_bar omega/kT and the h_bar cancels:

u(omega) = omega^2 kT/(pi^2 c^3),

which is linear in T, exactly what Rubens measured, and which integrates to infinity. When h_bar omega is much greater than kT, the exponential dominates and u(omega) = (h_bar omega^3/pi^2 c^3) e^(-h_bar omega/kT), which is Wien's law. Planck's formula is not an interpolation between two theories; both limits come out of one expression.

Integrating: the Stefan-Boltzmann constant, digit by digit

Total energy density is the integral of u(omega) over all frequencies. Substitute x = h_bar omega/kT:

U/V = (h_bar/(pi^2 c^3)) (kT/h_bar)^4 times the integral from 0 to infinity of x^3/(e^x - 1) dx.

That integral is a standard one, equal to pi^4/15 = 6.49394, so

U/V = a T^4 with a = pi^2 k^4/(15 h_bar^3 c^3).

Evaluate a from the defined constants. k^4 = 3.63356 x 10^-92, pi^2 = 9.86960, so the numerator is 3.58618 x 10^-91. h_bar^3 = 1.17281 x 10^-102 and c^3 = 2.69440 x 10^25, so the denominator is 15 x 3.15993 x 10^-77 = 4.73990 x 10^-76. Dividing,

a = 7.5660 x 10^-16 J m-3 K-4, against the accepted 7.565733 x 10^-16.

The power leaving a small hole in the cavity per unit area is c/4 times the energy density, the quarter coming from averaging the outward component of an isotropic flux over a hemisphere. So

sigma = a c/4 = 2 pi^5 k^4/(15 h^3 c^2) = 7.5660 x 10^-16 x 2.99792458 x 10^8 / 4 = 5.6706 x 10^-8 W m-2 K-4.

NIST lists sigma = 5.670374419 x 10^-8. Every digit you can carry through by hand agrees. The Stefan-Boltzmann law was an empirical fit by Josef Stefan in 1879 and a thermodynamic argument by Boltzmann in 1884, and neither could produce the constant. Statistical mechanics plus one quantum hypothesis produces it exactly.

The upshot: a constant measured in a furnace is 2 pi^5 k^4/(15 h^3 c^2).

Where the spectrum peaks, and a trap

Rewrite Planck's law per unit wavelength and maximise. Setting the derivative to zero gives the transcendental equation x = 5(1 - e^(-x)) with x = hc/(lambda k T), whose root is x = 4.965114. Hence

lambda_max T = hc/(4.965114 k) = 2.897771 x 10^-3 m K,

Wien's displacement constant. For the Sun's effective temperature of 5772 K that gives 502 nm, in the green.

Now the trap. Maximise the same radiation law expressed per unit frequency instead and you get a different transcendental equation, x = 3(1 - e^(-x)), root 2.821439, and nu_max/T = 5.8789 x 10^10 Hz/K. For the Sun that is 3.39 x 10^14 Hz, whose wavelength is 883 nm, in the near infrared. The two peaks are 380 nm apart and both are correct: a spectrum is a density, and a density transforms with a Jacobian when you change variable. There is no single wavelength at which a blackbody emits most; the question is not well posed until you say per unit what.

Common misconceptions

  • "Classical physics predicted an observed catastrophe, and Planck rescued it." The historical order runs the other way. Rayleigh published his classical mode-counting result in June 1900, Planck's formula came in October, and the phrase ultraviolet catastrophe was coined by Ehrenfest in 1911. Planck was responding to Rubens's long-wavelength data, not to a divergence.
  • "Photon number is a conserved quantity like particle number." Walls absorb and emit photons freely, so the equilibrium number is whatever minimises the free energy, and mu = 0. Everything in this lesson depends on that.
  • "The peak wavelength and the peak frequency describe the same colour." They differ by a factor of 1.76 in wavelength for any blackbody, because lambda_max nu_max is not c. Quoting one as if it were the other is a common error in popular accounts of stellar colours.
  • "A blackbody has to be black." The Sun is a blackbody to good accuracy and is not black. What is required is thermal equilibrium between radiation and matter, which a cavity or an optically thick gas provides.

Three more numbers from the same integral

Photon number density. Drop the factor of energy from the integrand and the integral becomes 2 zeta(3) = 2.40411, giving n = 2 zeta(3)(kT)^3/(pi^2 h_bar^3 c^3). For the cosmic microwave background at 2.7255 K, (kT)^3 = 5.328 x 10^-68 and pi^2 h_bar^3 c^3 = 3.119 x 10^-76, so n = 4.107 x 10^8 per cubic metre, or 411 photons per cubic centimetre. That number appears in every cosmology text and you have just computed it.

Energy density of the microwave background. a T^4 = 7.566 x 10^-16 x (2.7255)^4 = 4.17 x 10^-14 J/m3, which is 0.26 eV per cubic centimetre. Compare that with the mean rest-mass energy density of ordinary matter in the universe, roughly 200 times larger, and you see that the universe is matter dominated now and was radiation dominated early.

Radiation pressure. For a gas whose particles satisfy E = pc, P = U/3V, in contrast to 2U/3V for a non-relativistic gas. So P_rad = aT^4/3. At the Sun's centre, T = 1.57 x 10^7 K, that gives 1.53 x 10^13 Pa. The total central pressure is 2.3 x 10^16 Pa, so radiation supplies about 0.07 percent of the support. In a star of sixty solar masses the same calculation makes radiation pressure comparable to gas pressure, which is what limits how massive a star can be.

One more result follows from the grand potential. For photons mu = 0, so F = Phi = -PV = -aVT^4/3, and S = -(dF/dT) = (4/3) a V T^3. An adiabatic expansion therefore keeps V T^3 constant, so as the universe expanded by a linear factor of 1100 since recombination, the radiation temperature fell from about 3000 K to 2.7 K. The cooling of the microwave background is a thermodynamic identity.

Try it: the luminosity of two stars

Exercise. Compute the Sun's total power output from its radius and surface temperature.

Solution. Take R = 6.957 x 10^8 m and T = 5772 K. Then T^4 = 1.1100 x 10^15 K4 and sigma T^4 = 6.294 x 10^7 W/m2. The surface area is 4 pi R^2 = 6.0816 x 10^18 m2, so

L = 6.294 x 10^7 x 6.0816 x 10^18 = 3.828 x 10^26 W,

which is the accepted solar luminosity to four figures.

Change one input. Sirius A has T = 9940 K and R = 1.711 solar radii. The temperature is 1.72 times higher, so the flux is 1.72^4 = 8.8 times larger, and the area is 1.711^2 = 2.93 times larger. Multiplying, L = 25.7 solar luminosities, against a measured 25.4. Two numbers and a fourth power account for a star being twenty-five times brighter than the Sun while being only a little larger.

What to remember

Cavity modes have density V omega^2/(pi^2 c^3), each mode holds photons with Bose-Einstein occupancy at mu = 0, and the product is Planck's law. It reduces to the linear-in-T classical form at low frequency, which is what Rubens measured in October 1900, and to Wien's exponential at high frequency. Integrating gives U/V = aT^4 with a = pi^2 k^4/(15 h_bar^3 c^3) = 7.566 x 10^-16, and the emitted flux gives sigma = ac/4 = 5.6706 x 10^-8 W m-2 K-4, matching NIST. The wavelength peak obeys lambda_max T = 2.8978 mm K, while the frequency peak obeys a different equation and lies elsewhere. The same integrals give 411 CMB photons per cubic centimetre, a radiation pressure of aT^4/3, and the entropy (4/3)aVT^3 whose constancy makes the background cool as the universe expands.

Photons were the easy case, because they are massless and their chemical potential is zero. The next lesson applies the same mode counting to the vibrations of a solid, where the modes are finite in number and the cutoff turns out to matter enormously.

Sources

  1. Wikipedia contributors. (n.d.). Planck's law. Wikipedia. en.wikipedia.org
  2. National Institute of Standards and Technology. (2019). Stefan-Boltzmann constant. CODATA fundamental physical constants. physics.nist.gov
  3. Wikipedia contributors. (n.d.). Cosmic microwave background. Wikipedia. en.wikipedia.org
  4. Planck, M. (1901). Ueber das Gesetz der Energieverteilung im Normalspectrum. Annalen der Physik, 4, 553-563.
  5. Bose, S. N. (1924). Plancks Gesetz und Lichtquantenhypothese. Zeitschrift fuer Physik, 26, 178-181.
Key terms
Mode density
V omega^2/(pi^2 c^3) electromagnetic modes per unit angular frequency, including both polarisations.
Planck's law
The spectral energy density h_bar omega^3/(pi^2 c^3 (exp(h_bar omega/kT) - 1)) of radiation in equilibrium.
Ultraviolet catastrophe
The divergence produced by giving kT to each of an unbounded number of high-frequency modes; named by Ehrenfest in 1911.
Radiation constant a
pi^2 k^4/(15 h_bar^3 c^3) = 7.566 x 10^-16 J m^-3 K^-4, the coefficient in U/V = aT^4.
Stefan-Boltzmann constant
sigma = ac/4 = 2 pi^5 k^4/(15 h^3 c^2) = 5.670374 x 10^-8 W m^-2 K^-4.
Wien displacement constant
lambda_max T = 2.897771 mm K, from the root 4.965114 of x = 5(1 - exp(-x)).
Radiation pressure
P = aT^4/3, one third of the energy density because photons obey E = pc.
Photon gas entropy
S = (4/3)aVT^3, whose constancy under adiabatic expansion makes the microwave background cool as the universe grows.

Einstein, Debye, and the Heat Capacity of Real Solids

  • Compare the Einstein and Debye models of lattice vibrations term by term and explain why only one of them produces a T cubed law at low temperature.
  • Derive the Debye cutoff frequency and heat capacity, and evaluate both limits.
  • Test the Debye model numerically against copper and diamond, and identify where a single Debye temperature stops working.

In 1875 Heinrich Weber measured the heat capacity of diamond over a wide temperature range and found it climbing from under 3 joules per mole per kelvin near 200 K to about 21 at 1000 K, still short of the 24.9 that Dulong and Petit's law demands at every temperature. Einstein took Weber's table, plotted it in his 1907 paper, and fitted it with one adjustable frequency by applying Planck's quantisation to the vibrations of the atoms themselves. The fit was good, the idea was right, and the low-temperature end was wrong in a way that took five more years and a better mode count to fix.

This lesson puts the two models side by side and settles the comparison with numbers.

Einstein's model

Treat a solid of N atoms as 3N independent quantum harmonic oscillators, all at the same angular frequency omegaE. Lesson 4 already gave the heat capacity of one such oscillator, so with Theta_E = h_bar omega_E/k,

C = 3Nk (Theta_E/T)^2 e^(Theta_E/T)/(e^(Theta_E/T) - 1)^2.

At high temperature this tends to 3Nk, recovering Dulong-Petit and hence the classical result of Lesson 3. At low temperature it falls as e^(-Theta_E/T), because every mode has an energy gap and none can be excited until kT reaches it.

By 1911 Walther Nernst's laboratory in Berlin had heat capacities down to about 20 K for several solids, and they did not fall exponentially. They fell as T^3. The discrepancy is not marginal: at 10 K, copper's measured heat capacity is 0.055 J/(mol K), while Einstein's formula with any reasonable Theta_E gives something like 10^-7, seven orders of magnitude too small.

Debye's repair

Peter Debye's 1912 insight was that a solid's low-frequency vibrations are not independent atomic oscillations at all. They are sound waves, with wavelengths long compared with the lattice spacing, and their frequency goes to zero as the wavelength grows. Those modes have no gap, so they can always be excited, however cold the crystal.

Count them exactly as Lesson 11 counted photons, replacing the speed of light by an average sound speed v_s and two polarisations by three, one longitudinal and two transverse:

g(omega) = 3 V omega^2/(2 pi^2 v_s^3).

A crystal of N atoms has exactly 3N vibrational modes and no more, so unlike the photon gas the spectrum must be cut off. Debye's prescription is the simplest one: integrate the continuum density up to whatever frequency exhausts the count.

the integral from 0 to omega_D of g(omega) d(omega) = 3N gives omega_D^3 = 6 pi^2 N v_s^3 / V,

and the Debye temperature is Theta_D = h_bar omega_D/k. Each mode is a phonon oscillator with mu = 0, so the energy is

U = the integral from 0 to omega_D of h_bar omega g(omega)/(e^(beta h_bar omega) - 1) d(omega),

and differentiating with respect to T and substituting x = h_bar omega/kT,

C = 9Nk (T/Theta_D)^3 times the integral from 0 to Theta_D/T of x^4 e^x/(e^x - 1)^2 dx.

High temperature. When Theta_D/T is small the integrand is approximately x^2, the integral is (Theta_D/T)^3/3, and C = 3Nk. Dulong-Petit again.

Low temperature. The upper limit runs to infinity, the integral becomes the standard 4 pi^4/15 = 25.976, and

C = (12 pi^4/5) Nk (T/Theta_D)^3 = 1943.8 (T/Theta_D)^3 J/(mol K).

There is the T^3, and it comes from exactly the integral that produced T^4 for the photon gas in Lesson 11, one derivative later.

The comparison, read off a table

Einstein (1907)Debye (1912)
Picture of the vibrations3N independent oscillators, all at one frequencyElastic continuum carrying sound waves up to a cutoff
Adjustable parameterThetaEThetaD
High-temperature limit3R, correct3R, correct
Low-temperature behaviourexp(-ThetaE/T), far too fast1943.8(T/ThetaD)3, correct
Copper at 298 K23.5 J/(mol K)23.4 J/(mol K)
Copper at 10 Kabout 10-7 J/(mol K)0.048 J/(mol K)
Measured copper at 10 K0.055 J/(mol K), of which 0.007 is electronic

The two models are indistinguishable above roughly Theta/2 and separated by seven orders of magnitude at 10 K. The physical difference is one word: gap. Einstein's every mode has one and freezes exponentially; Debye's acoustic branch reaches down to zero frequency and always has modes cheap enough to excite.

What matters here: the low-temperature exponent is not a detail of the fit. It reports whether the excitation spectrum has a gap, which is one of the most useful diagnostics in condensed matter physics. A superconductor's exponential heat capacity below T_c is how the energy gap was first detected.

Testing Debye against real solids

SolidThetaD (K)T3 coefficient, J/(mol K4)C at 298 K predictedCV at 298 K measured
Lead1051.68 x 10-324.625.4
Silver2251.71 x 10-424.324.3
Copper3434.82 x 10-523.423.7
Aluminium4282.48 x 10-522.623.5
Diamond22301.75 x 10-74.16.1

Four rows agree to within a few percent. Diamond is off by 50 percent, and the failure is instructive rather than embarrassing. If you fit diamond's room-temperature heat capacity instead of its low-temperature one, you need Theta_D = 1860 K, which then gives 6.13 J/(mol K) against the measured 6.11. Diamond has no single Debye temperature.

That is general. Theta_D extracted from data drifts with temperature by 10 to 20 percent for most solids, because Debye's linear dispersion and sharp cutoff are a caricature of a real phonon spectrum, which has optical branches, flattened zone-boundary acoustic branches, and van Hove singularities in its density of states. Debye's model is a one-parameter interpolation that is exact in both limits and approximate in between, which is precisely what makes it useful.

A check that has no business working as well as it does

The Debye temperature is fitted to calorimetry near a few kelvin. The sound speed is measured by timing an ultrasonic pulse. The two have nothing in common except the theory, so use one to predict the other.

Copper: Theta_D = 343 K gives omega_D = k Theta_D/h_bar = 4.491 x 10^13 rad/s. With N/V = 8.49 x 10^28 per cubic metre, omega_D^3 = 6 pi^2 (N/V) v_s^3 rearranges to v_s = omega_D [V/(6 pi^2 N)]^(1/3) = 4.491 x 10^13 x 5.84 x 10^-11 = 2620 m/s.

Now the elastic measurement. Copper's longitudinal sound speed is 4760 m/s and its transverse speed is 2325 m/s, and Debye's average is defined by 3/v_s^3 = 1/v_L^3 + 2/v_T^3, which gives v_s = 2612 m/s. The two agree to 0.4 percent. A heat capacity measured near absolute zero predicts the speed of sound in a metal.

Common misconceptions

  • "Debye's theory is exact and Einstein's is wrong." Both are approximations to the same true phonon density of states. Debye is exact in both limits and wrong in the middle; Einstein is exact only at high temperature. For optical phonon branches, which are nearly flat, an Einstein term is actually the better description, and modern fits often use a Debye term plus one or two Einstein terms.
  • "The Debye temperature is a fixed property like a melting point." It is a fitted parameter and depends on which part of the data you fit. Diamond needs 2230 K at low temperature and 1860 K at room temperature.
  • "The T cubed law holds up to the Debye temperature." It holds below roughly Theta_D/50. For copper that means below about 7 K, which is why low-temperature calorimetry is where the constants are extracted.
  • "Phonons are real particles." They are quantised vibrational modes of the lattice. Like photons in a cavity their number is not conserved, so mu = 0, and they exist only as excitations of a medium.

Try it: separating the electrons from the lattice

Exercise. Below about 10 K a metal's heat capacity is C = gamma T + A T^3, one term from the electron gas and one from the phonons. Given copper measurements of C = 9.50 mJ/(mol K) at 5 K and C = 55.4 mJ/(mol K) at 10 K, extract gamma and Theta_D.

Solution. Divide by T: C/T = gamma + A T^2, so plotting C/T against T^2 gives a straight line with intercept gamma and slope A. At 5 K, C/T = 1.900 mJ/(mol K2) with T^2 = 25. At 10 K, C/T = 5.540 with T^2 = 100. The slope is (5.540 - 1.900)/(100 - 25) = 0.04853 mJ/(mol K4), and the intercept is 1.900 - 25 x 0.04853 = 0.687 mJ/(mol K2).

So gamma = 0.687 mJ/(mol K2), against the accepted 0.695, and A = 4.853 x 10^-5 J/(mol K4). Inverting A = 1943.8/Theta_D^3 gives Theta_D = (1943.8/4.853 x 10^-5)^(1/3) = (4.005 x 10^7)^(1/3) = 342 K, against the accepted 343.

Change one input. Try the same on an insulator, where gamma is zero because there are no free electrons. Two data points then overdetermine one parameter, and any departure from a line through the origin in the C/T against T^2 plot signals something other than acoustic phonons: a magnetic contribution, a Schottky anomaly from impurities, or in the case of a superconductor below T_c, an exponential term hiding in the data.

Where this leaves us

Einstein's 1907 model puts all 3N vibrational modes at one frequency, gets Dulong-Petit right at high temperature, and fails at low temperature because every mode is gapped. Debye's 1912 model counts sound waves as Lesson 11 counted photons, cuts the spectrum off at 3N modes to fix omega_D, and produces C = 1943.8 (T/Theta_D)^3 J/(mol K) at low temperature and 3R at high temperature. It matches lead, silver, copper and aluminium at room temperature to a few percent, and it requires two different Debye temperatures to fit diamond at low and high temperature, because a single elastic cutoff is a caricature of a real phonon spectrum. A calorimetric Theta_D of 343 K predicts copper's Debye sound speed to within 0.4 percent of the ultrasonic value.

One term in that low-temperature heat capacity has been taken on trust: the electronic gamma T. The next lesson derives it, and finds that the same calculation gives the pressure of the electron gas in a metal and, with one number changed, the pressure holding up a white dwarf.

Sources

  1. Wikipedia contributors. (n.d.). Debye model. Wikipedia. en.wikipedia.org
  2. Weisstein, E. W. (n.d.). Debye model. Eric Weisstein's World of Physics. scienceworld.wolfram.com
  3. Debye, P. (1912). Zur Theorie der spezifischen Waermen. Annalen der Physik, 39(4), 789-839.
  4. Kittel, C. (2005). Introduction to solid state physics (8th ed.), Chapter 5. John Wiley and Sons.
Key terms
Einstein model
3N independent oscillators at a single frequency, giving heat capacity that falls exponentially at low temperature.
Phonon
A quantised lattice vibration, with no conserved number and therefore zero chemical potential.
Debye frequency
The cutoff omega_D fixed by requiring the continuum mode count to equal 3N, giving omega_D^3 = 6 pi^2 N v_s^3/V.
Debye temperature
Theta_D = h_bar omega_D/k; 105 K for lead, 343 K for copper, above 1800 K for diamond.
Debye T cubed law
C = (12 pi^4/5)Nk(T/Theta_D)^3 = 1943.8(T/Theta_D)^3 J/(mol K), valid below roughly Theta_D/50.
Acoustic branch
The vibrational modes whose frequency tends to zero with wave number, which is why they never freeze out.
Debye sound speed
The average defined by 3/v_s^3 = 1/v_L^3 + 2/v_T^3, which for copper is 2612 m/s.
Electronic heat capacity coefficient
The gamma in C = gamma T + AT^3, extracted as the intercept of a C/T against T^2 plot.

Module 5: Degenerate and Condensed Matter

Push the quantum distributions until the temperature stops mattering. Fill the Fermi sea of a metal and find the pressure inside a white dwarf, then run the same logic for bosons and watch a gas of rubidium atoms collapse into a single quantum state.

The Electron Gas in Copper, and the Pressure Inside Sirius B

  • Construct the Fermi sphere at zero temperature and compute the Fermi wavevector, energy, temperature and velocity of the conduction electrons in copper.
  • Derive the Sommerfeld result that only a fraction of order T/T_F of the electrons can absorb heat, and check the predicted heat capacity against calorimetry.
  • Compute the degeneracy pressure of a degenerate electron gas and apply it to Sirius B, then identify where the non-relativistic treatment fails and why that failure produces the Chandrasekhar limit.

Copper conducts because about one electron per atom is free to wander: 8.49 x 1028 of them in every cubic metre, a gas some three thousand times denser than the air in this room. Equipartition, from Lesson 7, says a classical gas of point particles carries (3/2)Nk, so a mole of those electrons should add 12.47 J/(mol K) to copper's heat capacity. The measured molar heat capacity of copper at 298 K is 24.44 J/(mol K), and the phonons of Lesson 12 already account for 23.4 of that. The electrons are contributing roughly 0.21. They are short by a factor of sixty.

Paul Drude's 1900 electron gas was born with this problem and never solved it. Arnold Sommerfeld solved it in 1927 by changing exactly one thing: he filled the states with the Fermi-Dirac distribution of Lesson 10 instead of the Boltzmann distribution of Lesson 4. The consequences are not a small correction. With two numbers changed, the same arithmetic explains how a burnt-out star the mass of the Sun sits stably in a volume the size of the Earth.

Filling the box at absolute zero

Put N electrons in a cube of side L with periodic boundary conditions. The allowed wavevectors sit on a lattice of spacing 2 pi/L, so k space holds one state per volume 8 pi^3/V, and each takes two electrons, spin up and spin down. At T = 0 the Fermi-Dirac occupancy is a step, and the filled region is a sphere, the Fermi sphere, of radius k_F.

Count what fits inside it. N = 2 x (4/3) pi k_F^3 / (8 pi^3/V) = V k_F^3/(3 pi^2), so

k_F = (3 pi^2 n)^(1/3), and E_F = h_bar^2 k_F^2/(2 m_e).

Everything follows from the number density n and nothing else about the metal enters, which is the model's virtue or its weakness depending on what you ask it for.

Copper, worked. Density 8960 kg/m3, molar mass 63.546 g/mol, one conduction electron per atom:

n = (8960/0.063546) x 6.02214 x 10^23 = 8.491 x 10^28 per cubic metre.

3 pi^2 n = 29.6088 x 8.491 x 10^28 = 2.5141 x 10^30, and the cube root is k_F = 1.360 x 10^10 m-1.

That corresponds to a wavelength 2 pi/k_F = 0.462 nm, roughly twice copper's nearest-neighbour spacing of 0.256 nm: the electron density and the atom density are the same density, so their length scales match.

E_F = (1.11212 x 10^-68)(1.360 x 10^10)^2/(1.8219 x 10^-30) = 1.1288 x 10^-18 J = 7.04 eV.

Convert that to a temperature and the point appears at once:

T_F = E_F/k = 1.1288 x 10^-18/1.380649 x 10^-23 = 8.18 x 10^4 K.

In short: room temperature is 300 K and the Fermi temperature of copper is 82,000 K. On the scale that matters to the electrons, a copper wire at room temperature, at the boiling point of water, or in a furnace at 1300 K, is at absolute zero to within half a percent. Melting the copper barely disturbs the Fermi sea.

The Fermi velocity is v_F = h_bar k_F/m_e = 1.574 x 10^6 m/s, which is 0.53 percent of the speed of light. Those electrons are moving at fifteen hundred kilometres per second at absolute zero, and no amount of cooling will slow them down. Exclusion, not thermal agitation, puts them there.

Why the heat capacity is sixty times too small

Now the original puzzle. Raise the temperature from 0 to T. Which electrons can absorb energy? Only one that can find an empty state within about kT of where it started, and every state well below E_F is surrounded by full states. Only the shell of thickness kT around the Fermi surface is free to move, a fraction of roughly T/T_F of the total, so the energy rises by about N(T/T_F) x kT and the heat capacity is of order Nk(T/T_F) rather than (3/2)Nk. The Sommerfeld expansion makes the coefficient exact:

C_el = (pi^2/2) N k (T/T_F), so per mole C_el = (pi^2/2) R T/T_F = gamma T with gamma = pi^2 R/(2 T_F).

For copper at 300 K, T/T_F = 300/81758 = 3.67 x 10^-3, giving C_el = 0.151 J/(mol K) and gamma = 0.502 mJ/(mol K2).

Lesson 12 extracted gamma = 0.687 mJ/(mol K2) from copper's low-temperature calorimetry, and the accepted value is 0.695. The ratio 0.695/0.502 = 1.38 is the thermal effective mass: copper's band structure makes its electrons respond as though 38 percent heavier than free electrons. That is not a failure to apologise for; it is how the effective mass is measured.

The point: the classical prediction was 12.47 J/(mol K) and the truth is 0.21. The suppression factor is (pi^2/3)(T/T_F) = 1.2 percent, and it is small for exactly one reason: T_F is enormous, because the electron density is enormous, because exclusion stacks electrons into states of very high kinetic energy whether they are hot or not.

The gas that pushes at four hundred thousand atmospheres

The same stacking has a mechanical consequence. Integrating over the filled sphere gives U = (3/5) N E_F, and since E_F scales as n^(2/3), U is proportional to V^(-2/3). Then P = -dU/dV = (2/5) n E_F.

For copper: P = 0.4 x 8.491 x 10^28 x 1.1288 x 10^-18 = 3.83 x 10^10 Pa, that is 38.3 GPa, or 378,000 atmospheres. The electron gas inside an ordinary copper wire pushes outward at nearly four hundred thousand atmospheres, held in by the electrostatic attraction of the ion cores.

That claim is testable. Squeeze the metal and the pressure resists, with bulk modulus B = -V dP/dV = (5/3)P = 63.9 GPa for copper against a measured 140 GPa. The free electron gas gets the order of magnitude and half the number, missing the rest because it ignores the ion cores, which resist compression by a mechanism the model does not contain.

Common misconceptions

  • "Degeneracy pressure is electrons repelling each other electrostatically." Set the electron charge to zero and nothing in this calculation changes: P = (2/5)nE_F follows from state counting and exclusion alone. The pressure exists because the only place to put the next electron is a state of higher kinetic energy.
  • "At absolute zero the electrons stop moving." Copper's Fermi velocity is 1.57 x 106 m/s at T = 0 however cold you go. For fermions the lowest allowed configuration is fast.
  • "A white dwarf is held up by the heat left over from the star." Sirius B has a surface temperature of 25,000 K and a Fermi temperature, from the numbers computed below, of about 3.7 x 109 K. Thermal pressure is a rounding error. The star will cool for billions of years and not shrink measurably.

Change two numbers: Sirius B

In 1862 Alvan Graham Clark, testing a new 18.5 inch refractor, saw a faint companion beside Sirius. Sirius B has a mass of 1.018 solar masses inside a radius of 5,634 km, slightly smaller than the Earth. Nothing above changes except the density and the identity of the nuclei.

M = 1.018 x 1.989 x 10^30 = 2.025 x 10^30 kg; V = (4/3) pi (5.634 x 10^6)^3 = 7.49 x 10^20 m3; mean density = 2.70 x 10^9 kg/m3. A sugar cube of it would weigh 2.7 tonnes.

The star is carbon and oxygen, so there are two nucleons per electron and n_e = rho/(2 m_u) = 2.70 x 10^9/(3.321 x 10^-27) = 8.14 x 10^35 per cubic metre, which is 9.6 million times the electron density in copper. Feed that into the same two formulas:

E_F = (6.104 x 10^-39)(3 pi^2 x 8.14 x 10^35)^(2/3) = 5.09 x 10^-14 J = 0.318 MeV,

P = (2/5) n E_F = 0.4 x 8.14 x 10^35 x 5.09 x 10^-14 = 1.66 x 10^22 Pa.

Is that enough? Gravity has to be balanced, and for a sphere of uniform density the central pressure required is P_c = 3GM^2/(8 pi R^4):

P_c = 3 x 6.674 x 10^-11 x (2.025 x 10^30)^2/(8 pi x (5.634 x 10^6)^4) = 3.24 x 10^22 Pa.

Two calculations sharing no input except the star's mass and radius, one from gravitation and one from the exclusion principle, land within a factor of two of each other. That factor is the price of assuming uniform density: a real white dwarf is centrally condensed, and integrating the same equation of state through a proper polytrope closes the gap.

Bottom line: degeneracy pressure does not care about temperature, so it does not go away when the star cools. That is why a white dwarf is stable rather than slowly collapsing.

Where this calculation breaks, and what breaks it

Look again at E_F = 0.318 MeV and compare it with the rest energy of an electron, m_e c^2 = 0.511 MeV. The ratio is 0.62. The electrons at the top of the Fermi sea in Sirius B are not slightly relativistic; they are substantially relativistic, and E = p^2/2m, which the whole derivation used, is wrong for them.

Take the opposite extreme, E = pc, and redo the integral. It gives U = (3/4) N E_F with E_F = h_bar c (3 pi^2 n)^(1/3), so P = (1/4) n E_F, proportional to n^(4/3) rather than n^(5/3).

That exponent decides the fate of the star. Gravity demands P of order GM^2/R^4. Non-relativistic degeneracy supplies P proportional to (M/R^3)^(5/3) = M^(5/3)/R^5, and setting them equal gives R proportional to M^(-1/3): a heavier white dwarf is a smaller white dwarf, which is strange and is observed. In the relativistic limit degeneracy supplies M^(4/3)/R^4, the radius cancels from both sides, and balance is possible at one mass only. Subrahmanyan Chandrasekhar did this arithmetic on the boat from Madras to England in 1930 and published the mass in 1931: about 1.44 solar masses for mu_e = 2, the Chandrasekhar limit. Arthur Eddington told the Royal Astronomical Society in 1935 that there ought to be a law of Nature preventing it. There is not.

Try it: sodium, and then something much worse

Exercise. Sodium has one conduction electron per atom and n = 2.65 x 10^28 per cubic metre. Compute E_F, T_F and the degeneracy pressure, and compare with copper.

Solution. 3 pi^2 n = 7.846 x 10^29, and raising to the two thirds power gives 8.507 x 10^19 m-2. Then E_F = 5.193 x 10^-19 J = 3.24 eV, T_F = 3.76 x 10^4 K, and P = 5.5 GPa. Sodium's Fermi energy is less than half copper's on less than a third the electron density, because the dependence is only n^(2/3); its degeneracy pressure is seven times smaller, and sodium is correspondingly soft enough to cut with a knife.

Change one input. Replace the electron with the neutron. A neutron star of 1.4 solar masses in a 12 km radius has n = 2.30 x 10^44 neutrons per cubic metre. The prefactor h_bar^2/(2m) falls by the neutron-to-electron mass ratio of 1839, but n has risen by 108, and E_F comes out at 74 MeV with a pressure of 1.1 x 10^33 Pa, eleven orders of magnitude above Sirius B. Same three lines of algebra, a different fermion.

What to carry forward

At T = 0 a gas of fermions fills a sphere in k space out to k_F = (3 pi^2 n)^(1/3), giving E_F = h_bar^2 k_F^2/2m_e: 7.04 eV for copper, a Fermi temperature of 82,000 K, and a Fermi velocity of 0.53 percent of c that survives to absolute zero. Because only the fraction T/T_F near the surface can absorb heat, the electronic heat capacity is (pi^2/2)Nk(T/T_F), 0.151 J/(mol K) for copper at 300 K against a classical 12.47, and the measured gamma of 0.695 against the predicted 0.502 mJ/(mol K2) gives a thermal effective mass of 1.38. The same filled sphere exerts P = (2/5)nE_F = 38 GPa in copper and 1.7 x 10^22 Pa in Sirius B, within a factor of two of the 3.2 x 10^22 Pa gravity demands there. Since Sirius B's Fermi energy is 62 percent of the electron rest energy, the equation of state softens toward n^(4/3), and at that exponent the radius drops out of the balance, which is the Chandrasekhar limit.

Fermions were forced apart by exclusion. The next lesson runs the same grand canonical machinery for bosons, where the statistics pull the other way.

Sources

  1. Wikipedia contributors. (n.d.). Free electron model. Wikipedia. en.wikipedia.org
  2. Wikipedia contributors. (n.d.). Sirius. Wikipedia. en.wikipedia.org
  3. Wikipedia contributors. (n.d.). Chandrasekhar limit. Wikipedia. en.wikipedia.org
  4. Kittel, C. (2005). Introduction to solid state physics (8th ed.), Chapter 6. John Wiley and Sons.
  5. Chandrasekhar, S. (1931). The maximum mass of ideal white dwarfs. Astrophysical Journal, 74, 81-82.
Key terms
Fermi sphere
The region of k space filled at zero temperature, of radius k_F = (3 pi^2 n)^(1/3).
Fermi energy
The energy of the highest filled state at T = 0; 7.04 eV in copper, 3.24 eV in sodium.
Fermi temperature
E_F/k, the temperature at which thermal energy would rival the Fermi energy; 8.18 x 10^4 K for copper.
Degenerate
Describes a fermion gas at T much less than T_F, where the occupancy is essentially the zero-temperature step.
Sommerfeld expansion
The systematic expansion in T/T_F that gives C_el = (pi^2/2)Nk(T/T_F) and the O((T/T_F)^2) shift of mu below E_F.
Thermal effective mass
The ratio of measured gamma to the free electron prediction; 1.38 for copper.
Degeneracy pressure
P = (2/5)nE_F in the non-relativistic limit and (1/4)nE_F in the ultrarelativistic limit, present at zero temperature.
Chandrasekhar limit
About 1.44 solar masses, the mass above which relativistic electron degeneracy cannot balance gravity.

Bose-Einstein Condensation, and Two Thousand Atoms in One State

  • Derive the saturation of the excited states of an ideal Bose gas and obtain the condensation temperature from the phase space density criterion n lambda cubed = 2.612.
  • Evaluate that temperature for liquid helium-4 and for rubidium-87 at the density reported by the 1995 JILA experiment, and account for the gap between calculation and measurement in each case.
  • Derive the condensation condition for atoms held in a harmonic trap rather than a box, and use it to check the JILA numbers for internal consistency.

The thermal de Broglie wavelength of a rubidium-87 atom at room temperature is 10.8 picometres, about a twentieth of the atom's own diameter. Cool the gas and the wavelength grows as one over the square root of the temperature. At 170 nanokelvin it is 454 nanometres: longer than the wavelength of the orange light used to photograph it, and some forty thousand times what it was at 300 K.

Somewhere in between, the wavelength passes the spacing between neighbouring atoms, and the gas stops being a collection of separate objects. The atoms overlap. Lesson 10 showed that when they overlap, the counting changes, and for bosons the counting says something that Einstein worked out in 1925 and nobody managed to arrange in a gas for the next seventy years.

The excited states can only hold so much

Take the grand canonical result from Lesson 10 for bosons, with the ground state energy set to zero:

n_bar(eps) = 1/(e^((eps - mu)/kT) - 1).

The occupancy must be positive, so mu can never exceed the lowest energy: mu <= 0, and mu can only approach zero from below. Now sum over states in the usual way, replacing the sum with an integral over the free-particle density of states g(eps) proportional to eps^(1/2). The number of particles in excited states works out to

N_exc = V g_(3/2)(z)/lambda^3, where z = e^(mu/kT) and lambda = h/sqrt(2 pi m kT).

Here is the whole phenomenon in one observation. The function g_(3/2)(z) is increasing in z, z cannot exceed 1, and g_(3/2)(1) = zeta(3/2) = 2.6124, a finite number. So the excited states of a Bose gas have a maximum capacity:

N_exc,max = 2.6124 V/lambda^3.

Cool the gas at fixed N and lambda grows, so that ceiling falls. When it drops below N, the leftover particles have nowhere to go except the one state the integral quietly threw away. The density of states vanishes as eps^(1/2) at eps = 0, so the integral assigns the ground state zero weight, which is harmless when the ground state holds one particle in 1020 and catastrophic when it holds half of them.

Setting N = 2.6124 V/lambda^3 and solving for T gives the Bose-Einstein condensation temperature:

T_c = (2 pi h_bar^2/(m k)) (n/2.6124)^(2/3),

and below it the condensate fraction is N_0/N = 1 - (T/T_c)^(3/2).

So what?: the criterion is not a temperature. It is n lambda^3 = 2.612, a pure number comparing the volume an atom's wave packet occupies with the volume it has to itself. Rewriting the same condition, condensation begins when the mean interparticle spacing n^(-1/3) falls to 0.73 of the thermal wavelength.

Helium-4, where the number is close and the physics is not

Fritz London made the obvious first test in 1938, weeks after Kapitza, Allen and Misener announced that liquid helium below 2.17 K flows without viscosity. Liquid helium-4 has a density of 145 kg/m3, so with m = 6.646 x 10^-27 kg,

n = 145/6.646 x 10^-27 = 2.18 x 10^28 per cubic metre,

T_c = (7.615 x 10^-19)(2.18 x 10^28/2.6124)^(2/3) = 7.615 x 10^-19 x 4.116 x 10^18 = 3.13 K.

The measured lambda point is 2.17 K. A formula that assumes no interactions whatsoever, applied to a liquid whose atoms are permanently within a hard-core diameter of each other, lands 44 percent high. London was right that the mechanism was Bose statistics, and the quantitative agreement was never going to be better than that.

A sharper warning comes from the condensate fraction. The ideal gas predicts that at T = 0 every atom is in the ground state. Neutron scattering from liquid helium finds the condensate fraction below 10 percent even at the lowest temperatures, because interactions deplete the ground state occupancy severely. Helium-4 is a superfluid that contains a condensate; it is not a condensate.

What it takes to do it properly

To test Einstein's prediction you need a gas dilute enough that the atoms genuinely do not interact, which means a density perhaps a billion times lower than helium's, which by the two thirds power law means a temperature roughly a million times lower. Two techniques made that possible. Laser cooling in a magneto-optical trap brings a rubidium vapour from 300 K to around 10 microkelvin by scattering photons that preferentially oppose the atoms' motion. That is nowhere near enough, and it stalls, because the same scattered photons heat the sample.

The second technique is older and cruder: evaporative cooling, which is what you do when you blow across hot soup. Hold the atoms in a magnetic trap, then use a radio frequency field as an adjustable knife that flips the spin of any atom energetic enough to reach a chosen height, letting it escape; the rest rethermalise cooler. Lower the knife and repeat. The cost is brutal: JILA threw away more than 99.9 percent of their atoms to cool the remainder by a factor of a hundred.

On 5 June 1995 Eric Cornell and Carl Wieman ran that sequence on rubidium-87 and switched off the trap. The cloud expands, and after tens of milliseconds an absorption image records where the atoms went, which is a picture of the velocity distribution rather than of the cloud. What appeared was a narrow spike sitting on a broad thermal pedestal: a population with almost no velocity spread, coexisting with a normal gas. About two thousand atoms, at 170 nanokelvin. Cornell, Wieman and Wolfgang Ketterle shared the 2001 Nobel Prize in Physics for it.

Now check the number, and watch it fail

The published paper reports the condensate first appearing near 170 nK at a number density of 2.5 x 10^12 per cubic centimetre. Put that into the formula.

n = 2.5 x 10^18 m-3, so n/2.6124 = 9.570 x 10^17, and the two thirds power is 9.712 x 10^11 m-2. With 2 pi h_bar^2/(m k) = 3.507 x 10^-20 K m2 for rubidium-87,

T_c = 3.507 x 10^-20 x 9.712 x 10^11 = 3.41 x 10^-8 K = 34 nanokelvin.

Thirty four, not one hundred and seventy. The other way round: at 170 nK the thermal wavelength is 454 nm, so n lambda^3 = 0.23, eleven times short of 2.612. Those two numbers say the gas should not have condensed.

The formula is not wrong and the measurement is not wrong. The box is wrong. A uniform gas in a box has one density; atoms in a magnetic trap sit in a harmonic potential and pile up at the centre, so the density varies by orders of magnitude across the cloud. Condensation starts where the phase space density is highest, which means the n in the formula is the peak density at the trap centre, not a figure averaged over a cloud. To reach n lambda^3 = 2.612 at 170 nK you need 2.8 x 10^13 per cubic centimetre at the centre.

Worth holding on to: before you put a density into a thermodynamic formula, find out what geometry it was measured in and what geometry the formula assumes. This is the single most common way a correct calculation produces a wrong answer.

Redo it for a trap

In a harmonic trap of angular frequencies omega_x, omega_y, omega_z, the density of states is no longer proportional to eps^(1/2). Counting the levels of a three-dimensional oscillator gives g(eps) = eps^2/(2 (h_bar omega_bar)^3), where omega_bar is the geometric mean of the three frequencies. Run the same argument: integrate the Bose occupancy against that density of states at mu = 0, and the capacity of the excited states is

N_exc,max = zeta(3) (kT/(h_bar omega_bar))^3 = 1.202 (kT/(h_bar omega_bar))^3,

so kT_c = h_bar omega_bar (N/1.202)^(1/3) = 0.94 h_bar omega_bar N^(1/3), and the condensate fraction below T_c is 1 - (T/T_c)^3, a cube rather than a three halves power. Confinement changes the exponents, because it changed the density of states.

Test it against JILA. Two thousand atoms condensing at 170 nK requires

omega_bar = kT_c/(0.94 h_bar N^(1/3)) = 2.347 x 10^-30/(9.918 x 10^-35 x 12.60) = 1.88 x 10^3 rad/s,

that is a mean trap frequency of about 300 Hz, which is squarely in the range that magnetic traps of that generation delivered. The trapped-gas formula reproduces the experiment from its own numbers; the box formula does not.

Common misconceptions

  • "The atoms condense because they attract one another." The derivation above contains no interaction of any kind. Condensation in an ideal Bose gas is a purely statistical effect: the excited states run out of capacity, and the surplus has nowhere else to go. Interactions modify T_c by a few percent in a dilute gas and matter enormously afterwards, but they do not cause the transition.
  • "A condensate is a lump of atoms sitting still." The atoms occupy one quantum state, the ground state of the trap, which has a finite spatial extent and therefore, by the uncertainty principle, a finite momentum width. That is why the expansion image shows a narrow spike rather than a delta function, and why the spike's width measures the trap.
  • "Condensation happens when a gas gets cold enough." It happens when n lambda^3 reaches 2.612. Helium reaches it at 2.17 K and rubidium at 170 nK, a difference of seven orders of magnitude in temperature, entirely because their densities differ by ten.
  • "Bose-Einstein condensation and superfluidity are the same thing." An ideal condensate is not a superfluid: Landau's criterion gives it a critical velocity of zero, so the slightest perturbation creates excitations. Superfluidity requires interactions to produce a phonon-like excitation spectrum. Helium-4 is superfluid with under 10 percent condensate; the connection is real but it is not an identity.

Try it: the density you are not allowed to have

Exercise. Suppose you wanted rubidium-87 to condense at a comfortable 1 microkelvin instead of 170 nanokelvin. What density would that take, and why does nobody do it?

Solution. At 1 microkelvin, lambda = h/sqrt(2 pi m kT) = 187 nm and lambda^3 = 6.57 x 10^-21 m3. The criterion needs n = 2.6124/6.57 x 10^-21 = 3.98 x 10^20 per cubic metre, that is 4 x 10^14 per cubic centimetre, about 160 times the JILA density.

Change one input. That density is forbidden by chemistry rather than by physics. Three atoms that meet can form a molecule, releasing binding energy that ejects all three from the trap, and the rate of that process scales as the square of the density. At 1014 per cubic centimetre a rubidium sample lives for milliseconds. At the JILA density it lived, as the paper reports, for more than fifteen seconds. Going colder is cheaper than going denser, which is why every subsequent experiment went colder.

Putting it together

The excited states of an ideal Bose gas hold at most 2.6124 V/lambda^3 particles, because mu cannot rise above the ground state energy and zeta(3/2) is finite. Cool past the point where that ceiling falls below N and the surplus occupies the single lowest state, which the density of states integral had silently discarded. The criterion is n lambda^3 = 2.612, giving T_c = (2 pi h_bar^2/mk)(n/2.6124)^(2/3): 3.13 K for liquid helium-4 against a measured lambda point of 2.17 K, and 34 nK for rubidium-87 at the density JILA reported, against a measured onset of 170 nK. The helium discrepancy is interactions; the rubidium discrepancy is geometry, since the box formula wants the peak density of a trapped cloud rather than an average. Redoing the count for a harmonic trap gives kT_c = 0.94 h_bar omega_bar N^(1/3) and a condensate fraction going as 1 - (T/T_c)^3, and it reproduces the JILA result with a mean trap frequency near 300 Hz.

Both of the last two lessons watched a system change character sharply as a parameter crossed a threshold. That is a phase transition, and the next module treats them directly, starting with the one van der Waals wrote down in 1873 and the mistake his equation makes if you read it literally.

Sources

  1. Wikipedia contributors. (n.d.). Bose-Einstein condensate. Wikipedia. en.wikipedia.org
  2. The Nobel Prize in Physics 2001. NobelPrize.org. nobelprize.org
  3. Wikipedia contributors. (n.d.). Superfluid helium-4. Wikipedia. en.wikipedia.org
  4. Anderson, M. H., Ensher, J. R., Matthews, M. R., Wieman, C. E., and Cornell, E. A. (1995). Observation of Bose-Einstein condensation in a dilute atomic vapor. Science, 269(5221), 198-201.
  5. Pethick, C. J., and Smith, H. (2008). Bose-Einstein condensation in dilute gases (2nd ed.), Chapters 2 and 10. Cambridge University Press.
Key terms
Thermal de Broglie wavelength
lambda = h/sqrt(2 pi m kT); 10.8 pm for rubidium-87 at 300 K and 454 nm at 170 nK.
Phase space density
n lambda^3, the number of atoms within a thermal wavelength cubed; condensation begins at 2.612.
Saturation of the excited states
The finite ceiling 2.6124 V/lambda^3 on how many bosons the excited states can hold, which is what forces condensation.
Condensate fraction
1 - (T/T_c)^(3/2) in a uniform box and 1 - (T/T_c)^3 in a harmonic trap.
Evaporative cooling
Removing the most energetic atoms with a radio frequency knife and letting the rest rethermalise cooler.
Lambda point
The 2.17 K transition at which liquid helium-4 becomes superfluid, against an ideal-gas prediction of 3.13 K.
Quantum depletion
The interaction-driven reduction of ground state occupancy, which keeps liquid helium's condensate fraction below 10 percent.
Three-body recombination
The loss process, scaling as density squared, that sets the practical upper limit on trapped atomic densities.

Module 6: Phase Transitions and the Arrow of Time

Watch a system change character as a parameter crosses a threshold. Repair the van der Waals isotherm with the Maxwell construction, solve the Ising model exactly in one dimension and quote Onsager in two, and then face the two objections that nearly destroyed Boltzmann's H-theorem.

The Isotherm That Cannot Be Right: van der Waals and the Maxwell Construction

  • Evaluate the van der Waals equation on a subcritical isotherm, locate the branch where dP/dv is positive, and explain why no substance can occupy it.
  • Derive the Maxwell equal-area construction from the equality of Gibbs free energy between coexisting phases, and apply it numerically to carbon dioxide at 280 K.
  • Obtain the critical point from the equation, show that Z_c = 3/8 is its one genuine prediction, and compare that with measured values for six fluids.

Set T = 280 K in the van der Waals equation for carbon dioxide, with a = 0.3640 Pa m6/mol2 and b = 4.267 x 10^-5 m3/mol, and evaluate the pressure at two molar volumes. At 100 cm3/mol it gives 4.21 MPa. At 180 cm3/mol it gives 5.72 MPa.

Read that again. The gas was allowed to expand by 80 cubic centimetres and its pressure went up by more than a third. A substance on that branch has a negative isothermal compressibility: squeeze a small region of it and the squeezed region pushes back less than its surroundings, so the squeeze continues, and the fluid tears itself apart. Nothing in the laboratory does this. Something in the reasoning is wrong, and the useful question is exactly which step.

Two corrections to the ideal gas

Johannes Diderik van der Waals defended his Leiden thesis in 1873 with an equation that adds two terms to Pv = RT. The first is geometric: molecules occupy space, so the volume available for motion is not v but v - b, where b is an excluded volume per mole. The second is attractive: a molecule at the wall has neighbours on one side only, so it strikes the wall more gently than a molecule in the bulk would. The number of such deficient collisions is proportional to the density, and the strength of each pull is proportional to the density again, so the pressure defect goes as 1/v^2. Together,

(P + a/v^2)(v - b) = RT, or P = RT/(v - b) - a/v^2.

Both constants are checkable against molecular size. Treating molecules as hard spheres gives b = 4 N_A (4/3) pi r^3, four times the volume they actually occupy. For carbon dioxide, b = 4.267 x 10^-5 m3/mol implies r = 0.162 nm, and the kinetic diameter of carbon dioxide measured by gas diffusion is 0.33 nm, so r = 0.165 nm. Two independent routes to a molecular radius agree to two percent. The van der Waals equation is not numerology.

Where the reasoning actually fails

Multiply out and you have a cubic in v:

P v^3 - (Pb + RT) v^2 + a v - ab = 0.

Above the critical temperature it has one real root and everything is simple. Below, it can have three, and the flaw in the opening calculation is now visible. It is not in the equation. It is in the silent assumption that whatever the equation returns at a given (T, P) is what the substance does.

A cubic with three roots is offering three candidate homogeneous states. Thermodynamics does not oblige the system to pick one. It obliges the system to minimise the Gibbs free energy at fixed T and P, and one of the options available is not a homogeneous state at all: split into two phases of different density, a liquid at the small root and a vapour at the large one, in whatever proportion fills the container. The equation of state supplies candidates. It does not adjudicate between them.

Remember: an equation of state is a statement about one homogeneous phase. Everything about coexistence has to be imposed from outside it, by minimising a free energy.

The equal-area rule, derived

Two phases coexist at the same T and P when their molar Gibbs free energies are equal, since otherwise molecules move to the cheaper phase. At fixed temperature, dg = v dP, so equality of g between the liquid point and the vapour point means

the integral from P_l to P_g of v dP = 0, taken along the van der Waals isotherm.

Integrating by parts turns that into a statement about the picture. On a plot of P against v, drawing the horizontal line at the coexistence pressure cuts two closed areas out of the S-shaped loop, one above the line and one below. The condition says those two areas are equal. That is the Maxwell construction, published by James Clerk Maxwell in 1875, and it replaces the whole loop with a flat line at a pressure the equation itself never singles out.

Carry it out for carbon dioxide at 280 K. Solving the cubic together with the equal-area condition numerically gives

P_sat = 5.3 MPa, with v_liquid = 81 cm3/mol and v_vapour = 263 cm3/mol.

The measured saturation pressure of carbon dioxide at 280 K is 4.16 MPa. The construction is exactly right and the equation it is applied to is approximate, so the answer is 27 percent high. In reduced units the comparison is cleaner: van der Waals puts P_sat/P_c at 0.71 where the real fluid puts it at 0.56, and that overestimate is systematic across substances.

The one number the equation genuinely predicts

At the critical point the two turning points of the cubic merge, so dP/dv = 0 and d^2P/dv^2 = 0 together. Solving those two conditions gives

v_c = 3b, T_c = 8a/(27 R b), P_c = a/(27 b^2).

For carbon dioxide: T_c = 8 x 0.3640/(27 x 8.3145 x 4.267 x 10^-5) = 304.0 K and P_c = 0.3640/(27 x 1.8207 x 10^-9) = 7.40 MPa, against measured values of 304.13 K and 7.377 MPa. Impressive, and entirely circular: a and b are two numbers, and they were fitted to those two measurements in the first place.

The critical point, however, has three coordinates. Having spent its two constants on T_c and P_c, the equation must predict the third, and the cleanest way to state that prediction is the critical compressibility factor:

Z_c = P_c v_c/(R T_c) = [a/(27b^2)](3b)/(R x 8a/(27Rb)) = 3/8 = 0.375, for every substance.

FluidTc (K)Measured ZcError in vc
Helium5.190.301+25 percent
Argon150.70.291+29 percent
Nitrogen126.20.289+30 percent
Methane190.60.286+31 percent
Carbon dioxide304.10.274+37 percent
Water647.10.229+64 percent

Not one fluid is near 0.375, and, more damning, not one is above it. Scattered errors mean noise; errors all in one direction mean a missing physical effect. What is missing is fluctuations: near the critical point the density varies wildly from place to place on scales far larger than a molecule, and van der Waals's derivation replaced the local density everywhere by its average. Lesson 16 shows what that costs.

The part of the loop that is real after all

Do not throw the whole S away. Its middle section, where dP/dv > 0, runs between the two turning points, which for carbon dioxide at 280 K are at 95 and 185 cm3/mol. That locus is the spinodal, and inside it no state can exist for even an instant.

But between the coexistence volumes, 81 and 263 cm3/mol, and the spinodal volumes, 95 and 185, lie two strips where dP/dv < 0 and the fluid is mechanically stable while still costing more free energy than the two-phase mixture. These are metastable: a liquid heated past its boiling point without boiling, and a vapour cooled past its dew point without condensing. Both are ordinary laboratory objects.

Donald Glaser filled a chamber with liquid hydrogen in 1952, dropped the pressure to put it in exactly that superheated strip, and found that a charged particle crossing it left a track of bubbles, because the ions it produced supplied the nucleation sites the pure liquid lacked. C. T. R. Wilson had done the mirror image in 1911 with supersaturated vapour in a cloud chamber. Two of the twentieth century's principal particle detectors operate in the region a careless reading of the isotherm discards.

The upshot: the Maxwell construction tells you where equilibrium is. It does not tell you the system will get there. Nucleation needs a seed, and a clean liquid in a smooth container can sit superheated for a long time, which is why water in a microwave sometimes erupts when you disturb it.

Common misconceptions

  • "The wiggle is an error in the equation, so cut it out." Only the section between the turning points is unphysical. The two flanking strips are metastable states you can make, photograph and use to detect particles.
  • "a and b are measured molecular properties." They are fitted constants that happen to correspond roughly to molecular quantities. The correspondence for b is good to a factor of four exactly, since b is four times the hard-sphere volume, not one times it.
  • "Coexistence happens where the loop crosses its own pressure twice." The loop crosses a horizontal line at many pressures. Only one of them equalises the two areas, and only that one equalises the Gibbs free energies.
  • "Getting T_c and P_c right within 0.1 percent shows the equation is accurate." Those two were used to fix a and b. The test is the third coordinate, and there the equation is 29 to 64 percent wrong.

Try it: water, and then argon

Exercise. Water has T_c = 647.1 K and P_c = 22.064 MPa. Invert the critical relations to get a and b, then predict v_c and compare it with the measured 55.9 cm3/mol.

Solution. From T_c = 8a/(27Rb) and P_c = a/(27b^2), dividing one by the other gives b = R T_c/(8 P_c) = 8.3145 x 647.1/(8 x 2.2064 x 10^7) = 3.048 x 10^-5 m3/mol, and then a = 27 R^2 T_c^2/(64 P_c) = 0.5535 Pa m6/mol2. Both match the tabulated van der Waals constants for water. The prediction is v_c = 3b = 91.4 cm3/mol against a measured 55.9: too large by 64 percent.

Change one input. Repeat for argon, T_c = 150.687 K and P_c = 4.863 MPa. You get b = 3.221 x 10^-5 m3/mol and v_c = 96.6 cm3/mol against a measured 74.6, an error of 29 percent. Argon is a sphere with no dipole and no hydrogen bonds; water is a bent molecule that forms a tetrahedral network. The equation's error tracks how badly a molecule fails to be a featureless attractive sphere, which tells you precisely what the model left out.

The takeaway

The van der Waals equation adds an excluded volume b and a mean attraction a/v^2 to the ideal gas, and below T_c it becomes a cubic whose three roots include a branch with dP/dv > 0 that no substance can occupy. The equation is not at fault; the assumption that a single homogeneous phase must exist at every point is. Minimising the Gibbs free energy allows a two-phase split, and equality of g gives the equal-area construction, which for carbon dioxide at 280 K returns P_sat = 5.3 MPa with coexisting volumes of 81 and 263 cm3/mol, against a measured saturation pressure of 4.16 MPa. The critical relations v_c = 3b, T_c = 8a/27Rb and P_c = a/27b^2 fix a and b from two measurements, so the equation's only free prediction is Z_c = 3/8, and every real fluid sits below it, from argon's 0.291 to water's 0.229. Between the coexistence curve and the spinodal lie the metastable strips that make bubble and cloud chambers work.

Expanding the equation about the critical point gives a coexistence curve with v_g - v_l proportional to (T_c - T)^(1/2). In 1945 Edward Guggenheim plotted the measured coexistence curves of eight fluids together and found them collapsing onto a single curve with an exponent near 1/3. The next lesson explains why that discrepancy is not a defect of van der Waals in particular.

Sources

  1. Wikipedia contributors. (n.d.). Van der Waals equation. Wikipedia. en.wikipedia.org
  2. Wikipedia contributors. (n.d.). Maxwell construction. Wikipedia. en.wikipedia.org
  3. National Institute of Standards and Technology. (n.d.). Carbon dioxide: phase change data. NIST Chemistry WebBook. webbook.nist.gov
  4. van der Waals, J. D. (1873). Over de continuiteit van den gas- en vloeistoftoestand [doctoral thesis]. Leiden University.
  5. Guggenheim, E. A. (1945). The principle of corresponding states. Journal of Chemical Physics, 13(7), 253-261.
Key terms
Excluded volume b
The van der Waals volume correction, equal to four times the hard-sphere volume of a mole of molecules.
Attractive constant a
The coefficient in the pressure defect a/v^2, arising from pairwise attraction and hence from the density squared.
Maxwell construction
The equal-area rule that replaces the unphysical loop with a horizontal coexistence line, equivalent to equating molar Gibbs free energies.
Binodal
The coexistence curve, the locus of the two volumes joined by the Maxwell line; 81 and 263 cm^3/mol for CO2 at 280 K.
Spinodal
The locus where dP/dv = 0, inside which a homogeneous fluid is mechanically unstable; 95 and 185 cm^3/mol for CO2 at 280 K.
Metastable state
A superheated liquid or supercooled vapour, stable against small fluctuations but not the global free energy minimum.
Critical compressibility factor
Z_c = P_c v_c/(RT_c); van der Waals predicts 3/8 for everything, and measured values run from 0.301 down to 0.229.
Law of corresponding states
In reduced variables the van der Waals equation contains no substance-specific constants, so all fluids should fall on one curve.

The Ising Model: One Dimension Exactly, Two Dimensions by Onsager, and Why the Exponents Repeat

  • Solve the one-dimensional Ising model exactly by transfer matrix and explain, by counting domain walls, why it orders only at absolute zero while two dimensions does not.
  • Derive the mean field self-consistency equation and the Landau expansion, extract the exponents alpha, beta, gamma and delta, and compare them with Onsager's exact two-dimensional results.
  • State what universality claims, verify the Rushbrooke and Widom relations across three exponent sets, and identify what a universality class actually depends on.

Ernst Ising submitted a doctoral thesis at Hamburg in 1924 on a model his supervisor Wilhelm Lenz had invented: spins on a line, each up or down, each pulled into line by its two neighbours. Ising solved it exactly and found no magnetism. Not weak magnetism, not magnetism below some low temperature: none at all above absolute zero. He concluded the model failed, guessed three dimensions would fail the same way, and left research.

The guess was wrong and the calculation was right. Working out why one dimension is barren and two is not explains something larger: why a boiling liquid and a magnet losing its magnetism obey the same numerical laws near their critical points.

The model, and the two systems it secretly describes

Put a variable s_i = +1 or -1 on each lattice site and write H = -J (sum over neighbouring pairs) s_i s_j - h (sum over i) s_i. With J > 0 neighbours prefer to agree, and h is an external field. That is the whole Ising model.

It is also, relabelled, a fluid. Let s_i = +1 mean the site holds a molecule and -1 mean it is empty; then J is the attraction between neighbouring molecules, h is the chemical potential, and the magnetisation is the density. The magnet's Curie point and the fluid's critical point are the same mathematics in different clothes.

One dimension, solved

For a ring of N spins the partition function factorises through a 2 by 2 transfer matrix with entries T(s, s') = exp(beta J s s' + beta h (s + s')/2), so Z = trace(T^N) and in the thermodynamic limit only the larger eigenvalue survives. At zero field the eigenvalues are 2 cosh(beta J) and 2 sinh(beta J), giving f = -kT ln(2 cosh(J/kT)).

That is smooth, with derivatives of every order, for every T > 0. A phase transition is a non-analyticity in the free energy; there is none, so there is no transition. The correlation is C(r) = [tanh(J/kT)]^r, decaying with length xi = -1/ln tanh(J/kT): 3.7 lattice spacings at kT = J, 2 x 10^8 at kT = J/10. It grows without bound as T falls, but reaches infinity only at T = 0.

Why dimension decides it

Rudolf Peierls gave the argument in 1936 and it takes three lines. Start a chain of N aligned spins and insert one domain wall. The cost is 2J, one unhappy bond; the wall can sit at any of about N bonds, so the entropy gain is k ln N, and Delta F = 2J - kT ln N is negative for any T > 0 once N > exp(2J/kT). At kT = J a chain of eight spins already prefers to break; at kT = J/10 it takes 5 x 108. Entropy scales with system size and the cost of a wall does not.

In two dimensions a domain wall is a closed loop. A loop of length L costs 2JL, and the number of self-avoiding loops of that length grows no faster than 3^L, since each step has at most three choices. So Delta F >= L(2J - kT ln 3), positive for kT < 2J/ln 3 = 1.82 J: below that, large domains cost free energy rather than saving it, so a transition exists and lies above 1.82 J/k.

Key idea: ordering is a contest between the energy cost of a defect and the entropy of where to put it. In one dimension the defect costs a constant and buys ln N; in two the two sides both scale with L, so a coefficient decides and there is a temperature where the decision flips.

Mean field, and the price of averaging

Solving in three dimensions is beyond anyone, so approximate. Single out one spin and replace each of its z neighbours by the average magnetisation m. The spin then feels an effective field zJm + h, and a single spin in a field is Lesson 4's problem: m = tanh((zJm + h)/kT).

At h = 0 this always has the solution m = 0. A second appears when the slope of the right-hand side at the origin exceeds 1, that is when zJ/kT > 1, so mean field theory predicts kT_c = zJ and nothing else about the lattice matters.

LatticezMean field kTc/JTruthError
Chain220 (Ising, 1924)infinite
Square442.269 (Onsager, 1944)+76 percent
Simple cubic664.511 (numerical)+33 percent

Mean field always overestimates T_c, and the error shrinks as the coordination number grows: replacing a neighbour by its average discards that neighbour's fluctuation, and the more neighbours there are the less that matters. On a chain the approximation does not merely err, it invents a transition that is not there.

Landau's version, and the exponents that come with it

Lev Landau's 1937 reformulation drops the microscopic model entirely. Write the free energy as a power series in the order parameter m, keeping only terms the symmetry allows; reversing all spins must not change the zero-field energy, so only even powers appear:

f(m) = f_0 + a(T - T_c) m^2 + b m^4 - hm, with a and b positive.

Above T_c the quadratic coefficient is positive and the minimum sits at m = 0; below, it turns negative and the minimum splits in two. Minimising gives 2a(T - T_c)m + 4bm^3 = h, and four results drop out:

  • At h = 0, T < T_c: m = sqrt(a(T_c - T)/2b), so beta = 1/2.
  • At small h above T_c: chi = 1/(2a(T - T_c)), so gamma = 1.
  • On the critical isotherm h = 4bm^3, so delta = 3.
  • The heat capacity jumps by a^2 T_c/2b without diverging, so alpha = 0.

The core of it: Landau theory never mentions spins, lattices or the value of J. Its exponents follow from two assumptions: that f is smooth in m near T_c, and that the symmetry is m to -m. That is why van der Waals gave beta = 1/2 in Lesson 15. Every mean field theory ever written yields the same four numbers, which is impressive until you check them.

Onsager, and the numbers that do not match

In 1944 Lars Onsager published the exact free energy of the zero-field square-lattice Ising model, one of the hardest calculations completed in the subject. The transition temperature satisfies sinh(2J/kT_c) = 1, giving kT_c/J = 2/ln(1 + sqrt 2) = 2.269185, comfortably above Peierls's bound of 1.82 and far below mean field's 4. The spontaneous magnetisation is m = [1 - sinh^-4(2J/kT)]^(1/8), whose exponent is 1/8, not 1/2. Onsager announced that from the floor of a conference in 1948 without a derivation, and C. N. Yang published the proof in 1952.

ExponentMeasuresMean field2D exact3D Ising
alphaC, heat capacity0 (jump)0 (log)0.110
betam, order parameter1/21/80.3265
gammachi, susceptibility17/41.2372
deltam against h at Tc3154.7898
nuxi, correlation length1/210.6300
etacorrelations at Tc01/40.0363

Three columns of very different numbers, and yet two identities hold in all of them. Rushbrooke's relation alpha + 2 beta + gamma = 2 gives 0 + 1 + 1, 0 + 1/4 + 7/4, and 0.110 + 0.653 + 1.237. Widom's relation gamma = beta(delta - 1) gives 0.5 x 2 = 1, (1/8) x 14 = 7/4, and 0.3265 x 3.7898 = 1.2372. The exponents depend on dimension; the relations between them do not.

Universality

Here is the claim the table points at. The critical exponents do not depend on the lattice, on J, on whether the interaction is between spins or molecules, or on anything else microscopic. They depend on the dimensionality of space, the number of components of the order parameter, and the interaction range.

So xenon at its liquid-gas critical point, a uniaxial ferromagnet at its Curie point and a binary alloy at its ordering temperature all sit in the three-dimensional Ising class, measured with beta near 0.33 and gamma near 1.24. Superfluid helium-4 has a two-component order parameter and belongs elsewhere; its alpha is slightly negative, and the most precise value, -0.0127 plus or minus 0.0003, was obtained aboard the Space Shuttle in 1992, because on Earth the pressure gradient in a column of helium smears the transition over a wider temperature range than the effect being measured.

The explanation is Leo Kadanoff's block spins and Kenneth Wilson's renormalisation group. Coarse-grain repeatedly and watch the couplings: almost every microscopic detail shrinks away, and the handful that grow are fixed by symmetry and dimension. Above four dimensions the fluctuations that ruin mean field are themselves irrelevant, which is why Landau is exactly right for d >= 4.

Common misconceptions

  • "The one-dimensional model fails because one dimension is unphysical." The domain-wall argument applies in every dimension. It simply wins in one, where a defect's cost is a constant while its entropy grows with system size.
  • "Mean field theory is just wrong." It is exact above four dimensions and excellent whenever each degree of freedom couples to many others. BCS superconductivity is a mean field theory, and the region around T_c where it fails is measured in microkelvin.
  • "Onsager solved the Ising model." He solved the square lattice at zero field. The two-dimensional model in a field has never been solved, and neither has three dimensions in any field, including none.

Try it: how fast the magnetisation grows

Exercise. For the square lattice in mean field theory, z = 4 and kT_c = 4J. Solve m = tanh(m T_c/T) at T/T_c = 0.9 and compare with the near-critical form m = sqrt(3t), t = 1 - T/T_c.

Solution. Iterating m = tanh(1.1111 m) from any nonzero start converges to m = 0.525. The near-critical form gives sqrt(0.3) = 0.548, high by 4 percent, which is what a leading-order expansion should do ten percent away from T_c.

Change one input. Put the same reduced temperature into Onsager's exact answer with the true kT_c = 2.269 J. Then 2J/kT = 0.9793, sinh(0.9793) = 1.1435, and m = [1 - 1.1435^-4]^(1/8) = 0.415^(1/8) = 0.896. Mean field says the square lattice is 53 percent magnetised at nine tenths of its Curie temperature; exactly, it is 90 percent. An exponent of 1/8 sends the order parameter up almost vertically below T_c, and no adjustment of J reconciles the two curves.

Looking back

In one dimension the transfer matrix gives f = -kT ln(2 cosh(J/kT)), analytic everywhere, so there is no transition above T = 0: a domain wall costs 2J and earns kT ln N, and entropy always wins in a long enough chain. In two dimensions a wall costs and earns amounts both proportional to its length, Peierls's argument puts T_c above 1.82 J/k, and Onsager's 1944 solution puts it at 2.269 J/k. Mean field predicts kT_c = zJ, 76 percent high on the square lattice, 33 percent high on the simple cubic lattice, and catastrophically wrong on a chain. Landau's expansion reproduces mean field's exponents from symmetry alone, 0, 1/2, 1, 3, against exact two-dimensional values of 0, 1/8, 7/4, 15 and three-dimensional values of 0.110, 0.3265, 1.2372, 4.7898. All three sets satisfy alpha + 2 beta + gamma = 2 and gamma = beta(delta - 1). The exponents depend on dimension, order parameter symmetry and interaction range, and on nothing else, which is why a boiling fluid and a magnet share them.

All of this rests on equilibrium: on the assumption that a system left alone reaches a state describable by a free energy and stays there. The last lesson asks how time-reversible microscopic equations produce that behaviour, and whether Boltzmann ever answered his critics.

Sources

  1. Wikipedia contributors. (n.d.). Ising model. Wikipedia. en.wikipedia.org
  2. Wikipedia contributors. (n.d.). Square lattice Ising model. Wikipedia. en.wikipedia.org
  3. Wikipedia contributors. (n.d.). Critical exponent. Wikipedia. en.wikipedia.org
  4. Onsager, L. (1944). Crystal statistics I: A two-dimensional model with an order-disorder transition. Physical Review, 65(3-4), 117-149.
  5. Goldenfeld, N. (1992). Lectures on phase transitions and the renormalization group, Chapters 3 to 9. Addison-Wesley.
Key terms
Transfer matrix
The 2 by 2 matrix whose largest eigenvalue gives the free energy per spin of the one-dimensional Ising chain.
Domain wall
A boundary between oppositely aligned regions; its energy cost against its entropy decides whether order survives.
Peierls argument
The 1936 proof that two-dimensional order survives below 2J/(k ln 3) = 1.82 J/k, because loop entropy grows only as the exponential of the loop length.
Mean field theory
Replacing each neighbour by the average magnetisation, giving m = tanh(zJm/kT) and kT_c = zJ.
Landau expansion
f = f_0 + a(T - T_c)m^2 + b m^4 - hm, whose exponents follow from analyticity and symmetry alone.
Onsager solution
The exact 1944 free energy of the zero-field square-lattice Ising model, with kT_c/J = 2/ln(1 + sqrt 2) = 2.269185.
Universality class
The set of systems sharing critical exponents, fixed by spatial dimension, order parameter components and interaction range.
Rushbrooke and Widom relations
alpha + 2 beta + gamma = 2 and gamma = beta(delta - 1), satisfied by mean field, 2D and 3D exponent sets alike.
Upper critical dimension
d = 4 for the Ising class; above it fluctuations are irrelevant and Landau's exponents are exact.

The H-Theorem on Trial: Loschmidt, Zermelo, and What Is Still Open

  • State the H-theorem and identify precisely which step in its derivation is not time-symmetric.
  • Reconstruct the reversibility and recurrence objections, give Boltzmann's replies, and evaluate the recurrence time numerically for a cubic centimetre of air.
  • Distinguish the parts of the dispute that are settled from the parts that remain open, including the initial condition problem and the dependence of H on a choice of coarse-graining.

Josef Loschmidt and Ludwig Boltzmann had shared a building in Vienna for years, and Loschmidt had taught the younger man, when in 1876 he published an objection that fitted in a sentence. Boltzmann's 1872 paper had derived, from Newtonian collisions alone, a quantity that could only decrease. Loschmidt pointed out that Newtonian collisions run backwards perfectly well. Reverse every velocity in the gas and you have another legal mechanical state, one in which the quantity must increase. So it cannot be a theorem of mechanics.

Twenty years later Ernst Zermelo attacked the same result from a different direction and Boltzmann answered him too, more sharply. Both objections were correct. Both replies were correct. The interesting question is what precisely each side established, and what neither of them settled.

The claim on trial

Boltzmann's 1872 paper tracks f(v, t), the number of molecules per unit volume with velocity near v. Collisions redistribute velocities, and to count them he assumed that the number of collisions between molecules of velocity v_1 and v_2 is proportional to f(v_1) f(v_2): the two molecules arrive uncorrelated. That is the Stosszahlansatz, the assumption of molecular chaos, and everything turns on it.

Define H = the integral of f ln f over velocity space. Boltzmann showed that

dH/dt <= 0, with equality only when f is the Maxwell-Boltzmann distribution.

Since S = -kH up to constants, that is the H-theorem: the second law, apparently derived from mechanics, with the Maxwell distribution emerging as the unique endpoint rather than being assumed.

The case for the objection

Loschmidt, 1876. Newton's equations are invariant under t to -t together with v to -v. Take a gas that has relaxed from some ordered start, so H has fallen. Now reverse every molecular velocity. The resulting configuration is a perfectly legitimate state of the same mechanical system, and by determinism it retraces the original trajectory backwards, so H climbs back to where it started. A quantity that increases along a genuine mechanical trajectory cannot obey a mechanical theorem forbidding it to increase. This is not an intuition or an appeal to plausibility; it is a construction.

Zermelo, 1896. Planck's assistant in Berlin brought a second, independent weapon: Henri Poincare's recurrence theorem of 1890. A Hamiltonian system confined to a bounded region of phase space returns arbitrarily close to its initial state, for almost every initial condition. So H, a function of that state, must eventually return arbitrarily close to its initial value. Monotonic decrease forever is impossible, and no coarse-graining changes the theorem.

Both arguments rest on results nobody disputes. Neither depends on any property of gases. That is what makes them serious.

The case for the reply

Boltzmann's answer to Loschmidt, developed in 1877, is the point at which statistical mechanics becomes statistical. The reversed state exists; the question is how many states like it there are. Among all the microstates compatible with a given macrostate, the overwhelming majority evolve toward equilibrium, and the anti-thermodynamic ones are a vanishing minority. Entropy is the logarithm of the number of microstates, S = k log W, and the second law is a statement about counting.

That reply produced the equation on Boltzmann's gravestone. It is worth appreciating that the foundation of the subject was written as a concession.

To Zermelo, in 1896 and 1897 in the Annalen der Physik, Boltzmann conceded the theorem and disputed its relevance. Recurrence happens. How long do you wait? Take a cubic centimetre of air at room temperature: N = 2.7 x 10^19 molecules. The probability that all of them are in the left half at any instant is 2^-N, and with configurations shuffling on a collision time of about 10^-10 s, the expected waiting time in seconds is

10^-10 x 2^(2.7 x 10^19), whose base-ten logarithm is 2.7 x 10^19 x 0.30103 - 10 = 8.1 x 10^18.

The age of the universe is about 4 x 10^17 seconds, a number with 18 digits. The recurrence time is a number with eight quintillion digits. Zermelo's theorem is exactly true and describes nothing that will ever be observed.

Why this matters: the dispute was not settled by finding an error. It was settled by changing what the theorem was claiming. After 1877 the second law is not a mechanical identity, it is a statement about overwhelming probability, and both objections become true statements about events of measure effectively zero.

Where the asymmetry actually enters

A sharper question survives. If mechanics is time-symmetric, a derivation from mechanics cannot produce an asymmetric conclusion. So which step is not symmetric?

It is the Stosszahlansatz, and specifically its tense. Boltzmann assumed molecules are uncorrelated before they collide. After they collide, they are correlated: their velocities carry a record of the encounter. Applying the same assumption to the outgoing pair would give the opposite arrow. The time asymmetry is not in Newton's laws and not in the algebra; it is smuggled in with a hypothesis about which molecules are independent of each other, and that hypothesis is about the past.

Oscar Lanford proved in 1975 that the Boltzmann equation really does follow from Hamiltonian dynamics for a dilute hard-sphere gas, rigorously, with molecular chaos propagating from an uncorrelated initial condition. His result holds for times of order one fifth of a mean free time, and half a century later nobody has extended it much further. The rigorous derivation of irreversibility from mechanics currently covers rather less than one collision.

What the laboratory can do

Loschmidt's reversal is not purely hypothetical. In 1950 Erwin Hahn found that nuclear spins in a magnetic field, precessing at slightly different rates and so fanning out until the net magnetisation vanishes, can be brought back: a radio pulse that rotates every spin by 180 degrees makes the fast ones fall behind and the slow ones catch up, and the magnetisation reappears as a spin echo. The apparent approach to equilibrium was reversed by flipping the sign of the evolution, exactly as Loschmidt described. Rhim, Pines and Waugh extended it in 1971 to spins coupled to each other rather than merely dephasing independently.

The echo is never perfect, and it degrades further with each repetition. That decay is where the irreversibility genuinely lives: not in an impossibility of reversal, but in the extreme sensitivity of the reversed trajectory to anything left unreversed.

The other laboratory result Boltzmann would have wanted is the fluctuation theorem, which quantifies the ratio of the probabilities of entropy-producing and entropy-consuming trajectories over a finite time. In 2002 a group in Canberra dragged a micrometre-scale latex bead through water with an optical trap and measured that ratio directly, observing intervals of up to about two seconds during which the bead did negative work on its surroundings. The second law is violated routinely at that scale, in exactly the proportion the theory predicts.

What remains open

Two things, and they are not small.

The initial condition. The counting argument is symmetric in time. If the overwhelming majority of microstates compatible with today's macrostate evolve toward higher entropy in the future, the same argument says they came from higher entropy in the past, which is flatly false. Every glass of melting ice in the universe had a colder past. Nothing in statistical mechanics explains that; it has to be imposed, as the assumption that the universe began in an extraordinarily improbable low-entropy state. Boltzmann's own proposal was that our region is a giant fluctuation in an eternal equilibrium, which fails badly: a fluctuation producing one confused observer and nothing else is vastly more probable than one producing a galaxy, so the hypothesis predicts that your memories are fabricated.

The choice of description. H is built from the one-particle distribution f. Gibbs's fine-grained entropy, computed from the full N-particle distribution, is exactly constant under Hamiltonian evolution, by Liouville's theorem from Lesson 5. Entropy increases only once you agree to stop tracking correlations, and a different agreement gives a different entropy. Whether that makes entropy increase a fact about the world or a fact about our bookkeeping is still argued, in good faith, by people who understand the mathematics perfectly well.

What matters here: the disagreement that survives is not about any calculation. Everyone computes the same numbers. It is about whether the low-entropy past is a fact to be explained or a boundary condition to be assumed, and no experiment yet proposed distinguishes the positions.

Common misconceptions

  • "Loschmidt and Zermelo were refuted." Neither was. Both arguments are valid, and Boltzmann accepted both. What he denied was that they mattered, and the reply to Loschmidt is the origin of S = k log W.
  • "Entropy always increases." It increases with overwhelming probability. In a micrometre-scale system observed for a second, decreases are common and measurable, which is what the 2002 optical-trap experiment showed.
  • "The arrow of time comes from the expansion of the universe." Expansion is time-asymmetric but does not by itself supply a low-entropy start. A contracting universe would not run entropy backwards.
  • "The H-theorem is wrong." It is a correct theorem about the Boltzmann equation. The question is what the Boltzmann equation is a theorem about, and the answer involves an assumption regarding correlations before collisions, not after.

Try it: how many molecules does it take

Exercise. Estimate how long you would wait to see every molecule in the left half of a box, for N = 20, N = 100 and N = 2.7 x 10^19, taking one independent configuration per 10^-10 s.

Solution. For N = 20, 2^20 = 1.05 x 10^6, so the wait is 1.05 x 10^-4 s: about ten thousand times a second. For N = 100, 2^100 = 1.27 x 10^30 gives 1.3 x 10^20 s, roughly 4 x 1012 years, some three hundred times the age of the universe. For a cubic centimetre of air the exponent alone has nineteen digits. Between one hundred molecules and Avogadro's number, a possibility becomes an impossibility without any law changing.

Change one input. Stop demanding perfection. The number on the left has standard deviation sqrt(N)/2, so a fractional density excess d is d sqrt(N) standard deviations out. For a cubic centimetre of air, an excess of one percent is 5 x 10^7 standard deviations and will never be seen; an excess of one part in a billion is 5.2 standard deviations and happens constantly. Those tiny fluctuations scatter sunlight, and they are why the sky is blue rather than transparent.

Pulling it together

Boltzmann's 1872 H-theorem shows dH/dt <= 0 for the Boltzmann equation, with the Maxwell distribution as the unique stationary point. Loschmidt objected in 1876 that reversing every velocity produces a legal mechanical state along which H rises, and Zermelo objected in 1896 that Poincare recurrence forbids permanent decrease. Both are correct. Boltzmann answered the first by making the law statistical and writing S = k log W, and the second by computing the recurrence time for a cubic centimetre of air: a number of seconds whose logarithm is 8 x 10^18, against a universe some 4 x 10^17 seconds old. The asymmetry entered the derivation through the Stosszahlansatz, which assumes molecules are uncorrelated before colliding and not after, and Lanford's 1975 theorem makes that rigorous for about one fifth of a mean free time. Spin echoes perform Loschmidt's reversal in the laboratory, and the fluctuation theorem, verified with an optically trapped bead in 2002, measures how often entropy runs backwards in a small system. What is unresolved is why the initial condition had low entropy, and whether an entropy that depends on a chosen coarse-graining describes the world or our description of it.

That is where this course ends, and it is the honest place to end it. Everything in the preceding sixteen lessons, from counting the microstates of twenty coins to the pressure inside Sirius B, rests on a postulate about equal a priori probabilities and an assumption about which correlations to ignore. Both work superbly. Neither has been derived.

Sources

  1. Wikipedia contributors. (n.d.). H-theorem. Wikipedia. en.wikipedia.org
  2. Wikipedia contributors. (n.d.). Loschmidt's paradox. Wikipedia. en.wikipedia.org
  3. Uffink, J. (n.d.). Boltzmann's work in statistical physics. Stanford Encyclopedia of Philosophy. plato.stanford.edu
  4. Boltzmann, L. (1877). Ueber die Beziehung zwischen dem zweiten Hauptsatze der mechanischen Waermetheorie und der Wahrscheinlichkeitsrechnung. Sitzungsberichte der Kaiserlichen Akademie der Wissenschaften, 76, 373-435.
  5. Lanford, O. E. (1975). Time evolution of large classical systems. In Dynamical systems, theory and applications, Lecture Notes in Physics 38, 1-111. Springer.
Key terms
H-theorem
Boltzmann's 1872 result that H, the integral of f ln f, never increases under the Boltzmann equation.
Stosszahlansatz
The assumption that colliding molecules are uncorrelated before the collision; the one time-asymmetric step in the derivation.
Reversibility objection
Loschmidt's 1876 point that reversing all velocities gives a legal mechanical state along which H increases.
Recurrence objection
Zermelo's 1896 point that Poincare recurrence forbids any state function from decreasing forever.
Recurrence time
The waiting time for a bounded system to return near its initial state; for a cubic centimetre of air its logarithm is about 8 x 10^18.
Past hypothesis
The assumed low-entropy initial condition of the universe, needed because the counting argument is itself time-symmetric.
Coarse-graining
The choice of which correlations to stop tracking; Gibbs's fine-grained entropy is exactly constant under Hamiltonian flow.
Spin echo
Hahn's 1950 reversal of apparent dephasing by a 180 degree pulse, a laboratory version of Loschmidt's reversal.
Fluctuation theorem
The relation giving the ratio of probabilities of entropy-producing and entropy-consuming trajectories over a finite time.

Open the interactive version with quizzes and progress →