🪐 Astronomy · Undergraduate · ASTR 260

Stars, Galaxies & the Milky Way

Every claim in stellar astronomy traces back to a small set of measurable quantities: how far away a star is, how much energy it emits, how hot its surface is, and how much mass it contains. This course builds those measurements from the ground up and then spends them. You will compute a distance from a parallax angle, convert an apparent magnitude into a luminosity, read a spectrum for…

Start the interactive course (quizzes, progress, videos) →

Free forever. No sign-up, no ads. 16 lessons. The full lesson text is below so you can read it right here.

Module 1: Measuring Starlight

The four numbers everything else rests on: how far, how bright, how hot, and what a star is made of. Parallax and the parsec, the magnitude system and luminosity, and the spectra that turned points of light into objects with chemistry.

The First Rung: Parallax and the Parsec

  • Explain the geometry of trigonometric parallax and why the baseline is one astronomical unit.
  • Convert between parallax in arcseconds, distance in parsecs, light years and metres, and carry the arithmetic correctly.
  • Describe what Hipparcos and Gaia changed, and state where parallax stops being reliable and why.

Bessel, 61 Cygni, and three tenths of an arcsecond

In December 1838 Friedrich Bessel announced that the star 61 Cygni shifted back and forth against the background of fainter stars by about 0.31 arcseconds over the course of a year. He had spent eighteen months at Konigsberg measuring the star's position against two faint neighbours with a heliometer, an instrument whose objective lens was cut in half so the two halves could be slid past each other and the offset read off a scale. The shift he found is roughly the angular width of a small coin seen from sixteen kilometres away.

That number was the first direct measurement of the distance to any star other than the Sun. It gave 61 Cygni a distance of about 10.4 light years. The modern value, from the Gaia spacecraft, is 11.4 light years, so Bessel was low by about nine per cent using an instrument you could carry. Within three months Thomas Henderson published a parallax for Alpha Centauri from observations made years earlier at the Cape of Good Hope, and Friedrich Georg Wilhelm Struve published one for Vega. Two centuries of failure ended in about ninety days.

Everything in this course stands on that measurement, because distance is the quantity that converts an appearance into a property. A star that looks bright might be a nearby ember or a remote furnace. Until you know how far away it is, you know nothing about the star itself, only about the light arriving at your eye. So this lesson works the first rung of the distance ladder carefully, in numbers, before anything else.

Key idea: Parallax is the only stellar distance method that is pure geometry, with no assumption about what stars are like, which is why every other rung of the ladder is ultimately calibrated against it.

The geometry, done slowly

Hold a finger up at arm's length and look at it with one eye, then the other. The finger jumps against the far wall. It jumps because you observed it from two positions separated by the distance between your pupils, roughly 6.5 centimetres, and that separation is called the baseline. The angle of the jump depends on the ratio of the baseline to the distance of the finger. Move the finger closer and the jump grows.

Now scale it up. Earth's orbit gives you a baseline that changes continuously and returns to where it started every year. Observe a star in January and again in July and you have viewed it from two points separated by two astronomical units, which is 2 x 1.496 x 10^11 metres, or about 300 million kilometres. Nearby stars trace out a small ellipse on the sky over the year in response.

By convention the parallax angle, written p, is defined as half the total annual swing, which makes it the angle subtended by one astronomical unit, not two. That convention exists so the arithmetic comes out clean, and the definition of the parsec is what makes it clean.

For angles this small, the tangent of the angle equals the angle in radians to better than one part in a billion, so the exact relation d = 1 AU divided by tan(p) collapses to d = 1 AU divided by p, with p in radians. Convert to arcseconds by noting that one radian is 206,265 arcseconds, and you get

d (in AU) = 206,265 / p (in arcseconds)

Astronomers then defined a unit that absorbs the 206,265. One parsec is the distance at which one astronomical unit subtends one arcsecond, which is 206,265 AU, or 3.086 x 10^16 metres, or 3.262 light years. With that unit the relation becomes the cleanest formula in astronomy:

d (in parsecs) = 1 / p (in arcseconds)

The name is a contraction of parallax and arcsecond. It looks like jargon invented to keep outsiders out. It is really just a unit chosen to delete a conversion factor from a calculation astronomers were going to do millions of times.

Worked example: four stars, four distances

Gaia reports parallaxes in milliarcseconds, so the first move is always to convert. One milliarcsecond, written mas, is 0.001 arcseconds.

Proxima Centauri. Parallax 768.07 mas, which is 0.76807 arcseconds. d = 1 / 0.76807 = 1.302 parsecs. In light years that is 1.302 x 3.262 = 4.25 light years. This is the nearest star to the Sun, and it is the only star whose parallax exceeds one arcsecond.

Barnard's Star. Parallax 546.98 mas = 0.54698 arcseconds. d = 1 / 0.54698 = 1.828 parsecs = 5.96 light years. In metres, 1.828 x 3.086 x 10^16 = 5.64 x 10^16 m.

Sirius. Parallax 379.21 mas. d = 1 / 0.37921 = 2.637 parsecs = 8.60 light years. Keep this one; it comes back in the next lesson when we work out how luminous Sirius really is.

Vega. Parallax 130.23 mas. d = 1 / 0.13023 = 7.679 parsecs = 25.0 light years.

Notice how quickly the angles shrink. A star at 100 parsecs has a parallax of 10 mas. At 1,000 parsecs it is 1 mas, which is the angular width of that small coin at about 5,000 kilometres. And 1,000 parsecs is a short walk in a galaxy 30,000 parsecs across. The whole reason stellar distances were an unsolved problem until 1838 is that the angles are absurd.

The point: Parallax and distance are reciprocals, so the measurement gets harder in exact proportion to how far away the interesting object is.

Why it took two hundred and ninety-five years

Copernicus published in 1543 and immediately owed the world a parallax. If Earth moves, nearby stars must shift. No shift was seen, and Tycho Brahe treated that as a decisive argument against the heliocentric model, which it would have been if the stars were as close as anyone then imagined. The correct conclusion, which Tycho found unbelievable, was that the stars are enormously farther away than the planets.

Three things had to be solved before 1838. The first was instrumental: measuring an angle of a third of an arcsecond requires a stable mount, a fine micrometer or divided lens, and a way to compare a star with a neighbour rather than with an absolute coordinate frame. Bessel's heliometer did exactly that.

The second was choosing the right star. Bessel picked 61 Cygni because it has the largest proper motion known at the time, more than 5 arcseconds per year of steady drift across the sky. High proper motion is a good bet for proximity, since nearby stars appear to move faster for the same real velocity, exactly as nearby cars do. That was a genuine piece of scientific judgment rather than luck.

The third was disentangling parallax from two larger effects that mimic it. James Bradley, hunting parallax in the 1720s, found a 20.5 arcsecond annual wobble in the position of Gamma Draconis that was ninety degrees out of phase with what parallax should produce. It was the aberration of starlight, caused by Earth's orbital velocity of 30 km/s combined with the finite speed of light, and Bradley announced it in 1729. He later found nutation, an 18.6 year nodding of Earth's axis, and announced that in 1748. Two failed parallax hunts produced two major discoveries and, more importantly, produced the corrections without which no parallax measurement could be trusted.

From one star at a time to 1.5 billion

By 1900 a few dozen stellar parallaxes were known, most of them poor. Ground-based astrometry is limited by the atmosphere, which smears a stellar image into a blob typically one arcsecond across and moves it around unpredictably. You can beat that blob down with long exposures and careful statistics, but not by the factor of a thousand that real progress needed.

The answer was to leave the atmosphere. ESA launched Hipparcos in August 1989, and although a failed apogee boost motor left it in the wrong orbit, the mission was salvaged and its 1997 catalogue delivered positions and parallaxes for 118,218 stars at a precision near one milliarcsecond. That single catalogue rewrote the distance scale of the solar neighbourhood.

Gaia, launched in December 2013, did it again by two more orders of magnitude. Gaia spins slowly and sweeps two fields of view across the sky, measuring the changing angles between stars rather than absolute positions, which is the same trick Bessel used, industrialised. Data Release 3, published in June 2022, contains astrometry for about 1.47 billion sources, with parallax uncertainties around 0.02 to 0.03 mas for the brighter ones. The spacecraft stopped collecting science data in January 2025; further releases from the accumulated archive were still to come when this course was written, so treat any specific catalogue number you read as a snapshot with a date attached.

What does 0.03 mas buy you? The fractional error in distance equals the fractional error in parallax, so a ten per cent distance requires a parallax measured to ten per cent, meaning p must be at least about 0.3 mas. That is a distance of roughly 3,300 parsecs. Gaia therefore gives good geometric distances across a substantial fraction of the Milky Way's disk, which is why nearly everything in this course now has a better number attached to it than it did in 2010.

Why this matters: Gaia converted parallax from a technique that reached the nearest few hundred stars into one that maps a large part of the Galaxy, which recalibrated every method built on top of it.

Where parallax fails, and the trap in 1/p

Three things break parallax, and it is worth knowing all three because press releases rarely mention them.

Distance. Beyond a few thousand parsecs the angle sinks into the noise. Nothing about the method is wrong; there is simply no signal.

The star itself. Betelgeuse is a red supergiant whose visible surface boils in enormous convective cells, so its apparent centre of light wanders by an amount comparable to its parallax. Hipparcos gave 6.55 mas with an uncertainty of 0.83 mas, which is a distance somewhere between about 135 and 175 parsecs. Radio observations of the star's atmosphere, which is less lumpy than the optical surface, give a similar value with different systematics. The honest statement is that we know the distance to one of the most famous stars in the sky to about ten per cent, and that limits everything else we can say about it, including its luminosity and its mass.

The reciprocal. This one catches careful people. Measurement errors are roughly symmetric in parallax, but d = 1/p is a nonlinear transformation, so symmetric errors in p become asymmetric errors in d. Worse, because there are far more distant stars than nearby ones, a star whose true parallax is small is more likely to be scattered up by noise than a large-parallax star is to be scattered down. Simply inverting a noisy parallax therefore gives a distance that is systematically too small on average, an effect known as Lutz-Kelker bias. Modern catalogues address it by publishing distances estimated with an explicit prior on how stars are distributed in space rather than by inversion. For a parallax measured to better than about five per cent, 1/p is fine. For a twenty per cent parallax, it is not, and you should use a published distance estimate instead.

Common misconceptions

"Parallax is how we measure distances to galaxies." No. Parallax reaches a few thousand parsecs. The nearest large galaxy, Andromeda, is about 780,000 parsecs away, where the parallax would be about 0.0013 mas, far below anything measurable. Galaxy distances come from methods calibrated by parallax, not from parallax itself, and you will build that chain in Module 5.

"The parallax angle is the whole annual shift." It is half of it, by definition, so that d = 1/p works. If you measure the full swing of a star as 0.6 arcseconds, its parallax is 0.3 arcseconds and its distance is 3.3 parsecs.

"A star with a big parallax is a bright star." Parallax measures distance only. Proxima Centauri has the largest parallax of any star and is invisible to the unaided eye, because it is a small red dwarf radiating about 0.0017 times the Sun's power. Nearness and brightness are different questions, and separating them is the job of the next lesson.

"Modern parallaxes are exact." They are extraordinarily good and they still carry systematic offsets. Gaia's parallaxes required a global zero-point correction of order 0.017 mas, which matters enormously when you are measuring parallaxes of 0.2 mas and not at all when you are measuring 300 mas.

Where this leaves us

Parallax is the shift in a star's apparent position caused by Earth's motion around the Sun, and the parallax angle is defined as the shift produced by a one astronomical unit baseline. Distance in parsecs is the reciprocal of parallax in arcseconds, a relation that holds because the parsec was defined to make it hold. Bessel's 0.31 arcseconds for 61 Cygni in 1838 was the first success after nearly three centuries of failure, and the failures along the way produced aberration and nutation as consolation prizes. Hipparcos took the method to space and 118,218 stars; Gaia took it to roughly 1.5 billion and to precisions near 0.02 mas, which yields useful geometric distances out to a few thousand parsecs. Beyond that, and for stars whose surfaces are lumpy or whose parallaxes are noisy, the method degrades in specific and predictable ways, including the systematic bias that makes naive inversion of a poor parallax give a distance that is too small. Every other distance technique in this course is eventually traced back to this one.

Sources

  1. European Space Agency. (n.d.). Gaia: Mapping a billion stars. ESA Science and Exploration. esa.int
  2. European Space Agency. (n.d.). Gaia data release documentation. ESA Gaia Cosmos. cosmos.esa.int
  3. Encyclopaedia Britannica. (n.d.). Parallax. britannica.com
  4. Fraknoi, A., Morrison, D., & Wolff, S. C. (2022). Astronomy 2e, Chapter 19: Celestial distances. OpenStax, Rice University. openstax.org
  5. Hirshfeld, A. W. (2001). Parallax: The race to measure the cosmos. W. H. Freeman. (A book-length account of Bradley, Bessel, Henderson and Struve; no open edition.)
Key terms
Trigonometric parallax
The apparent shift of a nearby star against distant background stars caused by Earth's orbital motion, measured as the angle subtended by one astronomical unit.
Baseline
The separation between the two observing positions used in a parallax measurement; for stellar parallax it is Earth's orbital radius.
Arcsecond
One 3,600th of a degree; one radian equals 206,265 arcseconds.
Parsec
The distance at which one astronomical unit subtends one arcsecond, equal to 206,265 AU, 3.086 x 10^16 m, or 3.262 light years.
Proper motion
A star's steady angular drift across the sky due to its real motion through space, distinct from the annual parallax ellipse.
Aberration of starlight
The apparent displacement of a star caused by the combination of Earth's orbital velocity and the finite speed of light, discovered by Bradley in 1729.
Milliarcsecond
One thousandth of an arcsecond, the working unit of modern space astrometry.
Lutz-Kelker bias
The systematic underestimate of distance that results from inverting a noisy parallax, because more stars are scattered inward from large distances than outward from small ones.

Brightness Twice Over: Magnitudes, Flux, and Luminosity

  • Convert between magnitude differences and flux ratios, and explain why the scale is logarithmic and inverted.
  • Apply the inverse square law and the distance modulus to move between apparent magnitude, absolute magnitude and luminosity.
  • Explain what a bolometric correction fixes, and identify when ignoring it produces a badly wrong luminosity.

Betelgeuse loses two thirds of its light

Through the winter of 2019 and into February 2020, Betelgeuse faded. It normally sits near visual magnitude 0.5, bright enough that it is one of the ten or so brightest stars in the sky. By the middle of February 2020 observers were measuring it near magnitude 1.6, the faintest in more than a century of records, and it was visibly demoted in Orion's shoulder to anyone who looked up.

A drop of 1.1 magnitudes sounds mild. Work out what it means physically and it is not. The magnitude scale is logarithmic, and the flux ratio corresponding to a magnitude difference is 10 raised to the power 0.4 times that difference:

flux ratio = 10^(0.4 x 1.1) = 10^0.44 = 2.75

Betelgeuse was delivering less than 37 per cent of its usual light. Roughly two thirds of it had gone. That is the sort of thing that ought to be described in a physical unit, and this lesson is about the units, because almost every mistake beginners make in stellar astronomy is a unit mistake dressed up as a conceptual one.

Three quantities get confused constantly, so fix them now. Flux is energy arriving per second per square metre at your detector, measured in watts per square metre. Luminosity is energy leaving the star per second, in watts, and it belongs to the star alone. Magnitude is a logarithmic bookkeeping system for flux that astronomers inherited from antiquity and have never managed to abandon.

Remember: Flux is about you and your distance from the star; luminosity is about the star. Converting one into the other is the whole job, and it requires a distance, which is why the previous lesson came first.

Why the scale runs backwards, and why 2.512

Around 129 BC Hipparchus catalogued stars into six classes, calling the brightest ones of the first magnitude and the faintest visible ones of the sixth. Nothing about that is a physical scale. It is a ranking, and it runs the wrong way because it counts rank, not brightness.

In 1856 Norman Pogson noticed something useful about the old classes. A first magnitude star delivered roughly one hundred times the flux of a sixth magnitude star, which meant five magnitude steps corresponded to a factor of about one hundred. So he defined the step exactly: five magnitudes equals a factor of exactly 100 in flux. One magnitude is therefore the fifth root of 100, which is

100^(1/5) = 2.512

and the two working formulas follow at once. To turn a magnitude difference into a flux ratio:

F1 / F2 = 10^(0.4 x (m2 - m1))

and to go the other way:

m2 - m1 = 2.5 x log10(F1 / F2)

Note the order of the subscripts. The star with the larger magnitude number is the fainter one. Brighter objects run to negative numbers: Sirius is at -1.46, Venus at its best near -4.9, the full Moon near -12.7, and the Sun at -26.74.

Try the extreme. The Sun and Sirius differ by 25.28 magnitudes, so the flux ratio is 10^(0.4 x 25.28) = 10^10.11, about 1.3 x 10^10. The Sun delivers roughly thirteen billion times as much light to your eye as the brightest star in the night sky. That is precisely the compression the logarithmic scale exists to manage, and it is why astronomers keep a system that otherwise reads like a practical joke.

Flux, the inverse square law, and a luminosity you can check

A star radiates in all directions. At a distance d, all of that power is spread over the surface of a sphere of area 4 x pi x d^2, so

F = L / (4 x pi x d^2)

Nothing in that equation is subtle, and everything in observational stellar astronomy runs through it. Here is a case where you can check the answer, because both sides are measured independently.

Satellites measure the solar flux at Earth's mean distance at about 1,361 watts per square metre. Earth's mean distance is 1 AU = 1.496 x 10^11 m. So

L = 4 x pi x d^2 x F = 4 x 3.1416 x (1.496 x 10^11)^2 x 1,361

(1.496 x 10^11)^2 = 2.238 x 10^22 m^2

4 x pi x 2.238 x 10^22 = 2.812 x 10^23 m^2

L = 2.812 x 10^23 x 1,361 = 3.83 x 10^26 watts

The accepted solar luminosity is 3.828 x 10^26 W. Two measurements you could make from a rooftop and a planetary ephemeris give the total power output of a star, and that number is the yardstick for every other star in this course. One solar luminosity, written L_sun, is 3.828 x 10^26 W.

The inverse square law also carries a warning. It assumes nothing absorbs the light on the way. Interstellar dust does absorb and scatter starlight, dimming and reddening it, which is called extinction. In the plane of the Milky Way extinction can exceed one magnitude per kiloparsec, so a distant star in the disk looks farther away than it is if you forget to correct. Module 3 comes back to the dust; for now, hold the caveat.

The upshot: Flux falls as the inverse square of distance, so measuring flux tells you about the star only after you have supplied a distance and corrected for anything absorbing along the path.

Absolute magnitude: putting every star at the same distance

Comparing apparent magnitudes is comparing appearances. To compare stars, put them all at a standard distance and ask how bright they would then look. The convention is 10 parsecs, and the result is the absolute magnitude, written M.

Combine the inverse square law with Pogson's formula and you get the distance modulus:

m - M = 5 x log10(d) - 5, with d in parsecs

The quantity m - M is itself called the distance modulus, and once you can read it you can read half of observational astronomy. It is zero at 10 parsecs, negative for closer stars, and positive for anything farther.

Worked example: Sirius. Apparent magnitude m = -1.46, distance 2.637 pc from the previous lesson.

log10(2.637) = 0.4212

m - M = 5 x 0.4212 - 5 = 2.106 - 5 = -2.894

M = m - (m - M) = -1.46 + 2.894 = +1.43

Now convert to luminosity. The Sun's absolute visual magnitude is +4.83, so

L / L_sun = 10^(0.4 x (4.83 - 1.43)) = 10^(0.4 x 3.40) = 10^1.36 = 22.9

Sirius radiates about 23 times the Sun's power in visible light. It is the brightest star in the night sky partly because it is genuinely luminous and largely because it is close.

Worked example in reverse: an unknown star. Suppose you identify a star as a type whose absolute magnitude you know to be M = -3.5, and you measure m = +11.5. Then m - M = 15.0, so 15.0 = 5 log10(d) - 5, giving log10(d) = 4.0 and d = 10,000 parsecs. That inversion is the engine of the entire cosmic distance ladder: find an object whose absolute magnitude you can determine some other way, measure how bright it looks, and read off the distance. Cepheid variables in Module 5 are exactly this trick with a very good M.

Where the visual magnitude lies to you

Return to Betelgeuse and try the same arithmetic. Take m = 0.5 at its usual brightness and a distance of about 168 parsecs.

log10(168) = 2.225, so m - M = 5 x 2.225 - 5 = 6.13, and M = 0.5 - 6.13 = -5.63

L / L_sun = 10^(0.4 x (4.83 + 5.63)) = 10^(0.4 x 10.46) = 10^4.18 = 15,200

Fifteen thousand solar luminosities. But every reference you consult will tell you Betelgeuse radiates somewhere around 100,000 times the Sun's power. The calculation is not wrong; the question was.

Visual magnitudes measure flux through a filter centred near 550 nanometres, roughly where the human eye peaks. Betelgeuse has a surface temperature near 3,600 K and radiates most of its energy in the infrared, where a V filter simply does not look. A star that hot radiates its peak near 800 nanometres, well outside the visible band.

The fix is the bolometric magnitude, M_bol, which accounts for radiation at all wavelengths, and the bolometric correction, BC, which is the number you add to convert:

M_bol = M_V + BC

BC is defined to be near zero for stars like the Sun, whose output peaks in the visible, and becomes strongly negative for both very cool stars, which dump their energy into the infrared, and very hot stars, which dump it into the ultraviolet. For a red supergiant like Betelgeuse, BC is about -1.9. So

M_bol = -5.63 - 1.9 = -7.5

and using the IAU's adopted solar bolometric magnitude of +4.74,

L / L_sun = 10^(0.4 x (4.74 + 7.5)) = 10^(0.4 x 12.24) = 10^4.90 = 79,000

Now the answer lands in the right neighbourhood. The remaining spread in published values, roughly 90,000 to 150,000 solar luminosities, comes almost entirely from the ten per cent uncertainty in the distance you met in the last lesson, since luminosity scales as distance squared.

What matters here: A visual magnitude is a measurement through a specific window. For any star much hotter or much cooler than the Sun, the visual band misses most of the energy, and a luminosity computed without a bolometric correction can be wrong by a factor of five or more.

Colours, filters, and one number that carries a temperature

Since a magnitude depends on the filter, astronomers standardised the filters. The Johnson-Morgan UBV system defines U in the near ultraviolet near 365 nm, B in the blue near 445 nm, and V in the visual near 551 nm, later extended into the red and infrared with R, I, J, H and K bands.

Measure a star in two bands and subtract, and you have a colour index. The most common is B - V. Because a hot star emits proportionally more blue light, a hot star has a small or negative B - V, while a cool star has a large one. The Sun sits at B - V = 0.65. Vega, at about 9,600 K, is defined near 0.00 by historical convention. Betelgeuse is around 1.85.

StarApproximate B - VApproximate surface temperatureApparent visual magnitude
Rigel-0.0312,000 K0.13
Vega0.009,600 K0.03
Sun0.655,772 K-26.74
Arcturus1.234,300 K-0.05
Betelgeuse1.853,600 K0.5 (variable)

Two magnitudes, one subtraction, and you have a temperature estimate without ever taking a spectrum. That is why colours are the workhorse of large surveys, which measure billions of stars in a handful of filters. The catch, again, is dust: interstellar extinction removes blue light preferentially, so a reddened star masquerades as a cooler one. Correcting for that is a permanent chore in Galactic astronomy.

The range you are working in

It helps to know the ends of the scale you are using. The unaided eye under a genuinely dark sky reaches about magnitude 6. A pair of binoculars gets you to about 10. Gaia's survey is complete to about G = 20.7. The deepest Hubble exposures reach objects near magnitude 30, which is about 250,000 times fainter than the naked-eye limit in flux terms, since a 24 magnitude difference is 10^(0.4 x 24) = 10^9.6, or about four billion. Work that out yourself and you will see the compression clearly: nine and a half orders of magnitude in flux fit into 24 tick marks on a scale invented by someone sorting stars by eye in the second century BC.

Common misconceptions

"A magnitude 1 star is twice as bright as a magnitude 2 star." It is about 2.512 times as bright, not twice, and the factor is exact only in the sense that five steps make exactly one hundred.

"The brightest stars in the sky are the most luminous stars." Of the twenty brightest stars, several are ordinary nearby stars. Sirius is 23 solar luminosities at 2.6 parsecs; Deneb is tens of thousands of solar luminosities at hundreds of parsecs. Apparent brightness mixes distance and luminosity irreversibly until you separate them.

"Absolute magnitude is the star's real brightness." It is the apparent magnitude the star would have at 10 parsecs, in a particular filter. It is still a magnitude, still filter-dependent, and still needs a bolometric correction before it becomes a total energy output.

"Magnitude is an outdated system we should drop." It survives because it is well matched to the way detectors and errors behave. Photometric uncertainties are naturally multiplicative, and a logarithmic scale turns multiplicative errors into additive ones, which is exactly what you want when combining thousands of measurements.

Putting it together

Magnitude is a logarithmic flux scale, running backwards because it inherited Hipparchus's ranking, with five magnitudes defined as exactly a factor of one hundred and one magnitude therefore a factor of 2.512. Flux relates to luminosity through the inverse square law, which turns a rooftop measurement of 1,361 watts per square metre into the Sun's output of 3.828 x 10^26 watts. Absolute magnitude standardises every star to 10 parsecs, and the distance modulus m - M = 5 log10(d) - 5 moves you between apparent brightness, absolute brightness and distance, in whichever direction you have two of the three. Visual magnitudes see only a narrow band, so bolometric corrections are mandatory for anything much hotter or cooler than the Sun; forgetting one turns Betelgeuse's 100,000 solar luminosities into 15,000. Colour indices such as B - V give a cheap temperature, provided you have corrected for the interstellar dust that reddens everything in the Galactic plane. With distance and luminosity in hand, the next lesson takes the spectrum apart and extracts temperature, composition and surface gravity.

Sources

  1. NASA Science. (n.d.). Stars. National Aeronautics and Space Administration. science.nasa.gov
  2. Encyclopaedia Britannica. (n.d.). Magnitude. britannica.com
  3. Fraknoi, A., Morrison, D., & Wolff, S. C. (2022). Astronomy 2e, Chapter 17: Analyzing starlight. OpenStax, Rice University. openstax.org
  4. European Southern Observatory. (2021, 16 June). Mystery of Betelgeuse's dip in brightness solved. ESO. eso.org
  5. Carroll, B. W., & Ostlie, D. A. (2017). An introduction to modern astrophysics (2nd ed.). Cambridge University Press. (Chapter 3 derives the magnitude system and bolometric corrections in full.)
Key terms
Flux
Radiant energy arriving per second per square metre at the observer, measured in watts per square metre.
Luminosity
Total radiant energy leaving a star per second, in watts; one solar luminosity is 3.828 x 10^26 W.
Apparent magnitude
A logarithmic measure of the flux received from an object, with smaller and more negative numbers meaning brighter.
Absolute magnitude
The apparent magnitude an object would have at a standard distance of 10 parsecs.
Distance modulus
The quantity m - M, equal to 5 log10(d) - 5 with d in parsecs, used to convert between apparent brightness and distance.
Bolometric magnitude
A magnitude accounting for radiation at all wavelengths rather than through a single filter.
Bolometric correction
The offset added to a visual magnitude to obtain a bolometric one; strongly negative for very hot and very cool stars.
Colour index
The difference between magnitudes in two filters, such as B - V, which encodes surface temperature.
Interstellar extinction
Dimming and reddening of starlight by intervening dust, which must be corrected before distances or temperatures are trusted.

Reading a Spectrum: Temperature, Composition, and the Harvard Sequence

  • Apply Wien's law and the Stefan-Boltzmann law to obtain a star's temperature and radius from its spectrum and luminosity.
  • Explain why absorption line strengths peak at particular temperatures rather than tracking abundance directly.
  • Order the OBAFGKM sequence by temperature and interpret an MK type such as G2V or M2Iab.

Fraunhofer counts 574 dark lines

Joseph von Fraunhofer was an optician, not an astronomer. Around 1814, in Munich, he was testing glass for telescope lenses and needed a pure colour to measure how strongly each type of glass bent light. Passing sunlight through a slit and a prism, he found the spectrum was interrupted by hundreds of narrow dark gaps. He mapped 574 of them, labelled the strongest with the letters A through K, and used their fixed positions as the precise wavelength standards he had been looking for.

He also noticed something he could not explain. The pair of dark lines he called D sat at exactly the wavelength of the bright yellow light a sodium flame gives off in a laboratory. Forty-five years later Gustav Kirchhoff and Robert Bunsen worked out why, and in doing so turned every star in the sky into a chemistry sample. A cool gas absorbs precisely the wavelengths that the same gas, when hot, emits.

That is the whole basis of what follows. A stellar spectrum is a continuous rainbow with a pattern of dark lines cut into it, and the two parts carry different information. The shape of the continuum gives you temperature and, with a distance, radius. The lines give you composition, pressure, velocity, rotation and magnetic field. This lesson takes both apart.

Key idea: The continuum tells you how hot the star is; the lines tell you what it is and what it is doing. Confusing the two is the most common error in reading a spectrum.

Kirchhoff's three rules, and why a star shows dark lines at all

Kirchhoff summarised the laboratory behaviour in three statements. A hot dense object produces a continuous spectrum. A hot thin gas produces bright emission lines at wavelengths characteristic of its atoms. A cool thin gas in front of a continuous source produces dark absorption lines at those same wavelengths.

A star gives you the third case because of the temperature gradient in its atmosphere. Light emerges from a dense, hot layer, the photosphere, and then passes outward through progressively cooler, thinner gas. At the wavelength of a strong atomic transition the gas is opaque, so the photons you finally see at that wavelength escaped from higher up, where it is cooler and therefore dimmer. The line looks dark not because the light was destroyed but because it came from a cooler layer. This is why the Sun's spectrum shows the same lines whether you look at the centre of the disk or the limb, with different strengths.

The continuum: two laws and a radius you cannot measure directly

To a good approximation a star radiates as a blackbody, and two laws follow.

Wien's displacement law gives the wavelength of peak emission:

lambda_max = 2.898 x 10^-3 / T, with lambda in metres and T in kelvin

For the Sun, T = 5,772 K, so lambda_max = 2.898 x 10^-3 / 5,772 = 5.02 x 10^-7 m = 502 nanometres, which is green. For Betelgeuse at 3,600 K it is 805 nm, in the near infrared. For Rigel at 12,000 K it is 242 nm, in the ultraviolet. Two of those three stars radiate their peak where the human eye cannot see at all, which is exactly the bolometric correction problem from the previous lesson, stated in a different currency.

The Stefan-Boltzmann law gives the power radiated per square metre of surface:

F_surface = sigma x T^4, with sigma = 5.670 x 10^-8 W m^-2 K^-4

Multiply by the surface area of a sphere and you get the luminosity:

L = 4 x pi x R^2 x sigma x T^4

This is the single most useful equation in stellar astronomy, because it links three quantities of which you can usually measure two. Rearranged into solar units it becomes

R / R_sun = (L / L_sun)^0.5 x (T_sun / T)^2

Worked example: how big is Betelgeuse? Nobody can measure a stellar radius with a ruler, and only a handful of stars have been resolved as disks at all. But we have L and T. Take L = 100,000 L_sun and T = 3,600 K, with T_sun = 5,772 K.

(100,000)^0.5 = 316

(5,772 / 3,600)^2 = (1.603)^2 = 2.57

R / R_sun = 316 x 2.57 = 813

About 800 solar radii. The Sun's radius is 6.96 x 10^8 m, so Betelgeuse is roughly 5.7 x 10^11 m across in radius, which is about 3.8 astronomical units. Put it where the Sun is and its surface would swallow the orbit of Mars and reach most of the way to the asteroid belt. Notice how the answer was assembled: a parallax gave a distance, the distance and apparent magnitude gave a luminosity, a bolometric correction fixed the luminosity, a spectrum gave a temperature, and Stefan-Boltzmann converted the pair into a size. Nothing was measured directly except angles and fluxes.

So what?: Luminosity, temperature and radius are locked together by L = 4 pi R^2 sigma T^4, so any two of them determine the third, and that is how stellar sizes are known at all.

Why hydrogen lines are strongest in stars that are not mostly special

Here is the trap that took astronomy fifty years to escape. The Balmer lines of hydrogen, the ones in the visible spectrum, are strongest in A-type stars at about 9,500 K. They are weak in the Sun and almost absent in the hottest O stars. The obvious reading is that A stars are rich in hydrogen and O stars are not. The obvious reading is wrong.

Balmer absorption requires a hydrogen atom already in its first excited state, with its electron in the n = 2 level, ready to be kicked up to n = 3 or higher. Two competing effects control how many atoms are in that state.

Raise the temperature and collisions excite more atoms out of the ground state into n = 2, so Balmer absorption strengthens. Raise it further and collisions start stripping the electron off entirely. An ionised hydrogen atom is a bare proton with no electron to absorb anything, so the lines vanish. The two effects peak against each other at about 9,500 K, which is exactly where A stars sit.

Meghnad Saha wrote down the equation governing the ionisation balance in 1920, and it, together with the Boltzmann distribution over excitation states, converts a measured line strength into an abundance rather than leaving it as a raw appearance. That machinery is why a modern spectrum yields a number for how much iron a star contains relative to the Sun, written as the metallicity [Fe/H], where zero means solar and -2 means one hundredth of solar iron.

In short: Line strength depends on temperature at least as much as on abundance, and untangling the two requires the physics of excitation and ionisation, not eyeballing the spectrum.

OBAFGKM: a sequence that was alphabetical before it was physical

The Harvard College Observatory took on the classification of stellar spectra in the 1880s under Edward Pickering, funded by a bequest from the widow of Henry Draper. Williamina Fleming sorted spectra into classes A through Q, ordered by the strength of the hydrogen lines, which seemed the natural criterion. Antonia Maury reorganised the sequence and added a second parameter describing line width, which Pickering disliked and which turned out to be the seed of the luminosity classification.

Annie Jump Cannon then made the decision that stuck. She dropped most of the letters, kept a few, reordered the survivors, and produced the sequence O, B, A, F, G, K, M, subdivided by digits from 0 to 9. It looks arbitrary because it is a historical residue of an alphabetical scheme, but the order is now known to be a temperature order running from hottest to coolest. Cannon classified the spectra of about 225,300 stars for the Henry Draper Catalogue, published in nine volumes between 1918 and 1924, and several hundred thousand in total across her career. Generations of students have memorised the order with the mnemonic Oh Be A Fine Girl, Kiss Me, which is a period piece in itself.

ClassSurface temperatureColourDominant spectral featuresExampleShare of main-sequence stars
OAbove 30,000 KBlueIonised helium; weak hydrogenZeta OphiuchiFar under 0.01 per cent
B10,000 to 30,000 KBlue whiteNeutral helium; hydrogen strengtheningRigel, SpicaAbout 0.1 per cent
A7,500 to 10,000 KWhiteHydrogen Balmer lines at maximumSirius A, VegaAbout 0.6 per cent
F6,000 to 7,500 KYellow whiteHydrogen weakening; ionised metals appearingProcyon AAbout 3 per cent
G5,200 to 6,000 KYellowIonised calcium H and K dominant; many metalsThe Sun, Alpha Centauri AAbout 8 per cent
K3,700 to 5,200 KOrangeNeutral metals strong; molecular bands beginningArcturus, Epsilon EridaniAbout 12 per cent
M2,400 to 3,700 KRedTitanium oxide bands dominateProxima Centauri, BetelgeuseRoughly three quarters

Read the last column twice. About three quarters of the stars in the solar neighbourhood are M dwarfs, and not one of them is visible to the unaided eye. The stars you can see at night are an extreme and unrepresentative sample, biased toward the rare luminous classes because those are the ones bright enough to notice from far away. Below M the sequence continues into L, T and Y for brown dwarfs and cool substellar objects, added from the late 1990s onward as infrared surveys found them.

Cecilia Payne and the composition of everything

In 1925 Cecilia Payne submitted a doctoral thesis at Radcliffe College applying Saha's ionisation theory to stellar spectra properly, for the first time on a large scale. Correcting line strengths for temperature and ionisation, she concluded that hydrogen and helium are overwhelmingly the most abundant elements in stars, with hydrogen roughly a million times more abundant than the metals by number.

This contradicted the settled view, which held that the Sun had broadly the same composition as Earth. Henry Norris Russell, the leading American astronomer of the day and the referee on her work, told her the result could not be right. She published, but inserted a sentence saying the hydrogen and helium abundances were almost certainly not real. Four years later Russell reached the same conclusion by an independent route and published it, and the result is often credited to him.

Two things are worth taking from this beyond the injustice. First, the correction that made the difference was a temperature correction, exactly the physics of the previous section: the metals in the Sun looked abundant because their lines are strong at 5,800 K, and the hydrogen looked scarce because most of it is in the ground state and invisible in the visible band. Second, that composition is the reason the rest of this course works the way it does. Stars are hydrogen with a helium contaminant and a trace of everything else, and the story of stellar evolution is the story of what happens to that hydrogen.

The second dimension: luminosity classes

Two stars can share a spectral class and differ by a factor of ten thousand in luminosity. Betelgeuse and Proxima Centauri are both M stars. The spectra are not identical, and Antonia Maury had spotted the difference decades before anyone systematised it: the lines in luminous stars are narrower.

The reason is pressure. A supergiant's atmosphere is enormously extended and therefore extremely low in density, so atoms are rarely disturbed by neighbours as they absorb. In a dwarf's compact, high gravity atmosphere, frequent collisions perturb the energy levels and smear each line out, an effect called pressure broadening. Measure the line widths and you measure surface gravity, and surface gravity separates dwarfs from giants from supergiants.

William Morgan, Philip Keenan and Edith Kellman formalised this in 1943 as the MK system, adding a Roman numeral to Cannon's letter and digit:

ClassNameTypical example
Ia, IbLuminous and less luminous supergiantsRigel (B8Ia), Betelgeuse (M1-2Iab)
IIBright giantsSargas
IIIGiantsArcturus (K1.5III), Aldebaran (K5III)
IVSubgiantsProcyon A (F5IV-V)
VMain sequence, historically called dwarfsThe Sun (G2V), Vega (A0V)

So the Sun's full type is G2V: a G star two tenths of the way toward K, on the main sequence. That short string encodes a temperature of about 5,772 K, a surface gravity typical of a main-sequence star, and by implication a mass, a radius and a luminosity. Getting a spectral type is the cheapest way to learn almost everything about a star, which is why surveys collect millions of them.

What else the lines carry

Four more measurements come free with a spectrum. A uniform wavelength shift of every line gives the radial velocity through the Doppler effect, which is how binary orbits and Galactic rotation are measured. Symmetric broadening beyond what pressure explains gives the projected rotation speed, since one limb approaches while the other recedes. Splitting of certain lines into components gives the magnetic field strength through the Zeeman effect, which is how solar magnetism was first measured, by George Ellery Hale in 1908. And the relative strengths of metal lines, once the temperature is known, give metallicity, which turns out to be a rough age and population indicator, as Module 5 will use.

Common misconceptions

"Strong hydrogen lines mean a hydrogen-rich star." All main-sequence stars are hydrogen-rich. Balmer line strength peaks near 9,500 K because that is where the population of atoms in the n = 2 state is largest, and it falls off on both sides for reasons that have nothing to do with abundance.

"OBAFGKM is a random order somebody should fix." It is a historical residue, and renaming it now would orphan a century and a half of literature. The order itself is physically meaningful: it is monotonic in temperature.

"Most stars are like the Sun." The Sun is in the top ten per cent or so by mass and luminosity. About three quarters of stars are M dwarfs, cooler, smaller and vastly fainter, and none is visible without a telescope. Every naked-eye impression of the stellar population is distorted by this selection effect.

"A star's colour tells you its composition." Colour tells you temperature. Composition comes from line strengths interpreted through ionisation and excitation physics, and the differences in composition among ordinary stars are small, at the level of the trace elements rather than the bulk.

What to carry forward

A stellar spectrum splits into a continuum and a line pattern. The continuum obeys Wien's law, giving a peak wavelength of 2.898 x 10^-3 divided by the temperature, and the Stefan-Boltzmann law, whose rearrangement lets you compute a radius from luminosity and temperature, which is how Betelgeuse is known to be roughly 800 solar radii. Absorption lines arise because a cooler outer atmosphere sits above a hotter photosphere, and their strengths depend on temperature through excitation and ionisation, so hydrogen Balmer lines peak in A stars at about 9,500 K for reasons unrelated to abundance. The Harvard sequence OBAFGKM, assembled by Fleming, Maury and Cannon and applied to more than 225,000 stars by Cannon alone, is an alphabetical accident that turned out to be a temperature order, and roughly three quarters of all stars fall in its last, invisible class. Cecilia Payne's 1925 thesis showed that stars are overwhelmingly hydrogen and helium, which is the premise the rest of the course runs on. The MK luminosity class added surface gravity as a second dimension through pressure broadening, and with temperature and luminosity in hand you are ready for the diagram that organises them both.

Sources

  1. NASA Science. (n.d.). Stars. National Aeronautics and Space Administration. science.nasa.gov
  2. Encyclopaedia Britannica. (n.d.). Stellar classification. britannica.com
  3. Fraknoi, A., Morrison, D., & Wolff, S. C. (2022). Astronomy 2e, Chapter 17: Analyzing starlight. OpenStax, Rice University. openstax.org
  4. NOIRLab. (n.d.). Education and public outreach resources. NSF NOIRLab. noirlab.edu
  5. Payne, C. H. (1925). Stellar atmospheres: A contribution to the observational study of high temperature in the reversing layers of stars. Harvard College Observatory Monograph No. 1. (Doctoral thesis, Radcliffe College.)
  6. Morgan, W. W., Keenan, P. C., & Kellman, E. (1943). An atlas of stellar spectra, with an outline of spectral classification. University of Chicago Press.
Key terms
Photosphere
The layer of a star from which visible light escapes, and the surface whose temperature spectral types describe.
Wien's displacement law
The relation lambda_max = 2.898 x 10^-3 / T, giving the wavelength at which a blackbody radiates most strongly.
Stefan-Boltzmann law
L = 4 pi R^2 sigma T^4, linking a star's luminosity, radius and surface temperature.
Saha equation
The 1920 relation describing how the ionisation state of a gas depends on temperature and pressure, essential for converting line strengths into abundances.
Harvard spectral sequence
The classification O, B, A, F, G, K, M, subdivided 0 to 9, ordered from hottest to coolest.
Luminosity class
The Roman numeral in an MK type, from Ia supergiants to V main-sequence stars, inferred from pressure broadening of spectral lines.
Pressure broadening
Widening of spectral lines caused by collisions in a dense, high gravity atmosphere, which distinguishes dwarfs from giants.
Metallicity
The abundance of elements heavier than helium relative to the Sun, usually written [Fe/H], with 0 meaning solar.

Module 2: Structure and Energy

The diagram that organises every star, the mechanical balance that keeps one from collapsing or exploding, and the nuclear reactions that pay the energy bill for ten billion years.

The Hertzsprung-Russell Diagram

  • Construct and orient an HR diagram, including the reason the temperature axis runs backwards.
  • Identify the main sequence, giant branch, supergiants and white dwarfs, and explain what physically distinguishes each region.
  • Infer a star's mass, radius and approximate remaining lifetime from its position on the diagram.

Hertzsprung plots the Hyades

In 1911 Ejnar Hertzsprung, a Danish chemical engineer who had taught himself astronomy, did something that looks obvious now and was not. He took two star clusters, the Pleiades and the Hyades, and plotted each star's colour against its apparent magnitude. Because every star in a cluster sits at essentially the same distance, apparent magnitude works as a stand-in for luminosity within that cluster, and the distance problem cancels out.

The points did not scatter. They fell along a narrow band. Two years later Henry Norris Russell, working with the parallax distances then available for field stars, presented a version at a meeting of the Royal Astronomical Society and published it in 1914, plotting spectral class against absolute magnitude. His plot showed the same band, plus a scatter of luminous stars sitting well above it and a few faint ones far below.

That plot is the most useful diagram in astronomy. It is not a picture of anything, it is a scatter graph, and it organises stellar physics the way the periodic table organises chemistry. This lesson builds it from the quantities you spent Module 1 measuring, and then reads it.

Why this matters: Once you can place a star on the HR diagram, you can usually name its mass, its radius, what it is burning, and roughly how long it has left, from two measured numbers.

The axes, and their inherited perversities

The vertical axis is luminosity, usually plotted logarithmically in solar units, running from about 10^-4 L_sun at the bottom to about 10^6 at the top. Ten orders of magnitude. Equivalently it is absolute magnitude, running from about +16 at the bottom to -10 at the top, and since magnitudes already run backwards, that axis is plotted with the most negative value at the top so that brighter is still up.

The horizontal axis is surface temperature, or a stand-in for it: spectral class or a colour index such as B - V. And it runs backwards. Hot stars are on the left, cool stars on the right, 40,000 K at one end and 2,500 K at the other. This is not a convention anyone would choose today. It survives because Russell plotted spectral class in the Harvard order O, B, A, F, G, K, M from left to right, and that order happens to run hot to cool. Every HR diagram published since has kept it so that older diagrams remain readable.

When a diagram uses colour and apparent magnitude instead of temperature and luminosity, which is what you do for a cluster, it is called a colour-magnitude diagram. It is the same object with cheaper axes.

Two ways to fill the diagram, and only one of them is honest

Plot the brightest stars in the sky and you get a diagram dominated by luminous giants and supergiants, with the main sequence represented only at its upper end. Plot every star within, say, 10 parsecs of the Sun, and you get something entirely different: a dense clump of faint red points in the lower right, the Sun sitting unusually high, and almost nothing else.

Both diagrams are correct and they describe different samples. The first is magnitude limited and therefore biased toward intrinsically luminous stars, which can be seen from far away and so sweep up a huge volume. The second is volume limited and tells you what the stellar population actually is. Of the roughly sixty known stars within about five parsecs of the Sun, the overwhelming majority are M dwarfs; there is one A star, Sirius; and there are several white dwarfs. The Sun is in the top few per cent by luminosity of its own neighbourhood.

This distinction is not pedantry. Almost every wrong intuition in stellar astronomy comes from reasoning about a magnitude-limited sample as though it were a fair one. When you read that some class of star is common or rare, the first question is always which sample the claim came from.

The main sequence is a mass sequence

About ninety per cent of stars at any moment lie on the diagonal band running from hot and luminous at the upper left to cool and faint at the lower right. That band is the main sequence, and a star sits on it for as long as it is fusing hydrogen into helium in its core.

The band is narrow, and the reason is that position along it is set almost entirely by one parameter: mass. Give a star a mass and the physics of Module 2's next two lessons fixes its radius, its core temperature, its luminosity and its surface temperature to within a modest spread. Composition and age shift a star slightly, which is why the band has width rather than being a line, but mass does the work.

Spectral typeMass (solar)Luminosity (solar)Radius (solar)Surface T (K)Main-sequence lifetime
O5V40500,0001242,000About 1 million years
B0V1720,000730,000About 10 million years
A0V2.9542.49,800About 500 million years
F0V1.66.51.57,300About 3 billion years
G2V (Sun)1.01.01.05,772About 10 billion years
K0V0.80.40.855,200About 20 billion years
M5V0.20.0080.33,200Longer than the age of the universe

Look at the second and third columns together. Multiply the mass by 40 and the luminosity goes up by a factor of 500,000. That steep dependence is the mass-luminosity relation, which for stars in the broad middle of the range is approximately

L / L_sun = (M / M_sun)^3.5

Check it on the A0 star: 2.9^3.5. Take logs: 3.5 x log10(2.9) = 3.5 x 0.4624 = 1.618, so L = 10^1.618 = 41. The table says 54, so the rule is good to a few tens of per cent, which is what you want from a one-line approximation. The exponent is not universal: it is closer to 4 at the low-mass end and flattens toward 3 for the most massive stars, because the internal physics changes.

The core of it: Position along the main sequence is essentially a mass axis, and luminosity climbs roughly as the 3.5 power of mass, so a small range of masses spans an enormous range of brightness.

Reading a radius straight off the plot

Because L = 4 pi R^2 sigma T^4, any point on the diagram implies a radius. Lines of constant radius run as diagonals from lower left to upper right: at fixed R, a hotter star is more luminous as T^4.

That gives you an immediate physical reading of the diagram's geography. Take two stars at the same temperature, 5,000 K, one at 1 L_sun and one at 1,000 L_sun. The luminosity ratio is 1,000 and the temperatures are equal, so the radius ratio is the square root of 1,000, about 32. The luminous one is 32 times larger. It is not hotter; it is bigger. That is what a red giant is.

Now go the other way. A white dwarf at 25,000 K, four times the Sun's temperature, with a luminosity of 0.003 L_sun. Radius ratio equals (0.003)^0.5 x (5772/25000)^2 = 0.0548 x 0.0533 = 0.0029. About 0.003 solar radii, or roughly 2,000 kilometres, comparable to the Earth. A solar mass or so of material in an Earth-sized volume gives a mean density around 10^9 kg per cubic metre, which is a tonne per cubic centimetre. Module 4 explains what holds it up.

The four neighbourhoods

The main sequence runs diagonally across the middle. Hydrogen core fusion. About ninety per cent of stars, and the Sun.

The giant branch and red clump sit above and to the right. Cool, large, luminous. These are post-main-sequence stars whose cores have run out of hydrogen and whose envelopes have swollen. Arcturus at 25 solar radii and Aldebaran at about 44 live here.

The supergiants occupy a horizontal band across the top, at luminosities from 10^4 to over 10^6 solar. Betelgeuse at roughly 800 solar radii and Rigel at about 70 are both here, at opposite ends in temperature, which tells you the band is populated by massive stars in several different evolutionary states.

The white dwarfs form a separate sequence in the lower left: hot, and yet faint, which is only possible if they are tiny. Sirius B, at 25,000 K and about 0.056 solar luminosities, was the first one recognised, and its position on this diagram was the clue that something extraordinary was going on inside it.

Between the main sequence and the giant branch lies a relatively empty region called the Hertzsprung gap, and its emptiness is itself a measurement: stars cross it quickly, so few are caught in the act.

The diagram as a clock

Here is the move that makes the HR diagram indispensable rather than merely tidy. Massive stars are luminous, and luminous stars burn through their fuel fast, as the next lessons will quantify. So in any group of stars born together, the most massive ones leave the main sequence first, then the next most massive, and so on down.

Plot a cluster's colour-magnitude diagram and you will see the main sequence intact at the faint end and truncated at the bright end. The point where it bends away is the main-sequence turnoff, and the mass of the star at that point has a known lifetime, which is therefore the cluster's age. This single trick dates open clusters at tens of millions of years and globular clusters at twelve to thirteen billion, and it is worked properly in Module 5.

Bottom line: The HR diagram is not a snapshot of stars sitting still. It is a map of evolutionary states, and the population of each region tells you how long stars spend there.

A worked identification

Four stars, each with a measured luminosity and temperature. Place each one and say what it is.

Star 1: L = 0.0006 L_sun, T = 3,000 K. Cool and very faint. Radius = (0.0006)^0.5 x (5772/3000)^2 = 0.0245 x 3.70 = 0.091 solar radii. Small, cool, faint: an M dwarf on the lower main sequence, roughly 0.1 solar masses, and it will still be fusing hydrogen long after the Sun is a white dwarf.

Star 2: L = 0.001 L_sun, T = 15,000 K. Hot and faint is the impossible combination unless the star is tiny. Radius = (0.001)^0.5 x (5772/15000)^2 = 0.0316 x 0.148 = 0.0047 solar radii, about 3,300 km. A white dwarf.

Star 3: L = 200 L_sun, T = 4,500 K. Radius = (200)^0.5 x (5772/4500)^2 = 14.1 x 1.645 = 23 solar radii. Cool and large and luminous: a red giant, probably around one to two solar masses, currently past the main sequence.

Star 4: L = 20,000 L_sun, T = 30,000 K. Radius = (20,000)^0.5 x (5772/30000)^2 = 141 x 0.037 = 5.2 solar radii. Hot, luminous, moderately sized: a massive main-sequence B star of roughly 17 solar masses, with a life expectancy around ten million years.

Common misconceptions

"Stars move along the main sequence as they age." They do not travel down the band toward cooler types. A star of a given mass sits at essentially one place on the main sequence for its entire hydrogen-burning life, brightening slowly and shifting slightly, and then leaves the band sideways toward the giant region. The main sequence is a mass axis, not an age axis.

"The HR diagram shows where stars are in space." Both axes are physical properties of the star. Position on the diagram says nothing about location, and two stars adjacent on the plot may be on opposite sides of the Galaxy.

"Red giants are hot because they are bright." They are bright because they are enormous. A red giant's surface is cooler than the Sun's; it wins on area, and area beats temperature here because the radius grows by a factor of tens while the temperature falls by less than a factor of two.

"The gap in the diagram means no stars have those properties." The Hertzsprung gap is sparsely populated because stars cross it in a geologically short time, not because the region is forbidden. Sparse regions on the diagram usually mean fast evolutionary phases.

The short version

The HR diagram plots luminosity against surface temperature, with temperature running backwards because Russell inherited the Harvard letter order in 1913. About ninety per cent of stars fall on the main sequence, where position is set almost entirely by mass and luminosity scales roughly as mass to the 3.5 power, spanning ten orders of magnitude for masses spanning a factor of a few hundred. Because L = 4 pi R^2 sigma T^4, every point on the diagram implies a radius, which is how you can tell at a glance that a cool luminous star is a giant tens of times the Sun's size and a hot faint star is a white dwarf the size of Earth. The four regions correspond to distinct physical states, and their relative populations are a direct measure of how long each state lasts. Crucially, the point at which a cluster's main sequence bends away gives the cluster's age. What the diagram does not explain by itself is why mass should determine everything, and that is the subject of the next two lessons: hydrostatic balance, and the nuclear reactions that pay for it.

Sources

  1. Encyclopaedia Britannica. (n.d.). Hertzsprung-Russell diagram. britannica.com
  2. European Space Agency. (n.d.). Gaia mission science highlights. ESA Science and Exploration. esa.int
  3. Fraknoi, A., Morrison, D., & Wolff, S. C. (2022). Astronomy 2e, Chapter 18: The stars, a celestial census. OpenStax, Rice University. openstax.org
  4. NASA Science. (n.d.). Stars. National Aeronautics and Space Administration. science.nasa.gov
  5. Russell, H. N. (1914). Relations between the spectra and other characteristics of the stars. Popular Astronomy, 22, 275-294 and 331-351.
Key terms
Hertzsprung-Russell diagram
A plot of stellar luminosity against surface temperature, with temperature increasing to the left.
Colour-magnitude diagram
An HR diagram using a colour index and apparent magnitude, used for clusters where all stars share a distance.
Main sequence
The diagonal band occupied by stars fusing hydrogen in their cores, containing about ninety per cent of stars.
Mass-luminosity relation
The approximate rule L / L_sun = (M / M_sun)^3.5 for main-sequence stars.
Magnitude-limited sample
A sample selected by apparent brightness, which over-represents intrinsically luminous stars.
Volume-limited sample
A sample containing everything within a given distance, which represents the true stellar population.
Main-sequence turnoff
The point where a cluster's main sequence bends toward the giant branch, whose stellar mass gives the cluster's age.
Hertzsprung gap
The sparsely populated region between the main sequence and the giant branch, thinly occupied because stars cross it quickly.

Inside a Star: Hydrostatic Equilibrium and Energy Transport

  • State hydrostatic equilibrium and use it to estimate a star's central pressure and temperature from its mass and radius.
  • Distinguish radiative and convective energy transport and predict which operates where in stars of different masses.
  • Compare the dynamical, thermal and nuclear timescales and explain why they make stars mechanically stable.

The Sun rings, and we listen

In 1960 Robert Leighton pointed a spectroheliograph at the Sun from Mount Wilson and watched the Doppler shift of a single patch of the surface. It did not sit still. The patch rose and fell with vertical speeds of a few hundred metres per second and a period close to five minutes. Leighton, Robert Noyes and George Simon published the result in 1962, and for a decade nobody was sure what it meant.

It turned out to be sound. The Sun's interior traps acoustic waves, which bounce between the surface and the depth where increasing sound speed refracts them back up. Millions of distinct modes are predicted and many thousands have been individually measured, each one sampling a different depth. Because the frequency of a mode depends on the sound speed along its path, and the sound speed depends on temperature and composition, measuring thousands of frequencies and inverting them maps the interior.

The results are startlingly precise. Helioseismology puts the base of the Sun's convective envelope at 0.713 of the solar radius, with an uncertainty of about 0.003. It measures the sound speed profile through most of the interior to a fraction of a per cent. It fixed the helium abundance of the outer layers, and it revealed that the convection zone rotates differentially while the radiative interior below rotates nearly as a solid body.

So this lesson is not a hypothetical. Stellar interiors are inaccessible in the sense that you cannot go there, and they are measurable in the sense that matters. What follows is the physics those measurements confirm.

The point: The interior of a star is not guesswork. For the Sun it has been mapped acoustically, and every model in this course is checked against that map.

The balance that defines a star

A star is a ball of gas that gravity is trying to crush. It does not collapse, and it does not fly apart, which means the inward pull is balanced at every depth by an outward push. Consider a thin shell of gas at radius r inside the star. Gravity pulls it inward with a force proportional to the mass inside r and the shell's own mass. Pressure pushes it outward, because the pressure below the shell is greater than the pressure above it. Balance those and you have hydrostatic equilibrium:

dP/dr = -G x M(r) x rho(r) / r^2

In words: pressure must fall outward, and it must fall at exactly the rate needed to hold up the weight of everything above. Nothing in that statement is specific to stars. It is the same equation that governs the pressure at the bottom of the ocean, applied to a self-gravitating sphere.

What is specific to stars is the size of the numbers. Let us estimate the central pressure. Dimensionally, integrating the equation across the star gives P_c of order G M^2 / R^4. For the Sun, with M = 1.989 x 10^30 kg and R = 6.957 x 10^8 m:

M^2 = 3.956 x 10^60, so G M^2 = 6.674 x 10^-11 x 3.956 x 10^60 = 2.64 x 10^50

R^4 = (6.957 x 10^8)^4 = 2.34 x 10^35

P_c is of order 2.64 x 10^50 / 2.34 x 10^35 = 1.1 x 10^15 pascals

Detailed solar models give 2.3 x 10^16 pascals, about twenty times larger, because the Sun is strongly centrally concentrated and the crude estimate treats it as uniform. Be honest about what the estimate is worth: it gets the scaling right, not the number. And the scaling is the useful part, because P_c goes as M^2 / R^4. Double a star's mass at fixed radius and the central pressure quadruples. That single dependence is why mass controls everything about a star, which is the fact the HR diagram displayed without explaining.

From pressure to temperature

In an ordinary star the gas is a fully ionised plasma that behaves as an ideal gas, so

P = rho x k x T / (mu x m_H)

where k is Boltzmann's constant, m_H is the mass of a hydrogen atom, and mu is the mean molecular weight in units of m_H, about 0.62 for fully ionised gas of solar composition. Combine that with the hydrostatic estimate and the central temperature comes out as roughly

T_c is of order G M mu m_H / (k R)

Put the solar numbers in. G M = 6.674 x 10^-11 x 1.989 x 10^30 = 1.328 x 10^20. Multiply by mu m_H = 0.62 x 1.673 x 10^-27 = 1.037 x 10^-27, giving 1.377 x 10^-7. Divide by k R = 1.381 x 10^-23 x 6.957 x 10^8 = 9.61 x 10^-15.

T_c is of order 1.43 x 10^7 kelvin

Fourteen million kelvin. The value from detailed solar models, confirmed by helioseismology, is 15.7 million kelvin. A one-line dimensional argument lands within about ten per cent, which is better than the argument deserves and is the reason astronomers trust order-of-magnitude reasoning as far as they do.

That temperature is the crucial output. Fifteen million kelvin is the threshold at which hydrogen nuclei can fuse at a useful rate, as the next lesson will show. Gravity, acting on a solar mass of hydrogen, produces exactly the conditions that ignite the fuel. Nobody arranged that; it is a consequence of the numbers.

Two more solar numbers worth memorising: the mean density is 1,408 kg per cubic metre, slightly more than water, while the central density is about 150,000 kg per cubic metre, or 150 times water. The Sun is a gas throughout, even at 150 times the density of water, because at 15 million kelvin it is far too hot for atoms to bind.

The four equations, in words

Modelling a star means solving four coupled differential equations from the centre to the surface.

Hydrostatic equilibrium fixes how pressure changes with radius, as above.

Mass conservation says the mass enclosed grows with radius as dM/dr = 4 pi r^2 rho, which is bookkeeping.

Energy generation says the luminosity flowing outward grows with radius wherever nuclear reactions are running: dL/dr = 4 pi r^2 rho epsilon, where epsilon is the energy released per kilogram per second.

Energy transport relates the temperature gradient to the energy flux, and takes a different form depending on whether radiation or convection is doing the carrying.

To close the system you also need three pieces of input physics: an equation of state relating pressure to density, temperature and composition; an opacity describing how hard it is for radiation to get through the material; and nuclear reaction rates. Supply those, specify a mass and a composition, and the structure follows. The claim that mass and composition alone determine a star's structure and evolution is called the Vogt-Russell theorem, which is not actually a theorem, since rotation, magnetic fields, binarity and mass loss all break it. It is a very good first approximation and you should hold it loosely.

Worth holding on to: Stellar models are not fitted to data. Give the equations a mass and a composition and they produce a luminosity, a radius and a temperature that can then be compared with a measured star. That is why agreement means something.

Getting the energy out: a photon's very slow walk

Energy is made in the core and must reach the surface. In most of the Sun's interior it travels by radiative diffusion: a photon is emitted, travels a short distance, is absorbed or scattered, is re-emitted in a random direction, and repeats.

How short is that distance? In the deep interior the mean free path is of order one centimetre. A random walk covering a net distance R requires roughly (R divided by the step length) squared steps. The Sun's radius is 7 x 10^10 cm, so

N is of order (7 x 10^10)^2 = 4.9 x 10^21 steps

Each step takes 1 cm divided by 3 x 10^10 cm/s, so the total time is about 4.9 x 10^21 x 3.3 x 10^-11 = 1.6 x 10^11 seconds, roughly 5,000 years. More careful calculations, which account for the mean free path changing enormously with depth, give figures from about ten thousand to a few hundred thousand years. Any specific number you see quoted depends on assumptions, so treat the honest answer as tens to hundreds of thousands of years. The point survives: the energy warming your face left the core before agriculture, and the neutrinos from the same reactions arrived here in eight and a half minutes.

What blocks the photons is opacity, and several processes contribute. Free electrons scatter photons. Ions absorb photons and eject electrons, which is bound-free absorption. Electrons passing near ions absorb photons in free-free transitions. Near the solar surface, where it is cool enough for the negative hydrogen ion to exist, that ion dominates the opacity, and it is the reason the Sun has a sharp visible edge at all.

When radiation is not enough: convection

If radiation cannot carry the flux without demanding an impossibly steep temperature gradient, the gas starts moving instead. A blob that is displaced upward and finds itself less dense than its new surroundings keeps rising. That is the Schwarzschild criterion for convective instability, and where it is satisfied the star boils.

You can see this on the Sun. The visible surface is covered in granules, cells roughly 1,000 kilometres across with bright rising centres and dark sinking edges, each lasting something like ten minutes. They are the tops of convection cells, and their existence is direct evidence that the outer third of the Sun is convective, exactly where helioseismology places the boundary at 0.713 solar radii.

Where convection occurs depends on mass, and the pattern reverses across the main sequence.

Stellar massCoreEnvelopeWhy
Below about 0.35 solarConvective throughoutHigh opacity at low temperature; no radiative zone survives
About 0.35 to 1.3 solarRadiativeConvectiveGentle proton-proton energy generation spreads over a large core; cool outer layers are opaque
Above about 1.3 solarConvectiveRadiativeCNO burning is so temperature sensitive that energy is generated in a tiny central region, forcing convection there

Two consequences matter later. A fully convective low-mass star mixes its entire hydrogen supply into the core, so it can burn nearly all of it and lives for hundreds of billions of years. A massive star with a convective core continually replenishes the core with fresh fuel from a larger region, which changes how its evolution proceeds. And the Sun, with a convective envelope over a radiative core, has a shear layer between them called the tachocline, which is where its magnetic dynamo is thought to operate.

Three timescales, and why stars sit still

Three characteristic times control stellar behaviour, and their ratios explain the stability you observe.

The dynamical timescale is how long the star would take to react mechanically if pressure support vanished, essentially a free-fall time. For the Sun it is about half an hour.

The thermal, or Kelvin-Helmholtz, timescale is how long the star could shine on its stored gravitational energy alone, roughly G M^2 divided by (R L). For the Sun that comes to about 30 million years. In the 1860s Kelvin and Helmholtz computed exactly this and concluded that the Sun could be no older, which put physics in direct conflict with the geologists and with Darwin, both of whom needed vastly longer. They were right about the physics and wrong about the energy source.

The nuclear timescale is how long the fuel lasts: about 10 billion years for the Sun.

The ratios are roughly 1 to 10^12 to 10^14. A star is therefore in essentially perfect mechanical equilibrium at all times, because any imbalance is corrected within an hour while the structure evolves over billions of years. This is why hydrostatic equilibrium is not an approximation you apologise for. It is exact to a precision no other assumption in the subject can match.

The thermostat, and what happens when it is removed

Stars are stable for a second reason, and it is the reason a star does not behave like a bomb. Suppose fusion in the core speeds up slightly. The extra energy raises the pressure, the core expands, and expansion cools it. Cooling slows the reactions back down. The feedback is negative, and it is strong, because reaction rates depend on temperature to a very high power.

Push on it the other way and the same logic runs in reverse: a core that produces too little energy contracts, and contraction heats it, restoring the rate. This self-regulation is why the Sun has varied its output by only a few tens of per cent over four and a half billion years rather than flickering.

The thermostat depends on one thing: that pressure rises with temperature. Remove that link and it fails. In a degenerate gas, where quantum mechanics rather than thermal motion supplies the pressure, pressure barely depends on temperature at all. Ignite fusion in degenerate material and the core heats without expanding, the reactions accelerate, and the result is a runaway. That is precisely what happens in the helium flash in a low-mass red giant and, far more violently, in a Type Ia supernova. Both are in Module 4, and both are consequences of this single missing feedback.

Remember: A normal star is stable because pressure rises with temperature. Every stellar explosion in this course happens where that link has been broken.

Common misconceptions

"Radiation pressure holds up the Sun." Gas pressure does, by a wide margin. Radiation pressure contributes well under a per cent in the Sun. It becomes important only in very massive stars, where it can dominate and set an upper limit on stellar mass.

"The Sun is burning, like a fire." Chemical burning of the Sun's entire mass would last a few thousand years. The energy source is nuclear, and the word burning survives only as jargon.

"The interior is unknowable, so models are speculation." Helioseismology measures the internal sound speed to a fraction of a per cent and locates the base of the convection zone to better than half a per cent, and solar neutrinos measure the core reaction rate directly. Stellar interiors are among the better tested regions in physics.

"Energy takes a million years to escape the Sun, and this is a precise figure." Estimates range over more than an order of magnitude depending on how the varying mean free path is handled. Tens to hundreds of thousands of years is the defensible statement.

Pulling it together

A star is a self-gravitating ball held up by a pressure gradient, and hydrostatic equilibrium sets that gradient exactly: dP/dr = -G M(r) rho / r^2. Applying it to the Sun gives a central pressure scaling as M^2 / R^4 and, through the ideal gas law, a central temperature of about 14 million kelvin from a one-line estimate against 15.7 million from full models. That temperature is the ignition condition for hydrogen fusion, which is why mass determines everything else about a star. Four coupled equations, plus an equation of state, an opacity and reaction rates, determine the whole structure from mass and composition alone. Energy escapes by radiative diffusion, a random walk taking tens to hundreds of thousands of years, and by convection wherever the required temperature gradient becomes too steep, which is the outer third of the Sun, the entire volume of an M dwarf, and the core of any star above about 1.3 solar masses. The dynamical, thermal and nuclear timescales differ by roughly twelve and fourteen orders of magnitude, so stars are always in mechanical balance while evolving slowly, and the temperature dependence of fusion supplies a thermostat that fails only where matter becomes degenerate. All of this rests on an energy source we have not yet examined, and that is next.

Sources

  1. NASA Science. (n.d.). The Sun. National Aeronautics and Space Administration. science.nasa.gov
  2. NASA. (n.d.). Imagine the Universe: The life cycles of stars. Goddard Space Flight Center. imagine.gsfc.nasa.gov
  3. Fraknoi, A., Morrison, D., & Wolff, S. C. (2022). Astronomy 2e, Chapter 16: The Sun, a nuclear powerhouse. OpenStax, Rice University. openstax.org
  4. Encyclopaedia Britannica. (n.d.). Star: Stellar structure. britannica.com
  5. Leighton, R. B., Noyes, R. W., & Simon, G. W. (1962). Velocity fields in the solar atmosphere. I. Preliminary report. Astrophysical Journal, 135, 474-499.
  6. Kippenhahn, R., Weigert, A., & Weiss, A. (2012). Stellar structure and evolution (2nd ed.). Springer.
Key terms
Hydrostatic equilibrium
The balance at every depth between the inward pull of gravity and the outward pressure gradient, expressed as dP/dr = -G M(r) rho / r^2.
Mean molecular weight
The average mass per particle in a gas, in units of the hydrogen atom mass; about 0.62 for fully ionised solar-composition material.
Opacity
A measure of how strongly stellar material absorbs and scatters radiation, controlling how easily energy diffuses outward.
Radiative diffusion
Energy transport by photons performing a random walk through opaque material, taking tens to hundreds of thousands of years to cross the Sun.
Convection
Energy transport by bulk motion of gas, which takes over wherever the temperature gradient required for radiative transport becomes too steep.
Granulation
The pattern of convection cells about 1,000 km across visible on the solar photosphere.
Kelvin-Helmholtz timescale
The time a star could shine on gravitational contraction alone, about 30 million years for the Sun.
Helioseismology
The study of the Sun's interior through its trapped acoustic oscillations, which fixes the sound speed profile and the depth of the convection zone.
Degenerate gas
Matter whose pressure comes from quantum mechanical exclusion rather than thermal motion, and which therefore lacks the temperature-pressure feedback that stabilises normal stars.

The Furnace: Fusion, the Proton-Proton Chain, and the CNO Cycle

  • Explain how the mass defect between four hydrogen nuclei and one helium nucleus supplies stellar energy, and compute the energy released.
  • Describe the proton-proton chain and CNO cycle step by step and account for their very different temperature sensitivities.
  • Explain the solar neutrino problem and how its resolution vindicated stellar models rather than revising them.

Eddington in Cardiff, August 1920

Arthur Eddington gave the presidential address to the physics section of the British Association at Cardiff in 1920 and made a claim he could not prove. Francis Aston had recently measured atomic masses precisely enough to show that a helium atom weighs slightly less than four hydrogen atoms. On modern values, four hydrogen atoms come to 4.031300 atomic mass units and one helium-4 atom to 4.002603, a shortfall of 0.028697 units, about 0.7 per cent. Eddington argued that if stars could convert hydrogen into helium, that missing mass would appear as energy by Einstein's relation, and the Sun could shine for billions of years rather than the 30 million that gravitational contraction allowed.

He had no mechanism. The temperature he could calculate for the solar centre, around fifteen million degrees, was far too low for protons to overcome their mutual electrical repulsion by classical physics, and his critics said so. Eddington's reply was essentially that the critics should find a hotter place. He was right, and the resolution arrived from quantum mechanics eight years later.

Convert the mass defect and you get the number this whole lesson turns on:

E = 0.028697 u x 931.494 MeV per u = 26.73 MeV per helium nucleus formed

Compare that with a chemical reaction, which releases a few electron volts. Fusion is roughly ten million times more energetic per event, and that factor is the difference between a Sun that lasts a few thousand years and one that lasts ten billion.

Key idea: Stars run on a 0.7 per cent conversion of rest mass into energy, and that small percentage is enormous because c^2 is enormous.

Why fusion should be impossible, and why it happens anyway

Two protons repel. To fuse, they must approach within about 10^-15 metres, where the strong nuclear force can act. The electrical potential energy at that separation is

U = k e^2 / r = (8.99 x 10^9) x (1.602 x 10^-19)^2 / 10^-15 = 2.3 x 10^-13 joules, which is about 1.4 MeV

Now compare with the thermal energy available. At the solar core temperature of 15.7 million kelvin,

kT = 1.381 x 10^-23 x 1.57 x 10^7 = 2.17 x 10^-16 joules, which is about 1.35 keV

The barrier is about a thousand times higher than the typical particle energy. Classically, nothing fuses. This is exactly the objection Eddington could not answer.

Two effects rescue it. The first is statistical: in a Maxwell-Boltzmann distribution a small fraction of particles have energies far above the mean, and it is that high-energy tail that does all the work. The second is quantum mechanical. George Gamow showed in 1928 that a particle can tunnel through a potential barrier it cannot climb, with a probability that falls steeply with barrier height and rises with energy. Multiply the falling number of fast particles by the rising tunnelling probability and you get a narrow band of energies, well above kT but well below the barrier, where nearly all fusion occurs. That band is called the Gamow peak, and its narrowness is why reaction rates are so ferociously temperature sensitive.

There is a third obstacle, and it is the reason the Sun is not a bomb. The very first step of the chain, two protons combining, produces a deuterium nucleus, which contains one proton and one neutron. That means one of the protons has to turn into a neutron during the collision, which requires the weak nuclear force, and the weak force is called weak for good reason. Even at solar core densities and temperatures, an individual proton waits on average something like nine billion years before it fuses. The Sun's stability comes from this bottleneck. If the first step were fast, a solar mass of hydrogen would burn out in an eyeblink.

The proton-proton chain, step by step

Below about 18 million kelvin, hydrogen fusion proceeds mainly by the proton-proton chain. The dominant branch, called ppI, has three steps.

Step 1. Two protons fuse to deuterium, emitting a positron and an electron neutrino. Energy released: 1.44 MeV, of which the neutrino carries an average 0.27 MeV straight out of the star. Timescale for a given proton: about 9 billion years.

Step 2. The deuterium captures another proton, forming helium-3 and emitting a gamma ray. Energy released: 5.49 MeV. Timescale: about one second. Deuterium never accumulates in a star because this step is so fast.

Step 3. Two helium-3 nuclei fuse to helium-4, returning two protons to the gas. Energy released: 12.86 MeV. Timescale: about 400 years.

Because step 3 consumes two helium-3 nuclei, steps 1 and 2 must each run twice for every completed cycle. Add it all up and the net reaction is

4 protons goes to helium-4 + 2 positrons + 2 neutrinos + gamma rays, releasing 26.73 MeV

The positrons annihilate immediately with electrons, contributing their energy locally. The neutrinos do not. They interact so weakly that they leave the Sun in about two seconds and reach Earth eight and a half minutes later, carrying roughly two per cent of the total. So about 26.2 MeV per helium nucleus actually heats the star, and 0.5 MeV escapes.

Two minor branches exist. In the Sun, ppI accounts for most completions; ppII, which routes through beryllium-7 and lithium-7, handles most of the rest; and ppIII, through boron-8, is rare but produces the highest energy neutrinos, which matters enormously for detection because those are the easiest neutrinos to catch.

Near solar conditions the proton-proton rate rises roughly as the fourth power of temperature. That is already steep. The alternative is far steeper.

The CNO cycle: carbon as a catalyst

Hans Bethe and, independently, Carl Friedrich von Weizsacker worked out in 1938 and 1939 that hydrogen can also be converted to helium using carbon, nitrogen and oxygen as catalysts. The nuclei are consumed and regenerated, so a trace of carbon suffices to process an entire core of hydrogen. Bethe received the Nobel Prize for this work in 1967.

The cycle runs: carbon-12 captures a proton to give nitrogen-13; nitrogen-13 decays to carbon-13, emitting a positron and a neutrino; carbon-13 captures a proton to give nitrogen-14; nitrogen-14 captures a proton to give oxygen-15; oxygen-15 decays to nitrogen-15, emitting a positron and a neutrino; and nitrogen-15 captures a proton and splits into carbon-12 plus helium-4. The carbon-12 you started with is back, and four protons have become one helium nucleus. The energy released is the same 26.73 MeV, because the reactant and product are the same; only the route differs.

The slowest step is nitrogen-14 capturing a proton, which is why nitrogen builds up in CNO-processed material. That surplus nitrogen is an observable signature: the surfaces of certain evolved stars show enhanced nitrogen and depleted carbon, which is direct evidence that CNO burning happened and that the material was subsequently dredged to the surface.

The crucial difference is temperature sensitivity. Because the cycle involves capturing protons onto nuclei with charge 6, 7 and 8, the Coulomb barriers are six to eight times higher, so tunnelling is far harder and the rate depends far more sharply on temperature: roughly as T^18 to T^20 near solar core conditions.

Proton-proton chainCNO cycle
FuelHydrogen onlyHydrogen, with C, N, O as catalysts
Temperature dependenceAbout T^4About T^18 to T^20
Dominant below or aboveBelow about 18 million KAbove about 18 million K
Stellar mass where it dominatesBelow about 1.3 solar massesAbove about 1.3 solar masses
Consequence for structureEnergy spread over a broad core; core stays radiativeEnergy concentrated in a tiny core; core becomes convective
Share of the Sun's outputAbout 99 per centAbout 1 per cent

The last row was a prediction for eighty years and became a measurement in 2020, when the Borexino experiment under the Gran Sasso mountain in Italy reported the first detection of neutrinos from the CNO cycle in the Sun, at a rate consistent with the roughly one per cent contribution standard models had assumed.

The upshot: The same net reaction runs by two routes with wildly different temperature sensitivities, and which route dominates decides whether a star's core is radiative or convective, which shapes everything about how it evolves.

Working the Sun's fuel consumption

Numbers make this concrete. The Sun radiates 3.828 x 10^26 watts, and each completed chain yields 26.73 MeV.

26.73 MeV = 26.73 x 1.602 x 10^-13 = 4.28 x 10^-12 joules

Reactions per second = 3.828 x 10^26 / 4.28 x 10^-12 = 8.9 x 10^37

Each completion consumes four protons, so protons consumed per second = 3.6 x 10^38. Multiply by the proton mass of 1.673 x 10^-27 kg:

Hydrogen consumed = 6.0 x 10^11 kg per second, which is 600 million tonnes every second

Of that, 0.7 per cent becomes energy. Check it directly with E = mc^2:

mass converted = L / c^2 = 3.828 x 10^26 / (9.0 x 10^16) = 4.3 x 10^9 kg per second

Roughly 4.3 million tonnes of matter per second, permanently. Over the Sun's 4.6 billion year life, which is 1.45 x 10^17 seconds, that comes to 6.2 x 10^26 kg, about one ten-thousandth of the Sun's mass, or a hundred Earths. The Sun has lost the mass of a hundred Earths to sunlight, and its structure has barely noticed.

Now the lifetime. Only the core, roughly the inner ten per cent of the mass, gets hot enough to fuse, and hydrogen burning stops when the core hydrogen is gone. Take about 10 per cent of the Sun's mass, of which about 71 per cent is hydrogen, giving roughly 1.4 x 10^29 kg of usable fuel. At 6 x 10^11 kg per second that lasts 2.4 x 10^17 seconds, about 7.5 billion years. Adding the slow brightening the Sun undergoes as it ages, standard models give a total main-sequence lifetime near 10 billion years, of which about 4.6 billion are spent. This arithmetic generalises in the next module.

The experiment that looked like it falsified all of this

Photons take tens of thousands of years to leak out of the Sun, so sunlight tells you about the core only indirectly. Neutrinos leave immediately. If you can catch them, you are measuring the nuclear reaction rate in the core right now.

Raymond Davis Jr. tried, starting in the late 1960s, using a tank holding 615 tonnes of perchloroethylene cleaning fluid nearly 1,500 metres underground in the Homestake gold mine in South Dakota. A neutrino occasionally converts a chlorine-37 atom to argon-37, and Davis extracted and counted the argon atoms, a few per week from 10^30 chlorine atoms. It is one of the most difficult measurements ever attempted.

He found about one third of the neutrinos that John Bahcall's solar models predicted. The deficit held for three decades and became known as the solar neutrino problem. Two explanations were on the table: the solar models were wrong, or something happened to the neutrinos on the way. Gallium experiments in Russia and Italy and the water Cherenkov detector Kamiokande in Japan confirmed a deficit, which narrowed the possibilities without settling them.

The answer arrived from the Sudbury Neutrino Observatory in Ontario, a kilotonne of heavy water two kilometres underground, which could do what earlier detectors could not: measure both the electron neutrino flux and the total flux of all three neutrino flavours separately. In 2001 and 2002 SNO reported that the electron neutrino flux was indeed about a third of predictions, while the total flux across all flavours matched Bahcall's models. The neutrinos were not missing. Two thirds of them had changed flavour in transit, which requires that neutrinos have mass, contradicting the Standard Model of particle physics as it then stood.

Davis and Masatoshi Koshiba shared the 2002 Nobel Prize in Physics; Takaaki Kajita and Arthur McDonald shared the 2015 prize for the oscillation result. The episode is worth remembering as a case where a thirty-year discrepancy between a model and an experiment was resolved in favour of the astrophysics, and the surprise turned out to be in particle physics instead. Solar models are now among the best tested calculations in science, checked independently by helioseismology and by neutrino counting.

In short: Neutrinos let us watch the Sun's core in real time, and when they seemed to contradict stellar theory, the fault lay with what was known about neutrinos.

What comes after hydrogen

Hydrogen fusion is only the first stage. When the core hydrogen runs out and the core contracts and heats to about 100 million kelvin, helium begins to fuse by the triple-alpha process: two helium-4 nuclei form beryllium-8, which is unstable and falls apart in about 10^-16 seconds, but at sufficient density a third helium nucleus occasionally arrives in time to make carbon-12.

That reaction should be far too slow to matter. In 1953 Fred Hoyle argued that since carbon plainly exists, there must be an excited state of carbon-12 at just the right energy, near 7.65 MeV, to make the reaction resonant. William Fowler's group at Caltech looked, and it was there. It remains one of the few genuinely successful predictions made from the existence of the predictor.

Beyond helium come carbon, neon, oxygen and silicon burning in massive stars, each stage hotter, faster and less productive than the last, ending at iron. Module 4 works through why the sequence stops there and what happens when it does.

Common misconceptions

"The Sun's core is hot enough for protons to overcome their repulsion." It is about a thousand times too cool. Fusion happens because of the high-energy tail of the velocity distribution combined with quantum tunnelling, and both are essential.

"The CNO cycle uses carbon as fuel." Carbon, nitrogen and oxygen are catalysts. They are consumed and regenerated in each turn of the cycle, and the fuel is hydrogen, exactly as in the proton-proton chain. The net reaction and the energy released are identical.

"Fusion converts hydrogen entirely into energy." Only 0.7 per cent of the mass is converted. The other 99.3 per cent remains as helium, which is why stars end their lives full of ash rather than empty.

"The solar neutrino problem showed stellar models were wrong." The opposite. When the measurement was finally done in a way that counted all neutrino flavours, the models were right to within their uncertainties, and the new physics was neutrino mass.

"Fusion in a star is a chain reaction that could run away." In non-degenerate material the pressure-temperature feedback from the previous lesson prevents it, and the weak-interaction bottleneck in the first proton-proton step makes the whole process extraordinarily slow to begin with.

What you now know

Four hydrogen nuclei weigh 0.7 per cent more than the helium nucleus they can become, and that mass defect releases 26.73 MeV per helium formed, which Eddington identified as the stellar energy source in 1920 without knowing how it could happen. It happens because the Maxwell-Boltzmann tail supplies fast particles and quantum tunnelling lets them through a Coulomb barrier a thousand times higher than their typical energy, with the first step further throttled by the weak interaction so that a proton waits billions of years to react. Below about 18 million kelvin the proton-proton chain dominates, with a rate going as roughly T^4; above it the CNO cycle takes over, with a rate going as roughly T^18 to T^20 and carbon acting purely as a catalyst. That switch, at around 1.3 solar masses, decides whether a star's core is radiative or convective. The Sun consumes 600 million tonnes of hydrogen a second and converts 4.3 million tonnes of matter into energy in the process, a rate that gives it a main-sequence life near ten billion years. And the whole picture has been checked at the core itself: after thirty years in which solar neutrinos appeared to be two thirds missing, SNO showed they had merely changed flavour, and Borexino has since detected the one per cent of the Sun's energy that comes from CNO burning.

Sources

  1. NASA Science. (n.d.). The Sun. National Aeronautics and Space Administration. science.nasa.gov
  2. Encyclopaedia Britannica. (n.d.). Nuclear fusion. britannica.com
  3. Fraknoi, A., Morrison, D., & Wolff, S. C. (2022). Astronomy 2e, Chapter 16: The Sun, a nuclear powerhouse. OpenStax, Rice University. openstax.org
  4. Nobel Prize Outreach. (2002). The Nobel Prize in Physics 2002. nobelprize.org
  5. Bethe, H. A. (1939). Energy production in stars. Physical Review, 55(5), 434-456.
  6. Clayton, D. D. (1983). Principles of stellar evolution and nucleosynthesis. University of Chicago Press.
Key terms
Mass defect
The difference in mass between reactants and products of a nuclear reaction, which appears as energy; 0.7 per cent for hydrogen to helium.
Coulomb barrier
The electrical repulsion two nuclei must overcome to fuse, about 1.4 MeV for two protons.
Quantum tunnelling
The finite probability that a particle passes through a potential barrier it lacks the energy to climb, without which stars would not burn.
Gamow peak
The narrow band of particle energies, above thermal but below the barrier, in which nearly all fusion reactions occur.
Proton-proton chain
The dominant hydrogen fusion route below about 18 million K, with a rate scaling roughly as T^4.
CNO cycle
Hydrogen fusion catalysed by carbon, nitrogen and oxygen, dominant above about 18 million K, with a rate scaling roughly as T^18 to T^20.
Triple-alpha process
The fusion of three helium-4 nuclei into carbon-12 at around 100 million K, made possible by a resonance Hoyle predicted in 1953.
Solar neutrino problem
The three-decade discrepancy between the measured and predicted flux of solar electron neutrinos, resolved by neutrino flavour oscillation.

Module 3: Birth and the Main Sequence

How a cold cloud of molecular hydrogen becomes a star, and why the mass it ends up with decides, to within a factor of two, how long it will live.

From Cold Cloud to Protostar

  • Describe the physical conditions in molecular clouds and use the Jeans criterion to decide whether a region will collapse.
  • Trace the collapse sequence from cloud fragment through protostar and disk to the zero-age main sequence.
  • Explain the initial mass function and the physical limits that set the smallest and largest stellar masses.

A hole in the sky in Ophiuchus

Barnard 68 is a small dark cloud in Ophiuchus, about 125 parsecs away. In visible light it is a clean silhouette: an ink blot roughly a third of a light year across with almost no background star showing through. In the near infrared, where dust is far more transparent, hundreds of background stars appear right through it, each one dimmed and reddened by a specific amount that depends on how much dust sits along its particular line of sight.

Joao Alves, Charles Lada and Elizabeth Lada realised in 2001 that this was a measurement waiting to be made. Count the reddening of every background star and you map the cloud's column density point by point, with no assumptions about temperature or chemistry. The result, published in Nature, was a density profile that matched a Bonnor-Ebert sphere: a self-gravitating ball of isothermal gas held in balance by its own pressure, sitting right at the edge of stability. Barnard 68 holds about two solar masses of gas at a temperature near 10 kelvin, and it is close to collapsing.

That is where every star begins. Not in fire but in something a few degrees above absolute zero, dense enough to be opaque and cold enough for hydrogen to exist as molecules. This lesson follows the material from there to the moment fusion starts.

So what?: Star formation is a competition between gravity and everything that resists it, and the outcome is decided while the gas is still colder than liquid nitrogen.

The raw material

The space between stars is not empty. Roughly ten to fifteen per cent of the Milky Way's baryonic mass sits in the interstellar medium, in several phases at wildly different temperatures and densities. Introduction to Astronomy covers the full phase inventory; what matters here is the coldest, densest one.

Molecular clouds have temperatures of 10 to 20 kelvin and densities from about 100 to 10^6 particles per cubic centimetre. That density sounds large until you notice the best laboratory vacuums on Earth are far emptier than the air in this room and still denser than this by many orders of magnitude. What makes molecular clouds special is not that they are dense in absolute terms but that they are dense compared with everything around them, and cold enough that gravity can win.

The gas is mostly molecular hydrogen, and there is a problem with that: H2 is a symmetric molecule with no permanent dipole moment, so it barely radiates at the temperatures involved and is effectively invisible. Astronomers therefore trace molecular gas with carbon monoxide, which is about ten thousand times less abundant but radiates strongly at 2.6 millimetres, and convert CO brightness to H2 mass using a calibration factor that carries real uncertainty. Whenever you read a molecular cloud mass, remember there is a conversion factor inside it.

About one per cent of the mass is dust: silicate and carbonaceous grains typically a tenth of a micrometre across. Dust is a minority component that runs the show. It absorbs starlight and re-radiates in the infrared, which is what cools the cloud; it shields molecules from ultraviolet radiation that would otherwise destroy them; and its surfaces are where molecular hydrogen actually forms, since two hydrogen atoms in free space cannot easily shed the energy needed to bind.

Giant molecular clouds run from 10^4 to 10^6 solar masses and span tens of parsecs. The Orion complex, at about 400 parsecs, is the nearest large example, and the Orion Nebula is the illuminated face of one small part of it.

The Jeans criterion, worked

James Jeans asked in 1902 when a cloud's self-gravity beats its internal pressure. The answer, for a uniform isothermal sphere, is that collapse occurs above a critical mass

M_J = (5kT / (G mu m_H))^1.5 x (3 / (4 pi rho))^0.5

which is worth putting numbers into once. Take T = 10 K and a number density of 100 molecules per cubic centimetre, so with mu = 2.3 for molecular gas the mass density is rho = 3.85 x 10^-19 kg per cubic metre.

5kT / (G mu m_H) = (5 x 1.381 x 10^-23 x 10) / (6.674 x 10^-11 x 3.85 x 10^-27) = 6.91 x 10^-22 / 2.57 x 10^-37 = 2.69 x 10^15

Raised to the power 1.5: 1.39 x 10^23

(3 / (4 pi x 3.85 x 10^-19))^0.5 = (6.21 x 10^17)^0.5 = 7.88 x 10^8

M_J = 1.39 x 10^23 x 7.88 x 10^8 = 1.10 x 10^32 kg, about 55 solar masses

Different textbooks use slightly different geometric coefficients, so treat 55 as the scale rather than a precise threshold. What matters is the scaling, which is

M_J proportional to T^1.5 divided by the square root of density

Cold helps and dense helps. Warm the cloud to 100 K and the critical mass rises by a factor of 31. Compress it a hundredfold and the critical mass falls by a factor of ten. In a dense core at 10^4 molecules per cubic centimetre, still at 10 K, the Jeans mass is about 5.5 solar masses, which is the right order for the cores that actually make stars.

The free-fall time for material that has decided to collapse is

t_ff = (3 pi / (32 G rho))^0.5

At 10^4 molecules per cubic centimetre that comes to about 1.1 x 10^13 seconds, roughly 340,000 years. Fast on a stellar timescale, slow on a human one.

Why the Galaxy is not converting itself into stars right now

Add up the mass of molecular gas in the Milky Way, divide by the free-fall time, and you predict a star formation rate of a few hundred solar masses per year. The measured rate is one to two solar masses per year. Something is holding the clouds up, and there are three candidates, all of them real.

Turbulence. Molecular clouds show supersonic internal motions, with line widths far broader than thermal. Turbulent pressure supports the cloud on large scales, while simultaneously creating dense compressed regions on small scales where collapse can proceed. It both prevents and triggers.

Magnetic fields. The gas is weakly ionised, and the ions are tied to field lines, so a collapsing cloud must drag its magnetic field inward against its own resistance. Neutral gas slowly slips past the ions in a process called ambipolar diffusion, which lengthens collapse times considerably in the denser cores.

Feedback. Once massive stars form, their ultraviolet radiation, winds and eventual supernovae disrupt the remaining cloud. This is why star formation efficiency, the fraction of a cloud's mass that ends up in stars, is typically only a few per cent. Most of the gas is blown away before it can be used.

Something has to tip the balance, and the usual triggers are external: a supernova blast wave, the compression as gas passes through a spiral arm, a collision between clouds, or the expanding ionised bubble around a young massive star squeezing its neighbours. The pillars in the Eagle Nebula that Hubble photographed in 1995, and the similar structures JWST has since imaged in far more detail, are exactly this process caught in progress: dense knots resisting erosion by the radiation of nearby hot stars, with young stars forming inside some of them.

Bottom line: Left alone, molecular clouds mostly do not collapse. Star formation is inefficient, externally triggered, and self-limiting.

Collapse, fragmentation, and why clouds make clusters

Once a region exceeds its Jeans mass, it falls inward. Early on the collapse is roughly isothermal: the gas is optically thin, so compressional heating is radiated away by dust and the temperature stays near 10 kelvin.

Now look again at the scaling. If T stays constant while rho rises, M_J falls as the inverse square root of density. So as the cloud collapses, smaller and smaller sub-regions become individually unstable, and the cloud breaks into pieces, which break into pieces. This is fragmentation, and it is why a 10,000 solar mass cloud does not produce one 10,000 solar mass star. It produces a cluster of hundreds or thousands of stars spanning a wide range of masses.

Fragmentation stops when the gas becomes opaque to its own infrared radiation. At that point compression starts raising the temperature, T^1.5 in the numerator beats the density term, and the Jeans mass rises again. That opacity limit is at roughly 0.01 solar masses, and it sets the smallest fragment nature makes.

Inside a collapsing fragment, the centre becomes opaque first and a small hydrostatic core forms, surrounded by material still raining in. That is a protostar. Its luminosity comes from accretion, not fusion: infalling matter converts gravitational energy into heat as it lands. The protostar is buried in dust and invisible in optical light, which is why the study of star formation is largely an infrared and millimetre subject.

Angular momentum, disks, and jets

Every cloud rotates a little, and angular momentum is conserved. Collapse a rotating cloud a millionfold in radius and its rotation rate would rise catastrophically. If nothing intervened, the material would spin so fast that it could never fall onto the star.

Two things intervene. The material flattens into a rotating accretion disk, which allows angular momentum to be transported outward through the disk while mass moves inward. And a fraction of the infalling gas is thrown back out along the rotation axis in a pair of collimated jets, driven by the twisting magnetic field, which carries angular momentum away entirely.

These jets are conspicuous. Where they slam into surrounding gas they produce bright shock-excited knots called Herbig-Haro objects, and hundreds are catalogued. They are one of the few directly visible signposts of an otherwise buried process. The disks themselves are now imaged routinely at millimetre wavelengths, and what they become is the subject of Planetary Science, not this course.

Observationally, young stellar objects are sorted into classes by the shape of their infrared spectrum, which tracks how much cold dusty envelope remains. Class 0 objects are deeply embedded and radiate almost entirely in the far infrared. Class I retain a substantial envelope. Class II are T Tauri stars: optically visible pre-main-sequence stars with disks, strong emission lines, violent variability and powerful magnetic activity. Class III have lost most of the disk and are approaching the main sequence.

The approach to the main sequence

A protostar is not yet a star, because it is not yet fusing hydrogen. It shines by contracting, on the Kelvin-Helmholtz timescale from the previous lesson, and on the HR diagram it moves along a well-defined path.

Low-mass pre-main-sequence stars are fully convective, and a fully convective star of a given mass has a nearly fixed surface temperature. So it descends almost vertically on the diagram, becoming fainter at nearly constant temperature, a path called the Hayashi track. When a radiative core develops, the star turns left and moves at roughly constant luminosity toward higher temperature along the Henyey track, until the core reaches about 15 million kelvin and hydrogen ignites.

The contraction time depends steeply on mass, and the contrast is startling.

Mass (solar)Approximate time from protostar to main sequenceComment
15About 100,000 yearsReaches the main sequence while still embedded in its natal cloud
3About 2 million yearsVisible as a Herbig Ae star
1About 30 to 50 million yearsThe Sun spent this long as a T Tauri star
0.5About 100 million yearsStill contracting when much of a young cluster has settled
0.1Several hundred million yearsApproaches the main sequence very slowly

One consequence is worth noting. In a young cluster, the massive stars are already on the main sequence, and may already have died, while the low-mass stars are still contracting above it. Fitting the pre-main-sequence portion of a young cluster's diagram is one of the standard ways of dating it.

Along the way, at around a million kelvin, a protostar burns whatever deuterium it inherited. Deuterium burning is brief and unimportant energetically, but it draws a line: an object above roughly 13 Jupiter masses can fuse deuterium, and one below cannot, which is the most commonly used boundary between a brown dwarf and a planet.

The two ends of the mass range

Below about 0.075 to 0.08 solar masses, roughly 80 Jupiter masses, the core never gets hot enough for sustained hydrogen fusion because electron degeneracy pressure halts the contraction first. These are brown dwarfs, and they simply cool and fade forever. The first confirmed one, Gliese 229B, was announced in 1995, and the spectral classes L, T and Y were created to accommodate them.

At the top end, radiation pressure eventually resists further accretion, and very massive stars shed mass in fierce winds. Where exactly the limit falls is genuinely uncertain, and published masses for the most extreme stars in clusters such as R136 in the Large Magellanic Cloud have been revised downward as observations improved, because what looks like one enormous star can turn out to be several. The defensible statement is that stars above roughly 150 solar masses exist but are rare and hard to weigh.

Between the two ends, the distribution of masses at birth is called the initial mass function. Edwin Salpeter measured it in 1955 and found that above about half a solar mass the number of stars per unit mass falls as M to the power -2.35. Later work by Kroupa, Chabrier and others found it flattens below that. The practical meaning of the Salpeter slope is stark: for every star born above 10 solar masses, several hundred are born below one solar mass. Massive stars dominate the light and the chemical enrichment of a galaxy while being a tiny minority of its stars.

What matters here: Fragmentation plus the Salpeter slope means star formation makes many small stars and few large ones, and the few large ones do most of the visible work.

Common misconceptions

"Stars form when a cloud gets hot enough." Backwards. Cold clouds collapse; warm ones resist, because pressure support scales with temperature. Heating a molecular cloud is one of the ways star formation is shut down.

"A giant cloud makes a giant star." It makes a cluster. Because the Jeans mass falls as the density rises at constant temperature, a collapsing cloud fragments repeatedly, and the final mass distribution follows the initial mass function rather than the cloud mass.

"Protostars shine by fusion." They shine by accretion and contraction. Fusion begins only at the end of the process, and the transition defines the zero-age main sequence.

"Most of a molecular cloud turns into stars." A few per cent does. Radiation, winds and supernovae from the first massive stars disperse the rest, which is the main reason galaxies still contain gas after billions of years.

Where this leaves us

Stars form in molecular clouds at 10 to 20 kelvin and densities of 100 to a million particles per cubic centimetre, traced by carbon monoxide because molecular hydrogen is invisible, and cooled and shielded by the one per cent of the mass that is dust. Collapse begins where self-gravity beats pressure, which the Jeans criterion quantifies: about 55 solar masses at 10 kelvin and 100 particles per cubic centimetre, falling as the square root of density. Because the Jeans mass drops as an isothermal collapse proceeds, clouds fragment and produce clusters rather than single enormous stars, with fragmentation halting near 0.01 solar masses when the gas becomes opaque. Turbulence, magnetic fields and stellar feedback keep the Galaxy's star formation rate near one to two solar masses per year rather than the hundreds that free fall alone would give. Angular momentum forces the material through a disk and drives bipolar jets, and the resulting protostar shines by accretion and contraction along the Hayashi and Henyey tracks for anywhere from 100,000 years to several hundred million, depending on mass, before hydrogen ignites. Below about 0.08 solar masses it never does. The mass a star ends up with is the single number that determines its future, and the next lesson works out exactly how.

Sources

  1. NASA Science. (n.d.). Stars: How stars form. National Aeronautics and Space Administration. science.nasa.gov
  2. European Southern Observatory. (n.d.). Star formation images and releases. ESO. eso.org
  3. Fraknoi, A., Morrison, D., & Wolff, S. C. (2022). Astronomy 2e, Chapters 20 and 21: Between the stars, and the birth of stars. OpenStax, Rice University. openstax.org
  4. Encyclopaedia Britannica. (n.d.). Interstellar medium. britannica.com
  5. Alves, J. F., Lada, C. J., & Lada, E. A. (2001). Internal structure of a cold dark molecular cloud inferred from the extinction of background starlight. Nature, 409(6817), 159-161.
  6. Salpeter, E. E. (1955). The luminosity function and stellar evolution. Astrophysical Journal, 121, 161-167.
Key terms
Molecular cloud
Cold dense interstellar gas at 10 to 20 K where hydrogen exists as H2 and star formation occurs.
Jeans mass
The critical mass above which a cloud's self-gravity overcomes its internal pressure, scaling as T^1.5 divided by the square root of density.
Free-fall time
The characteristic collapse time of a self-gravitating cloud, given by (3 pi / 32 G rho)^0.5.
Fragmentation
The repeated breakup of a collapsing isothermal cloud into smaller unstable pieces, which produces clusters rather than single massive stars.
Protostar
A hydrostatic core accreting from a surrounding envelope, shining by accretion and contraction rather than fusion.
T Tauri star
An optically visible pre-main-sequence star with a disk, strong emission lines and violent magnetic activity.
Hayashi track
The nearly vertical path a fully convective pre-main-sequence star follows on the HR diagram as it contracts at almost constant temperature.
Brown dwarf
An object below about 0.08 solar masses whose core never reaches hydrogen fusion temperature, supported by electron degeneracy and cooling forever.
Initial mass function
The distribution of stellar masses at birth; Salpeter's 1955 measurement gives numbers falling as M^-2.35 above about half a solar mass.

How Long a Star Lives, and Why Mass Decides

  • Derive the main-sequence lifetime scaling from the mass-luminosity relation and apply it across the stellar mass range.
  • Explain why the simple M^-2.5 rule breaks down at both ends of the mass range and in which direction.
  • Use turnoff masses to date stellar populations, and state the Sun's past and future timeline with numbers.

Betelgeuse is younger than the last dinosaurs

Betelgeuse formed roughly ten million years ago. The non-avian dinosaurs died out sixty-six million years ago. That star is a sixth the age of the extinction that ended the Cretaceous, and it is already near the end of its life, an evolved supergiant that will explode within a period that is astronomically brief even if it is unlikely to be within your lifetime.

Meanwhile Proxima Centauri, a red dwarf so faint you cannot see it without a telescope from four light years away, has burned perhaps a per cent of its fuel and will still be fusing hydrogen when the universe is a hundred times its present age.

Both statements come from one calculation, and this lesson is that calculation. It is short, it is the single most useful piece of arithmetic in stellar astronomy, and its consequences run through the rest of the course.

The derivation, in three lines

How long a star lasts on the main sequence is the fuel available divided by the rate at which it is consumed.

Fuel is proportional to mass, because a fixed fraction of a star's hydrogen sits in the core where it can burn. Consumption rate is proportional to luminosity, because luminosity is exactly the rate at which the star spends energy. So

t proportional to M / L

From Module 2, the mass-luminosity relation for main-sequence stars is roughly L proportional to M^3.5. Substitute:

t proportional to M / M^3.5 = M^-2.5

Anchor it on the Sun, whose main-sequence lifetime is about 10 billion years, and you have a formula you can use anywhere:

t = 10^10 years x (M / M_sun)^-2.5

Sit with the sign for a moment. More mass means more fuel, which by itself would give a longer life. But luminosity climbs so much faster than mass that the extra fuel is overwhelmed by the extra spending. A star with ten times the Sun's mass has ten times the fuel and burns it more than three thousand times faster.

Why this matters: Mass and lifetime run in opposite directions, and steeply. This one inversion explains why the brightest stars in the sky are all young, why star-forming regions look the way they do, and why we can date star clusters at all.

The numbers, and where the rule starts lying

Run the formula across the mass range and compare it with what detailed stellar models give.

Mass (solar)Simple rule: 10^10 x M^-2.5Detailed models (approximate)Agreement
0.13 x 10^12 yearsAround 10^13 yearsRule too short by a factor of a few
0.55.7 x 10^10 yearsAround 6 x 10^10 yearsGood
1.01.0 x 10^10 years1.0 x 10^10 yearsAnchor point
2.01.8 x 10^9 yearsAround 1.2 x 10^9 yearsFair
5.01.8 x 10^8 yearsAround 1 x 10^8 yearsFair
151.1 x 10^7 yearsAround 1.1 x 10^7 yearsGood, by coincidence
253.2 x 10^6 yearsAround 7 x 10^6 yearsRule too short by a factor of two
603.6 x 10^5 yearsAround 3.5 x 10^6 yearsRule too short by a factor of ten

The rule works well from about half a solar mass to fifteen, and fails in the same direction at both ends: it underestimates real lifetimes. Two separate effects cause this, and both are worth understanding because they are physics, not fudge factors.

The exponent in the mass-luminosity relation is not constant. It is close to 4 below about half a solar mass, close to 3.5 in the middle range, and flattens toward 3 and below for the most massive stars. In very luminous stars the opacity is dominated by electron scattering, which does not depend on temperature, and radiation pressure becomes a significant fraction of the total. Both effects make luminosity climb less steeply with mass than the simple relation predicts, so massive stars are less extravagant than M^3.5 suggests and live proportionally longer.

The fraction of the star that is usable fuel is not constant either. The Sun burns roughly ten per cent of its total hydrogen on the main sequence, because only the inner core is hot enough and there is no mixing to bring in fresh fuel from outside. A star of 20 solar masses has a large convective core that continuously stirs fresh hydrogen into the burning region, raising the usable fraction. And a fully convective M dwarf below 0.35 solar masses mixes its entire interior, so nearly all of its hydrogen is available, which is why the rule badly underestimates the smallest stars.

There is a third effect at the extreme top end that runs the other way: very massive stars lose a substantial fraction of their mass in radiation-driven winds, so the star that reaches the end of the main sequence is lighter than the one that started it. That shortens the star's life relative to its birth mass while complicating any simple formula.

Remember: Use t = 10^10 x M^-2.5 years as a working estimate and expect it to be right to a factor of two over most of the range, with real lifetimes longer than the rule at both extremes.

What the arithmetic forces you to conclude

Four consequences follow directly, and each one is a fact about the sky you can check.

Every O and B star you can see is young. An O5 star lives about a million years. It cannot have travelled far from where it formed. So the bright blue stars trace out current and recent star formation, which is why they outline the spiral arms of galaxies and cluster in regions like Orion. The Orion Belt stars and the Trapezium are a few million years old at most.

No low-mass star has ever died. A star of 0.8 solar masses lives about 17 billion years by the simple rule, and detailed low-metallicity models put the oldest stars near 13 billion. The universe is roughly 13.8 billion years old. Every red dwarf ever formed, anywhere, is still on the main sequence. The lower main sequence is a permanent record of every generation of star formation that has ever occurred, which is one reason astronomers care about M dwarfs out of all proportion to their brightness.

The bright stars in the night sky are a biased sample twice over. Module 2 pointed out that they are magnitude limited and therefore luminous. Now add that luminous means massive, and massive means short-lived. The naked-eye sky is dominated by stars that will not exist in a hundred million years.

Any group of stars formed together carries a clock. If a cluster is 100 million years old, every star more massive than about 5 solar masses has already left the main sequence, and everything below it is still there. Find the mass at the break point and you have the age.

Working a cluster age

Here is the method that Module 5 will use repeatedly. Plot a cluster's colour-magnitude diagram. Find where the main sequence bends away toward the giant branch. Determine the mass of a main-sequence star at that point from its position. That mass has a lifetime, and because every star in the cluster formed at essentially the same time, that lifetime is the cluster's age.

Example 1. A cluster's turnoff sits at a star of about 5 solar masses. Lifetime = 10^10 x 5^-2.5. Since 5^2.5 = 55.9, that gives 1.8 x 10^8 years. Detailed models say closer to 1 x 10^8. So the cluster is roughly 100 million years old, which is about the age of the Pleiades.

Example 2. A cluster's turnoff sits at 1.3 solar masses. 1.3^2.5 = 1.93, so the rule gives 5.2 x 10^9 years. Detailed models for solar composition give roughly 4 billion years. That is the age of M67, an old open cluster whose stars are close enough to the Sun in mass and composition to be a useful comparison sample for solar-type stars.

Example 3, where the rule breaks. A globular cluster's turnoff is at about 0.8 solar masses. The rule gives 10^10 x 0.8^-2.5 = 10^10 x 1.75 = 1.75 x 10^10 years, which is older than the universe. The rule fails here because globular cluster stars are metal poor, typically a hundredth of the solar iron abundance, and low metallicity means low opacity, which makes a star of given mass hotter and more luminous and therefore shorter lived than the solar-composition anchor assumes. Full models with the correct composition give ages near 12 to 13 billion years, comfortably inside the age of the universe. This is a good example of when a scaling rule must be handed back to a proper calculation.

The Sun, before and after

Apply the reasoning to the one star whose history we can check against rocks.

The Sun has been on the main sequence for about 4.6 billion years, dated by radiometric ages of meteorites, and has roughly 5 billion left. But it is not constant. As hydrogen becomes helium, the mean molecular weight of the core rises, and to maintain the same pressure at higher mean particle mass the core must contract and heat, which raises the fusion rate. The Sun therefore brightens steadily.

Models put the young Sun at about 70 per cent of its present luminosity 4.5 billion years ago. That creates the faint young Sun problem: with today's atmosphere, an Earth receiving 70 per cent of the current solar flux would have been frozen solid, and yet there is sedimentary evidence for liquid water going back more than 3.8 billion years. The generally accepted resolution involves a very different early atmosphere with far more greenhouse gases, though the details remain argued over.

Going forward, the Sun brightens by roughly one per cent every 100 million years. That number has an uncomfortable consequence. Long before the Sun leaves the main sequence, rising luminosity will accelerate the weathering of silicate rocks, drawing down atmospheric carbon dioxide below what plants can use, and eventually drive a runaway loss of the oceans. Standard estimates put the end of complex surface life on Earth at roughly one billion years from now, not five. The star has a long future; the biosphere does not share all of it.

The core of it: A main-sequence star is not static. It brightens by tens of per cent over its life, and for the Sun that slow brightening matters more to Earth than the eventual red giant phase does.

Common misconceptions

"Bigger stars live longer because they have more fuel." They have more fuel and spend it far faster. Luminosity rises as roughly the 3.5 power of mass while fuel rises only linearly, so lifetime falls as roughly the 2.5 power.

"A star burns all its hydrogen." The Sun will use only about ten per cent, the fraction in the hot central region, because there is no mechanism to mix the outer nine tenths inward. Only fully convective low-mass stars come close to using it all.

"The Sun will boil the oceans when it becomes a red giant." The oceans are expected to be gone roughly a billion years from now, from ordinary main-sequence brightening, some four billion years before the red giant phase begins.

"Main-sequence lifetime formulas are exact." The M^-2.5 rule is a scaling anchored on the Sun. It is good to a factor of two over the middle of the range and worse at both ends, and it says nothing about composition, which matters enough to change a globular cluster's age by billions of years.

The takeaway

Main-sequence lifetime is fuel divided by burn rate, so t is proportional to M divided by L, and with L proportional to M^3.5 that becomes t proportional to M^-2.5. Anchored on the Sun's ten billion years, this gives a million years for the most massive stars and more than 10^12 years for the least massive, a range of six orders of magnitude from a mass range of three. The rule is good to a factor of two from about 0.5 to 15 solar masses and underestimates lifetimes at both ends, because the mass-luminosity exponent flattens for massive stars and because convective mixing raises the usable fuel fraction for both massive stars and fully convective dwarfs. The consequences are direct: every luminous blue star is young and sits near where it formed, no star below about 0.8 solar masses has ever finished its main-sequence life anywhere in the universe, and the mass at a cluster's main-sequence turnoff dates that cluster, provided you use full models rather than the scaling rule when the composition is not solar. The Sun itself is 4.6 billion years into a ten billion year run, brightening about one per cent per hundred million years, which will end Earth's habitability long before the star's own hydrogen does.

Sources

  1. NASA Science. (n.d.). Stars. National Aeronautics and Space Administration. science.nasa.gov
  2. Encyclopaedia Britannica. (n.d.). Stellar evolution. britannica.com
  3. Fraknoi, A., Morrison, D., & Wolff, S. C. (2022). Astronomy 2e, Chapter 22: Stars from adolescence to old age. OpenStax, Rice University. openstax.org
  4. NASA. (n.d.). Imagine the Universe: The life cycles of stars. Goddard Space Flight Center. imagine.gsfc.nasa.gov
  5. Prialnik, D. (2009). An introduction to the theory of stellar structure and evolution (2nd ed.). Cambridge University Press.
Key terms
Main-sequence lifetime
The time a star spends fusing hydrogen in its core, approximately 10^10 years times (M / M_sun)^-2.5.
Fuel fraction
The proportion of a star's hydrogen that reaches temperatures high enough to fuse, about ten per cent for the Sun and near one hundred per cent for fully convective dwarfs.
Turnoff mass
The mass of the most massive star still on the main sequence in a coeval population, whose lifetime equals the population's age.
Faint young Sun problem
The conflict between a Sun 30 per cent fainter 4.5 billion years ago and geological evidence for liquid water on the early Earth.
Mean molecular weight
Average particle mass in the gas; it rises as hydrogen becomes helium, forcing the core to contract and the star to brighten.
Radiation-driven wind
Mass loss from luminous stars driven by radiation pressure on spectral lines, significant enough in the most massive stars to change their evolution.

Module 4: How Stars End

The endgames. Red giants and white dwarfs for stars like the Sun, core collapse and nucleosynthesis for the massive ones, and the neutron stars and black holes left behind, with the evidence that they are really there.

Red Giants, Planetary Nebulae, and White Dwarfs

  • Explain why exhausting core hydrogen makes a star expand rather than shrink, and trace the giant branch and asymptotic giant branch.
  • Describe electron degeneracy pressure, the helium flash and the Chandrasekhar limit, and state what each explains.
  • Lay out the Sun's future as a numbered timeline with masses, radii and durations.

A wobble, a point of light, and an impossible density

In 1844 Friedrich Bessel, the same astronomer who measured the first stellar parallax, noticed that Sirius does not move across the sky in a straight line. Its proper motion wobbled with a period of about fifty years. He concluded that Sirius was being pulled by an unseen companion, and he made the same argument for Procyon. Eighteen years later, on 31 January 1862, Alvan Graham Clark was testing a new 18.5 inch lens in Cambridgeport, Massachusetts, pointed it at Sirius, and saw a faint point beside it, exactly where Bessel's calculation put it.

The companion was odd. It is about ten thousand times fainter than Sirius A in visible light. The obvious reading is that it must be cool and red, like other faint stars. In 1915 Walter Adams took its spectrum at Mount Wilson and found the opposite: Sirius B is hot, with a surface temperature we now put near 25,000 kelvin.

Run the Stefan-Boltzmann arithmetic from Module 1 and the conclusion is forced. A star that hot and that faint must be tiny: about 0.0084 solar radii, roughly 5,800 kilometres, slightly smaller than Earth. Orbital dynamics give it a mass of about 1.02 times the Sun's. Divide one by the other and the mean density comes to about 2.4 x 10^9 kilograms per cubic metre, two and a half tonnes in a sugar cube. Eddington wrote that the message from the star was absurd and that the sensible response was to stop dismissing absurdities.

Sirius B is a white dwarf, and it is what the Sun will be in about seven billion years. This lesson is the route from here to there.

The point: Every claim in this lesson about the far future of the Sun is anchored on objects we can already see, because the universe has been running this process for billions of years and the products are all around us.

Running out, and the counterintuitive response

A star leaves the main sequence when its core hydrogen is gone. What remains at the centre is helium, at a temperature far too low for helium fusion, which needs about 100 million kelvin against the roughly 15 million available.

With no nuclear energy source, the helium core cannot maintain its pressure against the weight above, so it contracts and heats. Contraction releases gravitational energy, and the hydrogen in a shell immediately surrounding the core is compressed and heated as well, so hydrogen fusion continues in that shell. Shell burning is more vigorous than core burning was, because the shell sits at higher temperature and density.

Now the surprise. The star's outer envelope expands enormously. The core is shrinking and heating while the envelope is swelling and cooling, the two moving in opposite directions at the same time. This behaviour is sometimes called the mirror principle, and it is worth being honest about it: it emerges reliably from the equations of stellar structure and there is no short physical argument for it that survives scrutiny. The usual hand-waving version, that the shell's extra energy pushes the envelope out, is at best incomplete. What you should take away is the observed and computed fact, not a false sense that it is obvious.

The star climbs the red giant branch. For a one solar mass star the endpoint of that climb is a radius near 170 solar radii, about 0.8 astronomical units, and a luminosity of a couple of thousand times the present Sun, with a surface temperature down near 3,000 kelvin. On the HR diagram the star has moved from the middle of the main sequence far up and to the right.

Ascending the giant branch also strips material away. A red giant's surface gravity is minute, since the same mass is spread over a radius a hundred times larger, and pulsations plus radiation pressure on dust drive a steady wind. A star arrives at the top of the giant branch noticeably lighter than it started.

Degeneracy, and the flash it causes

While the envelope swells, the helium core is being squeezed toward a state where ordinary gas physics stops applying. Electrons are fermions, and the Pauli exclusion principle forbids two of them from occupying the same quantum state. Compress a gas enough and the low-energy states fill up, forcing electrons into high-momentum states whether or not the gas is hot. The resulting electron degeneracy pressure depends on density and hardly at all on temperature.

That is precisely the broken thermostat from Module 2. In a degenerate core, adding heat does not cause expansion and cooling.

So when the core of a low-mass star finally reaches about 100 million kelvin, with a core mass near 0.45 to 0.5 solar masses, helium ignites through the triple-alpha process under degenerate conditions and the reaction runs away. Energy generation spikes to something like 10^11 solar luminosities for a few seconds, comparable to an entire galaxy, in an event called the helium flash.

Nothing visible happens. All of that energy is absorbed by the overlying envelope, and its only external effect is that the star readjusts over the following thousands of years. The flash lifts the degeneracy: once thermal energy exceeds the degeneracy energy the gas becomes an ordinary gas again, the core expands, and helium burning settles into a stable, thermostatted state.

Above roughly 2 solar masses this does not happen at all, because the core reaches 100 million kelvin before it becomes degenerate, and helium ignites quietly. The helium flash is a low-mass phenomenon and a direct consequence of degeneracy.

In short: Degeneracy pressure is what happens when quantum mechanics runs out of room, and it removes the feedback that keeps ordinary stellar burning gentle.

The horizontal branch and the second ascent

With helium burning steadily in the core to carbon and oxygen, and hydrogen still burning in a shell, the star settles into a phase lasting roughly 100 million years for a solar-mass star, about one per cent of its main-sequence life. In globular clusters these stars form a distinctive horizontal branch on the colour-magnitude diagram, and the variable stars of Module 5 live in a narrow strip crossing it.

When core helium runs out, history repeats one level up. The carbon-oxygen core contracts and becomes degenerate; helium burns in a shell around it; hydrogen burns in a shell around that; and the envelope expands again, further than before. This is the asymptotic giant branch, or AGB, and it is where several important things happen.

The two shells do not burn steadily. The helium shell ignites in pulses every 10,000 to 100,000 years, each one briefly dumping energy and driving convection that reaches down into processed material and dredges it to the surface. This is third dredge-up, and it is how carbon manufactured in the star's interior arrives at the surface, turning some AGB stars into carbon stars whose spectra show more carbon than oxygen.

The pulses also drive the s-process: slow neutron capture. Free neutrons, released by side reactions on carbon-13 and neon-22, are captured by iron-peak nuclei one at a time, slowly enough that unstable products decay before the next capture arrives. Over many pulses this builds elements up the periodic table along the valley of stability. Roughly half the elements heavier than iron in the universe were made this way, in the interiors of AGB stars, and the other half by the rapid process covered in the next lesson. Strontium, barium, lead and most of the zirconium in the world came from this route.

Finally, the AGB is where the star loses most of its mass. Pulsation lifts material to where dust can condense, radiation pressure pushes on the dust, and the dust drags the gas along. Loss rates reach 10^-4 solar masses per year in a final superwind, which means a star can shed several tenths of a solar mass in a few tens of thousands of years. A one solar mass star ends this phase with a bare core of about 0.54 solar masses and no envelope left.

The planetary nebula, which is neither

William Herschel coined the term planetary nebula in 1785 because these objects looked like small greenish disks through his telescope, resembling Uranus, which he had discovered four years earlier. The name is doubly wrong: there are no planets involved, and the objects are not diffuse clouds of the sort the word nebula originally described. It has stuck for 240 years.

What happens is that the exposed core, now at 30,000 kelvin and rising toward 100,000 or more, floods the departing envelope with ultraviolet light. The gas is ionised and fluoresces. The green colour Herschel saw comes mostly from a forbidden transition of doubly ionised oxygen at 500.7 nanometres, a line so unfamiliar to nineteenth century chemists that it was briefly attributed to a new element called nebulium, before Ira Bowen identified it in 1927 as ordinary oxygen radiating in conditions no laboratory could then reproduce.

Planetary nebulae are short-lived. The gas disperses in something like ten to twenty thousand years, which is why only a few thousand are known in the Galaxy despite most stars passing through the phase. Their shapes are wildly varied, from the near-circular ring of M57 to bipolar hourglasses and multi-lobed structures, and there is good evidence that binary companions shape many of them, since spherical mass loss struggles to produce the observed asymmetries.

The white dwarf, and the limit that made a career

What remains is a carbon-oxygen sphere of about 0.6 solar masses in a volume the size of Earth, supported entirely by electron degeneracy pressure, with no nuclear reactions at all. It has no energy source. It simply radiates away its stored heat, cooling for the rest of time.

White dwarfs have a property that catches everyone the first time: more massive white dwarfs are smaller. The radius scales roughly as the inverse cube root of the mass. Extra mass compresses the star, degeneracy pressure responds by rising steeply with density, and the equilibrium radius shrinks.

Extrapolate that and something must give. In 1930, on the voyage from India to England, the nineteen-year-old Subrahmanyan Chandrasekhar did the calculation including special relativity, and found that as the star shrinks the electrons approach the speed of light, which softens the equation of state. Beyond a critical mass, degeneracy pressure cannot win at any density. That critical mass is

M_Ch = 1.44 solar masses (for carbon-oxygen composition)

Eddington, then the most influential astrophysicist alive, rejected the result publicly at a Royal Astronomical Society meeting in 1935, dismissing the relativistic treatment as a mathematical trick. The physics community sided with Chandrasekhar within a few years. He received the Nobel Prize in 1983, nearly half a century later, and the limit that bears his name governs both the maximum mass of a white dwarf and the mechanism of Type Ia supernovae in the next lesson.

Cooling is slow. A newly exposed white dwarf can exceed 100,000 kelvin; the coolest ones found in the Galactic disk are near 4,000 kelvin. Because the cooling rate is calculable, the faintest white dwarfs in a population act as a clock. The observed drop-off in the white dwarf luminosity function of the Galactic disk gives an age of roughly 8 to 10 billion years, which is an entirely independent check on ages derived from main-sequence turnoffs. No white dwarf anywhere has yet had time to cool to blackness.

Worth holding on to: A white dwarf is not a star that has gone out. It is a star that has stopped and is now a slowly cooling ember whose temperature records how long ago it stopped.

The Sun's timeline, with numbers

Time from nowPhaseRadiusLuminosityDuration
NowMain sequence, 4.6 billion years in1 solar1 solarAbout 5 billion years remaining
About 1 billion yearsStill main sequence, but Earth's oceans lostAbout 1.1 solarAbout 1.1 solarOngoing
About 5 billion yearsSubgiant, hydrogen shell burning beginsGrowing past 2 solarA few solarAround 1 billion years
About 6.5 billion yearsRed giant branch tip, helium flashAbout 170 solarAbout 2,000 solarFlash lasts seconds
Following thatCore helium burningAbout 10 solarAbout 50 solarAround 100 million years
ThenAsymptotic giant branch, thermal pulses, superwindOver 200 solarThousands solarA few million years
ThenPlanetary nebulaNebula spans light yearsCore over 30,000 K10,000 to 20,000 years
ThereafterWhite dwarf, about 0.54 solar massesAbout 0.01 solarFalling foreverIndefinite

Whether Earth itself survives is genuinely uncertain and depends on a competition. The Sun loses about a third of its mass, which lets the planets spiral outward; tidal interaction with the vastly expanded envelope pulls them back in. Mercury and Venus are lost. Earth is on the boundary, and published models disagree. Either way the surface is sterilised billions of years earlier.

Common misconceptions

"The Sun will explode." It will not. Stars below about 8 solar masses end as white dwarfs after shedding their envelopes gently. Supernovae require either a much more massive star or a white dwarf pushed over the Chandrasekhar limit by a companion.

"A red giant is hotter because it is brighter." Its surface is cooler than the Sun's, near 3,000 kelvin. The luminosity comes from area: 170 solar radii is nearly 30,000 times the surface area.

"The helium flash blows the star apart." Nothing is visible from outside. The energy is absorbed by the envelope, and the flash's real effect is to lift degeneracy and let helium burning become stable.

"White dwarfs get bigger as they gain mass." They get smaller, roughly as the inverse cube root of mass, until the Chandrasekhar limit at 1.44 solar masses where degeneracy support fails entirely.

"Planetary nebulae have something to do with planets." Only Herschel's 1785 telescope and the resemblance to Uranus. The term is a historical accident preserved by inertia.

Summing up

When core hydrogen runs out, the helium core contracts and heats while a hydrogen shell ignites around it and the envelope expands, carrying the star up the red giant branch to about 170 solar radii and 2,000 solar luminosities. In stars below about 2 solar masses the core becomes degenerate on the way, so when helium finally ignites at 100 million kelvin it does so in a runaway helium flash that releases roughly 10^11 solar luminosities for seconds and is entirely invisible from outside. Stable core helium burning follows on the horizontal branch for about 100 million years, then the pattern repeats on the asymptotic giant branch with two burning shells, thermal pulses, dredge-up of carbon, s-process production of about half the elements heavier than iron, and mass loss reaching 10^-4 solar masses per year. The ejected envelope is ionised by the exposed core and glows for ten to twenty thousand years as a planetary nebula, misnamed by Herschel in 1785. What is left is a white dwarf of roughly 0.6 solar masses in an Earth-sized volume, held up by electron degeneracy pressure, shrinking as it gains mass, and unable to exist at all above the Chandrasekhar limit of 1.44 solar masses. Sirius B, seen by Clark in 1862 and understood only after 1915, is one of them. What happens to stars too massive for this route is the next lesson.

Sources

  1. NASA. (n.d.). Imagine the Universe: The life cycles of stars. Goddard Space Flight Center. imagine.gsfc.nasa.gov
  2. Encyclopaedia Britannica. (n.d.). White dwarf star. britannica.com
  3. Fraknoi, A., Morrison, D., & Wolff, S. C. (2022). Astronomy 2e, Chapters 22 and 23: Stars from adolescence to old age, and the death of stars. OpenStax, Rice University. openstax.org
  4. NASA Science. (n.d.). Hubble mission: Planetary nebulae. National Aeronautics and Space Administration. science.nasa.gov
  5. Chandrasekhar, S. (1931). The maximum mass of ideal white dwarfs. Astrophysical Journal, 74, 81-82.
  6. Wali, K. C. (1991). Chandra: A biography of S. Chandrasekhar. University of Chicago Press.
Key terms
Shell burning
Fusion in a layer surrounding an exhausted core, which is more vigorous than the earlier core burning and drives the envelope's expansion.
Red giant branch
The evolutionary track on which a post-main-sequence star expands to tens or hundreds of solar radii while its surface cools.
Electron degeneracy pressure
Pressure arising from the Pauli exclusion principle, depending on density but almost not at all on temperature.
Helium flash
The runaway ignition of helium in a degenerate core of a low-mass star, releasing about 10^11 solar luminosities for seconds with no external sign.
Asymptotic giant branch
The late phase with helium and hydrogen burning shells, thermal pulses, dredge-up and heavy mass loss.
s-process
Slow neutron capture in AGB stars, which builds roughly half the elements heavier than iron along the valley of stability.
Planetary nebula
The ionised, glowing ejected envelope of a dying intermediate-mass star, visible for ten to twenty thousand years.
Chandrasekhar limit
The maximum mass of a white dwarf, about 1.44 solar masses, above which electron degeneracy pressure cannot resist gravity.
White dwarf cooling sequence
The steadily fading population of white dwarfs whose faint end dates a stellar population independently of turnoff ages.

Supernovae and the Origin of the Elements

  • Explain why fusion stops at iron and describe the collapse and bounce of a massive stellar core.
  • Distinguish core-collapse from thermonuclear supernovae by their spectra, progenitors, energetics and remnants.
  • Attribute the elements of the periodic table to the specific astrophysical processes that produce them.

Twenty-four neutrinos, and then a new star

At 07:35 Universal Time on 23 February 1987, three particle detectors on three continents recorded a burst of events within about thirteen seconds of each other. Kamiokande-II, in a mine in Japan, logged around a dozen. The Irvine-Michigan-Brookhaven detector in an Ohio salt mine logged eight. The Baksan detector in the Caucasus logged five. Individually each would have been noise. Together, arriving simultaneously in instruments built to look for proton decay, they were a signal.

Hours later, at Las Campanas Observatory in Chile, Ian Shelton developed a photographic plate of the Large Magellanic Cloud and found a star that had not been there the night before. Oscar Duhalde, a telescope operator at the same site, had already noticed it by eye. So had Albert Jones, an amateur observer in New Zealand. Supernova 1987A, at about 51 kiloparsecs, was the closest supernova observed since the invention of the telescope.

Two facts from that night changed the subject permanently. First, archival plates identified the star that had exploded: Sanduleak -69 202, a blue supergiant of roughly 20 solar masses. Everyone had expected a red supergiant, and the models had to accommodate a blue one. Second, and far more important, the neutrino burst arrived before the light and carried the signature the theory had predicted: a core collapse releasing the overwhelming majority of its energy as neutrinos over a few seconds. A mechanism worked out on paper in the 1930s and refined for fifty years was confirmed by two dozen particle detections.

Key idea: A core-collapse supernova is a neutrino event that happens to be visible. About 99 per cent of the energy leaves as neutrinos, about 1 per cent as kinetic energy of the ejecta, and roughly one part in ten thousand as light.

Why the fusion chain stops

Fusion releases energy because the product is more tightly bound than the reactants. Plot binding energy per nucleon against mass number and you get a curve that rises steeply from hydrogen, climbs through helium and carbon, and peaks near mass 56 to 62, with iron-56 and nickel-62 both close to 8.8 MeV per nucleon. Past that peak the curve turns down. Fusing iron into anything heavier costs energy rather than releasing it.

A massive star therefore works its way up the curve in stages, each one hotter, faster and less profitable than the last. For a star of about 25 solar masses the sequence runs roughly like this.

StageCore temperatureApproximate durationMain products
Hydrogen burning40 million KAbout 7 million yearsHelium
Helium burning200 million KAbout 700,000 yearsCarbon, oxygen
Carbon burning800 million KAround 600 yearsNeon, sodium, magnesium
Neon burning1.6 billion KAround a yearOxygen, magnesium
Oxygen burning1.8 billion KAround six monthsSilicon, sulfur
Silicon burning2.5 billion KAbout a dayIron-peak nuclei

Read the duration column again. The star spends millions of years on hydrogen and one day on silicon. The acceleration has a specific cause: above about a billion kelvin the dominant energy loss is not photons diffusing out over thousands of years but neutrinos, produced by thermal processes in the core, which leave immediately. The star is pouring energy directly into space, and it must burn fuel faster and faster to keep up. The last stages are essentially a race the star cannot win.

The result is an onion structure: an iron core surrounded by shells of silicon, oxygen, neon, carbon, helium and hydrogen, each still burning at its base. The iron core grows as silicon burning feeds it, and it is supported by electron degeneracy, exactly like a white dwarf. So it has the same limit. When it reaches roughly 1.4 solar masses, compressed into a volume about the size of Earth, it cannot hold.

One second

Two processes remove the support at once. At these temperatures, gamma rays are energetic enough to break iron nuclei apart into helium and free neutrons, a process called photodisintegration, which consumes energy and lowers the pressure. Simultaneously the density is high enough that electrons are forced into protons, producing neutrons and neutrinos, which removes the electrons that were providing the degeneracy pressure.

The core collapses. The inner portion falls inward at speeds approaching a quarter of the speed of light, and the collapse takes well under a second. It halts when the central density reaches nuclear density, around 2.3 x 10^17 kilograms per cubic metre, where the strong force makes the material abruptly stiff. The infalling inner core overshoots slightly, rebounds, and drives a shock wave outward.

The shock does not immediately succeed. It runs into the still-infalling outer core, loses energy photodisintegrating iron, and stalls within milliseconds. What revives it, in the standard picture developed by Hans Bethe, James Wilson and others in the 1980s, is neutrinos. The newborn neutron star is radiating something like 3 x 10^46 joules of neutrinos in a few seconds, and although neutrinos interact feebly, the densities behind the stalled shock are high enough that a per cent or so of that energy is deposited, reviving the shock and blowing off the star's outer layers. Multi-dimensional simulations show that convection and instabilities behind the shock are essential to making this work, which is one reason the problem resisted a full solution for decades.

The energy budget is worth writing out, because it is so lopsided.

ChannelEnergyShare
NeutrinosAbout 3 x 10^46 JAbout 99 per cent
Kinetic energy of ejectaAbout 10^44 JAbout 1 per cent
Visible light over monthsAbout 10^42 JAbout 0.003 per cent

The visible supernova, which can briefly outshine its entire host galaxy, is an afterthought in the energy accounting. That is why the 1987A neutrino detection mattered so much: it measured the main event.

What remains at the centre is a neutron star, or if the core is massive enough, a black hole. For SN 1987A no compact object was seen for decades, and its absence was a genuine puzzle. Observations with ALMA and Chandra found increasing evidence for a hot compact source in the debris, and in 2024 a team reported JWST spectroscopy showing emission from highly ionised argon and sulfur at the centre, consistent with ionisation by a neutron star. Treat that as a recent result with a date on it rather than a settled textbook fact.

Two kinds of explosion under one name

Supernovae were classified before anyone understood them, purely by spectrum. Type I shows no hydrogen lines; Type II does. That division turns out to cut across the physics, so the modern scheme requires both the observational label and the mechanism.

TypeSpectral signatureMechanismProgenitorRemnant
IaNo hydrogen, strong siliconThermonuclear disruptionCarbon-oxygen white dwarf in a binaryNothing
IbNo hydrogen, helium presentCore collapseMassive star stripped of its hydrogen envelopeNeutron star or black hole
IcNeither hydrogen nor heliumCore collapseMassive star stripped of hydrogen and heliumNeutron star or black hole
IIHydrogen presentCore collapseRed or blue supergiant with envelope intactNeutron star or black hole

So Ib, Ic and II are the same physical event seen through progressively more stripped envelopes, while Ia is something else entirely. The observational division is historical, and you have to hold both layers in mind when reading the literature.

Type Ia: the white dwarf that cannot say no

A carbon-oxygen white dwarf sitting alone will cool forever. Put it in a binary with a companion that can transfer mass, and it approaches the Chandrasekhar limit. As it does, the central density and temperature rise until carbon ignites, and because the material is degenerate the thermostat is absent. The burning front sweeps through the star in seconds, fusing much of it to nickel-56, and the energy released exceeds the star's gravitational binding energy. The white dwarf is completely destroyed. There is no remnant at all.

The light curve is powered not by the explosion itself but by radioactivity: nickel-56 decays to cobalt-56 with a half-life of about six days, and cobalt-56 to iron-56 with a half-life of about 77 days. The characteristic decline of a Type Ia light curve follows those half-lives, which is a satisfying piece of evidence that the model is right.

Because the explosion happens at a nearly fixed mass by a nearly fixed mechanism, Type Ia supernovae have nearly the same peak luminosity: absolute magnitude about -19.3, which is roughly five billion times the Sun's output. Mark Phillips showed in 1993 that the residual scatter is not random: intrinsically brighter events decline more slowly, and correcting for that relation standardises them to about 0.1 to 0.15 magnitudes. That precision made them the tool that measured cosmic acceleration in 1998 and earned the 2011 Nobel Prize, which is Cosmology's subject rather than ours.

One honest caveat. Which binary configuration actually produces Type Ia events is still argued. In the single-degenerate picture the companion is an ordinary star transferring mass; in the double-degenerate picture two white dwarfs merge. Searches for surviving companions and for signatures of accretion have not settled it, and the field currently accepts that both channels may contribute. That the standardisation works empirically does not mean the progenitor question is closed.

Bottom line: Type Ia supernovae are precise enough to have measured the expansion history of the universe while their progenitor systems remain unidentified. Precision and understanding are different things.

Where the elements come from

Assemble everything and you can attribute the periodic table.

ElementsSourceProcess
Hydrogen, most helium, a little lithiumThe first few minutes of the universeBig Bang nucleosynthesis
Lithium, beryllium, boronInterstellar mediumCosmic ray spallation of heavier nuclei
Carbon, nitrogen, much oxygenLow and intermediate mass starsHelium burning, CNO processing, dredge-up
Oxygen through calciumMassive stars and their supernovaeAdvanced fusion stages and explosive burning
Iron-peak elementsType Ia supernovae mostly, core collapse partlyNickel-56 production and decay
About half of everything past ironAsymptotic giant branch starss-process, slow neutron capture
The other half, including gold, platinum, uraniumNeutron star mergers, possibly rare supernovaer-process, rapid neutron capture

The last row was theory until 17 August 2017. On that date the LIGO and Virgo detectors recorded gravitational waves from the merger of two neutron stars, a gamma-ray burst followed 1.7 seconds later, and telescopes across the world found the optical counterpart in the galaxy NGC 4993. The fading source, called a kilonova, reddened over days in exactly the way models predicted for a cloud of freshly synthesised heavy elements with high opacity. Estimates of the mass of r-process material ejected run to a few hundredths of a solar mass, thousands of Earth masses of the heaviest elements, made in a fraction of a second.

The r-process needs a neutron flux so intense that a nucleus captures many neutrons before it has time to beta decay, which requires conditions that essentially only occur when neutron star material is torn apart. Core-collapse supernovae were the long-standing candidate and may still contribute in rare magnetorotational cases, but the 2017 event is the first direct evidence of the site.

The practical version of this table: the calcium in your bones was made in massive stars, the carbon in every organic molecule in your body came largely from dying intermediate-mass stars, the iron in your blood came mostly from exploding white dwarfs, and the iodine in your thyroid and any gold you own were made when two neutron stars collided.

So what?: Chemical abundance is not a fixed property of the universe. It has been built up over thirteen billion years by specific, identifiable stellar processes running at different rates.

How often, and what we have watched

Estimates of the Galactic core-collapse rate cluster around two to three per century, based on counting massive stars, supernova remnants and rates in similar galaxies. We have not seen one with the naked eye since Kepler's star of 1604. Tycho's star of 1572 preceded it, and the guest star of 1054 recorded by Chinese astronomers left the Crab Nebula. The remnant Cassiopeia A dates from around 1680 and was barely noticed, if at all.

The discrepancy between two per century and none in four hundred years is not a paradox. Most of the Galaxy's massive stars sit in the disk behind dozens of magnitudes of dust extinction. We are almost certainly missing them optically. Neutrino detectors, which dust does not affect, now stand watch for exactly this reason, and a network exists to alert optical observers within minutes if a burst is detected. If a Galactic supernova occurs tomorrow, the neutrinos will announce it hours before the light arrives, as they did in 1987.

Common misconceptions

"A supernova is a very large explosion of the whole star." For core collapse it is an implosion that rebounds. The core falls in and becomes a neutron star; the envelope is ejected by a shock revived by neutrinos. Only Type Ia genuinely destroys the whole object.

"Iron cannot fuse because it is too heavy." It can fuse; the reaction simply costs energy instead of releasing it, because iron sits at the peak of the binding energy per nucleon curve. That is a thermodynamic statement, not a mechanical one.

"Supernovae make all the heavy elements." They make the intermediate-mass elements and a large share of the iron peak. Roughly half the elements beyond iron come from slow neutron capture in AGB stars, and much of the rest from neutron star mergers.

"The Crab Nebula supernova was seen only in China." The 1054 event appears in Chinese and Japanese records and in Arabic sources, and possible references exist elsewhere. Its absence from European chronicles is a genuine and much-discussed gap rather than proof of nothing.

"Type Ia supernovae are perfectly standard candles." They are standardisable, not standard. The Phillips relation corrects a real intrinsic spread, and residual scatter of about 0.1 magnitudes remains, which is the dominant systematic in some cosmological measurements.

What to remember

Fusion pays out only up to the binding energy peak near iron, so a massive star burns through hydrogen, helium, carbon, neon, oxygen and silicon in stages that shorten from millions of years to a single day as neutrino losses take over the energy budget. The resulting iron core is supported by electron degeneracy and collapses once it passes about 1.4 solar masses, falling at a quarter of light speed, bouncing at nuclear density, and driving a shock that stalls until neutrino heating revives it. Ninety-nine per cent of the 3 x 10^46 joules released leaves as neutrinos, which is why two dozen detections on 23 February 1987 confirmed the mechanism more convincingly than the visible supernova that followed. Types Ib, Ic and II are core collapse seen through progressively stripped envelopes; Type Ia is a white dwarf pushed to the Chandrasekhar limit that detonates and leaves nothing, powered afterwards by nickel-56 decay and standardisable through the Phillips relation despite an unresolved progenitor question. Between them, plus AGB stars and neutron star mergers, these processes built every element heavier than lithium, and the 2017 detection of a kilonova identified the site of the r-process after decades of inference. What is left behind at the centre of a core collapse is the subject of the next lesson.

Sources

  1. NASA Science. (n.d.). Supernovae. National Aeronautics and Space Administration. science.nasa.gov
  2. Encyclopaedia Britannica. (n.d.). Supernova. britannica.com
  3. Chandra X-ray Observatory. (n.d.). Supernovas and supernova remnants. Harvard-Smithsonian Center for Astrophysics. chandra.harvard.edu
  4. LIGO Scientific Collaboration. (n.d.). GW170817: Observation of gravitational waves from a binary neutron star merger. LIGO Caltech. ligo.caltech.edu
  5. Fraknoi, A., Morrison, D., & Wolff, S. C. (2022). Astronomy 2e, Chapter 23: The death of stars. OpenStax, Rice University. openstax.org
  6. Phillips, M. M. (1993). The absolute magnitudes of Type Ia supernovae. Astrophysical Journal Letters, 413, L105-L108.
Key terms
Binding energy per nucleon
The energy holding each nucleon in a nucleus, peaking near mass 56 to 62 at about 8.8 MeV, which is why fusion stops at iron.
Photodisintegration
The breakup of heavy nuclei by energetic gamma rays at billions of kelvin, which removes pressure support from a collapsing core.
Core bounce
The abrupt halt of collapse when the inner core reaches nuclear density and stiffens, sending a shock wave outward.
Delayed neutrino mechanism
The revival of a stalled supernova shock by absorption of about one per cent of the neutrino energy emitted by the newborn neutron star.
Type Ia supernova
The thermonuclear disruption of a carbon-oxygen white dwarf pushed toward the Chandrasekhar limit, leaving no remnant.
Phillips relation
The 1993 finding that intrinsically brighter Type Ia supernovae decline more slowly, which standardises them as distance indicators.
r-process
Rapid neutron capture producing roughly half the elements heavier than iron, with neutron star mergers confirmed as a site in 2017.
Kilonova
The reddening optical and infrared transient powered by radioactive decay of freshly made heavy elements after a neutron star merger.

Neutron Stars, Pulsars, and Black Holes

  • Describe the physical properties of neutron stars and explain how rotation and magnetic field produce pulsar emission.
  • Compute a Schwarzschild radius and state what an event horizon is and is not.
  • Assemble the observational evidence for stellar-mass and supermassive black holes, including gravitational waves and the Event Horizon Telescope images.

A bit of scruff on a chart recorder

In 1967 Jocelyn Bell was a graduate student at Cambridge who had spent two years helping to build a radio telescope out of poles, wire and 2,048 dipole antennas spread over four and a half acres of field. It produced about thirty metres of chart paper a day, and she analysed it by eye. In August she noticed a smudge she called a bit of scruff, appearing in the same patch of sky and not behaving like the interference she had learned to recognise.

On 28 November 1967 she ran the recorder fast enough to resolve it. The signal was a train of pulses, each about 0.3 seconds long, repeating every 1.3373 seconds, steady to a precision better than any clock in the observatory. Nothing astronomical was known to do that. The group briefly labelled the source LGM-1, for little green men, and then found a second source in a different part of the sky, which settled the question: whatever this was, it was natural and there was more than one.

The explanation came quickly. Walter Baade and Fritz Zwicky had proposed in 1934 that a supernova leaves behind a star made of neutrons. Thomas Gold and Franco Pacini argued in 1968 that a rapidly rotating, strongly magnetised neutron star would sweep a beam of radio emission across the sky like a lighthouse. Within a year a pulsar was found at the centre of the Crab Nebula, the remnant of the supernova of 1054, spinning thirty times a second, and the connection between supernovae, neutron stars and pulsars was complete.

The 1974 Nobel Prize in Physics went to Antony Hewish and Martin Ryle. Bell Burnell was not included, a decision that has been argued over ever since and that she herself has discussed with more equanimity than most observers manage.

What matters here: Neutron stars were predicted in 1934 as an exotic possibility and found in 1967 by someone reading chart paper carefully. Both halves of that sentence are how physics usually works.

What survives a core collapse

Take about 1.4 solar masses and compress it into a sphere 20 to 24 kilometres across. The numbers that follow are all consequences of that one sentence.

The mean density exceeds 10^17 kilograms per cubic metre, comparable to an atomic nucleus. A cubic centimetre weighs a hundred million tonnes.

Surface gravity is g = GM/R^2 = (6.674 x 10^-11 x 2.8 x 10^30) / (10^4)^2 = 1.9 x 10^12 metres per second squared, about two hundred billion times Earth's.

Escape velocity is v = (2GM/R)^0.5 = (3.72 x 10^20 / 10^4)^0.5 = 1.9 x 10^8 metres per second, roughly 0.6 times the speed of light. A neutron star is close enough to being a black hole that general relativity is required to describe it correctly.

Rotation follows from angular momentum conservation. Collapse the Sun, which turns once in about 25 days, from 7 x 10^8 metres to 10^4 metres. Angular velocity goes as 1/R^2, so it rises by a factor of (7 x 10^4)^2 = 4.9 x 10^9, giving a period of about 4 x 10^-4 seconds. Real neutron stars are born spinning at tens of milliseconds rather than fractions of one, because the collapsing core is not the whole star and considerable angular momentum is lost along the way. But the mechanism is unambiguous: the spin is inherited and concentrated, not generated.

Magnetic field follows the same logic. Magnetic flux through a surface is conserved in a highly conducting plasma, so field strength goes as 1/R^2 and gains the same factor of five billion. Observed pulsar fields of order 10^8 tesla require the progenitor core to have been more strongly magnetised than the Sun's surface, which is expected, and magnetars reach 10^10 to 10^11 tesla, the strongest magnetic fields known in the universe.

What lies inside is genuinely unsettled. The outer crust is a lattice of nuclei; below that neutrons drip out of nuclei and form a superfluid; and the core may contain hyperons, a condensate of mesons, or deconfined quark matter. Which of these is right determines the equation of state, and therefore the maximum mass a neutron star can have, currently believed to be somewhere around 2.2 to 2.5 solar masses. The heaviest well-measured neutron stars come in near 2.1 solar masses, which already rules out the softest proposed equations of state. This is one of the few places where astronomical observation directly constrains nuclear physics that no accelerator can reach.

The lighthouse, and the energy budget that proves it

A pulsar's beam comes from its magnetic poles, which in general are not aligned with its rotation axis. Charged particles are accelerated along the field lines and radiate in a narrow cone. If that cone sweeps across Earth, you see a pulse once per rotation. If it does not, you see nothing, which means the pulsars we detect are a minority of those that exist.

Pulsars slow down, because the rotating magnetic field radiates energy. The Crab pulsar's period of 33 milliseconds lengthens by about 36 nanoseconds per day. That sounds negligible until you convert it into power. The rotational kinetic energy of a neutron star is enormous, and the rate at which the Crab is losing it works out to roughly 5 x 10^31 watts.

Now measure the Crab Nebula. Its total luminosity across the spectrum is about 10^5 solar luminosities, which is about 4 x 10^31 watts. The nebula shines because the pulsar's spin-down feeds it. Two completely independent measurements, a timing measurement of a period derivative and a photometric measurement of a nebula, agree. That agreement is the strongest single piece of evidence that the lighthouse model is correct.

Because they slow down at a measurable rate, pulsars carry an approximate age: divide the period by twice the period derivative and you get a characteristic age, which for the Crab gives about 1,240 years against a true age, from the 1054 supernova, of nearly a thousand. Close enough to be useful, wrong enough to remind you it is an estimate.

The upshot: A pulsar is a flywheel losing energy to a magnetic brake, and the energy it loses is visible as the nebula around it.

Millisecond pulsars, and clocks good enough to weigh spacetime

Some pulsars spin hundreds of times a second. The fastest known, in the globular cluster Terzan 5, turns 716 times per second, which means a point on its equator moves at a substantial fraction of the speed of light. These are old pulsars that were spun up by accreting matter from a binary companion, a process called recycling, and they have unusually weak magnetic fields, so they brake slowly and stay fast.

Their stability is extraordinary. A millisecond pulsar's pulse arrival times can be predicted years in advance to within microseconds, comparable to the best atomic clocks. That turns them into instruments.

The most spectacular use came from PSR B1913+16, discovered by Russell Hulse and Joseph Taylor at Arecibo in 1974: a pulsar in a 7.75 hour orbit with another neutron star. General relativity predicts that such a system radiates gravitational waves and therefore loses orbital energy, so the orbit should shrink and the period shorten by a specific, calculable amount: about 76 microseconds per year. Decades of timing showed exactly that decay, agreeing with the prediction to better than a fraction of a per cent. Hulse and Taylor received the Nobel Prize in 1993 for what was, for forty years, the only evidence that gravitational waves existed.

More recently, arrays of millisecond pulsars monitored for decades have been used as a Galaxy-sized detector. In 2023 several international collaborations reported evidence for a background of nanohertz gravitational waves, most plausibly from the combined signal of supermassive black hole binaries across the universe. That result was still being strengthened when this course was written, so treat it as an active area rather than a settled one.

Black holes, and how small a radius has to be

If the collapsing core exceeds the maximum neutron star mass, nothing known can stop it. The result is a region from which no signal can escape, bounded by a surface called the event horizon. For a non-rotating black hole its radius is the Schwarzschild radius:

R_s = 2GM / c^2

which in convenient form is

R_s = 2.95 km x (M / M_sun)

Put some numbers through it. The Sun would need to be squeezed to a radius of 2.95 kilometres. Earth would need a radius of 8.9 millimetres. A 10 solar mass black hole has a horizon 29.5 kilometres across in radius. The black hole at the centre of the Milky Way, at 4.3 million solar masses, has a Schwarzschild radius of 1.27 x 10^7 kilometres, about 0.085 astronomical units, roughly a fifth of Mercury's orbital radius.

Three points cause more confusion than the rest of the subject combined.

First, the event horizon is not a surface made of anything. It is a boundary in spacetime. There is nothing there to touch, and an infalling observer crossing the horizon of a very large black hole would notice nothing locally remarkable at the moment of crossing.

Second, black holes are described by only three numbers: mass, angular momentum and electric charge. Everything else about whatever formed them is lost. Astrophysical black holes have negligible charge, so in practice mass and spin describe them completely. This simplicity is why they are, in a strange sense, the easiest objects in astrophysics to model.

Third, a black hole's gravity at a distance is exactly the gravity of the same mass in any other form. If the Sun were replaced by a 2.95 kilometre black hole of one solar mass, Earth's orbit would not change at all. It would get very cold, and nothing else would happen. Black holes do not draw in material from a distance any more effectively than an ordinary star of the same mass.

The evidence, in four independent kinds

X-ray binaries. Cygnus X-1 was found in a 1964 rocket flight as one of the brightest X-ray sources in the sky. In 1971 it was matched to a blue supergiant, HDE 226868, orbiting an invisible companion. The visible star's orbit gives a lower limit on the companion's mass, and that limit exceeds anything a neutron star can be. A 2021 study using very long baseline radio astrometry to pin the distance revised the compact object's mass upward to about 21 solar masses. X-rays come from an accretion disk heated to millions of kelvin by friction as material spirals in, which is how a black hole becomes conspicuous despite emitting nothing itself.

Stellar orbits at the Galactic centre. Since the early 1990s two groups, led by Reinhard Genzel in Germany and Andrea Ghez in the United States, have tracked individual stars orbiting an invisible mass at the centre of the Milky Way, using infrared observations that see through the dust. The star S2 completes an orbit every 16.05 years and passes within about 120 astronomical units of the centre at some 7,650 kilometres per second. Kepler's laws applied to that orbit give a mass of about 4.3 million solar masses inside a region smaller than the solar system. Nothing but a black hole fits. In 2018 the GRAVITY instrument measured the gravitational redshift of S2's light at closest approach, matching general relativity. Genzel and Ghez shared the 2020 Nobel Prize.

Gravitational waves. On 14 September 2015 the two LIGO detectors registered a signal lasting about 0.2 seconds, sweeping upward in frequency and then cutting off. It was the merger of black holes of about 36 and 29 solar masses into one of about 62. The missing 3 solar masses were radiated as gravitational waves, and for a fraction of a second the power output exceeded the combined light of every star in the observable universe. The announcement came on 11 February 2016 and the Nobel Prize followed in 2017. Since then the detector network has accumulated a substantial and growing catalogue of compact object mergers, so any count you read carries a date.

Direct imaging. On 10 April 2019 the Event Horizon Telescope, an array of radio dishes across the planet operating as a single instrument, released an image of the centre of the galaxy M87: a ring of emission about 42 microarcseconds across, with a dark centre, around a mass of 6.5 billion suns. On 12 May 2022 the same collaboration released an image of Sagittarius A star, far closer but far smaller and far harder because the gas around it changes on timescales of minutes. Both images show the shadow the horizon casts on the emission behind and around it, at the size general relativity predicts.

Remember: No single one of these four would be conclusive alone. Together, using entirely different physics and entirely different instruments, they converge on the same objects with the same masses.

The gap between the two endpoints

Neutron stars top out somewhere near 2.2 to 2.5 solar masses. The lightest black holes found in X-ray binaries sit near 5. For decades that stretch looked genuinely empty, and it was called the mass gap, with explanations proposed involving how supernova explosions eject material.

Gravitational wave detections have complicated the picture. Events have been recorded involving objects of about 2.6 solar masses, sitting squarely in the gap, whose nature is ambiguous: too heavy for a confident neutron star, too light for a comfortable black hole. Whether the gap is real, an artefact of how X-ray binaries are selected, or somewhere in between remains open. It is a good example of a question that only a new kind of instrument could reopen.

Common misconceptions

"Black holes suck things in." They exert exactly the gravity their mass implies. Material spirals in only if it is on a trajectory that takes it close, which is true of any compact object. Most matter in a galaxy passes black holes by, exactly as it passes stars.

"The event horizon is a physical surface." It is a boundary defined by which trajectories can escape. Nothing is located there, and for a supermassive black hole the tidal forces at the horizon are mild.

"Pulsars are pulsing stars." They rotate at a constant rate and sweep a beam. The pulsation is a geometric effect, and pulsars whose beams miss Earth are undetectable as pulsars altogether.

"The Event Horizon Telescope photographed a black hole." It imaged the bright emission surrounding one and the dark region its horizon casts within that emission. The black hole itself emits nothing, and the image is a reconstruction from interferometric data rather than a photograph in the ordinary sense.

"Neutron stars are made of neutrons and nothing else." The crust is ordinary nuclei and electrons, there is a proton fraction throughout, and the composition of the core is one of the open questions in nuclear physics.

Looking back

A core collapse that stops at nuclear density leaves a neutron star: roughly 1.4 solar masses in a 20 kilometre sphere, with a surface gravity of 10^12 metres per second squared, an escape velocity around 0.6c, an inherited rotation concentrated by a factor of five billion, and a magnetic field concentrated by the same factor. When the resulting beam sweeps across Earth we call it a pulsar, and the Crab pulsar's measured spin-down power of about 5 x 10^31 watts matches the Crab Nebula's luminosity, which confirms the mechanism. Recycled millisecond pulsars are clocks accurate enough that the orbital decay of PSR B1913+16 confirmed gravitational radiation forty years before it was detected directly, and pulsar timing arrays are now being used to search for a nanohertz gravitational wave background. Above the maximum neutron star mass, collapse continues to a black hole with a Schwarzschild radius of 2.95 kilometres per solar mass, characterised entirely by mass and spin, with a horizon that is a boundary rather than a surface and a distant gravitational field identical to any other mass. Four independent lines of evidence establish that these objects exist: X-ray binaries such as Cygnus X-1, the stellar orbits around Sagittarius A star that earned the 2020 Nobel Prize, the gravitational wave detections beginning with GW150914 in 2015, and the Event Horizon Telescope images of 2019 and 2022. The mass range between the heaviest neutron stars and the lightest black holes is still being mapped.

Sources

  1. NASA Science. (n.d.). Black holes. National Aeronautics and Space Administration. science.nasa.gov
  2. LIGO Scientific Collaboration. (n.d.). Detection of gravitational waves. LIGO Caltech. ligo.caltech.edu
  3. Event Horizon Telescope Collaboration. (n.d.). Images and results. EHT. eventhorizontelescope.org
  4. Encyclopaedia Britannica. (n.d.). Pulsar. britannica.com
  5. Hewish, A., Bell, S. J., Pilkington, J. D. H., Scott, P. F., & Collins, R. A. (1968). Observation of a rapidly pulsating radio source. Nature, 217, 709-713.
  6. Fraknoi, A., Morrison, D., & Wolff, S. C. (2022). Astronomy 2e, Chapters 23 and 24: The death of stars, and black holes and curved spacetime. OpenStax, Rice University. openstax.org
Key terms
Neutron star
A collapsed stellar core of roughly 1.4 solar masses in a 20 kilometre sphere, supported by neutron degeneracy and the strong force.
Pulsar
A rotating magnetised neutron star whose beamed emission sweeps across the observer once per rotation.
Spin-down power
The rate at which a pulsar loses rotational kinetic energy, which for the Crab matches the luminosity of its nebula.
Millisecond pulsar
An old pulsar spun up by accretion from a companion, with a weak field and timing stability rivalling atomic clocks.
Equation of state
The relation between pressure and density in neutron star matter, which sets the maximum neutron star mass and is not yet determined.
Schwarzschild radius
The event horizon radius of a non-rotating black hole, equal to 2GM/c^2, or 2.95 km per solar mass.
Event horizon
The boundary beyond which no signal can escape a black hole; a feature of spacetime rather than a material surface.
No-hair theorem
The result that a black hole is fully described by mass, angular momentum and charge, with all other information about its origin lost.
Mass gap
The sparsely populated range between the heaviest neutron stars near 2.2 solar masses and the lightest confirmed black holes near 5.

Module 5: Systems, Variables, and Populations

Binaries, which are the only direct route to a stellar mass; pulsating stars, which are the rung of the distance ladder that reaches other galaxies; and clusters, which let you date a whole population at once.

Binaries: The Only Direct Way to Weigh a Star

  • Apply Kepler's third law and the mass ratio to obtain individual stellar masses from a visual binary orbit.
  • Compare visual, astrometric, spectroscopic and eclipsing binaries by what each measurement can and cannot deliver.
  • Explain Roche lobe overflow, mass transfer and the Algol paradox, and identify the exotic systems binarity makes possible.

Fifty years, seven and a half arcseconds, three solar masses

Sirius A and Sirius B orbit their common centre of mass once every 50.1 years. On the sky the orbit spans a semimajor axis of about 7.5 arcseconds. Those two numbers, plus the parallax distance of 2.637 parsecs from the first lesson of this course, are enough to weigh both stars.

Convert the angular size to a physical one. An angle of 7.5 arcseconds at 2.637 parsecs corresponds to a separation of 7.5 x 2.637 = 19.8 astronomical units, since one arcsecond at one parsec is one astronomical unit by construction. Now use Kepler's third law in the form Newton derived, expressed in solar units:

M1 + M2 = a^3 / P^2, with a in astronomical units, P in years, and the answer in solar masses

a^3 = 19.8^3 = 7,762

P^2 = 50.1^2 = 2,510

M1 + M2 = 7,762 / 2,510 = 3.09 solar masses

That is the total. To split it, watch how far each star moves from the barycentre: the more massive star traces the smaller ellipse, and the mass ratio is the inverse of the ratio of the two semimajor axes. For Sirius the ratio is about 2 to 1, giving 2.06 solar masses for Sirius A and 1.02 for Sirius B.

Those two numbers matter far beyond Sirius. The 1.02 is the mass that, combined with the Earth-sized radius from Module 4, gives a white dwarf its absurd density and anchors the Chandrasekhar limit against observation. And every stellar mass you have seen quoted in this course, every entry in every mass-luminosity table, traces back to measurements of this kind. Gravity is the only thing that responds to mass, so a star in isolation cannot be weighed at all.

Why this matters: Mass is the master variable of stellar astronomy, and binaries are the only place it can be measured directly. Everything else is calibration against them.

How common is a companion

Multiplicity is not exotic. Surveys of solar-type stars in the solar neighbourhood find that roughly forty to fifty per cent have at least one companion. The fraction rises steeply with mass: most O and B stars are in multiple systems, and many are in systems close enough to interact. It falls for the lowest masses, with M dwarf multiplicity nearer a quarter to a third.

Since M dwarfs dominate by number, the often-repeated claim that most stars are in binaries turns out to be false for the population as a whole, and true for the stars you can see. It is another instance of the sampling problem from Module 2, and it is worth pausing on because textbooks written before large M dwarf surveys stated the opposite.

Four ways a binary reveals itself

TypeHow it is detectedWhat you getWhat you do not get
VisualBoth stars resolved and tracked over years or decadesFull orbit, total mass with a distance, mass ratioAnything with a period longer than the observing record
AstrometricOne star wobbles against the background; the companion is unseenEvidence of a companion and constraints on its massA spectrum or luminosity for the companion
SpectroscopicSpectral lines shift periodically by the Doppler effectPeriod, orbital velocity amplitudes, and the combination M sin^3 iThe inclination, so masses remain ambiguous
EclipsingBrightness dips as one star passes in front of the otherInclination near 90 degrees, and stellar radii from eclipse durationsAnything, if the orbit is not aligned with our line of sight

Sirius is visual: Bessel inferred it astrometrically in 1844 and Clark resolved it in 1862, so it has been both.

Spectroscopic binaries come in two kinds. If both sets of lines are visible, shifting in opposite directions, the system is double-lined and you can measure both stars' velocity amplitudes and hence the mass ratio directly. If only one star is bright enough to contribute lines, the system is single-lined and you get less.

The great limitation of the spectroscopic method is that Doppler shifts measure only the component of velocity along the line of sight. A system seen face-on shows no shift at all, even with enormous orbital speeds. So spectroscopy delivers masses multiplied by sin^3 i, where i is the unknown inclination, and gives only lower limits.

This is exactly why eclipsing systems are precious. An eclipse can only happen if the orbit is nearly edge-on, so i is close to 90 degrees and sin i is close to 1. A system that is both eclipsing and double-lined spectroscopic yields absolute masses with no assumptions, and because eclipse durations measure how long each star takes to pass in front of the other, it yields radii too. Such systems give masses and radii accurate to a few per cent, and they are the fundamental calibrators against which all stellar structure models are tested.

Worked example: an eclipsing double-lined binary

Suppose you have observed a system with a circular orbit, a period of 2.00 days, radial velocity amplitudes of 100 km/s for the primary and 150 km/s for the secondary, and clear eclipses so that i is essentially 90 degrees.

First the mass ratio. Both stars orbit the same centre of mass, so the more massive one moves more slowly:

M1 / M2 = K2 / K1 = 150 / 100 = 1.5

Now the total mass. For a circular edge-on orbit,

M1 + M2 = P x (K1 + K2)^3 / (2 pi G)

with everything in SI units. P = 2.00 days = 1.728 x 10^5 s, and K1 + K2 = 250 km/s = 2.50 x 10^5 m/s.

(2.50 x 10^5)^3 = 1.5625 x 10^16

Numerator = 1.728 x 10^5 x 1.5625 x 10^16 = 2.700 x 10^21

Denominator = 2 x 3.1416 x 6.674 x 10^-11 = 4.194 x 10^-10

M1 + M2 = 2.700 x 10^21 / 4.194 x 10^-10 = 6.44 x 10^30 kg = 3.24 solar masses

Split it by the velocity amplitudes: the primary takes the share proportional to K2, so

M1 = 3.24 x (150/250) = 1.94 solar masses, and M2 = 3.24 x (100/250) = 1.30 solar masses

Two stars weighed to a few per cent from a light curve and a series of spectra, with no distance required and no model of stellar interiors assumed. Note that last point: nothing in this calculation used anything about how stars work. That independence is why eclipsing binaries are the benchmark.

The point: A double-lined eclipsing binary gives masses and radii from geometry and Doppler shifts alone, which is why a handful of such systems constrain the whole of stellar structure theory.

What binaries proved

The mass-luminosity relation you used in Module 2 to compute lifetimes is an empirical result from binaries. Once enough systems had been measured, plotting mass against luminosity produced the tight relation L proportional to roughly M^3.5, and Eddington showed in the 1920s that stellar structure theory predicted something close to it. Theory and measurement met, and that meeting is one of the foundations of the subject.

Binaries also supply a check on the HR diagram itself. Two stars in a binary formed at the same time from the same material, so they must lie on the same theoretical isochrone with only their masses differing. Systems where both components can be placed on the diagram are among the sharpest tests of stellar evolution models available.

When the two stars touch

Close binaries interact, and the geometry has a name. Around each star is a teardrop-shaped region called its Roche lobe, the volume within which material is gravitationally bound to that star rather than to its companion or to the pair. The two lobes meet at a saddle point called L1.

Systems are classified by whether the stars fill their lobes. Detached systems have both stars comfortably inside, and evolve as if they were single. Semi-detached systems have one star filling its lobe, so material flows through L1 onto the companion. Contact systems have both stars overflowing and sharing a common envelope.

Mass transfer produces a puzzle that took decades to resolve. Algol, whose eclipses John Goodricke timed in 1783 at a period he put near two days and twenty-one hours, consists of a hot main-sequence B star of about 3.2 solar masses and a cooler subgiant of about 0.7. The subgiant has evolved off the main sequence, but it is the less massive star. Since more massive stars evolve faster, the less massive star should still be sitting on the main sequence. This is the Algol paradox.

The resolution is mass transfer. The currently less massive star was originally the more massive one. It evolved first, expanded, filled its Roche lobe, and transferred most of its envelope to its companion, which is now the heavier and less evolved of the two. Evolutionary sequence and current mass ranking have simply been swapped by the transfer, and once you accept that, the paradox evaporates and becomes evidence.

The exotic objects that require a partner

A striking number of the phenomena in this course cannot happen to a single star at all.

Novae occur when hydrogen accreted from a companion accumulates on a white dwarf's surface until it ignites in a thermonuclear runaway, brightening the system by many magnitudes for weeks. The white dwarf survives, and the cycle can repeat.

Type Ia supernovae need a white dwarf pushed toward the Chandrasekhar limit, which requires either accretion from a companion or a merger with another white dwarf. No isolated white dwarf ever explodes.

X-ray binaries generate their emission from an accretion disk fed by a companion. Cygnus X-1 would be invisible without HDE 226868 supplying it with gas.

Millisecond pulsars exist because accretion from a companion spins an old pulsar back up. The recycling story from the previous lesson is entirely a binary story.

Blue stragglers are stars sitting above a cluster's main-sequence turnoff, apparently too massive to still be burning hydrogen given the cluster's age. They are best explained as products of mergers or mass transfer, which effectively gives a star a fresh supply of fuel and resets its clock.

Gravitational wave sources are, by definition, binaries. Every merger detected since 2015 is the end of a story that began with two massive stars orbiting each other.

Bottom line: Roughly half the interesting objects in stellar astrophysics are consequences of one star having a neighbour close enough to matter.

Common misconceptions

"Most stars are in binaries." Most stars are M dwarfs, and M dwarf multiplicity is only a quarter to a third. The claim is true for solar-type and massive stars and false for the population as a whole, which is a distinction older textbooks did not draw.

"A spectroscopic binary gives you the masses." It gives masses multiplied by sin^3 i. Without an eclipse or some other constraint on inclination, you have lower limits only.

"The two stars orbit the heavier one." Both orbit the common centre of mass, which lies between them, closer to the heavier star. The mass ratio is exactly the inverse of the ratio of their orbital sizes, and that is how the split in the Sirius calculation was made.

"The Algol paradox shows stellar evolution theory is wrong." It shows that the theory of single stars does not apply to interacting ones. Once mass transfer is included, Algol becomes a confirmation rather than a problem.

Recap

A binary orbit plus a distance gives a total mass through Kepler's third law in the form M1 + M2 = a^3 / P^2, and the ratio of the two stars' distances from the barycentre splits it. For Sirius that yields 2.06 and 1.02 solar masses from a 50.1 year period and a 7.5 arcsecond orbit. Binaries are detected visually, astrometrically, spectroscopically and through eclipses, and each method delivers something different: spectroscopy alone gives only M sin^3 i, but a system that both eclipses and shows two sets of Doppler-shifted lines gives absolute masses and radii to a few per cent using nothing but geometry. Those benchmark systems established the mass-luminosity relation empirically and remain the primary test of stellar structure models. Close binaries interact through Roche lobe overflow, and mass transfer explains the Algol paradox in which the less massive star is the more evolved one. Finally, novae, Type Ia supernovae, X-ray binaries, millisecond pulsars, blue stragglers and every gravitational wave source are phenomena that a single star cannot produce. With masses in hand from binaries and distances in hand from parallax, the next lesson extends the distance scale far beyond where parallax can reach.

Sources

  1. Encyclopaedia Britannica. (n.d.). Binary star. britannica.com
  2. Fraknoi, A., Morrison, D., & Wolff, S. C. (2022). Astronomy 2e, Chapter 18: Measuring stellar masses. OpenStax, Rice University. openstax.org
  3. NASA Science. (n.d.). Stars. National Aeronautics and Space Administration. science.nasa.gov
  4. Chandra X-ray Observatory. (n.d.). Binary and X-ray binary systems. Harvard-Smithsonian Center for Astrophysics. chandra.harvard.edu
  5. Torres, G., Andersen, J., & Gimenez, A. (2010). Accurate masses and radii of normal stars: Modern results and applications. Astronomy and Astrophysics Review, 18, 67-126.
Key terms
Barycentre
The common centre of mass about which both stars in a binary orbit, located closer to the more massive star.
Visual binary
A pair resolved into two points whose relative orbit can be traced directly on the sky.
Spectroscopic binary
A system detected by periodic Doppler shifts of its spectral lines, giving masses only in the combination M sin^3 i.
Eclipsing binary
A system whose orbit is nearly edge-on, so the stars periodically block one another, fixing the inclination and revealing radii.
Roche lobe
The teardrop-shaped region around each star within which material remains gravitationally bound to that star.
Mass transfer
The flow of material from a star that fills its Roche lobe onto its companion through the L1 point.
Algol paradox
The observation that the less massive star in Algol is the more evolved one, resolved by mass transfer having reversed the original ratio.
Blue straggler
A cluster star lying above the main-sequence turnoff, explained by mergers or mass transfer supplying fresh hydrogen.

Pulsating Stars and the Distance Ladder

  • Explain the kappa mechanism and why pulsating stars occupy a narrow instability strip on the HR diagram.
  • Use a period-luminosity relation to convert a Cepheid's period and apparent magnitude into a distance.
  • Lay out the rungs of the distance ladder, state how each is calibrated, and explain why errors compound.

Twenty-five variable stars on a photographic plate

Henrietta Swan Leavitt worked at the Harvard College Observatory measuring the brightnesses of stars on glass photographic plates, many of them taken at Harvard's southern station in Arequipa, Peru. Comparing plates of the Small Magellanic Cloud taken at different times, she catalogued more than a thousand variable stars, and she noticed something in the subset whose periods she could determine.

The brighter ones took longer to cycle. In 1912, in a Harvard Circular reporting work that was entirely hers, she published a table of 25 variables and plotted period against apparent magnitude. The points fell on a line.

Here is why that line was worth more than anything else on the plate. All 25 stars are in the Small Magellanic Cloud, which means all of them are at essentially the same distance. A relation between period and apparent magnitude within a single system is therefore a relation between period and absolute magnitude, with only an unknown constant offset. Measure the period of any such star anywhere, and you know its intrinsic luminosity. Compare with how bright it looks, and you have its distance.

Leavitt had found the first standard candle that reached beyond our own Galaxy. Within thirteen years it would be used to prove that other galaxies exist.

The core of it: A period is a clock, and a clock can be measured through any amount of distance without degrading. That is why a period-luminosity relation is a more powerful distance tool than anything based on apparent size or brightness alone.

Why a star pulsates at all

Delta Cephei, the star that gives the class its name, was found to vary by John Goodricke in 1784, the same observer who timed Algol. It brightens and fades by about one magnitude every 5.366 days, rapidly rising and slowly declining, and its spectral type and radial velocity vary in step. The star is physically swelling and shrinking, changing radius by around ten per cent and surface temperature by hundreds of kelvin.

The mechanism is a heat engine, and Eddington identified its principle in the 1920s before the details were understood. The key is a layer inside the star where helium is partially ionised, meaning it is in the process of losing its second electron.

Follow one cycle. The star contracts, compressing that layer. Normally compression makes gas more transparent, but in a partial ionisation zone the extra energy goes into further ionisation rather than into heating, and the opacity rises. Radiation is dammed up beneath the layer. Pressure builds and pushes the envelope outward. As the layer expands and cools, the ions recombine, the opacity drops, the trapped radiation escapes, the support fails, and the envelope falls back under gravity. Then it compresses again.

This is the kappa mechanism, named for the symbol conventionally used for opacity. It only works if the partial ionisation zone sits at the right depth: too shallow and there is too little mass above it to matter, too deep and convection carries the energy instead. That condition is met only in a narrow, nearly vertical band on the HR diagram called the instability strip, running from the supergiants down through the horizontal branch. A star does not choose to pulsate. It pulsates while it is crossing that strip and stops when it leaves.

The families of pulsator

ClassPeriod rangeTypical M_VMass and populationWhere they are useful
Classical Cepheids1 to 100 days-2 to -74 to 20 solar masses, young Population IDistances to spiral galaxies out to tens of megaparsecs
Type II Cepheids1 to 50 daysAbout 1.5 magnitudes fainter than classical at the same periodAround 0.5 solar masses, old Population IIDistances within the halo and to globular clusters
RR Lyrae stars0.2 to 1 dayAbout +0.6, nearly independent of period0.5 to 0.8 solar masses, horizontal branch, oldGlobular clusters and the Galactic halo
Mira variables100 to 1,000 daysVariable and less regularAsymptotic giant branch starsRough distances; large amplitude but poor precision
Delta Scuti starsUnder 0.3 daysAround +21.5 to 2.5 solar massesAsteroseismology of individual stars

The distinction in the first two rows caused one of the largest single corrections in the history of astronomy, and it is worth the paragraph it takes.

Classical Cepheids are massive young stars. Type II Cepheids are low-mass old stars that happen to cross the same strip on their way through late evolution. They obey different period-luminosity relations, offset by roughly 1.5 magnitudes, which is a factor of four in luminosity and a factor of two in distance. In 1952 Walter Baade, using the new 200 inch telescope at Palomar, showed that the two types had been conflated, and that the extragalactic distance scale calibrated on one had been applied to the other. Every distance beyond the Milky Way roughly doubled overnight. Among other things, this removed an embarrassment: the previous scale made the universe younger than the age of Earth.

Using the relation: a worked distance

A commonly used form of the classical Cepheid period-luminosity relation in the V band is

M_V = -2.81 x log10(P) - 1.43, with P in days

Different calibrations give slightly different coefficients, which matters at the few per cent level and is one of the reasons the last section of this lesson exists.

Test it on a known star. Delta Cephei has a period of 5.366 days and a mean apparent magnitude of about 3.95.

log10(5.366) = 0.7297

M_V = -2.81 x 0.7297 - 1.43 = -2.05 - 1.43 = -3.48

m - M = 3.95 - (-3.48) = 7.43

d = 10^((7.43 + 5)/5) = 10^2.486 = 306 parsecs

The distance measured directly by parallax is about 273 parsecs. Our answer is 12 per cent high, which is a genuinely useful result to sit with. The discrepancy comes mostly from two things: we ignored interstellar extinction, which makes the star look fainter and therefore farther, and the coefficients in the relation depend on which calibration you adopt and on the star's metallicity. Real Cepheid distance work is done in the near infrared precisely because extinction is much weaker there, and it uses metallicity-corrected relations.

Now use it where parallax cannot reach. A Cepheid in a distant galaxy has a period of 30 days and a corrected apparent magnitude of 24.2.

log10(30) = 1.477, so M_V = -2.81 x 1.477 - 1.43 = -4.15 - 1.43 = -5.58

m - M = 24.2 + 5.58 = 29.78

d = 10^((29.78 + 5)/5) = 10^6.956 = 9.0 x 10^6 parsecs, or about 9 megaparsecs

Twenty-nine million light years, from one period and one magnitude. No other stellar property scales like this.

The plate that ended a debate

In October 1923 Edwin Hubble photographed the Andromeda Nebula with the 100 inch telescope on Mount Wilson, looking for novae. On one plate he marked a star as a nova, then compared it with earlier plates, found it varying periodically, crossed out his note and wrote VAR with an exclamation mark. It was a Cepheid.

Applying Leavitt's relation put Andromeda far outside any plausible size for the Milky Way. Hubble announced the result at the start of 1925, and the question of whether the spiral nebulae were objects inside our Galaxy or separate galaxies of their own, argued publicly by Harlow Shapley and Heber Curtis in 1920, was settled. The universe acquired other galaxies. Module 6 picks up what followed.

So what?: A single variable star on a single photographic plate multiplied the known size of the universe by a large factor, because a period-luminosity relation converts a timing measurement into a distance.

The ladder, and why it is a ladder

No single method spans the range from the nearest star to the most distant galaxy. Each technique is calibrated using the one below it and used to calibrate the one above, which is why the whole structure is called the cosmic distance ladder.

RungMethodApproximate reachCalibrated by
1Trigonometric parallaxA few thousand parsecsNothing; it is pure geometry
2Main-sequence fitting to clustersTens of thousands of parsecsParallax distances to nearby clusters
3RR Lyrae starsWithin the Galaxy and the nearest satellitesParallax and cluster distances
4Classical CepheidsTens of megaparsecs with space telescopesParallax to Galactic Cepheids, plus the Large Magellanic Cloud
5Tip of the red giant branchComparable to Cepheids, in any galaxy typeCepheids and geometric anchors
6Type Ia supernovaeHundreds of megaparsecs and beyondCepheids or the red giant tip in galaxies that have hosted one
7Redshift and the Hubble lawThe observable universeEverything below it

Two structural features follow from this design and you should never forget either.

First, errors compound. A systematic error in the Cepheid calibration propagates directly into every Type Ia distance and therefore into the expansion rate. This is not a hypothetical worry: Baade's 1952 correction was exactly such an error, and it moved every extragalactic distance by a factor of two.

Second, the ladder has recently produced a discrepancy nobody has resolved. Measuring the present expansion rate by climbing the ladder gives a value near 73 kilometres per second per megaparsec, with an uncertainty around one. Inferring the same quantity from the cosmic microwave background, using the standard cosmological model, gives about 67.4 with an uncertainty near 0.5. The two disagree by more than their combined error bars. Enormous effort has gone into finding a flaw in the ladder, including JWST observations testing whether crowded fields bias Cepheid photometry, and so far the ladder has held up. Whether the resolution is an unrecognised systematic or new physics is open, and Cosmology takes the question further than we will here.

Common misconceptions

"Cepheids vary because something passes in front of them." That is an eclipsing binary. A Cepheid physically expands and contracts, changing its radius by about ten per cent and its temperature along with it, driven by the kappa mechanism.

"Any variable star can be used as a standard candle." Only classes with a tight relation between an observable timing property and luminosity. Mira variables pulsate with huge amplitudes and are poor distance indicators because their relation is loose.

"Leavitt's law gives distances directly." It gives a luminosity for a given period only up to a constant, which has to be fixed by measuring the distance to at least one Cepheid some other way. Leavitt provided the shape of the relation; the zero point came from parallaxes and from other methods, and refining it remains active work.

"The distance ladder is obsolete now that we have Gaia." Gaia strengthened the bottom rung enormously, which improves everything above it. It did not replace the upper rungs, because parallax still stops at a few thousand parsecs and galaxies start at hundreds of thousands.

"The Hubble tension means somebody made a mistake." It may. It may also mean the cosmological model is incomplete. Treating it as obviously one or the other is the error; it is currently an open measurement disagreement between two careful and independent methods.

The short version

Leavitt's 1912 table of 25 variables in the Small Magellanic Cloud established that a Cepheid's pulsation period predicts its luminosity, because a single galaxy puts all its stars at one distance and turns apparent magnitudes into absolute ones. Stars pulsate when a partial ionisation zone sits at the right depth to act as a valve, dumping opacity when compressed and releasing it when expanded, which confines pulsators to a narrow instability strip. Classical Cepheids are young massive stars; Type II Cepheids are old low-mass stars obeying a relation about 1.5 magnitudes fainter, and conflating the two led to a factor-of-two error in extragalactic distances that Baade corrected in 1952. Applying a period-luminosity relation to Delta Cephei gives 306 parsecs against a measured 273, a discrepancy that comes from ignoring extinction and from the choice of calibration, and applying it to a 30 day Cepheid at magnitude 24.2 gives 9 megaparsecs. Hubble did exactly this in 1923 for a Cepheid in Andromeda and proved that other galaxies exist. Because each rung of the distance ladder is calibrated on the one below it, systematic errors propagate all the way up, which is the context in which the current five-sigma disagreement between the ladder's expansion rate near 73 and the microwave background's near 67 should be understood.

Sources

  1. Encyclopaedia Britannica. (n.d.). Cepheid variable. britannica.com
  2. NASA Science. (n.d.). Hubble mission: The distance scale. National Aeronautics and Space Administration. science.nasa.gov
  3. Fraknoi, A., Morrison, D., & Wolff, S. C. (2022). Astronomy 2e, Chapter 19: Celestial distances. OpenStax, Rice University. openstax.org
  4. NOIRLab. (n.d.). Variable stars and the cosmic distance ladder. NSF NOIRLab. noirlab.edu
  5. Leavitt, H. S., & Pickering, E. C. (1912). Periods of 25 variable stars in the Small Magellanic Cloud. Harvard College Observatory Circular, 173, 1-3.
  6. Baade, W. (1956). The period-luminosity relation of the Cepheids. Publications of the Astronomical Society of the Pacific, 68, 5-16.
Key terms
Period-luminosity relation
The tight correlation between a Cepheid's pulsation period and its intrinsic luminosity, discovered by Leavitt in 1912.
Kappa mechanism
The pulsation driver in which a partial ionisation zone becomes more opaque under compression, trapping radiation and pushing the envelope outward.
Instability strip
The narrow, nearly vertical band on the HR diagram where the ionisation zone sits at the depth required for pulsation.
Classical Cepheid
A young, massive Population I pulsator with periods of 1 to 100 days, used for distances to tens of megaparsecs.
Type II Cepheid
An old, low-mass pulsator obeying a period-luminosity relation about 1.5 magnitudes fainter than the classical one.
RR Lyrae star
A horizontal branch pulsator with a period under a day and an absolute magnitude near +0.6, used for halo and globular cluster distances.
Standard candle
An object whose intrinsic luminosity can be determined independently, so its apparent brightness yields a distance.
Cosmic distance ladder
The chain of methods in which each technique is calibrated by the one below it, so systematic errors propagate upward.

Star Clusters and Stellar Populations

  • Contrast open and globular clusters by size, age, location, metallicity and dynamical fate.
  • Read a cluster colour-magnitude diagram to obtain an age, and explain why old open clusters are rare.
  • Describe the Population I and II scheme, its modern refinement, and how alpha element ratios date a population's enrichment history.

A blackout over Los Angeles

Walter Baade was a German citizen living in California when the United States entered the Second World War. Classified as an enemy alien, he was barred from the war research that took most of his colleagues at Mount Wilson away, and found himself with unusual access to the 100 inch telescope. He also had something no astronomer there had enjoyed for years: a dark sky, because wartime dimout regulations had turned down the lights of Los Angeles.

Baade used the combination to attempt something that had defeated everyone: resolving individual stars in the crowded central region of the Andromeda Nebula and its two small companions. In 1943 he succeeded, using red-sensitive plates, and what he found was not what anyone expected. In the spiral arms, the brightest stars are blue. In the nucleus and in the companions, the brightest stars are red giants, and there are no luminous blue stars at all.

He published in 1944, dividing stars into two populations. Population I is what the solar neighbourhood and the spiral arms contain: young stars including hot blue ones, mixed with gas and dust. Population II is what fills the nucleus, the halo, and globular clusters: old stars with no gas, no dust, and no blue giants because everything massive died long ago.

That distinction is where the study of stellar populations begins, and it is inseparable from star clusters, which are the cleanest examples of each type.

Key idea: A cluster is a laboratory sample in which everything shares an age, a distance and a composition, leaving mass as the only variable. Nothing else in astronomy holds so much fixed.

Two kinds of cluster, and they are not variations on a theme

Open clustersGlobular clusters
Number of starsTens to a few thousand10^4 to more than 10^6
DiameterA few to about 20 parsecsTens of parsecs, strongly concentrated at the centre
LocationThe Galactic disk, close to the planeThe halo and bulge, distributed roughly spherically
AgeA few million years to a few billionTypically 11 to 13 billion years
MetallicityRoughly solarA tenth to a thousandth of solar iron
Gas and dustOften still present in young onesEssentially none
Number known in the Milky WayThousands, and rising sharply with GaiaAround 150 to 160
Shape and bindingIrregular, loosely bound, dissolvingSpherical, tightly bound, long lived
ExamplesPleiades, Hyades, M67, NGC 3532Omega Centauri, 47 Tucanae, M13, M15

The two are not the same kind of object at different ages. Globular clusters are a hundred to a thousand times more massive than typical open clusters, they formed under conditions that no longer exist in the Galaxy, and they occupy a different part of it.

Two objects deserve individual mention. The Hyades, at about 47 parsecs, is the nearest open cluster, close enough that its stars are spread across a large patch of sky and its distance was the anchor for the whole cluster distance scale before Hipparcos. Omega Centauri is the most massive globular cluster in the Galaxy and behaves oddly: unlike a normal globular, it contains stars of markedly different metallicity, which is difficult to explain if the whole thing formed in one burst. The leading interpretation is that it is the stripped nucleus of a dwarf galaxy the Milky Way swallowed, which makes it a fossil of a merger rather than a cluster in the usual sense.

Reading a cluster diagram

Plot every star in a cluster on a colour-magnitude diagram and the structure is immediately legible, because a shared distance means apparent magnitude is a fair proxy for luminosity.

The main sequence appears as a band, intact at the faint red end and terminated at the bright blue end. The break point is the turnoff.

The subgiant and giant branches extend to the upper right from the turnoff, populated by stars that have already exhausted core hydrogen.

The horizontal branch, prominent in old metal-poor clusters, runs left from the giant branch at roughly constant luminosity; these are core helium burning stars. A gap in it marks the instability strip, where the stars in transit are RR Lyrae variables rather than steady points.

A white dwarf sequence appears at the bottom left in deep exposures, and its faint end dates the cluster independently, as Module 4 described.

Blue stragglers sit above the turnoff, where nothing should still be burning hydrogen, and are products of mergers or mass transfer.

Turning the turnoff into an age is the calculation from Module 3. Determine the mass of a main-sequence star at the turnoff and quote its lifetime.

ClusterApproximate turnoff massApproximate ageType
NGC 2264Above 10 solar masses; low-mass stars still contractingA few million yearsVery young open
PleiadesAbout 5 solar massesAround 100 million yearsYoung open
HyadesAbout 2.5 solar massesAround 650 million yearsIntermediate open
M67About 1.3 solar massesAround 4 billion yearsOld open
47 TucanaeAbout 0.85 solar massesAround 12 billion yearsGlobular

Notice the very young cluster in the first row. In NGC 2264 the massive stars are already on the main sequence while the low-mass stars are still contracting toward it and sit above the main sequence at the faint end. Fitting the pre-main-sequence part of the diagram dates young clusters where there is no turnoff yet, and the two methods overlap in the range where both work.

Worth holding on to: One diagram of one cluster gives its age, its distance, its composition and a census of every evolutionary stage its stars have reached. That is why clusters are the primary test bed for stellar evolution theory.

Why old open clusters are rare

Look at the table again and notice that M67 at four billion years is described as old, while every globular is three times older. Open clusters do not usually survive. They dissolve, and there are three mechanisms.

Evaporation. In a bound system, stars exchange energy in close encounters. Over time some acquire enough speed to escape. A loosely bound cluster of a few hundred stars leaks members steadily, and once enough have gone the remainder is even less bound.

Galactic tides. An open cluster orbiting in the disk feels a differential gravitational pull from the Galaxy that stretches it along its orbit. Beyond a certain radius, the tidal radius, a star is more strongly held by the Galaxy than by the cluster.

Encounters with molecular clouds. A giant molecular cloud passing nearby delivers a gravitational shock. Since open clusters live in the disk where the clouds are, this happens repeatedly.

Typical open cluster lifetimes run from a few hundred million years to about a billion, which is why old ones are so few. Globular clusters escape all three fates: they are far more massive and therefore far more tightly bound, and they spend most of their orbits well out of the disk where clouds and tides are weaker.

The dissolved clusters do not vanish. Their stars join the general disk population, and Gaia has identified stellar streams and moving groups that are the dispersed remains of former clusters, still sharing a common velocity long after any spatial grouping has gone. The Sun was almost certainly born in a cluster that dispersed billions of years ago, and searches for its lost siblings are an active if difficult undertaking.

Populations, and the scheme that replaced them

Baade's two populations were the right first cut, and they have since been refined into a picture with more components.

The thin disk holds young metal-rich stars on nearly circular orbits close to the plane, with a scale height of a few hundred parsecs. This is where the Sun lives and where all current star formation happens.

The thick disk holds older, more metal-poor stars on more inclined orbits, with a scale height near a kiloparsec, probably heated or accreted early in the Galaxy's history.

The halo holds the oldest and most metal-poor stars on randomly oriented, often highly elongated orbits, together with the globular clusters.

The bulge is a mix, containing both very old stars and a range of metallicities.

Population III is the name given to the hypothetical first stars, formed from gas with no elements heavier than lithium. With no metals to provide cooling, such gas cannot fragment as efficiently, so theory expects the first stars to have been very massive and very short lived. None has ever been observed. Searches continue, including with JWST, and the current situation is that Population III is a well-motivated prediction with no direct confirmation.

Reading a chemical clock

Populations differ chemically because each generation of stars enriches the gas from which the next one forms, and the enrichment is not uniform. This gives a second, independent clock.

Recall from Module 4 that core-collapse supernovae come from massive stars and therefore begin within a few million years of a burst of star formation, producing large amounts of the alpha elements: oxygen, neon, magnesium, silicon, calcium. Type Ia supernovae come from white dwarfs in binaries and therefore require a delay of hundreds of millions to billions of years, and they produce mostly iron.

So gas that is enriched quickly is rich in alpha elements relative to iron, and gas enriched over a longer period has proportionally more iron. Plot the ratio [alpha/Fe] against [Fe/H] for a population and you get a characteristic shape: a flat, elevated plateau at low iron abundance, then a downward bend as Type Ia supernovae begin contributing. The location of that bend, called the knee, tells you how rapidly the population formed its stars.

Halo stars sit on the plateau, indicating rapid early formation. Thin disk stars sit below it, indicating prolonged formation with plenty of time for Type Ia enrichment. Dwarf galaxies show knees at lower iron abundance, indicating slower, less efficient star formation. A single spectrum thus carries information about the star formation history of the system the star was born in, which is a remarkable thing for a spectrum to know.

The upshot: Chemistry is a clock as well as a composition. The ratio of alpha elements to iron encodes how quickly a population built its stars, because the two element groups arrive on different schedules.

The oldest things we can date

Globular cluster ages, from turnoff fitting with full models at the correct low metallicity, come out at roughly 12 to 13 billion years. Since the universe is 13.8 billion years old, these clusters formed within the first billion or two years, which makes them direct evidence about the early Galaxy.

It has not always fit so neatly. Through the 1990s some globular cluster ages came out at 16 to 18 billion years, against an expansion age that then looked closer to 12. Stars older than the universe is not a small problem, and it was taken seriously. The resolution came from several directions at once: Hipparcos parallaxes revised cluster distances, improved opacities and better treatment of helium diffusion changed the models, and the discovery of cosmic acceleration in 1998 raised the expansion age. The discrepancy closed from both ends. It is a useful case study, because the crisis was real, the resolution was not a single clever fix, and nobody had to abandon either stellar physics or cosmology.

Common misconceptions

"Globular clusters are old open clusters." They are a different class of object: hundreds of times more massive, in a different part of the Galaxy, and formed under conditions that no longer exist. Open clusters do not survive long enough to become globulars, and they were never massive enough anyway.

"All the stars in a cluster have the same brightness or type." They share an age, distance and composition, and differ enormously in mass, which is exactly what makes them useful. The spread in the colour-magnitude diagram is the point.

"Population II means the second generation of stars." The numbering is historical and runs backwards relative to time: Population I is younger. Population III, the genuinely first generation, was named later and has never been observed.

"A cluster's age comes from its oldest stars." Every star in the cluster has the same age. The age comes from identifying which mass is just now leaving the main sequence, because that mass has a known lifetime.

"Metal-poor means the star is short of iron because it used it up." Stars do not consume iron. A metal-poor star formed from gas that had not yet been enriched, so its composition is a record of when and where it was born.

Putting it together

Baade's wartime observations of Andromeda's nucleus separated stellar populations into a young, blue, gas-rich Population I and an old, red, gas-free Population II, and star clusters are the cleanest examples of each. Open clusters hold tens to thousands of stars in the disk at roughly solar metallicity and dissolve within a few hundred million to a billion years through evaporation, Galactic tides and encounters with molecular clouds. Globular clusters hold 10^4 to more than 10^6 stars in the halo at a tenth to a thousandth of solar metallicity, and survive for 12 to 13 billion years because they are massive and spend little time in the disk. A cluster colour-magnitude diagram displays the main sequence, its turnoff, the giant and horizontal branches, the white dwarf sequence and the blue stragglers all at once, and the turnoff mass converts directly into an age, from a few million years for NGC 2264 to about 12 billion for 47 Tucanae. The two-population scheme has since been refined into thin disk, thick disk, halo and bulge components, with an unobserved Population III of metal-free first stars still predicted. Independently of any diagram, the ratio of alpha elements to iron dates how rapidly a population formed, because core-collapse and Type Ia supernovae deliver their products on different schedules. The next module takes these populations and assembles them into the Galaxy they belong to.

Sources

  1. Encyclopaedia Britannica. (n.d.). Star cluster. britannica.com
  2. NASA Science. (n.d.). Star clusters. National Aeronautics and Space Administration. science.nasa.gov
  3. European Space Agency. (n.d.). Gaia and Galactic archaeology. ESA Science and Exploration. esa.int
  4. Fraknoi, A., Morrison, D., & Wolff, S. C. (2022). Astronomy 2e, Chapters 22 and 25: Stars from adolescence to old age, and the Milky Way galaxy. OpenStax, Rice University. openstax.org
  5. Baade, W. (1944). The resolution of Messier 32, NGC 205, and the central region of the Andromeda Nebula. Astrophysical Journal, 100, 137-146.
Key terms
Open cluster
A loosely bound group of tens to a few thousand stars in the Galactic disk, roughly solar in composition and dissolving within about a billion years.
Globular cluster
A tightly bound spherical system of 10^4 to over 10^6 old, metal-poor stars in the halo, typically 11 to 13 billion years old.
Turnoff
The point where a cluster's main sequence bends toward the giant branch; the mass there has a lifetime equal to the cluster's age.
Tidal radius
The distance from a cluster's centre beyond which the Galaxy's gravitational pull on a star exceeds the cluster's.
Population I
Young, metal-rich stars associated with gas and dust in the Galactic disk, including all hot blue stars.
Population II
Old, metal-poor stars in the halo, bulge and globular clusters, with no luminous blue members remaining.
Population III
The predicted but never observed first generation of stars, formed from gas containing no elements heavier than lithium.
Alpha element ratio
The abundance of oxygen, magnesium, silicon and similar elements relative to iron, which records how rapidly a population formed its stars.
Stellar stream
The dispersed remains of a cluster or accreted satellite, identifiable by shared velocities long after any spatial grouping has gone.

Module 6: The Galaxy and Beyond

The Milky Way from the inside: its structure, its rotation, and the mass that does not shine. Then the galaxies outside it, the Local Group we belong to, and the nuclei that outshine everything around them.

The Milky Way: Structure, Rotation, and the Missing Mass

  • Describe the Galaxy's disk, bar, bulge and halo components with their characteristic sizes and stellar populations.
  • Compute enclosed mass from a rotation curve and explain why a flat curve implies mass beyond the visible material.
  • Compare the dark matter and modified gravity accounts, naming the specific evidence each rests on.

Shapley moves the Sun off centre

In 1918 Harlow Shapley published the result of a project nobody else had attempted. He had measured distances to dozens of globular clusters, using RR Lyrae stars whose absolute magnitudes he estimated, and plotted their three-dimensional positions.

They were not distributed around the Sun. They formed a roughly spherical swarm centred on a point far away in the direction of Sagittarius, tens of thousands of light years distant. Shapley drew the obvious conclusion: the centre of that swarm is the centre of the Galaxy, and the Sun is nowhere near it.

His number was too large, around 15 kiloparsecs, because he had not corrected for interstellar dust dimming the more distant clusters, which he had no way of knowing about in 1918. The modern value, from tracking stars orbiting the black hole at the centre, is about 8.2 kiloparsecs, or roughly 27,000 light years. But the qualitative result stood, and it was the second great demotion in astronomy: Copernicus moved Earth out of the centre of the solar system, and Shapley moved the solar system out of the centre of the Galaxy.

The point: We live inside the object we are trying to map, at an unremarkable location, behind a great deal of dust. Everything difficult about Galactic astronomy follows from those three facts.

The anatomy of a barred spiral

ComponentSizeStellar contentApproximate mass
Thin diskAbout 30 kpc across, scale height around 300 pcYoung to intermediate age, near solar metallicity, all current star formationAround 4 x 10^10 solar masses
Thick diskSame radial extent, scale height around 900 pcOlder, more metal-poor, higher velocity dispersionAround 10^10 solar masses
Bar and bulgeBar half-length 3 to 5 kpc, boxy bulgeMostly old, wide metallicity rangeAround 10^10 solar masses
Stellar haloExtends beyond 100 kpcVery old, very metal-poor, plus about 150 globular clustersAround 10^9 solar masses
Interstellar gasConcentrated in a thin layer in the diskAtomic and molecular hydrogen, dustAround 10^10 solar masses
Dark haloExtends to 200 kpc or moreNothing that emits lightAround 10^12 solar masses

Read the first column and the last together. Everything we can see, in every wavelength, adds up to something under 10^11 solar masses. The rotation of the Galaxy requires roughly ten times that. The rest of this lesson is about how that conclusion was reached and how seriously to take it.

Some further specifics. The Sun sits about 8.2 kiloparsecs from the centre and about 20 parsecs above the midplane, orbiting at roughly 230 kilometres per second. The Galaxy holds somewhere between one and four hundred billion stars, a range that is wide because it depends almost entirely on how many undetected M dwarfs and brown dwarfs you assume.

The spiral arms are the part we know least well, for the embarrassing reason that we are inside them. Named structures include the Perseus, Sagittarius-Carina, Scutum-Centaurus and Norma arms, with the Sun sitting in a smaller feature called the Orion Spur. Whether the Galaxy has two major arms with substantial spurs or four comparable arms is still argued, and different tracers give different answers. Nobody has ever seen the Milky Way from outside, and every artist's impression you have seen is a model.

Seeing through the dust

Toward the Galactic centre the visual extinction exceeds thirty magnitudes, which is a factor of 10^12 in flux. Optical astronomy simply cannot look there. Four workarounds built the modern map.

The 21 centimetre line. Hendrik van de Hulst predicted in 1944 that neutral atomic hydrogen should emit a radio line at 21 centimetres from a spin-flip transition, and Ewen and Purcell detected it at Harvard in 1951, with Dutch and Australian confirmations following within months. Radio waves are indifferent to dust. Because the line is narrow, its Doppler shift gives the velocity of each cloud, and combining velocities with a rotation model places the gas in the Galaxy. This is how spiral structure was first traced.

Molecular tracers. Carbon monoxide at 2.6 millimetres does the same job for the cold molecular gas where stars form.

The infrared. Extinction falls steeply with wavelength, so surveys such as 2MASS and the Spitzer mid-infrared surveys see straight through the dust to the stellar distribution. The clearest evidence that the Milky Way is barred came from infrared star counts.

Precise parallaxes to masers and stars. Very long baseline radio interferometry measures parallaxes to water and methanol masers in star-forming regions with accuracies of microarcseconds, giving direct geometric distances across much of the disk, and Gaia has done the same optically for over a billion stars where the dust permits.

The rotation curve, and the number it will not give up

The Galaxy does not rotate as a solid body. Objects at different radii take different times to go around, which is called differential rotation. Measure the orbital speed as a function of radius and you have the rotation curve, which is the single most informative measurement in Galactic astronomy, because orbital speed is set entirely by the mass inside the orbit.

For a circular orbit, the gravitational force supplies the centripetal acceleration:

G M(R) / R^2 = v^2 / R, so M(R) = v^2 R / G

Try it at the Sun's radius. R = 8.2 kpc = 2.53 x 10^20 metres, v = 230 km/s = 2.3 x 10^5 m/s.

v^2 = 5.29 x 10^10

M = 5.29 x 10^10 x 2.53 x 10^20 / 6.674 x 10^-11 = 1.338 x 10^31 / 6.674 x 10^-11 = 2.0 x 10^41 kg

Divide by the solar mass: about 1.0 x 10^11 solar masses inside the Sun's orbit. That is already interesting, and it is not yet a problem.

The problem appears further out. Nearly all the visible material is inside about 15 kiloparsecs. Beyond that, if the visible mass were all there is, an orbiting object would behave like a planet outside the Sun, with speed falling as the inverse square root of radius. The prediction is a declining curve.

What is measured, using neutral hydrogen clouds and satellite objects at large radii, is a curve that stays roughly flat out to twenty kiloparsecs and beyond. Redo the calculation at R = 20 kpc = 6.17 x 10^20 metres with v = 220 km/s:

M = (2.2 x 10^5)^2 x 6.17 x 10^20 / 6.674 x 10^-11 = 4.47 x 10^41 kg, about 2.2 x 10^11 solar masses

The enclosed mass has doubled between 8 and 20 kiloparsecs, in a region containing very little visible matter. Extend the measurement using satellite galaxies and globular clusters at 100 kiloparsecs and the total climbs toward 10^12 solar masses. Something is out there that does not emit, absorb or scatter light at any wavelength we can detect, and there is roughly ten times as much of it as of everything else.

Why this matters: A flat rotation curve is not a subtle statistical effect. M(R) rising linearly with R in a region that looks empty is a direct, arithmetic consequence of Newtonian gravity applied to a measured velocity.

The same result everywhere else

The Milky Way alone would be a weak case, since we measure it from inside. The decisive work was done on other galaxies.

Vera Rubin and Kent Ford began measuring rotation curves of spiral galaxies in the late 1960s, starting with Andromeda, using a spectrograph sensitive enough to reach the faint outer regions. By 1980 they had published curves for around sixty spirals, and they were all flat. Not approximately flat, not flat in some cases: flat, systematically, in galaxies of different sizes, types and environments. Rubin's papers turned an oddity into a general property of spiral galaxies.

The mass discrepancy had appeared even earlier and in a different setting. In 1933 Fritz Zwicky measured the velocities of galaxies in the Coma cluster and found that the cluster would fly apart unless it contained far more mass than its galaxies could account for. He called the missing component dunkle Materie, dark matter. The result was largely ignored for forty years.

Three further independent lines now support the same conclusion. Gravitational lensing measures the mass of clusters from the bending of light from background galaxies, with no assumption about dynamics, and finds the same excess. The pattern of temperature fluctuations in the cosmic microwave background is fitted by a model in which non-baryonic matter is about five times as abundant as ordinary matter, a topic Cosmology develops properly. And simulations of structure formation fail to produce galaxies at all in the available time without a component that clumps gravitationally without radiating.

The alternative, and why it has not won

The inference from a flat rotation curve to unseen mass assumes Newtonian gravity holds at galactic accelerations, which are around 10^-10 metres per second squared, some eleven orders of magnitude weaker than anything ever tested in a laboratory. That is not a trivial assumption, and it deserves the strongest version of its challenge.

In 1983 Mordehai Milgrom proposed Modified Newtonian Dynamics, in which gravitational acceleration departs from the Newtonian form below a characteristic scale of about 1.2 x 10^-10 metres per second squared. Choosing that one constant reproduces flat rotation curves without any dark matter.

MOND's strongest evidence is that it makes sharp, correct predictions about individual galaxies. Given only the distribution of visible matter, it predicts the detailed shape of the rotation curve, including bumps and wiggles that follow the visible structure. It predicted the tight relation between a galaxy's total baryonic mass and the fourth power of its rotation speed, the baryonic Tully-Fisher relation, which holds over about five orders of magnitude with remarkably little scatter. Explaining that tightness in a dark matter framework, where the dark halo and the visible galaxy are largely independent, requires the two to be finely coordinated.

Its strongest evidence against comes from systems where mass and light can be separated. In the Bullet Cluster, two galaxy clusters that have collided, the hot X-ray emitting gas, which is most of the ordinary matter, was slowed by the collision and sits between the two galaxy groups, while gravitational lensing shows the mass concentrations passed through and now sit with the galaxies rather than with the gas. In a theory where gravity simply responds differently to ordinary matter, the lensing mass should follow the gas. It does not. MOND also does not account for cluster dynamics without adding some unseen matter anyway, and it has no accepted account of the cosmic microwave background power spectrum or of how structure formed.

The current position is that dark matter is the strongly favoured explanation and that no dark matter particle has ever been detected despite decades of increasingly sensitive direct searches. That absence is a genuine and unresolved tension, and the galaxy-scale regularities MOND highlights remain a real challenge for the standard picture to explain naturally rather than by fitting.

What the centre holds, and where the Galaxy came from

At the very centre sits Sagittarius A star, the 4.3 million solar mass black hole whose evidence Module 4 laid out, surrounded by an extremely dense nuclear star cluster and, further out, a reservoir of molecular gas called the central molecular zone. In 2010 the Fermi gamma-ray telescope revealed two enormous lobes of gamma-ray emission extending tens of thousands of light years above and below the plane, the Fermi bubbles, plausibly the relic of a past outburst from the centre.

The Galaxy was also assembled rather than simply formed. Analysis of Gaia data in 2018 identified a large population of halo stars sharing distinctive orbits and chemistry, interpreted as the debris of a substantial galaxy that merged with the Milky Way around ten billion years ago, now called Gaia-Enceladus or the Gaia Sausage. The Sagittarius dwarf galaxy is being disrupted in front of us right now, its stars strung out in a stream that wraps around the sky. Reading the Galaxy's assembly history from the chemistry and orbits of individual stars is called Galactic archaeology, and Gaia turned it from an aspiration into a working field.

Bottom line: The Milky Way is a barred spiral about 30 kiloparsecs across, assembled over billions of years from smaller pieces, in which everything that shines accounts for under a tenth of the mass its own rotation requires.

Common misconceptions

"Dark matter is just gas and dust we have not found." Ordinary matter, whatever its form, absorbs, emits or scatters light and shows up in X-ray, infrared or radio observations. Extensive searches have found it and counted it, and the total falls far short. The cosmic microwave background independently constrains the ordinary matter density, and dark matter exceeds it by about five to one.

"The Sun orbits the Galaxy the way Earth orbits the Sun." Earth orbits a central mass. The Sun orbits inside a distributed mass, so only the material within its orbit matters, and the rotation curve is flat rather than Keplerian.

"We have photographs of the Milky Way from outside." Every full-galaxy image of the Milky Way is a model or an artist's rendering. All our data comes from inside the disk, which is why arm structure remains contested.

"Dark matter is a fudge factor invented to save a theory." It was inferred independently from cluster dynamics in 1933, from spiral rotation curves in the 1970s, from lensing, from the microwave background and from structure formation. Those five lines use different physics and agree on the amount, which is not the signature of a fudge factor. That said, the failure to detect a particle is real and should be stated whenever the case is made.

What to carry forward

Shapley's 1918 globular cluster survey put the Sun about 8.2 kiloparsecs from the Galactic centre, off to one side of a barred spiral roughly 30 kiloparsecs across, with a thin disk of scale height 300 parsecs, a thicker older disk, a bar and bulge, a metal-poor stellar halo containing about 150 globular clusters, and a dark halo extending far beyond all of it. Because the disk is opaque in visible light, the map was built from the 21 centimetre hydrogen line detected in 1951, from carbon monoxide, from infrared surveys that revealed the bar, and from radio and optical parallaxes. The rotation curve, obtained from those tracers, converts directly into enclosed mass through M(R) = v^2 R / G, giving 1.0 x 10^11 solar masses inside the Sun's orbit and about 2.2 x 10^11 by 20 kiloparsecs even though almost nothing visible lies between. Rubin and Ford showed the same flatness in around sixty other spirals, Zwicky had found a similar discrepancy in the Coma cluster in 1933, and lensing, the microwave background and structure formation agree on the amount. MOND accounts for galaxy rotation curves and the baryonic Tully-Fisher relation from visible matter alone, but does not explain the Bullet Cluster, cluster dynamics or cosmological structure, and dark matter remains the favoured explanation despite never having been detected directly.

Sources

  1. Encyclopaedia Britannica. (n.d.). Milky Way Galaxy. britannica.com
  2. NASA Science. (n.d.). Galaxies and the Milky Way. National Aeronautics and Space Administration. science.nasa.gov
  3. European Space Agency. (n.d.). Gaia and the structure of the Milky Way. ESA Science and Exploration. esa.int
  4. Chandra X-ray Observatory. (n.d.). Dark matter and the Bullet Cluster. Harvard-Smithsonian Center for Astrophysics. chandra.harvard.edu
  5. Rubin, V. C., Ford, W. K., & Thonnard, N. (1980). Rotational properties of 21 Sc galaxies with a large range of luminosities and radii. Astrophysical Journal, 238, 471-487.
  6. Milgrom, M. (1983). A modification of the Newtonian dynamics as a possible alternative to the hidden mass hypothesis. Astrophysical Journal, 270, 365-370.
Key terms
Thin disk
The star-forming, metal-rich component of the Galaxy with a scale height near 300 parsecs, containing the Sun.
Stellar halo
A roughly spherical distribution of very old, metal-poor stars and globular clusters extending beyond 100 kiloparsecs.
Differential rotation
Rotation in which orbital period varies with radius, as it must in a system whose mass is distributed rather than central.
Rotation curve
Orbital speed plotted against galactocentric radius; its shape determines the enclosed mass through M(R) = v^2 R / G.
21 centimetre line
Radio emission from a spin-flip transition in neutral atomic hydrogen, predicted in 1944 and detected in 1951, which maps gas through dust.
Dark matter
Mass inferred gravitationally that neither emits nor absorbs light, required at roughly ten times the visible mass in the Milky Way.
MOND
Modified Newtonian Dynamics, Milgrom's 1983 proposal that gravity departs from the Newtonian form below about 1.2 x 10^-10 m/s^2.
Baryonic Tully-Fisher relation
The tight empirical relation between a galaxy's total ordinary matter and the fourth power of its rotation speed.
Galactic archaeology
Reconstructing the Galaxy's assembly history from the chemistry and orbits of individual stars, transformed by Gaia data.

Galaxies, the Local Group, and Active Nuclei

  • Classify galaxies on the Hubble sequence and state what each morphological type corresponds to physically.
  • Describe the membership and dynamics of the Local Group, including the Milky Way's future encounter with Andromeda.
  • Explain the accretion engine powering active galactic nuclei and evaluate the unified model against the observed variety.

Something coming toward us

In 1912, at Lowell Observatory in Arizona, Vesto Slipher took a spectrum of the Andromeda Nebula, a faint smudge nobody yet knew the nature of. Photographing a spectrum of such a diffuse object took many hours of exposure. When he measured the wavelengths of its absorption lines, he found them shifted toward the blue by an amount corresponding to about 300 kilometres per second of approach, a larger velocity than had been measured for any astronomical object up to that point.

Over the following years Slipher measured more spirals, and almost all of them were receding, some at over a thousand kilometres per second. That pattern would eventually become the expansion of the universe, which is Cosmology's subject. But Andromeda is one of the exceptions, and the reason is the subject of this lesson. Andromeda is not participating in the expansion relative to us, because it is gravitationally bound to the Milky Way. The two galaxies are falling toward each other.

This final lesson zooms out one last time: what kinds of galaxy exist, which ones we live among, and what happens at the centres of a small minority of them that makes those centres outshine everything else in the universe.

Hubble's sequence, and what it does not mean

Edwin Hubble classified several hundred extragalactic nebulae in a 1926 paper, and later drew the scheme as the tuning fork diagram that still appears in every textbook.

Ellipticals, labelled E0 to E7, are smooth featureless ovals, the number giving how flattened they appear. Lenticulars, S0, have a disk but no spiral arms. Spirals split into two prongs, ordinary S and barred SB, each running a, b, c according to how tightly the arms wind and how prominent the central bulge is. Irregulars have no organised structure.

Hubble called the ellipticals early type and the spirals late type, and those terms are still used. They are misleading and he knew it. There is no evolutionary sequence running from left to right on the diagram; ellipticals do not turn into spirals or the reverse. If anything, the modern picture runs the other way, with spirals merging to produce ellipticals. Read early and late as arbitrary labels.

What the shapes correspond to physically

TypeStarsGas and dustSupportTypical massWhere found
EllipticalOld, red, little current star formationVery little cold gasRandom stellar motions10^8 to 10^13 solar massesCluster cores
LenticularMostly oldDepletedRotating disk plus dispersion10^10 to 10^11Groups and clusters
SpiralMixed ages; young blue stars in armsAbundant, concentrated in the diskOrdered rotation10^9 to 10^12Field and group environments
IrregularOften young, metal-poorGas richDisordered10^8 to 10^10Often satellites of larger galaxies
Dwarf spheroidalOld, extremely metal-poor, very fewEssentially noneRandom motions10^5 to 10^8 in starsSatellites; the most numerous type by count

Two entries deserve emphasis. The largest ellipticals, the cD galaxies sitting at the centres of rich clusters, are the most massive galaxies known and appear to have grown by consuming their neighbours. At the other extreme, dwarf spheroidals are the most common galaxies in the universe by number and the hardest to find, since some contain only a few hundred thousand stars spread thinly. They also have the highest mass-to-light ratios known, in some cases requiring a hundred times more mass than their stars provide, which makes them among the cleanest laboratories for dark matter.

Modern surveys have added a complementary way of sorting galaxies that is physical rather than visual. Plot colour against luminosity for a large sample and galaxies separate into a red sequence of passive, mostly elliptical systems and a blue cloud of star-forming, mostly spiral ones, with a sparsely populated green valley between them. The sparseness of the valley says that galaxies cross it quickly, which means star formation shuts down rapidly rather than fading away, and explaining that shutdown is one of the central problems in the field.

Remember: Morphology tracks star formation history and dynamical support, not age or an evolutionary sequence along the tuning fork.

Relations that hold across whole galaxies

Galaxies are messy, so it is striking that they obey tight empirical relations.

The Tully-Fisher relation links a spiral's luminosity to the fourth power of its rotation speed, holding across several orders of magnitude. It is used as a distance indicator: measure the width of the 21 centimetre line, get the rotation speed, get the luminosity, compare with apparent brightness.

For ellipticals, the Faber-Jackson relation plays the same role using the velocity dispersion of the stars, and it is one projection of a tighter three-parameter relation called the fundamental plane connecting size, surface brightness and dispersion.

The most surprising is the M-sigma relation, established around 2000. The mass of a galaxy's central black hole correlates tightly with the velocity dispersion of its bulge, roughly as the fourth or fifth power, and amounts to something like a tenth of a per cent of the bulge mass. The surprise is one of scale: a black hole's gravitational influence extends only a few parsecs, while the bulge spans thousands. The two should not know about each other. That they do is the main evidence that the growth of a black hole and the growth of its host galaxy regulate one another, which brings us to active nuclei.

The neighbourhood

The Milky Way belongs to a small, loose collection called the Local Group, containing roughly a hundred known members and gaining new ones regularly as surveys find faint dwarfs.

Two galaxies dominate. Andromeda, M31, lies about 785 kiloparsecs away, and the Milky Way is comparable in mass. The third-ranked member is M33, the Triangulum galaxy, a modest spiral. Everything else is a dwarf. The Large and Small Magellanic Clouds, at roughly 50 and 62 kiloparsecs, are the most conspicuous, visible to the unaided eye from the southern hemisphere, and the Large Cloud is massive enough that its gravity noticeably disturbs the Milky Way's outer halo. The rest of the membership is dwarf spheroidals and ultra-faint systems, many discovered only in the last two decades.

The Local Group has no dense centre and no hot intracluster gas. It is a group, not a cluster, and it sits on the outskirts of the much larger Virgo Cluster's sphere of influence.

Andromeda's approach at roughly 110 kilometres per second relative to the Milky Way has for years been reported as a certain future collision in about four to five billion years, producing a merged elliptical. That statement deserves a caveat. The radial velocity is well measured; the transverse velocity, which decides whether the encounter is head-on or a wide pass, is far harder and depends on proper motions at the limit of what current instruments can do, and on the gravitational influence of M33 and the Large Magellanic Cloud. More recent analyses find the outcome considerably less certain than the standard account implies. Treat the collision as likely but not settled, and check the date on any source that states it flatly.

Nuclei that outshine their galaxies

In 1943 Carl Seyfert published a study of a dozen spiral galaxies with unusually bright, starlike nuclei whose spectra showed broad emission lines, indicating gas moving at thousands of kilometres per second. The work sat largely unused for twenty years.

Then, in 1963, Maarten Schmidt examined the spectrum of 3C 273, a radio source that appeared to coincide with an ordinary looking star. Its spectrum made no sense: a set of strong emission lines matching nothing in any catalogue. Schmidt realised the lines were the familiar Balmer series of hydrogen, shifted redward by 16 per cent. At that redshift the object lay far outside the Galaxy, and to appear as bright as it did from that distance it had to be radiating something like 4 x 10^12 solar luminosities, over a hundred times the entire Milky Way.

Worse, quasars vary. Some change brightness substantially in days to weeks. An object cannot vary coherently faster than light crosses it, so the emitting region must be no more than light-days across, a scale comparable to the solar system. A hundred galaxies' worth of light from a volume smaller than the orbit of Neptune.

Nuclear fusion cannot do this. Fusion converts 0.7 per cent of rest mass to energy. Accretion onto a black hole converts something like 6 to 40 per cent, depending on the hole's spin, because material spiralling inward through a disk radiates away a large fraction of its gravitational binding energy before crossing the horizon. Ten per cent is the usual working figure, more than ten times better than fusion, and it is the most efficient sustained energy release known in nature.

How much matter does that require? A quasar radiating 10^39 watts at ten per cent efficiency needs L divided by (0.1 c^2) kilograms per second, which is 10^39 / (0.1 x 9 x 10^16) = 1.1 x 10^23 kg per second, roughly a solar mass per year. That is a modest appetite, and it is why quasars can shine for tens of millions of years.

There is also a ceiling. Radiation pushes on infalling gas, and above the Eddington luminosity the outward push exceeds gravity and accretion chokes itself off. Numerically, L_Edd is about 1.26 x 10^31 watts per solar mass of black hole. So a source radiating 10^39 watts needs a black hole of at least 10^39 / 1.26 x 10^31, about 8 x 10^7 solar masses. Quasar luminosities therefore imply supermassive black holes directly, without any need to see one.

The upshot: Take the observed luminosity, divide by the Eddington limit per solar mass, and you have a lower bound on the black hole mass. The argument requires nothing but conservation of momentum.

One object, many names

Active galactic nuclei were catalogued under a confusing variety of names before anyone realised how many were the same thing seen differently: Seyfert 1 and Seyfert 2 galaxies, radio-loud and radio-quiet quasars, radio galaxies, BL Lacertae objects, blazars.

The unified model proposes a common structure. A supermassive black hole is surrounded by an accretion disk radiating across the ultraviolet and X-ray. Close to the disk, fast-moving clouds produce broad emission lines. Further out, slower clouds produce narrow lines. Surrounding the whole inner region is a thick doughnut of dust and molecular gas. And in some fraction of cases, a relativistic jet is launched along the rotation axis, sometimes extending for hundreds of kiloparsecs.

Then the appearance depends on viewing angle. Look down the axis and you see straight into the broad line region: a Seyfert 1 or a broad-lined quasar. Look through the torus from the side and the broad lines are hidden while the narrow lines, produced outside the torus, remain: a Seyfert 2. Look directly down a jet and its emission is beamed toward you and dominates everything: a blazar, violently variable and often polarised.

The model has real successes. Polarised light from some Seyfert 2 nuclei shows broad lines in reflection, exactly as expected if the broad line region exists but is hidden from direct view, which is close to a decisive test. It also has limits. Some genuine differences are not orientation at all, such as the split between radio-loud and radio-quiet systems, or objects observed to change type over years. Orientation explains much of the variety, not all of it.

The jet of M87 was noticed in 1918 by Heber Curtis as a straight ray emerging from the nucleus, decades before anything about it could be understood. It is now known to be a relativistic outflow launched from the vicinity of the 6.5 billion solar mass black hole imaged by the Event Horizon Telescope in 2019, and it is the same object seen at two ends of a century of instrumental progress.

Why almost no galaxy is active now

Quasars were far more common in the past. Their number density peaks at redshifts around two to three, corresponding to roughly ten to eleven billion years ago, and has declined steeply since. Meanwhile, essentially every massive galaxy examined closely turns out to contain a central supermassive black hole, active or not.

Put those together and the conclusion is that a quasar is a phase, not a species. Most large galaxies hosted an active nucleus at some point, and the Milky Way's own Sagittarius A star is a dormant example, currently accreting so little that it is many orders of magnitude below its Eddington limit.

What turns them off is thought to be the activity itself. Radiation and jets from an accreting black hole can heat and expel the surrounding gas, cutting off both the black hole's own supply and the galaxy's star formation. This idea, called AGN feedback, is what the M-sigma relation hints at, and it is a standard ingredient in galaxy formation models. It is also not fully verified: the coupling between a parsec-scale engine and a kiloparsec-scale galaxy is difficult to observe directly and is modelled with parameters rather than derived from first principles.

What matters here: Quasars are not exotic objects in a separate category. They are ordinary galaxies during a brief phase when the black hole they all contain happened to be well fed.

Common misconceptions

"The Hubble sequence is an evolutionary sequence." Ellipticals do not become spirals or the reverse along the tuning fork. The early and late labels are Hubble's terminology, not a claim about time, and mergers of spirals producing ellipticals runs opposite to the naive reading.

"Quasars are a different kind of object from ordinary galaxies." A quasar is a galactic nucleus in an active phase. The same galaxies, quieter, are all around us, and the black holes are still there.

"AGN are powered by many stars or by fusion." No stellar process reaches the required luminosity from a light-day-sized region. Only accretion onto a supermassive black hole, at roughly ten per cent mass-energy efficiency, works.

"The Milky Way and Andromeda will certainly collide in about four billion years." The approach is well measured, but the sideways motion is not, and analyses including the Large Magellanic Cloud and M33 make the outcome less certain than the familiar statement. It is a likely merger, not a scheduled one.

"Most galaxies are big spirals like the ones in pictures." Most galaxies are dwarfs, and most of those are faint spheroidals that surveys struggle to detect at all. The picturesque spirals are a bright, atypical minority, exactly as bright stars are among stars.

Pulling it together

Slipher's 1912 measurement of Andromeda approaching at 300 kilometres per second was the first sign that our galaxy has close neighbours bound to it. Hubble's tuning fork sorts galaxies into ellipticals, lenticulars, ordinary and barred spirals and irregulars, a classification that tracks star formation history and dynamical support rather than any evolutionary sequence, and which modern surveys supplement with the red sequence, blue cloud and sparsely populated green valley. Galaxies obey tight scaling relations, including Tully-Fisher for spirals, the fundamental plane for ellipticals, and the M-sigma relation linking a central black hole to a bulge thousands of times larger than its sphere of influence. The Local Group holds about a hundred members dominated by the Milky Way and M31, with M33 third and everything else a dwarf; the two large spirals are approaching, though the certainty of a merger has been overstated. A small fraction of galactic nuclei are active, and Schmidt's 1963 identification of a 16 per cent redshift in 3C 273 revealed objects radiating over a hundred galaxies' worth of light from a region light-days across. Only accretion onto a supermassive black hole, at roughly ten per cent efficiency against fusion's 0.7, can do that, and the Eddington limit converts an observed luminosity directly into a minimum black hole mass. The unified model explains much of the observed variety as an orientation effect, quasars peaked in number around ten to eleven billion years ago, and essentially every massive galaxy including our own contains the engine, mostly switched off.

Sources

  1. NASA Science. (n.d.). Galaxies. National Aeronautics and Space Administration. science.nasa.gov
  2. Encyclopaedia Britannica. (n.d.). Quasar. britannica.com
  3. Encyclopaedia Britannica. (n.d.). Galaxy. britannica.com
  4. Event Horizon Telescope Collaboration. (n.d.). The M87 black hole and its jet. EHT. eventhorizontelescope.org
  5. Fraknoi, A., Morrison, D., & Wolff, S. C. (2022). Astronomy 2e, Chapters 26 and 27: Galaxies, and active galaxies, quasars and supermassive black holes. OpenStax, Rice University. openstax.org
  6. Schmidt, M. (1963). 3C 273: A star-like object with large red-shift. Nature, 197, 1040.
Key terms
Hubble sequence
The tuning fork classification of galaxies into ellipticals, lenticulars, ordinary and barred spirals, and irregulars.
Dwarf spheroidal
A faint, gas-poor, dark-matter-dominated satellite galaxy; the most numerous galaxy type by count.
Red sequence and blue cloud
The two populations galaxies fall into on a colour-magnitude plot, separated by a sparsely occupied green valley.
Tully-Fisher relation
The tight link between a spiral galaxy's luminosity and the fourth power of its rotation speed, usable as a distance indicator.
M-sigma relation
The correlation between a galaxy's central black hole mass and its bulge velocity dispersion, evidence of coupled growth.
Local Group
The bound collection of about a hundred galaxies containing the Milky Way, M31 and M33, dominated by dwarfs in number.
Active galactic nucleus
A galactic centre powered by accretion onto a supermassive black hole, radiating far more than its stars.
Eddington luminosity
The luminosity at which radiation pressure balances gravity, about 1.26 x 10^31 watts per solar mass, setting a floor on an active nucleus's black hole mass.
Unified model
The account in which Seyfert types, quasars and blazars are the same structure viewed at different angles relative to an obscuring torus and jet.

Open the interactive version with quizzes and progress →