Module 1: From an Empty File to a Program That Talks Back
Get Python running, put your first lines on the screen, store values and read what somebody types, find out how Python does arithmetic, and learn to plan a program before you write it.
One Line, and What the Machine Does With It
- Run a Python program from a file and read the output it produces.
- Name the parts of a call to print and say what each part does.
- Read the last lines of an error message and use them to find the mistake.
Open an empty file. Type one line into it. Save it as hello.py. Then run it. The line is this:
print("Hello, world!")
And here is what came back, on its own line, in the black window where the program ran:
Hello, world!
Now change the small p at the start to a capital P and run exactly the same file again. Nothing else is different. Same words, same quotation marks, same file, same computer. This is what the machine says now, and these are the last three lines of what it said:
Print("Hello, world!")
^^^^^
NameError: name 'Print' is not defined. Did you mean: 'print'?
One character. That is the whole difference between a program that works and a program that stops. Nothing has broken, nothing is your fault, and the computer is not annoyed with you. It simply does not know a thing called Print, it knows a thing called print, and it has said so as plainly as it can, including a guess about what you meant. Most of learning to program is learning to close that gap between what you meant and what you typed, and by the end of this lesson you will have closed it twice on purpose.
Getting Python on your machine, or in your browser
You need Python 3.14. The projects use its standard library, without third-party Python packages. Turtle drawing later in the course also needs Tk and a graphical desktop; some systems require installing Tk separately.
If the computer is yours, go to python.org/downloads and get the installer for your system. On Windows, follow the official installation instructions for the Python Install Manager. Older installers use different screens, so do not depend on finding an Add Python to PATH checkbox. After installation, open a new terminal and check the version as shown below. If the command is unavailable, use the troubleshooting section of those instructions.
If the computer belongs to your school and you are not allowed to install software, you are not stuck. Open pyodide.org/en/stable/console.html in a browser. It loads a complete Python inside the web page, prints a banner naming the version, and then shows three greater-than signs waiting for you. Type print("Hello, world!") there and press Enter and the same line comes back. It is genuinely Python, compiled to run inside a browser, so you can try the first print examples there. Browser consoles differ in how they accept input and access files. Use a local Python installation for the full course when one is available.
The window where you type python hello.py has several names depending on who is talking: terminal, command line, shell, console, Command Prompt, PowerShell. They all mean the same thing here, which is a place where you type a command, press Enter, and the machine does it. To check that the installation worked, type this and press Enter:
python --version
You should see a line starting with the word Python and then a number beginning with 3. Use Python 3.14 to match this course. Older Python 3 versions may lack features or display different error messages. On systems where the command is python3, use python3 --version and python3 hello.py instead.
A program is a list of orders, obeyed exactly
A program is a list of instructions a computer carries out, one after another, in the order written. The word doing the work in that sentence is "exactly". A computer is extraordinarily fast and completely obedient, and it is not clever in the way a person is clever. It does not guess. It does not round off what you meant. If you ask a friend to grab you a drink, they fill in fifty details you never said: which glass, from which fridge, not the one at the back that has been there since August. A program fills in nothing at all. Every step has to be there.
That is why your capital P stopped the machine. A person reading Print would have read straight past it. Python cannot, because Print and print are two different names to it, in the same way that Nadia and nadia would be two different filenames to a strict filing clerk. Case matters. It will matter for the whole course, so it is worth being annoyed by it now rather than in week six.
Underneath, the chip in your computer does not understand Python at all. It understands very short, very dull instructions like add these two numbers or move this value there, written as patterns of ones and zeros. You write Python source instead. A program called the interpreter checks and executes that source. CPython, the implementation from python.org, first compiles the file to an internal form called bytecode. It then executes the instructions. This is why a syntax error near the end of a file can prevent even its first print from running.
Reading one line of Python: name, parentheses, argument
Look at the working line again, slowly, because there are four separate things happening in this short call.
print("Hello, world!")
The word print is a name. It names a piece of work that Python already knows how to do, which is putting something on the screen. A named piece of work like this is a built-in function, and you will write functions of your own in Lesson 8.
The pair of round brackets after the name is not decoration. It is the instruction to do it, now. Writing print on its own tells Python nothing except that a thing with that name exists; writing print() tells Python to run it.
Inside the brackets sits the piece of information you are handing over, which is called the argument. Here it is a run of characters between quotation marks, and a run of characters in quotation marks is a string. The quotation marks are the whole point. They mean: treat what is inside as literal text, letter by letter, do not try to work out what it means. That is why Hello, world! survived the trip from your file to the screen unchanged, including the comma, the space and the exclamation mark.
Single quotes work the same way as double quotes in Python, so 'Hello' and "Hello" are the same string. This course uses double quotes for text you print and single quotes when a piece of text has to sit inside another one, which keeps the two kinds of quotation mark from colliding.
Several lines, and the order they happen in
Real programs are longer than one line. When the interpreter runs a file it works from the top down, and it never jumps around unless you explicitly tell it to, which is what Lessons 5 to 7 are about. Here is a three-line file:
# My second program. Three lines, run top to bottom.
print("Hello, world!")
print("My name is Nadia.")
print("I am fourteen and this is my second program.")
Running it gives three lines, in the order the lines appear in the file:
Hello, world!
My name is Nadia.
I am fourteen and this is my second program.
Each call to print produced its own line. That is print's own habit: when it has finished printing what you gave it, it moves down to a fresh line. You did not ask for that and you cannot see it in the code, which is exactly the sort of quiet built-in behaviour worth noticing early.
Commas, and printing a sum you did not work out yourself
You can hand print more than one thing at a time by putting commas between them. When you do, print puts them all on one line with a single space between each one. And the things you hand it do not have to be text. Here is a small program that mixes both, saved as card.py:
# A tiny report card. Some lines print text, some print arithmetic.
print("================================")
print(" SCHOOL DAY, BY THE NUMBERS")
print("================================")
print("Lessons today:", 6)
print("Minutes per lesson:", 55)
print("Minutes in lessons:", 6 * 55)
print("That is this many hours:", 6 * 55 / 60)
And its output, exactly as it came back:
================================
SCHOOL DAY, BY THE NUMBERS
================================
Lessons today: 6
Minutes per lesson: 55
Minutes in lessons: 330
That is this many hours: 5.5
Stop on the sixth line of output for a moment, because something happened there that did not happen on the first three. You typed 6 * 55. What appeared was 330. Python did the multiplication before printing, because 6 * 55 is not in quotation marks, so Python treated it as a sum to work out rather than as characters to copy. The star is the multiplication sign on a keyboard. The slash on the last line is division, and it gave 5.5 rather than 5, decimals included.
Compare that with the first three lines of output, which are rows of equals signs and a title. Those were in quotation marks, so they came through untouched, spaces and all. This single distinction, between quoted text that is copied and unquoted code that is worked out, is the one you will use most often in the whole course. If you ever print something and get the characters of your sum back instead of the answer, look for a stray quotation mark.
Comments, which Python throws away
The first line of card.py begins with a hash mark:
# A tiny report card. Some lines print text, some print arithmetic.
A # outside a string starts a comment that continues to the end of the line. In print("#1"), the hash is part of the text and gets printed. Python skips the comment in card.py; that comment never appears in the output, which you can confirm by looking above: no line of that output mentions a report card. Comments exist for people, and one of the people is you in three weeks, when you have forgotten why you wrote something.
A comment that says what the code obviously says is wasted. # add one to the score above a line that adds one to the score tells nobody anything. A comment that says why is worth keeping: # the scoreboard starts at 1, not 0, so shift everything up. While you are a beginner, err on the side of writing too many, because putting a line into English forces you to check that you know what it does.
While we are talking about how code looks rather than what it does: Python has a written style guide, PEP 8, that nearly every Python programmer follows. You do not need it yet. Two habits from it are worth starting with, because they cost nothing. Put one statement on one line. Put a space on each side of an operator, so 6 * 55 rather than 6*55. Your file will look like everybody else's, which matters the first time you paste a problem into a forum and ask for help.
Two mistakes, made on purpose
You are going to make mistakes constantly, forever, at every level of skill. The only thing that changes with experience is how quickly you read the message and fix it. So let us break things deliberately while nothing is at stake.
First, the capital P from the top of the lesson. When a program stops with an error, Python prints a traceback: a short report saying where it was and what went wrong. Its first line says Traceback (most recent call last). Then comes a line beginning with File, giving the full name of your file on your machine and the line number, and a line saying in <module>, which just means the mistake was in the file itself rather than inside a function. Then the offending line is quoted back at you with carets underneath pointing at the part that failed, and the last line names the error. Read the last line first. It is the one that tells you what happened:
NameError: name 'Print' is not defined. Did you mean: 'print'?
A NameError means Python met a name it has never been given. That is the whole message. The fix is to correct the spelling.
Second, take card.py and delete one quotation mark, the closing one on the sixth line, so the line reads print("Minutes per lesson:, 55). Run it. This time the output is different in a way worth noticing:
print("Minutes per lesson:, 55)
^
SyntaxError: unterminated string literal (detected at line 6)
Nothing at all was printed first. Not even the rows of equals signs, which come earlier in the file and are perfectly correct. A SyntaxError is different from the other kind: it means Python could not even read your file as Python, so it never started running any of it. An opened quotation mark that is never closed swallows the rest of the line, and Python reaches the end of the line still waiting for the partner. Unterminated means exactly that: you started something and did not finish it.
The practical difference between those two errors is worth holding on to. If you see some output and then a traceback, the program ran and fell over partway. No output by itself does not identify the problem: a valid program may print nothing, and a runtime error may happen before its first print. Read the error type. A SyntaxError in the file being started means Python could not parse that file; a NameError means execution reached a name it could not find.
Common misconceptions
- "The computer knew what I meant." It did not, and it never will. It followed what you wrote. When the output is wrong the instructions were wrong, and the error message is telling you which line it was standing on when it lost the thread.
- "An error means I am bad at this." Errors are what programming looks like from the inside, for everyone, at every level. They are the machine being helpful in the only way it can: naming the line and the problem.
- "Capital letters are just style." Not in code.
Print,printandPRINTare three different names to Python, and only one of them exists. - "Quotation marks are optional around text." They are what makes text text.
print("Nadia")prints the word.print(Nadia)makes Python hunt for something called Nadia, find nothing, and stop with a NameError.
What you now know
Key idea: A program is a list of orders carried out exactly and in order, and an error message is a location report, not a verdict.
You installed Python or opened one in a browser, checked its version, and ran code you wrote yourself. You can put text on the screen with print, hand it several things at once with commas, and mix literal text in quotation marks with arithmetic that Python works out before printing. You know that a hash mark outside a string starts a comment for human readers. You have read two errors: a NameError, which happens while the program is running and does not undo output already printed, and a SyntaxError, which stops the file before a single line runs.
The next lesson gives your programs a memory. Right now every value you have printed vanished the instant it appeared. Storing a value under a name, and asking the person at the keyboard for one, turns a program that recites into a program that answers.
Sources
- Python Software Foundation. (2026). The Python tutorial: An informal introduction to Python. docs.python.org
- Python Software Foundation. (2026). Built-in functions: print. Python 3.14 documentation. docs.python.org
- Python Software Foundation. (2026). Errors and exceptions, Sections 8.1 and 8.2. Python 3.14 documentation. docs.python.org
- Python Software Foundation. (2026). Lexical analysis, Section 2.1.3: Comments. Python 3.14 documentation. docs.python.org
- van Rossum, G., Warsaw, B., and Coghlan, N. (2001, revised 2023). PEP 8: Style guide for Python code. Python Software Foundation. peps.python.org
- Key terms
- Program
- A list of instructions a computer carries out, one after another, in the order written.
- Interpreter
- The program that checks and executes Python code; CPython compiles source to bytecode before executing it.
- Function
- A named piece of work you run by writing its name followed by round brackets.
- Argument
- A piece of information handed to a function inside its brackets, such as the text print puts on the screen.
- String
- A run of characters between quotation marks, treated as literal text rather than as code to work out.
- Comment
- Text after a hash mark outside a string, through the end of that line; a note for people.
- Traceback
- The report Python prints when a program stops, naming the file, the line and the kind of error.
- SyntaxError
- An error in the shape of the code that stops Python before any line runs, such as a quotation mark never closed.
Giving the Program a Memory, and Asking It a Question
- Store a value under a name and reuse it instead of repeating it.
- Tell the four basic types apart and explain why quotation marks change the answer.
- Read a value from the keyboard with input, convert it with int, and print it inside an f-string.
Here is a three-line program that works perfectly and is still badly written:
print("Good morning, Priya.")
print("Priya, you have 3 pieces of homework.")
print("See you tomorrow, Priya.")
Good morning, Priya.
Priya, you have 3 pieces of homework.
See you tomorrow, Priya.
Now suppose it is not Priya. Suppose you want to hand this program to someone else in your form. You have to find the name three times and change it three times, and if you miss one, the program greets Priya and says goodbye to Sam and nothing warns you, because nothing is wrong as far as Python is concerned. With three lines that is irritating. With three hundred it is a bug waiting to happen. The whole of this lesson comes out of that one problem: a program needs somewhere to put a value so it can be written once and used everywhere.
A name that stands for a value
A variable is a name that stands for a value. You make one by writing the name, an equals sign, and the value:
name = "Priya"
homework_count = 3
print("Good morning,", name)
print(name, "you have", homework_count, "pieces of homework.")
print("See you tomorrow,", name)
Good morning, Priya
Priya you have 3 pieces of homework.
See you tomorrow, Priya
The name Priya now appears once in the file. Change the first line to name = "Sam" and all three lines of output change together. That is the point of a variable and it is worth more than it looks: a value that lives in one place can only be wrong in one place.
Notice what got worse, though. The full stops have gone. In the first version the text after Priya was inside the quotation marks; now the comma in print puts a space there instead and there is nowhere to put a full stop without an awkward extra argument. Hold that thought, because it is what f-strings fix at the end of this lesson.
What the equals sign is actually doing
In maths, x = 5 is a statement about the world that is either true or false. In Python it is an order: take the value on the right, and from now on let the name on the left refer to it. It is called assignment, and it only ever runs in one direction, right to left. 5 = x is not a Python statement at all; it is a syntax error, because you cannot assign a new meaning to the number five.
This matters because of what comes next. Look at this program, which keeps a score:
score = 0
print("At the start, score is", score)
score = score + 10
print("After the first question, score is", score)
score = score + 10
print("After the second question, score is", score)
score = "nine hundred"
print("And now score is", score, "which is a", type(score))
At the start, score is 0
After the first question, score is 10
After the second question, score is 20
And now score is nine hundred which is a <class 'str'>
The line score = score + 10 would be nonsense in a maths lesson, because no number equals itself plus ten. As an order it is perfectly sensible: work out the right-hand side using whatever score means right now, which is 0, get 10, and make score mean that instead. The name score now refers to 10. Assignment does not keep a history for that name, although another variable could still refer to an earlier value.
The last two lines show something else. A variable in Python is not a box of a fixed shape. The name score held a whole number, then a different whole number, and then a piece of text, and Python never objected. That freedom is convenient and it is also how a certain kind of bug gets in, which is why the next section is about types.
Four kinds of value, and how to ask which is which
Every value in Python has a type, and the type decides what you are allowed to do with it. Four of them will carry you through most of this course. The built-in function type will tell you which one you have:
name = "Priya"
homework_count = 3
hours_left = 2.5
is_finished = False
print(name, type(name))
print(homework_count, type(homework_count))
print(hours_left, type(hours_left))
print(is_finished, type(is_finished))
print("3" + "4")
print(3 + 4)
Priya <class 'str'>
3 <class 'int'>
2.5 <class 'float'>
False <class 'bool'>
34
7
| Type | Python calls it | Holds | Example |
|---|---|---|---|
| String | str | Characters, in quotation marks | "Priya" |
| Integer | int | A whole number, no decimal point | 3 |
| Float | float | A floating-point number, often an approximation | 2.5 |
| Boolean | bool | True or False, and nothing else | False |
The last two lines of that output are the ones to stare at. "3" + "4" gave 34 and 3 + 4 gave 7. The plus sign did two completely different jobs, and the only thing that told it which job to do was whether the values had quotation marks round them. Between two numbers, plus adds. Between two strings, plus joins them end to end, an operation called concatenation. Python is not confused here and it has not made a mistake; it did precisely what the types asked for. The built-in types documentation lists what every operator does for each type, and it is a page worth bookmarking now and reading properly in a month.
Booleans look like the least useful type here because they only have two values. They turn out to run every decision your programs make from Lesson 5 onwards, so meet them now: True and False, capital letters, no quotation marks.
Asking the person at the keyboard
So far your programs say the same thing every time. The function input changes that. It prints the text you give it, waits for somebody to type a line and press Enter, and hands back what they typed. Here is a program that asks for an age and adds one to it:
age = input("How old are you? ")
print("Next year you will be", age + 1)
It does not work. Fed the answer 14, here is the end of the error report:
print("Next year you will be", age + 1)
~~~~^~~
TypeError: can only concatenate str (not "int") to str
A TypeError means you asked for an operation that does not exist for those types, and the squiggles underneath point at the exact part of the line that failed, which here is age + 1. Read the message literally, because it is a full explanation: Python can only concatenate a string to a string, and you handed it an int. Which means age is a string. Which means input handed back a string.
That is the rule, and it catches everyone once: input always gives you a string, even when the person typed digits. The characters 1 and 4 are not the number fourteen any more than the word "fourteen" is. To get a number you convert it, using int for whole numbers or float for numbers with decimal points.
f-strings, and getting the full stops back
Here is the fixed program, and it also fixes the punctuation problem from the top of the lesson:
name = input("What is your name? ")
age_text = input("How old are you? ")
age = int(age_text)
print(f"Hello, {name}.")
print(f"Next year you will be {age + 1}.")
print(f"In ten years you will be {age + 10}, and {name} will still be {name}.")
Fed the answers Priya and 14, this is exactly what came back:
What is your name? How old are you? Hello, Priya.
Next year you will be 15.
In ten years you will be 24, and Priya will still be Priya.
Two prompts are sitting on one line and the answers are nowhere to be seen, which looks wrong until you know why. The answers were fed in from a file rather than typed. When you run this yourself and type, your answer appears just after each prompt, because the terminal echoes the keys you press; Python never printed them. Run it both ways once so the difference stops being mysterious.
Now the new piece. Putting the letter f immediately before the opening quotation mark makes an f-string, short for formatted string. Inside it, anything in curly brackets is worked out and dropped into the text at that spot. Everything outside the curly brackets is ordinary literal text, including the full stops and commas you wanted. That is why f"Hello, {name}." printed Hello, Priya. with the full stop tight against the name, which the comma version could not do.
The curly brackets can hold a calculation, not just a name: {age + 1} printed 15. And the same name can appear as many times as you like. From here on this course uses f-strings for nearly all output, because they put the shape of the finished line in front of you while you are writing it.
Names worth reading
You may call a variable almost anything: letters, digits and underscores, not starting with a digit, and not a reserved keyword such as if or while. The built-in name print is legal as a variable name, but avoid using it: an assignment such as print = 3 hides the function under that name. Which means you can call it x, or thing2, or homework_count, and Python will not care in the slightest. Your future self will.
The Python style guide, PEP 8, asks for lowercase names with underscores between words: homework_count, hours_left, is_finished. Two habits worth copying from the programs above. Names of things are nouns, so score rather than calculate. Names holding True or False read like questions, so is_finished rather than finished_status, because then if is_finished reads as English. A good name is the cheapest comment you will ever write.
Common misconceptions
- "The equals sign means the two sides are equal." It means evaluate the right-hand side and bind the name on the left to the resulting value. That is why
score = score + 10is a sensible instruction rather than an impossible equation, and why5 = xis not valid Python at all. - "A variable is a box that keeps a copy of everything you put in it." It keeps only the current value. When you assign again, the old value is not stored underneath and there is no undo. If you need the old value later, put it in a second variable before you overwrite the first.
- "If somebody types 14, the program has the number 14." It has the two characters 1 and 4 as a string. Until you pass it through int, arithmetic on it will either fail with a TypeError or, worse, quietly do something you did not mean.
- "Quotation marks around a number are harmless." They change a number into text, so operations can behave differently.
"3" + "4"is 34 and3 + 4is 7, and neither is a bug.
Putting it together
The point: A variable is a name pointing at a value, the type of that value decides what the operators do to it, and everything input hands you is text until you convert it.
You can now write a program that remembers things. You store a value with a single equals sign, reuse the name as often as you like, and change what it refers to whenever you want, including to a value of a different type. You can tell a str from an int from a float from a bool, and you know why the plus sign does one thing to numbers and another to text. You can ask a question with input, remember that the answer arrives as a string, convert it with int or float, and print a tidy line with an f-string that puts the calculation exactly where you want it.
The 5.5 hours from the first lesson is correct: 330 minutes includes half an hour beyond five hours. Next you will choose between whole-number and fractional answers and see why some fractions can only be approximated by floats. The next lesson is about the two places where Python's numbers do something surprising, both of which will eventually cost you an evening if you do not meet them now on purpose.
Sources
- Python Software Foundation. (2026). The Python tutorial: An informal introduction to Python (numbers, strings, and the assignment operator). docs.python.org
- Python Software Foundation. (2026). Input and output: Formatted string literals. Python 3.14 documentation. docs.python.org
- Python Software Foundation. (2026). Built-in types (str, int, float and bool, and what the operators do to each). docs.python.org
- Severance, C. (2016). Python for Everybody: Exploring data in Python 3, Chapter 2: Variables, expressions and statements. py4e.com
- Python Software Foundation. (2026). Lexical analysis, Section 2.3: Names. Python 3.14 documentation. docs.python.org
- Key terms
- Variable
- A name that stands for a value, created by writing name, equals sign, value.
- Assignment
- The action of the single equals sign: work out the right-hand side, then make the left-hand name refer to it.
- Type
- What kind of value something is. The type decides what the operators are allowed to do with it.
- String (str)
- Text: a run of characters in quotation marks. Adding two strings joins them end to end.
- Integer (int)
- A whole number with no decimal point, such as 3 or -12.
- Float
- A floating-point number, such as 2.5 or 1e3. Many fractional values are stored approximately.
- Boolean (bool)
- A value that is either True or False, written with capital letters and no quotation marks.
- f-string
- A string with the letter f before the opening quote, where anything in curly brackets is worked out and inserted.
- TypeError
- The error Python raises when you ask for an operation that does not exist between those two types.
Why 0.1 Plus 0.2 Is Not 0.3, and Other Arithmetic You Should Not Trust
- Choose between / and // by asking what question the answer has to answer.
- Use the remainder operator to split a quantity into whole parts plus what is left over.
- Explain why floats are approximations and handle money by counting in whole units.
Seventeen slices of pizza, five people, and a rule that slices must stay whole. Here is a program that shares them out, and its answer is wrong in a way that is worth an entire lesson:
slices = 17
people = 5
print("Slices each:", slices / people)
Slices each: 3.4
Nothing has crashed. There is no traceback, no red text, no error. Python has done exactly what it was asked and returned 3.4, and the answer still does not satisfy our rule of serving only whole slices. This is the most dangerous kind of wrong answer: the kind that looks like an answer.
The mistake was not in the arithmetic. It was earlier, in deciding what question to ask. There are two different questions hiding inside the word divide, and Python has a separate operator for each of them.
Two divisions, because there are two questions
The first question is: if I could cut anything into any fraction, how much is each share? That is /, ordinary division, and the answer is 3.4. The second question is: how many whole slices does each person get, and how many are left on the plate? Those are //, called floor division or integer division, and %, the remainder or modulo operator. Here is the whole program:
slices = 17
people = 5
print("Slices each:", slices / people)
print("Whole slices each:", slices // people)
print("Slices left over:", slices % people)
print("Check:", people * (slices // people) + (slices % people))
Slices each: 3.4
Whole slices each: 3
Slices left over: 2
Check: 17
Three each and two arguments about the last two. The final line is a habit worth acquiring now: whenever you split something, multiply the parts back and add the remainder, and see whether you get what you started with. It got 17, so nothing was invented and nothing was lost. Checks like that catch more bugs than staring at code does.
The double slash rounds toward negative infinity. For these positive operands, that gives the number of whole groups. 19 // 5 is 3 and so is 17 // 5, because in both cases three whole fives fit. The remainder is what fell off the end, and for a positive integer divisor it is nonnegative and smaller than that divisor, which is a useful sanity check in itself: if x % 5 ever came out as 7, something is very wrong.
The seven operators, and which one happens first
Here is every arithmetic operator you need, run on the same two numbers so you can compare them side by side:
a = 17
b = 5
print("a + b =", a + b)
print("a - b =", a - b)
print("a * b =", a * b)
print("a / b =", a / b)
print("a // b =", a // b)
print("a % b =", a % b)
print("a ** 2 =", a ** 2)
print("type of a / b is", type(a / b))
print("type of a // b is", type(a // b))
print("2 + 3 * 4 =", 2 + 3 * 4)
print("(2 + 3) * 4 =", (2 + 3) * 4)
print("-7 // 2 =", -7 // 2)
print("-7 % 2 =", -7 % 2)
a + b = 22
a - b = 12
a * b = 85
a / b = 3.4
a // b = 3
a % b = 2
a ** 2 = 289
type of a / b is <class 'float'>
type of a // b is <class 'int'>
2 + 3 * 4 = 14
(2 + 3) * 4 = 20
-7 // 2 = -4
-7 % 2 = 1
Four things in that output deserve a second look.
The double star is a power, so 17 ** 2 is 289. The caret is the bitwise XOR operator, not a power operator; if you write 17 ^ 2 you get 19, because the caret means something entirely different that has nothing to do with powers.
A single slash applied to these integers gives a float, even when the division comes out even. 10 / 5 is 2.0, not 2. That decimal point is Python telling you the answer is of the float type, and it will show up in your output whether you wanted it or not.
Multiplication happens before addition, exactly as in a maths lesson, so 2 + 3 * 4 is 14 rather than 20. Brackets override it. When a line of arithmetic gets long enough that you have to think about the order, put the brackets in anyway. They cost nothing and they tell the next reader what you meant.
Floor division rounds down, not towards zero. -7 // 2 is -4, not -3, because -4 is the whole number below -3.5. Most of the time you are working with positive numbers and never notice. The day you are working out how many places a scoreboard has moved and the answer went negative, you will notice.
A mean that came out too low
Here is a bug in the wild. Three test marks, and a program that reports the average:
mark1 = 58
mark2 = 71
mark3 = 67
wrong_mean = (mark1 + mark2 + mark3) // 3
right_mean = (mark1 + mark2 + mark3) / 3
print("Total:", mark1 + mark2 + mark3)
print("Mean with // :", wrong_mean)
print("Mean with / :", right_mean)
print(f"Mean, one decimal place: {right_mean:.1f}")
print("Rounded to a whole mark:", round(right_mean))
Total: 196
Mean with // : 65
Mean with / : 65.33333333333333
Mean, one decimal place: 65.3
Rounded to a whole mark: 65
The version with // reports 65 and nothing complains. It rounds down to 65 rather than preserving the fractional part. Had the marks been 58, 71 and 69, the true mean would have been 66.0 and // would have given 66, agreeing with the right answer by luck. That is what makes this bug nasty: it can pass tests whose means are whole numbers and be too low by less than one mark on other cases, and only ever downwards.
The last two lines show the two ways to make a long decimal readable, and they are not the same thing. {right_mean:.1f} inside an f-string changes only how the number is printed this once; the variable still holds 65.33333333333333. The function round produces a new, genuinely rounded value you can go on calculating with. Format for display; round when the rounded number is the thing you actually mean.
Three cans at 1.10, and the change that was not 1.70
Now the second surprise, and this one is not Python's fault at all. Three cans at 1.10 each, paid for with a five pound note:
price = 1.10
total = price * 3
change = 5.00 - total
print("Total:", total)
print("Change:", change)
print("Is the total 3.30?", total == 3.30)
print("Price times 100:", price * 100)
Total: 3.3000000000000003
Change: 1.6999999999999997
Is the total 3.30? False
Price times 100: 110.00000000000001
Three times 1.10 is not 3.30. One pound ten times a hundred is not a hundred and ten. And when you ask Python directly whether the total equals 3.30, it says False. These tiny representation errors can affect comparisons and later calculations. The displayed values alone do not establish an actual cash loss; payment systems need explicit rounding rules.
Why this happens, and why it is not a bug
Python floats on usual desktop systems store numbers in binary, in halves and quarters and eighths, not in tenths. Some fractions fit perfectly: a half is 0.1 in binary, a quarter is 0.01. One tenth does not fit at all. In binary it runs 0.0001100110011001100... forever, exactly as one third runs 0.3333... forever in decimal and never finishes.
A float has a fixed amount of room, typically about 15 to 17 significant decimal digits, so the endless pattern gets cut off. What is stored is the nearest number that does fit. You can see the real stored value by asking for twenty decimal places:
print(0.1 + 0.2)
print(0.1 + 0.2 == 0.3)
print(1.1 * 3)
print(f"{0.1:.20f}")
0.30000000000000004
False
3.3000000000000003
0.10000000000000000555
That last line is the whole explanation in one number. The thing Python calls 0.1 is not one tenth. It is a hair over one tenth. Add two of these hairs together and the error is large enough to appear in the printed answer. The Python tutorial chapter on floating point gives the stored value to its full length, and it ends ...15625, because it is an exact binary fraction that happens not to be a tenth.
This is not a Python quirk. It is common for binary floating-point arithmetic on modern computers, described by a standard called IEEE 754. Other languages also offer binary floating point, and some provide decimal types. The displayed number depends on the type and formatting.
Two fixes, and when to use each
The first fix is the serious one: if a quantity is made of whole things, count the whole things. Money is made of pence. So work in pence, which are integers, and Python integers have arbitrary precision, limited in practice by available memory.
price_p = 110
total_p = price_p * 3
change_p = 500 - total_p
print("Total in pence:", total_p)
print("Change in pence:", change_p)
print(f"Change as money: {change_p // 100} pounds {change_p % 100} pence")
print(f"Or printed straight: {change_p / 100:.2f}")
Total in pence: 330
Change in pence: 170
Change as money: 1 pounds 70 pence
Or printed straight: 1.70
Exact, every time, and look at what pulled the pounds and the pence apart on the fourth line: floor division and remainder, the same two operators that shared out the pizza. That pattern, total // size for how many whole units and total % size for what is left, is one of the most reused two-liners in programming. Seconds into minutes and seconds, pence into pounds and pence, days into weeks and days.
The second fix is for display only. {value:.2f} in an f-string prints a float to two decimal places. Use it to make output readable. Formatting does not repair stored values. Exact equality is appropriate when you need exactly equal stored values; for approximate measurements, use a tolerance suited to the problem, such as the math.isclose function in Lesson 18. For this whole-pence calculation, compare integers.
Common misconceptions
- "0.1 + 0.2 giving 0.30000000000000004 is a bug in Python." It is expected with the binary floating-point representation used here. The numbers 0.1 and 0.2 cannot be stored exactly in binary, so their sum cannot be either.
- "// rounds to the nearest whole number." It rounds down, always.
19 // 5is 3 even though 3.8 is nearer to 4, and-7 // 2is -4 rather than -3. - "10 / 5 gives the integer 2." It gives 2.0, a float, because / on two integers produces a float. If you need a whole number for counting or indexing, use
//or wrap the result in int. - "Rounding for display fixes the underlying value." An f-string format changes the characters printed and nothing else. The variable still holds the long value, and the next calculation will use the long one.
Where this leaves us
What matters here: Ask what the answer has to be before you choose the operator, and never store a quantity of whole things in a float.
You have two divisions now and a reason to choose between them: / when a fraction of the thing makes sense, // with % when it does not. You know that the double star is a power, that multiplication binds tighter than addition, and that floor division rounds downwards even past zero. You have seen a mean lose its fractional part from one wrong slash, and change come out a fraction of a penny short because a price was written as a decimal. And you know the two repairs: count in the smallest whole unit, and format only at the moment of printing.
Everything so far runs straight down the file and does the same thing every time. The next lesson does not add any new Python at all. It is about what to do before you type, because the programs from here on are long enough that starting at line one and hoping stops working.
Sources
- Python Software Foundation. (2026). Floating-point arithmetic: Issues and limitations. Python 3.14 tutorial. docs.python.org
- Python Software Foundation. (2026). The Python tutorial: Using Python as a calculator (the division, floor division and power operators). docs.python.org
- Python Software Foundation. (2026). Built-in types: Numeric types int, float, complex (operator table and the behaviour of floor division). docs.python.org
- IEEE. (2019). IEEE Standard for Floating-Point Arithmetic (IEEE 754-2019). Institute of Electrical and Electronics Engineers.
- Severance, C. (2016). Python for Everybody: Exploring data in Python 3, Chapter 2: Variables, expressions and statements (operators and precedence). py4e.com
- Key terms
- Floor division (//)
- Division rounded toward negative infinity. Integer operands give an int; a float operand can give a float.
- Remainder (%)
- The remainder after floor division. For integer operands with a positive divisor d, the result lies from 0 through d - 1.
- Operator precedence
- The order operators are applied in. Powers first, then multiply and divide, then add and subtract; brackets override.
- Float
- A number stored as a binary fraction with fixed room, so most decimals are held as the nearest approximation.
- IEEE 754
- A standard for floating-point arithmetic, including the binary formats used by typical Python floats.
- Format specifier
- The part after a colon inside an f-string, such as .2f, which controls how a value is printed without changing it.
- round
- A built-in function returning a new rounded value, unlike a format specifier, which only changes the printing.
The Answer Was 85.5 and Nobody Could Tell It Was Wrong
- Break a programming task into named steps before writing any code.
- Write pseudocode and a test case whose answer you already know.
- Build a program in slices, running each slice before adding the next.
Marcus wants to know what he needs in the physics exam. Coursework is worth 40 percent of the final mark, the exam the other 60. He has 68 for coursework and he wants 75 overall. He knows enough Python now to do this, so he opens a file and types the answer straight out:
coursework = input("Coursework mark? ")
target = input("Target overall? ")
print("You need", (target - coursework * 0.6) / 0.4, "in the exam")
print("You need", (target - coursework * 0.6) / 0.4, "in the exam")
~~~~~~~~~~~^~~~~
TypeError: can't multiply sequence by non-int of type 'float'
He recognises that one from Lesson 2. Input hands back a string, and a string times 0.6 is meaningless. He wraps both answers in int and runs it again:
coursework = int(input("Coursework mark? "))
target = int(input("Target overall? "))
print("You need", (target - coursework * 0.6) / 0.4, "in the exam")
Coursework mark? Target overall? You need 85.5 in the exam
No error. A number. It even looks plausible: high, but not absurd, exactly the sort of mark you would expect to need. Marcus writes 85.5 in his planner and goes to make toast.
It is wrong. The right answer is 79.7. He has just talked himself into six extra marks of panic, and nothing in that session was ever going to tell him.
Why 85.5 could not be caught
The bug is that the weights are the wrong way round. Coursework is worth 40, not 60, and the exam 60, not 40. Swap them in the formula and you get a number that is too high, but only by six. Not ten times too big. Not negative. Not zero. Just quietly, believably wrong.
Every skill you have gathered so far is useless against this. The program ran. The types were right. The arithmetic was performed correctly on the numbers it was given. There is no message to read, because nothing went wrong as far as Python is concerned. The only defence is to have decided, before writing the formula, what the right answer would look like.
That is what this lesson is: not a new piece of Python, but the twenty minutes before you type. Jeannette Wing, in a 2006 article that named the idea for a lot of people, described computational thinking as reformulating a hard problem into one you already know how to solve, using abstraction and decomposition when a task is too big to hold in your head. Marcus tried to hold the whole thing in his head. It was only three lines, and it was still too big.
Step one: write down what you have and what you want
In English, on paper or in a comment, before anything else. Two lists.
What I have: a coursework mark out of 100; a target overall mark out of 100; the coursework weight, 40 percent; the exam weight, 60 percent.
What I want: the exam mark, out of 100, that would make the weighted total equal the target.
Notice what writing this out already did. The words "coursework weight, 40" and "exam weight, 60" are now sitting next to each other where you can see them, which is the exact place Marcus went wrong. Half the bugs you will ever write are prevented by putting the quantities in a list and looking at it.
Step two: pseudocode, which is not a language
Pseudocode is the steps of a program written in ordinary language, structured like code but not obeying any syntax rules. There is no correct pseudocode and no interpreter to please. It exists so you can get the order of operations right while mistakes are still free to fix.
ask for the coursework mark
ask for the target overall mark
coursework weight is 40, exam weight is 60
the coursework contributes coursework * 40 points out of 10000
the target needs target * 100 points out of 10000
the exam must supply the difference
the exam mark is that difference divided by the exam weight
print it
check it by working the sum forwards again
Nine lines, thirty seconds, and the mistake is now almost impossible: the coursework is multiplied by its own weight on the line that mentions coursework. Write your pseudocode as comments in the real file and fill the code in underneath them. The comments that survive are usually the good ones.
Step three: find a case whose answer you already know
Before running anything, decide what a right answer looks like. You want cases you can do in your head:
- Coursework 100, target 100. The exam must be 100. Anything else means the program is broken.
- Coursework 50, target 50. The exam must be 50.
- Coursework 68, target 75. The exam must be somewhere above 75, because 68 is dragging the average down, and not enormously above, because coursework is the smaller weight. So: high seventies or low eighties.
That third case alone would not have caught 85.5. The first two would have caught it instantly. Easy cases are better tests than realistic ones, precisely because you know the answer without trusting the program.
Slice one: read the numbers and say them back
Now build, and run after every slice. The first slice does no arithmetic at all. It exists to prove that the inputs arrive as the right types with the right values.
# Slice 1: read the three numbers and say them back. No maths yet.
coursework = int(input("Coursework mark out of 100? "))
target = int(input("Target overall mark out of 100? "))
coursework_weight = 40
print(f"Coursework: {coursework}, worth {coursework_weight} percent")
print(f"Exam: unknown, worth {100 - coursework_weight} percent")
print(f"Target overall: {target}")
Coursework mark out of 100? Target overall mark out of 100? Coursework: 68, worth 40 percent
Exam: unknown, worth 60 percent
Target overall: 75
Dull, and worth it. The weights are printed where you can read them, and the exam weight is worked out as 100 - coursework_weight rather than typed as a second number, so the two can never disagree with each other. That is a small design decision made in slice one that removes a whole family of future bugs.
Slice two: work the sum forwards, where you can check it
Here is the trick that makes this problem easy. Working out the exam mark you need is the hard direction. Working out the overall mark from an exam mark is the easy direction, and you can check it by hand. So build the easy direction first:
# Slice 2: work an overall mark FORWARDS from an exam mark I supply.
# If this is wrong, nothing built on top of it can be right.
coursework = 68
exam = 80
coursework_weight = 40
exam_weight = 60
overall = (coursework * coursework_weight + exam * exam_weight) / 100
print(f"Coursework {coursework}, exam {exam}, overall {overall}")
# Two cases I can check in my head:
print("Both 100 should give 100:", (100 * 40 + 100 * 60) / 100)
print("Both 50 should give 50: ", (50 * 40 + 50 * 60) / 100)
Coursework 68, exam 80, overall 75.2
Both 100 should give 100: 100.0
Both 50 should give 50: 50.0
Both known cases came out right, so the weighting is correct. And 68 with an exam of 80 gives 75.2, which is just over the target of 75. So the answer Marcus wants is just under 80. That single line has already ruled out 85.5, before the hard part was written.
The input lines were temporarily replaced by fixed values here. That is deliberate. While you are testing a calculation, typing answers at a prompt forty times is a waste of your evening; hard-code the values, get the sum right, then put the inputs back.
Slice three: turn it round
# Slice 3: turn the forward sum round, and check it against slice 2.
coursework = 68
target = 75
coursework_weight = 40
exam_weight = 60
points_from_coursework = coursework * coursework_weight
points_needed = target * 100
points_from_exam = points_needed - points_from_coursework
exam_needed = points_from_exam / exam_weight
print("Exam mark needed:", exam_needed)
# The check: feed that exam mark back through the forward sum from slice 2.
overall = (coursework * coursework_weight + exam_needed * exam_weight) / 100
print("Feeding it back gives overall:", overall)
Exam mark needed: 79.66666666666667
Feeding it back gives overall: 75.0
79.67, just under 80, exactly as slice two predicted. And the check line is the good bit: it takes the answer, runs it back through the calculation that was already verified, and gets 75.0, which is the target. Two independent routes agreeing is much stronger evidence than one route looking reasonable.
Notice the four separate named steps. points_from_coursework, points_needed, points_from_exam, exam_needed. Marcus had all four of those crammed into one expression, which is why he could not see the weights were swapped. Splitting an expression into named steps costs three lines and buys you the ability to print any one of them.
The finished program, with the check left in
# grade_target.py
# What do I need in the exam to reach a target overall mark?
# Coursework and exam are weighted, and the weights must add to 100.
coursework = int(input("Coursework mark out of 100? "))
target = int(input("Target overall mark out of 100? "))
coursework_weight = 40
exam_weight = 100 - coursework_weight
# Work in weighted points out of 10000 so the numbers stay whole until the end.
points_from_coursework = coursework * coursework_weight
points_needed = target * 100
exam_needed = (points_needed - points_from_coursework) / exam_weight
exam_rounded = round(exam_needed, 1)
print(f"Coursework {coursework} at {coursework_weight} percent gives you {points_from_coursework / 100} of the final mark.")
print(f"To finish on {target} you need {exam_needed:.1f} in the exam.")
# Check: put the mark you would actually aim for back through the forward sum.
overall = (coursework * coursework_weight + exam_rounded * exam_weight) / 100
print(f"Check: {coursework} and {exam_rounded} really give an overall of {overall}")
Coursework mark out of 100? Target overall mark out of 100? Coursework 68 at 40 percent gives you 27.2 of the final mark.
To finish on 75 you need 79.7 in the exam.
Check: 68 and 79.7 really give an overall of 75.02
The check says 75.02, not 75.00, and that is correct rather than a bug: aiming at the rounded 79.7 rather than the exact 79.6667 overshoots the target by two hundredths of a mark. A check that reported a suspiciously perfect 75.0 would have been hiding the rounding. This is the difference between a check that tests something and a check that agrees with you.
What the slices actually bought
Four things, and they are worth naming because you will be tempted to skip all of them.
A slice that runs is a place to stand. When slice three went wrong, the problem had to be in slice three, because slices one and two had already printed correct output. Marcus's version had no such place: a wrong answer could have come from anywhere in three lines.
Building the easy direction first gave a way to check the hard one. Almost every calculation has an easy direction. Find it.
Naming the intermediate steps made them printable. You cannot print the middle of an expression, only the whole thing.
And writing what you had before writing how to combine it put the two weights side by side, which is where the bug was and where it would have been seen.
Common misconceptions
- "Planning is for long programs. Mine is three lines." Marcus's was three lines. Length is not the measure; the measure is whether you can tell a right answer from a wrong one by looking at it.
- "If it runs without an error, it works." Running is the low bar. A program that runs and produces a wrong number is more dangerous than one that crashes, because a crash tells you.
- "Pseudocode has to be written properly or it is not worth doing." There are no rules. It can be bullet points on the back of an envelope. Its entire value is forcing the order of steps into words before syntax gets a vote.
- "Test with realistic numbers." Test with numbers whose answer you already know, which usually means silly ones: all zeros, all hundreds, one of everything. A realistic case tells you nothing if you cannot independently say what it should produce.
The short version
Why this matters: Decide what a right answer looks like before you write the code that produces one, and build in pieces small enough that a wrong answer has nowhere to hide.
The method is four steps and it does not change: list what you have and what you want; write the order of operations as pseudocode in ordinary words; pick test cases whose answers you already know; then build in slices, running each before adding the next. Along the way, name the intermediate values instead of nesting everything into one expression, derive numbers from each other rather than typing them twice, and leave the check in the finished program.
That closes the first module. You can get input, store it, calculate with it and print it, and you have a method for deciding what to write. Every program so far has run in a straight line from top to bottom. Module 2 breaks the straight line: first by letting the program choose between two paths, then by letting it go round the same path many times.
Sources
- Wing, J. M. (2006). Computational thinking. Communications of the ACM, 49(3), 33-35. (On abstraction and decomposition when attacking a large task.) cs.cmu.edu
- Wikipedia contributors. (2026). Pseudocode. Wikipedia. en.wikipedia.org
- Wikipedia contributors. (2026). Decomposition (computer science). Wikipedia. en.wikipedia.org
- Polya, G. (1945). How to Solve It: A New Aspect of Mathematical Method. Princeton University Press. (Understand the problem, devise a plan, carry it out, look back.)
- Severance, C. (2016). Python for Everybody: Exploring data in Python 3, Chapter 1: Why should you learn to write programs? py4e.com
- Key terms
- Decomposition
- Breaking a task into smaller named parts, each simple enough to write and check on its own.
- Pseudocode
- The steps of a program written in ordinary language, structured like code but obeying no syntax rules.
- Incremental building
- Writing a program in slices and running each slice before adding the next, so a new bug has only one place to be.
- Test case
- An input whose correct output you already know, used to judge whether the program is right rather than merely running.
- Intermediate value
- A named step in the middle of a calculation, which can be printed and checked; an expression cannot.
- Hard-coding
- Writing a fixed value into the program instead of reading it, useful while testing a calculation and usually removed afterwards.
Module 2: Choosing a Path and Going Round Again
Let a program pick between branches, repeat work with while and for, count correctly, and package a piece of work as a function you can name and reuse.
Four Separate Ifs Gave a Mark of 91 the Grade Six
- Write an if statement with the correct colon and indentation, and say which lines belong to it.
- Compare values with the six comparison operators and describe the True or False they hand back.
- Choose between a chain of elif and a sequence of separate ifs, and explain what goes wrong with the wrong one.
A cinema charges 4.50 for under sixteens and 9.00 for everybody else. Twelve lines of Python can run the till:
age = int(input("How old are you? "))
if age < 16:
print("Child ticket: 4.50")
else:
print("Adult ticket: 9.00")
print("Enjoy the film.")
Run with 15, then with 16:
How old are you? Child ticket: 4.50
Enjoy the film.
How old are you? Adult ticket: 9.00
Enjoy the film.
The same file, run twice, did two different things. Up to now every program you have written did the same thing every time it ran. This one looks at what it was given and picks a path. The last line printed both times, which is the other half of the idea: some lines belong to the choice, and some lines happen regardless. Which lines are which is decided entirely by how far they are indented, and Python is unforgiving about it.
The shape of an if: condition, colon, indented block
An if statement has three parts. The word if, then a condition, then a colon. Underneath, indented by four spaces, come the lines that run only when the condition is true. That indented group is called a block.
The colon is not decoration; leaving it off is a syntax error. Four spaces is what PEP 8 asks for, and any consistent amount works, but mixing tabs and spaces in one file causes trouble that is genuinely hard to see, since both look like blank space on the screen. Set your editor to insert spaces when you press Tab and forget about it.
What happens if you forget to indent at all:
mark = 91
if mark >= 90:
print("top grade")
print("top grade")
^^^^^
IndentationError: expected an indented block after 'if' statement on line 2
The message is exact: an if statement promised a block and no block arrived. In most languages indentation is a courtesy to readers. In Python it is the structure itself, which means your code cannot be laid out in a way that disagrees with what it does.
The six comparisons, and what they hand back
A condition is just an expression whose value is True or False. Here is every comparison operator, run against a single value so you can see them together:
age = 15
print("age == 15 :", age == 15)
print("age != 15 :", age != 15)
print("age < 16 :", age < 16)
print("age <= 15 :", age <= 15)
print("age > 16 :", age > 16)
print("age >= 15 :", age >= 15)
print("type of the answer:", type(age < 16))
print('"cat" == "Cat" :', "cat" == "Cat")
print('"cat" < "dog" :', "cat" < "dog")
print("15 == '15' :", 15 == "15")
age == 15 : True
age != 15 : False
age < 16 : True
age <= 15 : True
age > 16 : False
age >= 15 : True
type of the answer: <class 'bool'>
"cat" == "Cat" : False
"cat" < "dog" : True
15 == '15' : False
| Operator | Asks | With age = 15 |
|---|---|---|
| == | Are these the same? | True |
| != | Are these different? | False |
| < | Is the left smaller? | True, against 16 |
| <= | Smaller or the same? | True, against 15 |
| > | Is the left bigger? | False, against 16 |
| >= | Bigger or the same? | True, against 15 |
Three results in that output are worth pausing on. The type of a comparison is bool, which means a comparison is not a special thing that only lives inside if statements; it is an ordinary value you can print, store in a variable, or pass around. Comparing "cat" with "Cat" gives False, because capital letters are different characters. Comparing "cat" with "dog" gives True, because strings compare alphabetically, which is how you sort names. And 15 == "15" is False, because the number fifteen and the two characters 1 and 5 are different things, which is the Lesson 2 rule about input arriving as text, now with teeth: forget the int and your comparison silently fails instead of raising an error.
One equals sign is an order, two is a question
This is the single most common beginner mistake in any language with both operators:
mark = 91
if mark = 90:
print("exactly ninety")
if mark = 90:
^^^^^^^^^
SyntaxError: invalid syntax. Maybe you meant '==' or ':=' instead of '='?
One equals sign assigns: put the value on the right into the name on the left. Two equals signs compare: are these the same? Python's message here is unusually helpful, but read it carefully, because it offers two suggestions and only one of them is what you meant. A useful trick when reading code aloud: say "becomes" for = and "is equal to" for ==. If the sentence does not make sense, the operator is wrong.
Four separate ifs graded a 91 as a six
Here is a program that looks completely reasonable and is badly wrong:
mark = 91
if mark >= 90:
grade = "9"
if mark >= 80:
grade = "8"
if mark >= 70:
grade = "7"
if mark >= 60:
grade = "6"
else:
grade = "U"
print("Mark", mark, "gets grade", grade)
Mark 91 gets grade 6
Ninety-one, and it printed a six. Trace it and you can see exactly why, because nothing mysterious happened. A separate if is a separate question, and Python asks all four of them.
- Is 91 at least 90? Yes. grade becomes "9".
- Is 91 at least 80? Yes, still. grade becomes "8", overwriting the 9.
- Is 91 at least 70? Yes. grade becomes "7".
- Is 91 at least 60? Yes. grade becomes "6".
The last test to succeed wins, and for a high mark every test succeeds, so the lowest band always wins. Note also that the else belongs only to the fourth if, not to all of them, which is another thing four separate statements cannot express.
elif: one chain, one winner
What was needed is a single decision with several outcomes. That is elif, a contraction of "else if":
mark = int(input("Mark out of 100? "))
if mark >= 90:
grade = "9"
elif mark >= 80:
grade = "8"
elif mark >= 70:
grade = "7"
elif mark >= 60:
grade = "6"
else:
grade = "U"
print(f"Mark {mark} gets grade {grade}")
Mark out of 100? Mark 91 gets grade 9
Mark out of 100? Mark 59 gets grade U
An if/elif/else chain is one question with several answers. Python tests the conditions from the top and stops at the first one that is true, skipping everything below it. At most one block runs. If none of them is true the else runs, and if there is no else, nothing runs at all and the program carries on.
Because the chain stops at the first match, order matters enormously. Put mark >= 60 at the top and every mark above 60 gets a six, which is the same bug wearing a different hat. The rule for a chain of thresholds is to start at the most demanding and work down, and each later condition can then quietly assume the earlier ones failed. That is why elif mark >= 80 does not need to also say "and less than 90": if it were 90 or more, the chain would already have stopped.
and, or, not, and a comparison that chains
Conditions combine. Python uses ordinary words rather than symbols for this, which makes them pleasant to read aloud:
age = 15
has_permission = True
is_weekend = False
print("age >= 13 and age <= 17 :", age >= 13 and age <= 17)
print("13 <= age <= 17 :", 13 <= age <= 17)
print("is_weekend or has_permission :", is_weekend or has_permission)
print("not is_weekend :", not is_weekend)
print("age >= 18 and has_permission :", age >= 18 and has_permission)
print("age >= 18 or has_permission :", age >= 18 or has_permission)
if 13 <= age <= 17 and has_permission:
print("Allowed in, with a guardian's note.")
else:
print("Not allowed in today.")
age >= 13 and age <= 17 : True
13 <= age <= 17 : True
is_weekend or has_permission : True
not is_weekend : True
age >= 18 and has_permission : False
age >= 18 or has_permission : True
Allowed in, with a guardian's note.
| Operator | True when | Common mistake |
|---|---|---|
| and | Both sides are true | Using it where you meant or, which then never fires |
| or | At least one side is true | Reading it as "one or the other but not both" |
| not | The thing after it is false | Stacking two nots and losing track of which way round it is |
The second line of that output is a Python convenience you will not find in most languages: 13 <= age <= 17 means exactly what a mathematician would expect, and is a tidier way of writing the line above it. Use it for ranges.
One trap with or. If you want to know whether a letter is a vowel, if letter == "a" or "e" or "i" does not work, and worse, it does not fail. The or joins three separate values, and a non-empty string counts as true, so the whole condition is always true. Write each comparison out in full: if letter == "a" or letter == "e" or letter == "i". Lesson 9 gives you a much neater way with the word in.
Common misconceptions
- "= and == are basically the same and Python works out which I meant." One assigns and one compares. Python refuses to guess, which is why
if mark = 90is a syntax error rather than a silent disaster. - "Several ifs in a row are the same as an if/elif chain." They are different programs. Separate ifs all get asked, so the last true one overwrites the rest. A chain stops at the first true one. This is the bug that graded a 91 as a six.
- "The else after a chain of ifs covers all of them." An else belongs to the single if or elif directly above it. Four ifs and an else means three uncovered ifs and one pair.
- "Indentation is just tidiness." In Python it is the grammar. The indented lines are the block, and moving a line left or right changes which decision it belongs to, with no warning that anything has changed.
- "if letter == 'a' or 'e' checks two letters." It checks one comparison and one string, and a non-empty string is treated as true, so the condition is always true. Repeat the variable in every comparison.
Pulling it together
Remember: An if/elif/else chain asks one question and runs at most one block; separate if statements ask separate questions and all of them get asked.
You can now write a program that behaves differently depending on what it is given. You know the six comparison operators and that each produces a bool, that strings compare alphabetically and case-sensitively, and that a number never equals the string of its digits. You know that one equals sign is an order and two make a question. You can combine conditions with and, or and not, chain a range comparison, and you have seen the vowel trap that makes an always-true condition out of a sensible-looking line. Above all you have seen what happens when a chain is written as separate questions, and why the order of a chain of thresholds has to run from the most demanding downwards.
The cinema till still serves one customer per run. The next lesson keeps the program going: asking again, and again, until something tells it to stop.
Sources
- Python Software Foundation. (2026). More control flow tools: if statements. Python 3.14 tutorial. docs.python.org
- Python Software Foundation. (2026). Built-in types: Comparisons and boolean operations. docs.python.org
- Severance, C. (2016). Python for Everybody: Exploring data in Python 3, Chapter 3: Conditional execution. py4e.com
- Sweigart, A. (2019). Automate the Boring Stuff with Python (2nd ed.), Chapter 2: Flow control. No Starch Press. automatetheboringstuff.com
- van Rossum, G., Warsaw, B., and Coghlan, N. (2001, revised 2023). PEP 8: Style guide for Python code (indentation: four spaces per level). peps.python.org
- Key terms
- Condition
- An expression that works out to True or False, used to decide whether a block runs.
- Block
- The group of lines indented under an if, elif or else, which run together or not at all.
- elif
- A continuation of an if, tested only when every condition above it was false. At most one branch of a chain runs.
- Comparison operator
- One of ==, !=, <, <=, > and >=, each of which produces a bool.
- and
- True only when both sides are true.
- or
- True when at least one side is true, including when both are.
- not
- Flips a truth value: not True is False.
- IndentationError
- The error raised when a block is missing or its lines do not line up, since indentation is Python's structure.
The Loop That Printed Nine Numbers When You Asked for Ten
- Write a while loop with all three of its moving parts and trace how many times it runs.
- Diagnose an infinite loop and an off-by-one error from their output.
- Use an accumulator and a while True loop with break to keep asking until an answer is valid.
Here is a countdown, and it works:
seconds = 5
while seconds > 0:
print(seconds)
seconds = seconds - 1
print("Go.")
5
4
3
2
1
Go.
Now delete the fourth line, the one that subtracts, and run it again. Here are the first six lines of what came out, and there is no seventh line because there is no end:
5
5
5
5
5
5
It carried on printing fives until it was stopped by hand, with Ctrl and C held down together. Nothing crashed. Nothing was wrong with the syntax. The program did exactly what it said, forever, because the condition seconds > 0 was true at the start and nothing in the loop ever made it false. That is the central fact about loops: a loop only ends because something inside it changes the thing the condition is looking at.
The three parts, and what happens if one is missing
Every while loop has three moving parts, and they are not all in the same place, which is why they are easy to lose.
- Something set up before the loop.
seconds = 5. Miss this and you get a NameError on the first test. - A condition after the word while.
seconds > 0. It is checked before every pass, including the first, so a loop whose condition starts false runs zero times. - Something inside the block that changes the condition's answer.
seconds = seconds - 1. Miss this and the loop is infinite.
The word while is honest: while this stays true, keep doing the block. Python checks the condition, runs the whole block, comes back and checks again, and only leaves when a check comes out False. It does not notice a change partway through the block; it finishes the block first.
If you run a program that will not stop, hold Ctrl and press C in the terminal. The program is interrupted and you get your prompt back. Everybody writes an infinite loop, usually more than once a month, and it is not a disaster. Just do not close the whole terminal window in a panic; you will do it while your file is unsaved.
Nine numbers when you asked for ten
Here is a bug you will write, probably tomorrow. Three loops, each printing a run of numbers, differing by a single character:
print("Version A, while n < 10:")
n = 1
while n < 10:
print(n, end=" ")
n = n + 1
print("Version B, while n <= 10:")
n = 1
while n <= 10:
print(n, end=" ")
n = n + 1
print("Version C, starting at 0 with n < 10:")
n = 0
while n < 10:
print(n, end=" ")
n = n + 1
Version A, while n < 10:
1 2 3 4 5 6 7 8 9
printed 9 numbers
Version B, while n <= 10:
1 2 3 4 5 6 7 8 9 10
printed 10 numbers
Version C, starting at 0 with n < 10:
0 1 2 3 4 5 6 7 8 9
printed 10 numbers
Version A asked for the numbers up to ten and produced nine of them. This is the off-by-one error, and it is so common it has a name and its own jokes. The loop stops the moment the condition fails, so n < 10 means the pass where n is 10 never happens.
The trap is that the difference between A and B is invisible when you read the code and obvious when you count the output. Which is the lesson: count the output. Not the line of code. Take any loop you have just written, give it a tiny input, run it, and count what comes out against what you expected. It takes ten seconds and it catches the error every time, which is more than can be said for staring at the condition.
Versions B and C both produce ten numbers, from different starting points. That is the underlying arithmetic: a loop from a start value up to but not including an end value runs end - start times. From 0 up to 10 is ten passes. From 1 up to 10 is nine. Write that down somewhere.
The end=" " in those print calls is a keyword argument: it tells print what to put after the text instead of moving to a new line. Here it is a space, which is why the numbers came out in a row. A bare print() with nothing in it then produces the line break.
Why programmers count from zero
Version C looks peculiar to anyone who counts things normally, and it is the version you will meet most often in real code. Starting at 0 and running while the counter is less than the total gives you the right number of passes without any arithmetic in your head: n < 10 runs ten times, full stop.
There is a second reason, which arrives properly in Lesson 9: the positions in a Python list are numbered from 0, so a list of ten things has positions 0 to 9. A loop counting 0, 1, 2 up to 9 lines up with those positions exactly, and a loop counting 1 to 10 does not. Getting used to the zero-based habit now saves a whole category of confusion later.
An accumulator: adding up whatever somebody types
A loop becomes useful when it builds something. The pattern is called an accumulator: a variable set up before the loop and added to inside it.
total = 0
count = 0
entry = input("Mark (or done): ")
while entry != "done":
total = total + int(entry)
count = count + 1
entry = input("Mark (or done): ")
print(f"You entered {count} marks adding to {total}.")
print(f"Mean: {total / count:.1f}")
Fed 58, 71, 67 and then done:
Mark (or done): Mark (or done): Mark (or done): Mark (or done): You entered 3 marks adding to 196.
Mean: 65.3
Look at where the input lines are. There are two of them, one before the loop and one as the last line inside it, and that duplication bothers everybody the first time. It is necessary: the condition has to have something to test before the first pass, and it has to have something fresh to test before every pass after that. The value that ends the loop, here the word done, is called a sentinel, and notice it is never added to the total, because as soon as it arrives the condition fails and the block is skipped.
total = 0 and count = 0 are the other half of the pattern. Starting an accumulator at zero is obvious; forgetting to start it at all, so the first total = total + ... raises a NameError, is the usual way to get this wrong.
while True and break: asking until the answer is sensible
Sometimes there is nothing sensible to test before the first pass, because you have not asked the question yet. The idiom for that is a loop with no exit condition and a deliberate escape:
while True:
answer = input("Pick a number from 1 to 5: ")
number = int(answer)
if 1 <= number <= 5:
break
print(f"{number} is not between 1 and 5. Try again.")
print(f"Thank you. You picked {number}.")
Fed 9, then 0, then 3:
Pick a number from 1 to 5: 9 is not between 1 and 5. Try again.
Pick a number from 1 to 5: 0 is not between 1 and 5. Try again.
Pick a number from 1 to 5: Thank you. You picked 3.
while True is a condition that is always true, so the loop would run forever. break leaves the loop immediately, skipping the rest of the block and everything the loop would have done afterwards, and carries on at the first line after it. Here the break is inside an if, so the escape happens only when the answer is in range. Note the rejection message is never printed on the successful pass, because break jumped over it.
This pattern, ask, check, break if good, complain and go round again, is how nearly all input validation is written. It will run the game loop in Lesson 20. Two cautions: a while True with no break in it is an infinite loop with extra steps, so write the break first; and break only leaves the loop it is directly inside, not two loops at once.
Common misconceptions
- "The loop checks the condition all the way through the block." It checks once, before each pass. If the condition becomes false halfway down the block, the rest of that block still runs, and only then does the loop end.
- "An infinite loop is a crash." Nothing has crashed. The program is obediently doing what you wrote. Ctrl and C stops it, and the fix is almost always a missing update inside the block.
- "while n < 10 counts to ten." It counts to nine. A loop stops the moment the condition fails, so the pass at the boundary never happens. Count the output, not the condition.
- "An accumulator does not need setting up."
total = total + xreads total before it writes it, so total must already exist. Set it to 0 before the loop. - "break ends the program." It ends the innermost loop only. The lines after the loop still run, which is exactly what makes the validation pattern work.
What to carry forward
The upshot: A loop needs a setup before it, a condition on it, and a change inside it. Lose the change and it never ends; misjudge the condition by one and it runs one pass too few.
You can now repeat work. You can count down, count up, and count from zero, and you know that a loop from a start to an end runs end minus start times when the condition uses a plain less-than. You can build a running total with an accumulator, end a loop on a sentinel value typed by the user, and use while True with break when there is nothing to test until after you have asked. You have seen an infinite loop and stopped it with Ctrl and C without panicking, which is a genuine skill.
Every loop here counted by hand: set a variable, compare it, add one to it, three lines and three chances to get it wrong. The next lesson replaces all three with one line, and it turns out to be the loop you will use most.
Sources
- Python Software Foundation. (2026). More control flow tools: break and continue, and the else clause on loops. Python 3.14 tutorial. docs.python.org
- Python Software Foundation. (2026). The Python tutorial: First steps towards programming (the while statement and its condition test). docs.python.org
- Severance, C. (2016). Python for Everybody: Exploring data in Python 3, Chapter 5: Iteration (infinite loops, break, and the accumulator pattern). py4e.com
- Sweigart, A. (2019). Automate the Boring Stuff with Python (2nd ed.), Chapter 2: Flow control (while loops and break). No Starch Press. automatetheboringstuff.com
- Wikipedia contributors. (2026). Off-by-one error. Wikipedia. en.wikipedia.org
- Key terms
- while loop
- A block that repeats as long as its condition is true, checked once before each pass.
- Infinite loop
- A loop whose condition never becomes false, usually because nothing inside it changes the value being tested.
- Off-by-one error
- A loop that runs one pass too many or too few, typically from choosing < where <= was meant.
- Accumulator
- A variable created before a loop and added to inside it, used to build a running total or count.
- Sentinel
- A special value that signals the end of input, such as the word done, and which is itself never processed.
- break
- A statement that leaves the innermost loop immediately and continues at the first line after it.
- Keyword argument
- An argument given by name, such as end=' ' in a print call, which changes what the function does rather than what it acts on.
One Line Instead of Three: for and range
- Rewrite a counted while loop as a for loop over range and say why it cannot loop forever.
- Predict what range produces with one, two or three arguments, including a negative step.
- Build a table with a nested loop and iterate directly over the characters of a string.
Here is the same job done twice. Print the squares of 1 to 5:
print("The while version:")
n = 1
while n <= 5:
print(n, "squared is", n * n)
n = n + 1
print("The for version:")
for n in range(1, 6):
print(n, "squared is", n * n)
The while version:
1 squared is 1
2 squared is 4
3 squared is 9
4 squared is 16
5 squared is 25
The for version:
1 squared is 1
2 squared is 4
3 squared is 9
4 squared is 16
5 squared is 25
Identical output, and the second version has thrown away two of the three lines that can go wrong. There is no setup line to forget, no n = n + 1 to leave out, and therefore no way to write an infinite loop by accident. Everything about how many times the loop runs is now in one place, on the for line, where you can read it.
That is the trade. A for loop is for when you know in advance what you are going through: five numbers, the letters of a word, the rows of a table. A while loop is for when you do not: keep asking until they type done, keep dealing cards until somebody wins. Reach for for first, and fall back to while when the ending genuinely depends on something that happens inside the loop.
What range actually gives you
The best way to see range is to make it show its working. Wrapping it in list displays every value it will hand out:
print("range(5) :", list(range(5)))
print("range(1, 6) :", list(range(1, 6)))
print("range(0, 20, 5) :", list(range(0, 20, 5)))
print("range(10, 0, -1):", list(range(10, 0, -1)))
print("range(3, 3) :", list(range(3, 3)))
print("len(range(5)) :", len(range(5)))
for i in range(3):
print("pass number", i)
range(5) : [0, 1, 2, 3, 4]
range(1, 6) : [1, 2, 3, 4, 5]
range(0, 20, 5) : [0, 5, 10, 15]
range(10, 0, -1): [10, 9, 8, 7, 6, 5, 4, 3, 2, 1]
range(3, 3) : []
len(range(5)) : 5
pass number 0
pass number 1
pass number 2
| Written as | Means | Produces |
|---|---|---|
| range(5) | Five values from 0 | 0 1 2 3 4 |
| range(1, 6) | From 1, stopping before 6 | 1 2 3 4 5 |
| range(0, 20, 5) | From 0, stepping by 5, stopping before 20 | 0 5 10 15 |
| range(10, 0, -1) | From 10, stepping down, stopping before 0 | 10 9 8 ... 1 |
| range(3, 3) | Start equals stop | nothing at all |
One rule runs through all five rows: range stops before the number you give it, never on it. This is the off-by-one error from Lesson 6, moved somewhere you can see it. It is also why range(5) gives five values and why len(range(5)) is 5: the count is exactly the stop value when you start from zero.
So range(1, 6) for the numbers 1 to 5, and if you want 1 to 10, write range(1, 11). The plus one in the stop value will feel wrong for about a week and then stop being visible. range(0, 20, 5) never produces 20 for the same reason, which is genuinely useful for a timetable that runs to but not through a boundary.
Notice too that range(3, 3) is empty and the loop simply does not run. A for loop over an empty range is not an error and prints nothing, which is worth remembering when a loop mysteriously produces no output at all.
Looping over the thing itself, not over its positions
range is not the only thing a for loop can walk through. It can walk through a string, one character at a time, and later through lists and dictionaries. That is why it is written for x in something: the word in is doing real work.
word = "physics"
for letter in word:
print(letter, end=".")
print()
vowels = 0
for letter in word:
if letter == "a" or letter == "e" or letter == "i" or letter == "o" or letter == "u":
vowels = vowels + 1
print(f"letters in {word}: {len(word)}, vowels: {vowels}")
total = 0
for n in range(1, 101):
total = total + n
print("1 + 2 + ... + 100 =", total)
p.h.y.s.i.c.s.
letters in physics: 7, vowels: 1
1 + 2 + ... + 100 = 5050
The loop variable takes each character in turn, and no counting was involved anywhere. The accumulator pattern from Lesson 6 works unchanged: set total to 0 before the loop, add inside it. The last three lines add the numbers 1 to 100 and get 5050, which is the answer Carl Friedrich Gauss is supposed to have produced in his head as a schoolboy, and which your computer produced by doing it the slow way a hundred times in well under a millisecond.
The vowel counting has that ugly chain of four ors, exactly as warned about in Lesson 5. Lesson 9 replaces the whole condition with if letter in "aeiou". It is worth writing it the long way once so that the short way feels like a relief rather than magic.
A loop inside a loop
Put one for loop in another and the inner one runs completely for every single pass of the outer one. That is how you make anything with rows and columns:
# A grid: the outer loop makes rows, the inner loop makes what is in a row.
for row in range(1, 6):
for column in range(1, 6):
print(f"{row * column:4}", end="")
print()
print()
# A triangle, where the inner loop's length depends on the outer loop.
for row in range(1, 6):
for star in range(row):
print("*", end="")
print()
1 2 3 4 5
2 4 6 8 10
3 6 9 12 15
4 8 12 16 20
5 10 15 20 25
*
**
***
****
*****
Read the first loop from the inside out. The inner loop prints five numbers with end="", so they all land on one line. When it finishes, the bare print() ends that line. Then the outer loop starts its next pass and the inner loop runs all over again. Five outer passes times five inner passes is twenty-five numbers, and {row * column:4} pads each one into four characters, which is what makes the columns line up.
The second example is the interesting one. Its inner loop is range(row), so the number of stars depends on which row the outer loop is on: one, then two, then three. When a nested loop confuses you, and it will, the question to ask is always the same. What does the inner loop do for one single value of the outer variable? Answer that, then multiply.
The bare print() that ends each row is the line people forget. Leave it out and all twenty-five numbers arrive on one enormous line, which is a very recognisable symptom.
Common misconceptions
- "range(1, 10) gives the numbers 1 to 10." It gives 1 to 9. range always stops before the second number. For 1 to 10, write range(1, 11).
- "range(5) starts at 1." It starts at 0 and gives 0, 1, 2, 3, 4. That is five values, which is the point of writing it that way.
- "Changing the loop variable inside the loop changes how many passes there are." It does not. A for loop takes the next value from range regardless of what you did to the variable, so assigning to it only affects the current pass.
- "A for loop can get stuck forever." A for loop over a range always ends, because the number of passes is fixed before it starts. Infinite loops are a while problem.
- "A nested loop runs the inner loop once per program." The inner loop runs from the start, every time the outer loop goes round. Five outer passes with five inner passes is twenty-five inner passes altogether.
Looking back
In short: Use for when you already know what you are stepping through, and remember that range stops before the number you give it.
You can now write a counted loop in one line instead of three, which removes the setup you might forget and the increment you might leave out. You know what range does with one, two and three arguments, that the third can be negative to count down, and that a range whose start equals its stop is empty and runs no passes at all. You can loop over the characters of a string without counting, keep an accumulator, and nest one loop inside another to build a grid or a shape, ending each row with a bare print.
The programs are getting longer, and the same few lines keep reappearing: ask for a number and check it is in range, add up a list of things, print a row. The next lesson gives those chunks names, so that each one is written once and used by name from then on.
Sources
- Python Software Foundation. (2026). More control flow tools: for statements and the range function. Python 3.14 tutorial. docs.python.org
- Python Software Foundation. (2026). Built-in functions: range and len. Python 3.14 documentation. docs.python.org
- Severance, C. (2016). Python for Everybody: Exploring data in Python 3, Chapter 5: Iteration (definite loops with for). py4e.com
- Sweigart, A. (2019). Automate the Boring Stuff with Python (2nd ed.), Chapter 2: Flow control (for loops and the range function). No Starch Press. automatetheboringstuff.com
- Key terms
- for loop
- A loop that takes each item of something in turn, with the number of passes fixed before it starts.
- range
- A built-in that produces a run of whole numbers, stopping before the value you give as the stop.
- Step
- The third argument to range, the amount added each time, which may be negative to count downwards.
- Loop variable
- The name after for, which is given the next item at the start of each pass.
- Nested loop
- A loop inside another. The inner one runs completely for every single pass of the outer one.
- Iterate
- To go through the items of something one at a time, which is what a for loop does to a range or a string.
Three Copies of the Same Block, and One of Them Was Wrong
- Define a function with def, call it with arguments, and explain the difference from a parameter.
- Return a value rather than printing it, and say what a function that only prints hands back.
- Describe what happens to a name assigned inside a function when the function ends.
A report card for three subjects. Written the obvious way, which is to get one block right and then copy it twice:
maths = 62
english = 71
science = 55
print("Maths")
print("=" * 20)
print(f"Mark: {maths}")
print(f"Grade: {maths // 10}")
print()
print("English")
print("=" * 20)
print(f"Mark: {english}")
print(f"Grade: {english // 10}")
print()
print("Science")
print("=" * 20)
print(f"Mark: {science}")
print(f"Grade: {science / 10}")
print()
Maths
====================
Mark: 62
Grade: 6
English
====================
Mark: 71
Grade: 7
Science
====================
Mark: 55
Grade: 5.5
Science got grade 5.5. Nobody has ever been awarded a grade 5.5. The last block has a single slash where the others have two, which happened during the copying and is invisible until you read all three blocks character by character against each other. This is the entire argument for functions, and it is not about elegance or saving keystrokes. It is that the same logic written three times is three things that can be different, and the difference is silent.
Also notice "=" * 20, which is a useful trick in its own right: a string multiplied by a number repeats it, so that produces twenty equals signs without typing twenty equals signs.
def: giving a piece of work a name
Here is the whole program again, with the repeated block written once:
def report(subject, mark):
print(subject)
print("=" * 20)
print(f"Mark: {mark}")
print(f"Grade: {mark // 10}")
print()
report("Maths", 62)
report("English", 71)
report("Science", 55)
Maths
====================
Mark: 62
Grade: 6
English
====================
Mark: 71
Grade: 7
Science
====================
Mark: 55
Grade: 5
Twenty-one lines became ten, and the grade calculation now exists once. If it is right, it is right for every subject. If it is wrong, it is wrong for every subject, which sounds worse but is far better: a bug that always happens is a bug you will find today.
The shape of a function definition is: the word def, a name, brackets holding the parameters, a colon, and an indented block. Defining a function does not run it. Python reads those six lines, notes that a thing called report exists, and moves on. Nothing happens until a line calls it by writing its name and brackets, and the values in those brackets, the arguments, are matched to the parameters in order: "Maths" goes to subject, 62 goes to mark.
Because a definition must come before its first call, functions live at the top of a file and the program that uses them at the bottom. That is a house rule of Python files everywhere, and it is why real code often looks upside down at first: the interesting part is at the end.
return, and why printing is not the same as handing back
The report function prints. That is fine for a report. But most functions should produce a value for the rest of the program to use, and that is return. The difference is the single most important idea in this lesson:
def grade_printed(mark):
print(mark // 10)
def grade_returned(mark):
return mark // 10
print("What grade_printed gives back:", grade_printed(62))
best = grade_returned(62)
print("What grade_returned gives back:", best)
print("Adding two returned grades:", grade_returned(62) + grade_returned(71))
print("Adding two printed grades:", grade_printed(62) + grade_printed(71))
6
What grade_printed gives back: None
What grade_returned gives back: 6
Adding two returned grades: 13
6
7
Traceback (most recent call last):
print("Adding two printed grades:", grade_printed(62) + grade_printed(71))
~~~~~~~~~~~~~~~~~~^~~~~~~~~~~~~~~~~~~
TypeError: unsupported operand type(s) for +: 'NoneType' and 'NoneType'
Read that output in order, because it tells a story. The first 6 on its own is grade_printed doing its printing, which happens before the surrounding print can display anything. Then the surrounding print reports that grade_printed gave back None, which is Python's word for no value at all: a function with no return statement returns None, always.
grade_returned, by contrast, handed back 6, which could be stored in a variable, and two of them could be added to make 13. The last line tried to add two Nones together and got a TypeError naming NoneType twice.
This is the bug beginners hit hardest, because the printing version looks like it works. It puts the right number on the screen. But the number is gone: nothing in the program can use it, compare it, add it up or save it. Print puts characters on a screen for a human. Return hands a value back to the program. When in doubt, return; the caller can always print it.
A return statement also ends the function immediately, like break for a loop. Anything written after it in that block does not run.
Names inside a function belong to the function
total = 0
def add_bonus(mark):
total = mark + 5
print("inside the function, total is", total)
return total
result = add_bonus(62)
print("the returned value is", result)
print("outside the function, total is still", total)
def show_mark():
print("a function can read an outer name:", total)
show_mark()
inside the function, total is 67
the returned value is 67
outside the function, total is still 0
a function can read an outer name: 0
Inside the function, total was 67. Outside, it is still 0. Assigning to a name inside a function creates a local name that belongs to that call and disappears when the call ends, even when a name exactly like it exists outside. The outer total was never touched.
The last two lines show the other half: a function can read an outer name it never assigns to. That asymmetry catches people, and the practical rule is simpler than the theory. Take values in through parameters, send values out through return, and do not rely on reading outer names. A function written that way can be moved to another file and still work, which is exactly the test of whether it is a real piece of work or just a fragment of this program.
Defaults, docstrings, and three functions that build a program
Two finishing touches, and then the whole thing assembled:
def grade_for(mark):
"""Turn a mark out of 100 into a GCSE-style grade from 9 down to U."""
if mark >= 90:
return "9"
elif mark >= 80:
return "8"
elif mark >= 70:
return "7"
elif mark >= 60:
return "6"
return "U"
def ask_for_mark(subject, tries=3):
"""Ask until the answer is a whole number from 0 to 100, or give up."""
for attempt in range(tries):
answer = input(f"{subject} mark out of 100? ")
if answer.isdigit() and 0 <= int(answer) <= 100:
return int(answer)
print("That is not a mark between 0 and 100.")
return 0
def report(subject, mark):
print(f"{subject}: {mark}, grade {grade_for(mark)}")
print("Checks on grade_for:")
print(grade_for(91), grade_for(90), grade_for(89), grade_for(59), grade_for(0))
maths = ask_for_mark("Maths")
english = ask_for_mark("English")
report("Maths", maths)
report("English", english)
print(f"Mean mark: {(maths + english) / 2:.1f}")
Fed 62, then the word ninety, then 71:
Checks on grade_for:
9 9 8 U U
Maths mark out of 100? English mark out of 100? That is not a mark between 0 and 100.
English mark out of 100? Maths: 62, grade 6
English: 71, grade 7
Mean mark: 66.5
Several things arrived at once there, so take them one at a time.
The line of checks is the important one. grade_for(91), grade_for(90), grade_for(89), grade_for(59), grade_for(0) gave 9, 9, 8, U, U, and those are the boundaries: the lowest mark that still gets a 9, the highest that does not, and the bottom of the range. A function small enough to test on its boundaries in one line is a function you can trust. This is the first step towards Lesson 14, where those checks stop being print statements and start being tests.
tries=3 is a default value. Call ask_for_mark("Maths") and tries is 3; call ask_for_mark("Maths", 5) and it is 5. Defaults let a function have an easy common case and a fuller one, without two functions.
The string in triple quotes just under each def is a docstring. It is not a comment; it is stored with the function, and typing help(grade_for) in the interactive shell prints it. One line saying what the function does and what it gives back is enough.
And answer.isdigit() is a taste of Lesson 11: a question you can ask a string about itself, which is True only when every character is a digit. It is what stops the word ninety reaching int and raising an error.
When to write a function
Three signals, in order of how reliable they are.
You are about to copy a block. This is the strongest one, and it is what went wrong at the top of the lesson. The principle has a name in software, do not repeat yourself, and the reason is not typing but truth: two copies can disagree.
A chunk of your program has a name. If you find yourself writing a comment that says what the next eight lines do, that comment is a function name waiting to happen. Delete the comment, write def ask_for_mark(subject): and the code documents itself.
You want to test a piece on its own. You cannot check the grade boundaries of a calculation buried in the middle of a program. You can check them in one line when it is a function that takes a mark and returns a grade.
Two habits from PEP 8 while you are here: function names are lowercase with underscores, like variables, and they are usually verbs, because a function does something. grade_for, ask_for_mark, report. Two blank lines between functions at the top level of a file is the convention, which is why the assembled program above looks airier than the ones before it.
Common misconceptions
- "A function that prints the answer has given me the answer." It has given the screen the answer and given the program None. Nothing can add it, compare it or store it. Use return.
- "Defining a function runs it." Python only notes that it exists. Nothing in the block happens until a call, which is why a file of nothing but definitions produces no output at all.
- "Assigning to a variable inside a function changes the one outside." It creates a separate local name that vanishes when the call ends. To get a value out, return it.
- "Parameters and arguments are the same word." The parameter is the name in the definition, the argument is the value in the call. They meet when the function is called, in order.
- "Code after a return still runs." return leaves the function immediately. Lines below it in that block are dead code.
The takeaway
So what?: Write it once, give it a name, take values in through parameters and hand the answer back with return. Copies of a block are copies of a bug waiting to disagree.
You can now define a function with def, give it parameters, call it with arguments matched in order, and hand a value back with return. You know that a function without a return gives back None, and you have seen the TypeError that None causes when the caller tries to use it. You know that names assigned inside a function are local and disappear when it ends, that a default value gives a parameter an easy common case, and that a docstring under the def says what the function is for. And you have a test for when to write one: whenever you are about to copy a block, whenever a chunk of your program has a name, and whenever you want to check a piece on its own.
That closes Module 2. You have decisions, loops and functions, which is the machinery of every program ever written. What you do not yet have is anywhere to put more than a handful of values. Module 3 is about holding many things at once: a list of marks, a dictionary of names, a string taken apart letter by letter.
Sources
- Python Software Foundation. (2026). More control flow tools: Defining functions, default argument values, and keyword arguments. Python 3.14 tutorial. docs.python.org
- Goodger, D., and van Rossum, G. (2001). PEP 257: Docstring conventions. Python Software Foundation. peps.python.org
- van Rossum, G., Warsaw, B., and Coghlan, N. (2001, revised 2023). PEP 8: Style guide for Python code (function names, blank lines between definitions). peps.python.org
- Severance, C. (2016). Python for Everybody: Exploring data in Python 3, Chapter 4: Functions (parameters, return values and why functions exist). py4e.com
- Sweigart, A. (2019). Automate the Boring Stuff with Python (2nd ed.), Chapter 3: Functions (local and global scope, the None value). No Starch Press. automatetheboringstuff.com
- Key terms
- Function definition
- A block introduced by def that names a piece of work. Defining it does not run it.
- Parameter
- A name in the brackets of a definition, which receives a value when the function is called.
- Argument
- A value supplied in the brackets of a call, matched to the parameters in order.
- return
- Hands a value back to whatever called the function and ends the function at once.
- None
- Python's word for no value. A function with no return statement gives back None.
- Local name
- A name assigned inside a function, which belongs to that call and disappears when it ends.
- Default value
- A value written in the definition, used when the caller leaves that argument out.
- Docstring
- A string just under the def line describing the function, stored with it and shown by help().
Module 3: Holding Many Things at Once
Put a whole set of values under one name with a list, look things up by name with a dictionary, and take text apart character by character to build a cipher.
Six Marks, One Name: Lists
- Build a list, read and change items by position, and explain why positions start at zero.
- Take a slice of a list or string and say what the two numbers mean.
- Use the common list methods and predict what happens when two names refer to one list.
Six test marks. With what you know so far, that is six variables, six lines to add them up, and a fresh problem the moment a seventh test is set. With a list it is one name:
marks = [62, 71, 55, 88, 43, 67]
print("The whole list:", marks)
print("How many:", len(marks))
print("Total:", sum(marks))
print("Mean:", sum(marks) / len(marks))
print("Best:", max(marks))
print("Worst:", min(marks))
for mark in marks:
print(mark, end=" ")
print()
The whole list: [62, 71, 55, 88, 43, 67]
How many: 6
Total: 386
Mean: 64.33333333333333
Best: 88
Worst: 43
62 71 55 88 43 67
A list is written in square brackets with commas between the items. Add a seventh mark and every one of those lines still works, unchanged, which is the point: the program no longer knows or cares how many marks there are.
len, sum, max and min are built-in functions that take a list and hand back one value. And the for loop from Lesson 7 walks through a list exactly as it walked through a string, because for x in something works on anything that has items in an order.
Positions start at zero
Each item has a position, called its index, and the first one is 0:
marks = [62, 71, 55, 88, 43, 67]
print("marks[0] :", marks[0])
print("marks[1] :", marks[1])
print("marks[5] :", marks[5])
print("marks[-1]:", marks[-1])
print("marks[-2]:", marks[-2])
print("len :", len(marks))
marks[2] = 60
print("after marks[2] = 60:", marks)
print("marks[6] :", marks[6])
marks[0] : 62
marks[1] : 71
marks[5] : 67
marks[-1]: 67
marks[-2]: 43
len : 6
after marks[2] = 60: [62, 71, 60, 88, 43, 67]
Traceback (most recent call last):
print("marks[6] :", marks[6])
~~~~~^^^
IndexError: list index out of range
Six items, and their indexes run 0, 1, 2, 3, 4, 5. The last index is always len - 1, which is the off-by-one error from Lesson 6 in its final form, and marks[6] raised an IndexError exactly as you would hope: loudly, on the spot, rather than quietly giving something wrong.
Negative indexes count backwards from the end, so marks[-1] is the last item and marks[-2] the one before it. This is genuinely useful. Getting the last item of a list you did not build yourself would otherwise mean marks[len(marks) - 1], which is both ugly and an off-by-one waiting to happen.
A list is mutable: marks[2] = 60 changed the item in place. Strings are not. Try word[0] = "P" on a string and Python raises a TypeError saying a str does not support item assignment, which is why Lesson 11 builds new strings rather than editing old ones.
Slices: a piece of a list, as a new list
Two numbers in the brackets with a colon between them give you a run of items:
marks = [62, 71, 55, 88, 43, 67]
print("marks[0:3] :", marks[0:3])
print("marks[:3] :", marks[:3])
print("marks[3:] :", marks[3:])
print("marks[1:4] :", marks[1:4])
print("marks[-2:] :", marks[-2:])
print("marks[:] :", marks[:])
print("marks[2:2] :", marks[2:2])
print("original unchanged:", marks)
word = "programming"
print("word[0:7] :", word[0:7])
print("word[-4:] :", word[-4:])
marks[0:3] : [62, 71, 55]
marks[:3] : [62, 71, 55]
marks[3:] : [88, 43, 67]
marks[1:4] : [71, 55, 88]
marks[-2:] : [43, 67]
marks[:] : [62, 71, 55, 88, 43, 67]
marks[2:2] : []
original unchanged: [62, 71, 55, 88, 43, 67]
word[0:7] : program
word[-4:] : ming
The rule is the same one range obeys: start included, stop excluded. So marks[0:3] gives three items, at positions 0, 1 and 2, and never touches position 3. Leave the start out and it means from the beginning; leave the stop out and it means to the end. marks[2:2] is empty because the start and stop are the same, which is exactly what range(3, 3) did.
Two consequences worth keeping. A slice always makes a new list, which is why the original was unchanged on the line after; marks[:] is therefore the standard way to take a copy. And slicing works on strings too, giving a new string, which is how word[0:7] pulled "program" out of "programming".
Methods: things a list can do to itself
A method is a function that belongs to a value and is called by writing a full stop after it. Lists have a set of them for adding and removing items:
queue = ["Ana", "Ben", "Cleo"]
print("start :", queue)
queue.append("Dev")
print("append Dev :", queue)
queue.insert(0, "Zoe")
print("insert Zoe :", queue)
queue.remove("Ben")
print("remove Ben :", queue)
served = queue.pop(0)
print("pop(0) gave:", served, "leaving", queue)
print("where is Cleo? ", queue.index("Cleo"))
marks = [62, 71, 55, 88, 43, 67]
marks.sort()
print("sorted :", marks)
marks.reverse()
print("reversed :", marks)
print("sorted(marks) makes a new list:", sorted([3, 1, 2]), "from", [3, 1, 2])
start : ['Ana', 'Ben', 'Cleo']
append Dev : ['Ana', 'Ben', 'Cleo', 'Dev']
insert Zoe : ['Zoe', 'Ana', 'Ben', 'Cleo', 'Dev']
remove Ben : ['Zoe', 'Ana', 'Cleo', 'Dev']
pop(0) gave: Zoe leaving ['Ana', 'Cleo', 'Dev']
where is Cleo? 1
sorted : [43, 55, 62, 67, 71, 88]
reversed : [88, 71, 67, 62, 55, 43]
sorted(marks) makes a new list: [1, 2, 3] from [3, 1, 2]
| Method | Does | Hands back |
|---|---|---|
| append(x) | Adds x at the end | None |
| insert(i, x) | Puts x at position i, shifting the rest along | None |
| remove(x) | Deletes the first x it finds | None |
| pop(i) | Removes the item at i | That item |
| index(x) | Finds where x is | Its position |
| sort() | Reorders this list | None |
| reverse() | Flips this list round | None |
The right-hand column is where the trap is. sort and append change the list and return None, so marks = marks.sort() throws your marks away and leaves you with None. If you want a sorted copy and the original kept, use the built-in sorted(marks), which returns a new list, as the last line of that output shows. The distinction between a method that changes a thing and a function that makes a new one runs right through Python and is worth noticing every time.
in: the question that replaces a chain of ors
Remember the vowel counter from Lesson 7, with its four ors? Here is the whole thing:
queue = ["Ana", "Cleo", "Dev"]
print("is Cleo in the queue?", "Cleo" in queue)
print("is Ben in the queue? ", "Ben" in queue)
print("is 'e' a vowel?", "e" in "aeiou")
print("is 'z' a vowel?", "z" in "aeiou")
is Cleo in the queue? True
is Ben in the queue? False
is 'e' a vowel? True
is 'z' a vowel? False
in asks whether something is among the items and gives back True or False. It works on lists and on strings, where the items are characters. So if letter in "aeiou" replaces four comparisons, reads like English, and cannot be got wrong the way the chain of ors could.
This is the same word that appears in for letter in word, doing a related job: there it takes each item in turn, here it asks whether one is present.
Two names, one list
Now the thing that will bite you, and which is worth twenty minutes of your attention today rather than two hours in a month:
original = [62, 71, 55]
same_list = original
copy_of_it = original[:]
original.append(99)
print("original :", original)
print("same_list :", same_list)
print("copy_of_it:", copy_of_it)
def add_a_mark(some_marks):
some_marks.append(100)
scores = [10, 20]
add_a_mark(scores)
print("after the function call, scores is", scores)
original : [62, 71, 55, 99]
same_list : [62, 71, 55, 99]
copy_of_it: [62, 71, 55]
after the function call, scores is [10, 20, 100]
same_list = original did not copy anything. It gave a second name to the same list, so appending through one name shows up through the other. copy_of_it = original[:] took a slice, which makes a new list, and that one stayed as it was.
The function shows why this matters. Lesson 8 said that assigning inside a function does not affect the outside, and that is still true. But some_marks.append(100) is not an assignment; it changes the list that both names point at, and the change is visible to the caller. Some functions are written to do exactly this, and the rule is to be deliberate: either a function returns a new list and leaves its argument alone, or it changes the list it was given and returns None, like sort. Doing both, or neither on purpose, is how lists mysteriously grow.
Common misconceptions
- "The first item is number one." It is number zero, and the last is at
len - 1. A list of six items has no index 6, and asking for one raises an IndexError. - "marks[0:3] gives four items." It gives three. The stop is excluded, exactly as with range.
- "marks = marks.sort() sorts the list into marks." sort returns None, so that line replaces your list with None. Either call
marks.sort()on its own or usesorted(marks). - "copy = original makes a copy." It makes a second name for the same list. Use
original[:], or the list's owncopymethod, when you want an independent one. - "A function cannot change a variable belonging to the caller." It cannot reassign one, but it can change the contents of a list it was handed, because both names point at the same list.
What to remember
The core of it: A list is one name for many values in order, its positions start at zero, and a slice takes a new list while a method usually changes the old one.
You can build a list, read and replace items by index, count backwards with negative indexes, and recognise the IndexError that a position one past the end produces. You can slice a list or a string, knowing the stop is excluded, and you know that a slice is always a new object. You can append, insert, remove, pop, sort and reverse, and you know which of those hand back None and which hand back a value. You can ask whether something is present with in, which finally kills the chain of ors. And you know that two names can share one list, which is a convenience when you meant it and a bug when you did not.
A list is right when position is what identifies a thing: first mark, second mark, last in the queue. It is the wrong shape when you want to look something up by a name rather than a number, which is what the next lesson is about.
Sources
- Python Software Foundation. (2026). Data structures: More on lists (list methods, and the return values of sort and append). Python 3.14 tutorial. docs.python.org
- Python Software Foundation. (2026). The Python tutorial: Lists (indexing, slicing and mutability). docs.python.org
- Severance, C. (2016). Python for Everybody: Exploring data in Python 3, Chapter 8: Lists. py4e.com
- Sweigart, A. (2019). Automate the Boring Stuff with Python (2nd ed.), Chapter 4: Lists (references, and passing lists to functions). No Starch Press. automatetheboringstuff.com
- Key terms
- List
- An ordered collection of values under one name, written in square brackets with commas between the items.
- Index
- The position of an item, counted from 0. The last item is at len minus one.
- IndexError
- The error raised by asking for a position that does not exist, such as index 6 of a six-item list.
- Slice
- A run of items taken with two numbers and a colon. The start is included, the stop is not, and the result is a new list.
- Method
- A function belonging to a value, called by writing a full stop and its name after the value.
- Mutable
- Able to be changed in place. Lists are mutable; strings are not.
- Aliasing
- Two names referring to the same list, so a change made through one is visible through the other.
- in
- An operator asking whether a value is among the items of a list or string, giving True or False.
When a Position Is the Wrong Way to Find Something
- Say when a dictionary suits a problem better than a list, and why.
- Create, read, add to and delete from a dictionary, and handle a key that is not there.
- Loop over keys, values and pairs, and count occurrences with the get pattern.
Four subjects, four marks, kept in two lists that have to stay lined up:
subjects = ["Maths", "English", "Science", "History"]
marks = [62, 71, 55, 88]
# Someone drops History from the subject list but forgets the marks.
subjects.remove("History")
for i in range(len(subjects)):
print(subjects[i], marks[i])
print("Which subject got 88?", marks.index(88), "->", subjects[marks.index(88)])
Maths 62
English 71
Science 55
Traceback (most recent call last):
print("Which subject got 88?", marks.index(88), "->", subjects[marks.index(88)])
~~~~~~~~^^^^^^^^^^^^^^^^^
IndexError: list index out of range
Look at the first three lines of that output, because they are the dangerous part. They are correct. The program printed three sensible rows and gave no hint that a mark of 88 was now floating loose with no subject attached to it. The crash only came later, when something asked which subject scored 88 and looked for a position that no longer existed.
Two lists that have to be kept in step are called parallel lists, and they are a promise the language cannot enforce. Nothing connects subjects[3] to marks[3] except your intention to keep them together, and intentions do not survive contact with a program you edit six weeks later.
A dictionary keeps the pairing
A dictionary stores pairs: a key and the value it points at. Written with curly brackets, a colon between each key and value, and commas between the pairs:
marks = {"Maths": 62, "English": 71, "Science": 55, "History": 88}
print("The whole thing:", marks)
print("English mark :", marks["English"])
print("How many subjects:", len(marks))
marks["Art"] = 74
print("after adding Art:", marks)
marks["Maths"] = 65
print("after correcting Maths:", marks)
del marks["History"]
print("after deleting History:", marks)
print("Is there a French mark?", "French" in marks)
print("Is there an Art mark? ", "Art" in marks)
The whole thing: {'Maths': 62, 'English': 71, 'Science': 55, 'History': 88}
English mark : 71
How many subjects: 4
after adding Art: {'Maths': 62, 'English': 71, 'Science': 55, 'History': 88, 'Art': 74}
after correcting Maths: {'Maths': 65, 'English': 71, 'Science': 55, 'History': 88, 'Art': 74}
after deleting History: {'Maths': 65, 'English': 71, 'Science': 55, 'Art': 74}
Is there a French mark? False
Is there an Art mark? True
Now the pairing cannot come apart. Deleting History removes the subject and its mark in one action, because they were never separate things. And notice the fourth and fifth lines: marks["Art"] = 74 created a new pair while marks["Maths"] = 65 replaced an existing one, using identical syntax. A dictionary key either exists or does not, and assigning to it handles both cases the same way.
List or dictionary
Both hold many values under one name. The question that decides between them is: how will you find a particular item again?
| List | Dictionary | |
|---|---|---|
| Written with | Square brackets, [62, 71] | Curly brackets, {"Maths": 62} |
| Found by | Position, counted from 0 | A key you chose |
| Order | Is the meaning: first, second, last | Is kept, but is not what identifies anything |
| Missing item gives | IndexError | KeyError |
| Good for | Marks in the order taken, a queue, the lines of a file | Marks by subject, a phone book, counting how often things occur |
| Duplicates | Allowed anywhere | Keys are unique; assigning again replaces |
Two quick tests. If "the third one" is a sensible thing to say about your data, use a list. If you would naturally say "the one for History", use a dictionary. And if you ever find yourself writing two lists that must be kept the same length, that is a dictionary with the pairing pulled apart.
The key that is not there
Asking a dictionary for a key it does not have is an error, and there is a second way to ask that is not:
marks = {"Maths": 62, "English": 71, "Science": 55}
print("get with a key that exists :", marks.get("Maths"))
print("get with a key that does not:", marks.get("French"))
print("get with a fallback :", marks.get("French", 0))
print("now the square brackets:")
print(marks["French"])
get with a key that exists : 62
get with a key that does not: None
get with a fallback : 0
now the square brackets:
Traceback (most recent call last):
print(marks["French"])
~~~~~^^^^^^^^^^
KeyError: 'French'
Square brackets demand the key and raise a KeyError if it is missing. The get method asks politely: it gives back None when the key is absent, or whatever fallback value you supply as a second argument. Both behaviours are useful. Use the square brackets when the key must be there and a missing one means your program is already wrong; use get with a fallback when absent genuinely means zero or nothing.
Three ways to walk through a dictionary
marks = {"Maths": 62, "English": 71, "Science": 55, "History": 88}
print("keys :", list(marks.keys()))
print("values:", list(marks.values()))
print("items :", list(marks.items()))
for subject in marks:
print(f"{subject:10} {marks[subject]}")
print("--")
for subject, mark in marks.items():
print(f"{subject:10} {mark}")
print("Total:", sum(marks.values()))
print("Mean :", f"{sum(marks.values()) / len(marks):.1f}")
print("Best subject:", max(marks, key=marks.get))
keys : ['Maths', 'English', 'Science', 'History']
values: [62, 71, 55, 88]
items : [('Maths', 62), ('English', 71), ('Science', 55), ('History', 88)]
Maths 62
English 71
Science 55
History 88
--
Maths 62
English 71
Science 55
History 88
Total: 276
Mean : 69.0
Best subject: History
A plain for subject in marks walks over the keys, which surprises people who expect the values. If you want both, marks.items() hands out pairs and you can catch them in two names at once, which is the tidiest of the three forms and the one to reach for.
The items are shown in round brackets, like ('Maths', 62). That is a tuple, which is a list that cannot be changed after it is made. You do not need to do anything with tuples in this course beyond recognising the brackets; they turn up when a function wants to hand back two things at once.
sum(marks.values()) works because values() produces the numbers alone. And max(marks, key=marks.get) looks strange but reads sensibly once you know the trick: go through the keys, and judge each one by what marks.get says about it, so the answer is the key with the largest value rather than the alphabetically last one.
Since Python 3.7 a dictionary keeps its pairs in the order they were added, which is why the output above is not shuffled. Do not lean on it for anything that matters; if order is the point, use a list, or sort the keys when you print them.
Counting, which is what dictionaries are best at
Here is the single most reused dictionary pattern in all of programming: counting how often each thing appears.
line = "the rain in spain stays mainly in the plain"
counts = {}
for letter in line:
if letter == " ":
continue
counts[letter] = counts.get(letter, 0) + 1
print(counts)
for letter in sorted(counts):
print(letter, "*" * counts[letter])
best = max(counts, key=counts.get)
print(f"Commonest letter: {best}, appearing {counts[best]} times")
{'t': 3, 'h': 2, 'e': 2, 'r': 1, 'a': 5, 'i': 6, 'n': 6, 's': 3, 'p': 2, 'y': 2, 'm': 1, 'l': 2}
a *****
e **
h **
i ******
l **
m *
n ******
p **
r *
s ***
t ***
y **
Commonest letter: i, appearing 6 times
The line doing the work is counts[letter] = counts.get(letter, 0) + 1, and it handles both cases at once. If this letter has been seen, get returns the running count and one is added. If it has not, get returns the fallback 0 and the count becomes 1. Without get you would need an if statement to check whether the key exists first, which is four lines instead of one.
continue is the partner to break from Lesson 6: it abandons the rest of this pass and goes straight to the next one, so spaces are skipped without being counted. And sorted(counts) sorts the keys, which is how the report came out alphabetically even though the dictionary is in the order the letters first appeared. That little bar chart, made of asterisks, is the seed of the charting in Lesson 17.
What can be a key
Values can be anything: numbers, strings, lists, even other dictionaries. Keys are stricter. A key must be an unchangeable value, so strings, numbers and tuples are fine and a list is not: try it and Python says a list is unhashable. The reason is that a dictionary finds things by computing a number from the key, and a key that could change afterwards would make the entry impossible to find again.
In practice this is barely a restriction, because keys are almost always strings or numbers. The useful part is what values can be. A dictionary whose values are lists is how you store several things per key, and a dictionary whose values are dictionaries is how you store a record per person, which is what the text adventure in Lesson 21 uses to hold its rooms.
Common misconceptions
- "for x in my_dict gives the values." It gives the keys. Use
.values()for the values or.items()for both at once. - "A missing key gives None." Square brackets raise a KeyError. Only
.get()gives None, which is exactly why it exists. - "Dictionaries have no order, so printing one is random." Since Python 3.7 insertion order is kept. It is still not what identifies an item, so do not build anything on the position of a pair.
- "You need an if to check whether a key is there before counting."
counts.get(key, 0) + 1handles the first time and every later time in one line. - "Two lists side by side are the same as a dictionary." Nothing stops one list from being edited without the other. A dictionary makes the pairing impossible to break by accident.
Where this leaves us
Bottom line: Use a list when position is the meaning and a dictionary when a name is, and reach for get whenever a key might not be there.
You can build a dictionary, read a value by its key, add and replace pairs with the same square-bracket assignment, delete with del, and test for a key with in. You know that square brackets raise a KeyError and that get returns None or a fallback instead. You can loop over keys, over values, or over both at once with items, and you can count occurrences in one line with the get pattern, which is the single most reused dictionary idiom there is. And you can say why two parallel lists are a bug waiting to happen.
You now have both shapes for holding data. The next lesson goes back to the simplest type of all, the string, and finds that it has a whole toolkit of its own, enough to take a message apart, shift every letter along the alphabet, and put it back together again.
Sources
- Python Software Foundation. (2026). Data structures: Dictionaries, and looping techniques. Python 3.14 tutorial. docs.python.org
- Python Software Foundation. (2026). Built-in types: Mapping types, dict (the get method and the requirement that keys be hashable). docs.python.org
- Severance, C. (2016). Python for Everybody: Exploring data in Python 3, Chapter 9: Dictionaries (counting with a dictionary). py4e.com
- Sweigart, A. (2019). Automate the Boring Stuff with Python (2nd ed.), Chapter 5: Dictionaries and structuring data. No Starch Press. automatetheboringstuff.com
- Key terms
- Dictionary
- A collection of key and value pairs, written in curly brackets, where a value is found by its key rather than its position.
- Key
- The label a value is stored under. Keys are unique and must be unchangeable values such as strings or numbers.
- Value
- The thing stored under a key. It can be anything, including a list or another dictionary.
- KeyError
- The error raised by asking with square brackets for a key the dictionary does not have.
- get
- A method that returns the value for a key, or None, or a fallback you supply, instead of raising an error.
- items
- A method producing each key and value together as a pair, for looping over both at once.
- Tuple
- An ordered group of values that cannot be changed after it is made, written in round brackets.
- continue
- Abandons the rest of this pass of a loop and starts the next one, in the way break leaves the loop entirely.
wkh sdvvzrug lv vfduohw
- Use the common string methods and explain why none of them changes the original string.
- Convert between a character and its number with ord and chr, and wrap past z with the remainder operator.
- Write functions that encode and decode a Caesar cipher, and break one by trying every shift.
A note is passed across a classroom. It says:
wkh sdvvzrug lv vfduohw
It is not gibberish and it is not a language. Every letter has been moved three places along the alphabet: w came from t, k from h, h from e. Undo the shift and it says "the password is scarlet". This is a Caesar cipher, named for the Roman general who is reported to have used one for military dispatches, and by the end of this lesson you will have written a program that produces it, another that undoes it, and a third that breaks it without being told the shift.
Everything you need is in the string type you have been using since Lesson 1. It turns out to have a toolkit.
A string cannot be changed, only rebuilt
Strings index and slice exactly like lists. What they do not do is change:
word = "python"
print("word[0]:", word[0])
print("word[-1]:", word[-1])
print("word[2:5]:", word[2:5])
shouted = word.upper()
print("after word.upper():", shouted, "and word is still", word)
word[0] = "P"
word[0]: p
word[-1]: n
word[2:5]: tho
after word.upper(): PYTHON and word is still python
Traceback (most recent call last):
word[0] = "P"
~~~~^^^
TypeError: 'str' object does not support item assignment
Strings are immutable: fixed once made. That is why word.upper() did not shout the original into capitals but handed back a new string, leaving word exactly as it was. Every string method works this way, which leads directly to the commonest string mistake in Python: writing name.strip() on a line by itself and wondering why the spaces are still there. The tidy string was made and thrown away. You have to catch it: name = name.strip().
Because you cannot edit a string in place, you build a new one. The pattern is the accumulator from Lesson 6 with an empty string instead of zero: start with result = "" and add a character at a time. That is exactly how the cipher will work.
The string methods worth knowing now
messy = " Ada Lovelace, 1815 "
print("original :", repr(messy))
print("strip() :", repr(messy.strip()))
print("upper() :", messy.strip().upper())
print("replace :", messy.strip().replace("1815", "1852"))
print("find('Love') :", messy.find("Love"))
print("find('Byron') :", messy.find("Byron"))
print("startswith :", messy.strip().startswith("Ada"))
print("len :", len(messy.strip()))
print("count('a') :", messy.count("a"))
print("'2026'.isdigit() :", "2026".isdigit())
print("'20a6'.isdigit() :", "20a6".isdigit())
print("'Ada'.isalpha() :", "Ada".isalpha())
original : ' Ada Lovelace, 1815 '
strip() : 'Ada Lovelace, 1815'
upper() : ADA LOVELACE, 1815
replace : Ada Lovelace, 1852
find('Love') : 7
find('Byron') : -1
startswith : True
len : 18
count('a') : 2
'2026'.isdigit() : True
'20a6'.isdigit() : False
'Ada'.isalpha() : True
| Method | Gives back | Used for |
|---|---|---|
| strip() | A copy with the spaces trimmed off both ends | Cleaning up anything typed by a person |
| upper(), lower() | A copy in one case | Comparing answers without caring about capitals |
| replace(a, b) | A copy with every a swapped for b | Fixing or censoring text |
| find(x) | The position of x, or -1 if absent | Locating something inside a line |
| count(x) | How many times x appears | Tallying without a loop |
| startswith(x) | True or False | Checking a prefix, such as a command word |
| isdigit(), isalpha() | True or False | Checking before int() or before shifting a letter |
Two details. repr shows a string with its quotation marks, which is the only way to see that strip actually removed something; printed normally, spaces at the ends are invisible. And find returns -1 rather than raising an error when the text is absent, which is a nasty trap if you forget: -1 is a perfectly valid index meaning the last character, so text[text.find("x")] on a missing x quietly gives you the wrong character instead of complaining.
Methods can be chained, as in messy.strip().upper(): strip hands back a string, and upper is then called on that. Read a chain left to right as a sequence of steps.
split and join, which take text apart and put it back
names = "Ada,Grace,Katherine,Dorothy"
parts = names.split(",")
print("split :", parts)
print("join with ' and ':", " and ".join(parts))
sentence = "the rain in spain"
print("split on spaces:", sentence.split())
print("title() :", sentence.title())
split : ['Ada', 'Grace', 'Katherine', 'Dorothy']
join with ' and ': Ada and Grace and Katherine and Dorothy
split on spaces: ['the', 'rain', 'in', 'spain']
title() : The Rain In Spain
split turns a string into a list, cutting at whatever you give it, or at runs of whitespace when you give it nothing. It is how a line of a data file becomes usable values, which is the whole of Lesson 16. join goes the other way and is written backwards from what everybody expects: the separator is the string you call the method on, and the list is the argument. ", ".join(parts), not parts.join(", "). Write it wrong once, read the error, and it will stick.
Every character is a number underneath
To shift a letter along the alphabet you need arithmetic, and letters are not numbers. Except that underneath, they are:
print("ord('a') :", ord("a"))
print("ord('z') :", ord("z"))
print("ord('A') :", ord("A"))
print("ord(' ') :", ord(" "))
print("chr(97) :", chr(97))
print("chr(122) :", chr(122))
for letter in "abc":
print(letter, "->", ord(letter), "->", ord(letter) - ord("a"))
print("shifting a by 3:", chr(ord("a") + 3))
print("shifting z by 3 the wrong way:", chr(ord("z") + 3))
print("shifting z by 3 with wraparound:", chr((ord("z") - ord("a") + 3) % 26 + ord("a")))
ord('a') : 97
ord('z') : 122
ord('A') : 65
ord(' ') : 32
chr(97) : a
chr(122) : z
a -> 97 -> 0
b -> 98 -> 1
c -> 99 -> 2
shifting a by 3: d
shifting z by 3 the wrong way: }
shifting z by 3 with wraparound: c
Every character your computer stores has a number, fixed by a standard that began as ASCII in the 1960s and grew into Unicode, which covers every writing system in use. Lowercase a is 97 and z is 122, so the twenty-six letters are twenty-six numbers in a row. ord gives the number for a character and chr gives the character for a number. Capital A is 65, which is a different number entirely, and that is why the cipher below lowercases everything first.
ord(letter) - ord("a") converts a letter into its position in the alphabet: a becomes 0, b becomes 1, c becomes 2. Do the arithmetic there, in the range 0 to 25, then add 97 back at the end. That is much easier to think about than juggling numbers in the nineties.
The middle of the three shift lines is the bug you would have written. chr(ord("z") + 3) is 125, which is a closing curly bracket, because past z the numbers simply carry on into punctuation. The fix is the remainder operator from Lesson 3: (position + 3) % 26 keeps the answer between 0 and 25, so 25 plus 3 becomes 28 becomes 2, which is c. Wrapping round a fixed range is what % is for, and it turns up again in clocks, calendars and dice.
The cipher, in two functions
def shift_letter(letter, amount):
"""Move one lowercase letter along the alphabet, wrapping past z."""
if not letter.isalpha():
return letter
position = ord(letter.lower()) - ord("a")
moved = (position + amount) % 26
return chr(moved + ord("a"))
def caesar(text, amount):
"""Encode or decode a whole message by shifting every letter."""
result = ""
for letter in text:
result = result + shift_letter(letter, amount)
return result
message = "meet me by the bike sheds at four"
secret = caesar(message, 3)
back = caesar(secret, -3)
print("plain :", message)
print("secret:", secret)
print("back :", back)
print("Did it come back the same?", back == message)
print("shift 13 twice:", caesar(caesar(message, 13), 13) == message)
print("shift 26:", caesar(message, 26) == message)
plain : meet me by the bike sheds at four
secret: phhw ph eb wkh elnh vkhgv dw irxu
back : meet me by the bike sheds at four
Did it come back the same? True
shift 13 twice: True
shift 26: True
The decomposition from Lesson 4 is doing the work here. shift_letter handles exactly one character and nothing else; caesar knows nothing about the alphabet and only walks the string. Each is short enough to check on its own.
Three details worth having. The guard if not letter.isalpha(): return letter passes spaces and punctuation through untouched, which is why the words in the secret message still have gaps between them. There is no separate decode function, because decoding is encoding with a negative shift; a function that is its own inverse with a sign flip is a small piece of elegance worth noticing. And the last two lines are checks: shifting by 13 twice comes back to the start because 26 is a whole trip round, and shifting by 26 changes nothing at all. Those are the boundary cases from Lesson 8, applied to a cipher.
Breaking it in twenty-five guesses
The Caesar cipher has one fatal weakness: there are only twenty-five shifts worth trying. A computer can try all of them faster than you can read them.
intercepted = caesar("the password is scarlet", 3)
print("intercepted:", intercepted)
for guess in range(1, 26):
print(f"{guess:2}: {caesar(intercepted, -guess)}")
intercepted: wkh sdvvzrug lv vfduohw
1: vjg rcuuyqtf ku uectngv
2: uif qbttxpse jt tdbsmfu
3: the password is scarlet
4: sgd ozrrvnqc hr rbzqkds
5: rfc nyqqumpb gq qaypjcr
6: qeb mxpptloa fp pzxoibq
7: pda lwoosknz eo oywnhap
8: ocz kvnnrjmy dn nxvmgzo
9: nby jummqilx cm mwulfyn
10: max itllphkw bl lvtkexm
11: lzw hskkogjv ak kusjdwl
12: kyv grjjnfiu zj jtricvk
13: jxu fqiimeht yi isqhbuj
14: iwt ephhldgs xh hrpgati
15: hvs doggkcfr wg gqofzsh
16: gur cnffjbeq vf fpneyrg
17: ftq bmeeiadp ue eomdxqf
18: esp alddhzco td dnlcwpe
19: dro zkccgybn sc cmkbvod
20: cqn yjbbfxam rb bljaunc
21: bpm xiaaewzl qa akiztmb
22: aol whzzdvyk pz zjhysla
23: znk vgyycuxj oy yigxrkz
24: ymj ufxxbtwi nx xhfwqjy
25: xli tewwasvh mw wgevpix
Twenty-four lines of nonsense and one line of English, on guess 3. You did not need to be clever; you needed a loop. That is cryptanalysis at its crudest, called a brute-force attack, and it works here because the number of possible keys is tiny. Real encryption is not stronger because its method is more tangled; it is stronger because the number of keys is so large that trying them all would take longer than the age of the universe.
Notice something else about that list. A human spots the English line instantly, but the program does not know which line is right. Teaching it to know would mean counting letter frequencies, which is exactly the dictionary counting pattern from Lesson 10 pointed at a new job. That is a good extension project for Lesson 23.
Common misconceptions
- "name.strip() removes the spaces from name." It builds a tidy copy and hands it back. If you do not catch it with
name = name.strip(), nothing has changed. - "You can fix one letter with word[0] = 'P'." Strings are immutable; that is a TypeError. Build a new string, for example
"P" + word[1:]. - "find returns nothing when the text is absent." It returns -1, which is a valid index meaning the last character. Test the result before you use it.
- "join is written list.join(separator)." It is the other way round:
", ".join(parts). The separator owns the method. - "Shifting z by 3 gives c automatically." It gives a curly bracket, because the character numbers run straight on past z. Only
% 26makes it wrap.
Summing up
Worth holding on to: A string method never changes the string; it hands you a new one, so catch the result. And when letters need arithmetic, turn them into positions 0 to 25 with ord, work there, and turn them back with chr.
You can index and slice a string, and you know why assigning to one position raises a TypeError. You have the methods that matter for now: strip, upper and lower, replace, find, count, startswith, and the is-checks that guard a conversion. You can split a line into a list and join a list into a line, in the right order. You know that every character has a number, that lowercase a is 97, and that the remainder operator is what makes an alphabet wrap round instead of running into punctuation. And you have written a cipher in two small functions, checked it by decoding what it encoded, and broken it with a twenty-five pass loop.
That closes Module 3. You can hold data in lists, in dictionaries and in text. Module 4 is about what happens when a program meets something it did not expect, which is every program, every day.
Sources
- Python Software Foundation. (2026). Built-in types: Text sequence type str, and string methods. Python 3.14 documentation. docs.python.org
- Python Software Foundation. (2026). Built-in functions: ord and chr. docs.python.org
- Wikipedia contributors. (2026). Caesar cipher. Wikipedia. en.wikipedia.org
- Severance, C. (2016). Python for Everybody: Exploring data in Python 3, Chapter 6: Strings (slicing, methods and immutability). py4e.com
- Sweigart, A. (2019). Automate the Boring Stuff with Python (2nd ed.), Chapter 6: Manipulating strings. No Starch Press. automatetheboringstuff.com
- Key terms
- Immutable
- Unable to be changed once made. Strings are immutable, so every string method returns a new string.
- strip
- A method returning a copy of a string with the whitespace trimmed from both ends.
- split
- A method cutting a string into a list, at a given separator or at runs of whitespace.
- join
- A method called on the separator, joining a list of strings into one string.
- ord
- A built-in giving the number that stands for a character, such as 97 for lowercase a.
- chr
- A built-in giving the character for a number, the reverse of ord.
- Caesar cipher
- A substitution cipher that moves every letter a fixed number of places along the alphabet.
- Brute-force attack
- Breaking a cipher by trying every possible key, which works on a Caesar cipher because there are only twenty-five.
Module 4: When It Goes Wrong
Read a traceback properly, handle the failures you can predict, take a genuinely broken program apart by method, and write checks that catch your mistakes before anybody else does.
Reading a Traceback From the Bottom Up
- Read a multi-line traceback and name which line of your own code started the trouble.
- Recognise the common exception types and what each one means.
- Handle a predictable failure with try and except without hiding the ones you did not predict.
A program that prints the mean mark for each student in a dictionary. Two students work. The third stops it dead:
def mean(marks):
return sum(marks) / len(marks)
def report(name, marks):
print(f"{name}: mean {mean(marks):.1f}")
students = {"Ana": [62, 71], "Ben": [88], "Cleo": []}
for name in students:
report(name, students[name])
Here is everything it printed. The File lines have been shortened to just the file name; on your machine each one shows the full path to your file.
Ana: mean 66.5
Ben: mean 88.0
Traceback (most recent call last):
File "marks.py", line 12, in <module>
report(name, students[name])
~~~~~~^^^^^^^^^^^^^^^^^^^^^^
File "marks.py", line 6, in report
print(f"{name}: mean {mean(marks):.1f}")
~~~~^^^^^^^
File "marks.py", line 2, in mean
return sum(marks) / len(marks)
~~~~~~~~~~~^~~~~~~~~~~~
ZeroDivisionError: division by zero
That block of text is the single most useful thing Python ever prints at you, and most beginners skim it and start guessing instead. It is not a punishment and it is not noise. It is a map, and this lesson is about reading it.
Reading it from the bottom up
Start at the last line, always. ZeroDivisionError: division by zero is what went wrong, in five words. An exception is Python's way of saying it cannot carry out an instruction, and the name before the colon is its type.
Then move up to the frame just above it. File "marks.py", line 2, in mean, followed by the actual line, return sum(marks) / len(marks), with squiggles under the part that failed. So the division happened inside the function mean, on line 2. Something divided by len(marks) and len must have been zero.
Now keep going up, because the frames are a chain of who called whom, oldest first. The top frame, line 12, in <module>, is your own file at the outermost level, the for loop. It called report on line 6. Report called mean on line 2. Mean divided by zero. Read the whole thing upwards and it is a sentence: the loop asked for a report, the report asked for a mean, and the mean divided by nothing.
The words most recent call last on the first line are telling you this order. The bottom frame is where the explosion happened; the frames above it are how Python got there. In a long traceback most of the frames may be inside Python's own library, and then the frame you want is the lowest one naming a file you wrote.
And the two lines of correct output at the top matter too. Ana and Ben printed fine, so the failure depends on the data rather than on the code being broken everywhere. Cleo has an empty list. That is the whole diagnosis, and it took four sentences.
The errors you will actually meet
There are dozens of exception types. These eight cover almost everything a beginner hits. Each line below was raised on purpose and caught, so all eight could run in one program:
1 NameError - name 'undefined_name' is not defined
2 ValueError - invalid literal for int() with base 10: 'twenty'
3 TypeError - can only concatenate str (not "int") to str
4 IndexError - list index out of range
5 KeyError - 'b'
6 ZeroDivisionError - division by zero
7 FileNotFoundError - [Errno 2] No such file or directory: 'no_such_file.txt'
8 ValueError - list.remove(x): x not in list
| Exception | Means | Usually caused by |
|---|---|---|
| NameError | No such name exists | A typo, or using a variable before assigning it |
| ValueError | Right type, impossible value | int("twenty"), or removing something that is not in a list |
| TypeError | That operation does not exist for these types | Adding a str to an int, or using a value that turned out to be None |
| IndexError | No item at that position | An off-by-one, or an empty list |
| KeyError | No such key in the dictionary | A misspelled key, or data that did not contain what you assumed |
| ZeroDivisionError | Divided by zero | An average of an empty collection |
| FileNotFoundError | No file of that name here | A wrong name, or running from a different folder |
| IndentationError | The block is missing or misaligned | A forgotten indent, or tabs mixed with spaces |
Two of those entries are worth a second look. ValueError appears twice in the output, once from int and once from remove, because the two failures are the same kind of thing: the type was fine and the value was impossible. And notice what a KeyError prints: just the key, 'b', with no explanation. It is terse rather than unhelpful. The full list lives in the built-in exceptions documentation.
try and except: handling what you can predict
Some failures are not your fault and should not stop the program. A person typing the word fourteen when asked for an age is not a bug; it is Tuesday. The tool for this is try and except:
def ask_for_age():
"""Keep asking until the answer really is a whole number."""
while True:
answer = input("How old are you? ")
try:
age = int(answer)
except ValueError:
print(f"'{answer}' is not a whole number. Try again.")
else:
return age
age = ask_for_age()
print(f"Next year you will be {age + 1}.")
marks = [62, 71]
try:
print("Mean:", sum(marks) / len(marks))
print("Sixth mark:", marks[5])
except ZeroDivisionError:
print("There were no marks to average.")
except IndexError as problem:
print("There is no such mark:", problem)
finally:
print("This line runs whatever happened.")
Fed the word fourteen, then 14:
How old are you? 'fourteen' is not a whole number. Try again.
How old are you? Next year you will be 15.
Mean: 66.5
There is no such mark: list index out of range
This line runs whatever happened.
The shape is: put the risky line in the try block, and what to do about a particular failure in an except block naming the exception type. When the risky line succeeds, the except block is skipped entirely. When it fails, Python jumps straight to the matching except, skipping the rest of the try block.
Four details in that program.
Name the exception you expect. except ValueError catches a bad conversion and nothing else. If something genuinely unexpected happened inside that try block, it would still stop the program, which is what you want.
else runs when nothing went wrong. Putting return age in the else rather than inside the try makes the intention plain: try only the risky line, and do the rest when it worked.
You can have several excepts. Python tries them in order and uses the first that matches. Here the IndexError one fired, printed its message, and the ZeroDivisionError one was ignored.
finally always runs. Failure, success, even a return: the finally block happens. It is where you put tidying up, such as closing a file, which is exactly its job in Lesson 15.
except IndexError as problem captures the exception itself into a name, so you can print the message Python would have printed. Useful when you want to report a problem rather than just react to it.
Catching everything is worse than catching nothing
The most tempting misuse of try is to wrap everything and move on:
# The wrong way: catch everything and say nothing useful.
def mean_bad(marks):
try:
return sum(marks) / len(marks)
except:
return 0
# The right way: fix the case you understand, let the rest speak.
def mean_good(marks):
if len(marks) == 0:
return 0.0
return sum(marks) / len(marks)
print("mean_bad on [] :", mean_bad([]))
print("mean_bad on ['a', 'b'] :", mean_bad(["a", "b"]))
print("mean_good on [] :", mean_good([]))
print("mean_good on ['a','b']:", mean_good(["a", "b"]))
mean_bad on [] : 0
mean_bad on ['a', 'b'] : 0
mean_good on [] : 0.0
Traceback (most recent call last):
return sum(marks) / len(marks)
~~~^^^^^^^
TypeError: unsupported operand type(s) for +: 'int' and 'str'
Both functions gave 0 for an empty list, which is fine. Now look at the second line. Handed a list of words instead of numbers, mean_bad returned 0, cheerfully, as though that were an answer. A mark of 0 will now travel through the rest of your program, get averaged into a report and printed on a screen, and nothing anywhere will ever mention that the data was wrong.
mean_good crashed, and the crash is the good outcome. The traceback names the real problem, a string where a number should be, which is a bug somewhere else entirely that you now get to fix.
So: a bare except, or a try wrapped round a whole function, converts a loud bug into a silent one. Catch the specific failure you have thought about and can do something sensible about. Let everything else through. And when a condition can simply be checked, as with the empty list here, check it with an if rather than waiting for the exception.
Syntax errors are a different animal
Everything above is a runtime error: the file was valid Python and something went wrong while it ran. A SyntaxError, as in Lesson 1, means the file could not be read as Python at all, so nothing ran and there is no traceback of frames, only a pointer at the offending line. That distinction tells you where to look: if you got some output first, the program ran; if you got none at all, the file never started.
One quirk worth knowing because it wastes an hour the first time. Python often reports a syntax error on the line after the real mistake, because a missing closing bracket means it keeps reading, hoping. If line 20 looks perfect, check line 19.
Common misconceptions
- "The first line of the traceback is the important one." It is the least important. Read the last line for what went wrong, and the lowest frame naming your own file for where.
- "A traceback with four File lines means four errors." It means one error and four frames: who called whom, on the way to the thing that failed.
- "try/except fixes the error." It decides what happens instead. If the except block hides the problem rather than resolving it, you have swapped a crash for a wrong answer.
- "A bare except is a safe default." It catches everything, including typos in your own code, and turns them into silence. Name the exception you expect.
- "Errors mean the program is badly written." A program that raises ZeroDivisionError on an empty class list has told you something true about your data. The alternative was printing a meaningless number.
Recap
Key idea: Read a traceback from the bottom up: the last line says what, the lowest frame in your own file says where, and the frames above say how Python got there.
You can read a multi-frame traceback and turn it into a sentence. You know the eight exceptions that account for most beginner trouble and what each one is usually telling you. You can wrap a risky line in try and handle a named exception in except, use else for the success path and finally for the tidying up, and capture the exception with as when you want to report it. And you know the most important rule about exception handling, which is not to use it where an if would do, and never to catch what you have not thought about.
This lesson was about the errors that announce themselves. The next one is about the other kind: the program that runs perfectly, produces a number, and is wrong.
Sources
- Python Software Foundation. (2026). Errors and exceptions (tracebacks, try and except, else and finally). Python 3.14 tutorial. docs.python.org
- Python Software Foundation. (2026). Built-in exceptions (the exception hierarchy and what each type means). docs.python.org
- Severance, C. (2016). Python for Everybody: Exploring data in Python 3, Chapter 3: Conditional execution (try and except as a safety net around input conversion). py4e.com
- Sweigart, A. (2019). Automate the Boring Stuff with Python (2nd ed.), Chapter 3: Functions (exception handling). No Starch Press. automatetheboringstuff.com
- Key terms
- Exception
- Python's report that it cannot carry out an instruction. It stops the program unless something catches it.
- Traceback
- The list of frames printed when an exception is not caught, showing who called whom on the way to the failure.
- Frame
- One entry in a traceback, naming a file, a line number and the function that line was in.
- try
- A block holding the risky lines, which Python abandons at the first failure.
- except
- A block saying what to do about one named kind of exception.
- finally
- A block that runs whatever happened, used for tidying up such as closing a file.
- Bare except
- An except with no exception named, which catches everything including your own typos and hides them.
- Runtime error
- A failure while the program runs, as opposed to a SyntaxError, which stops the file being read at all.
Four Bugs in Twenty Lines, Found One at a Time
- State the expected answer before looking for a bug.
- Place trace prints that show a loop variable, the data it is using, and a running total.
- Fix one bug at a time, re-running after each change, and explain why fixing several at once is worse.
Here is a program that runs, prints four tidy lines, and is wrong in four separate ways:
names = ["Ana", "Ben", "Cleo", "Dev", "Eve"]
marks = [72, 50, 45, 91, 63]
def summary(names, marks):
passes = 0
total = 0
best_name = ""
best_mark = 0
for i in range(1, len(names)):
total = marks[i]
if marks[i] > 50:
passes = passes + 1
if marks[i] > best_mark:
best_name = names[i - 1]
best_mark = marks[i]
print("Students:", len(names))
print("Passes :", passes)
print("Average :", total / len(names))
print("Top :", best_name, best_mark)
summary(names, marks)
Students: 5
Passes : 2
Average : 12.6
Top : Cleo 91
No traceback. No red text. Four labelled lines that look exactly like the output of a program that works. Everything in Lesson 12 is useless here, because nothing went wrong as far as Python is concerned. This is a logic error, and finding it is a different skill with a different method.
Step zero: know the answer before you look for the bug
The method starts before any code is touched. Five marks: 72, 50, 45, 91, 63.
- Students: 5. That is the only line above that is right.
- Passes at 50 or more: Ana 72, Ben 50, Dev 91, Eve 63. That is 4.
- Total: 72 + 50 + 45 + 91 + 63 = 321. Average: 321 divided by 5 = 64.2.
- Top: Dev, 91.
Two minutes with a calculator, and now the four printed lines are not a mystery but a list of specific discrepancies. Without this step you are reading code hoping something looks wrong, which is the slowest debugging technique there is and the one everybody uses.
Notice which discrepancy is most useful: the average of 12.6. Not slightly off. Not off by one. Off by a factor of five, and smaller than every single mark, which is impossible for an average. The wildest wrong number is the best place to start, because a big error usually has a simple cause.
The method, in four words
Print. Reason. Isolate. Fix. And then run it again, after every single change.
- Print the values the suspect line depends on, inside the loop, where they change.
- Reason about what you see against what you predicted. The moment the trace stops matching your expectation is the moment of the bug.
- Isolate: narrow it to one line before changing anything.
- Fix that one thing, run, and confirm the number you predicted.
The hardest of these to obey is the last. When you have spotted three suspicious lines the urge is to correct all three and run once. Do not. If the result is still wrong you now have no idea which of your three changes helped, which did nothing, and which broke something new.
Bug one: an average smaller than every mark
The average uses total, so print total where it changes. One line, inside the loop:
total = marks[i]
print(f" TRACE i={i} name={names[i]} mark={marks[i]} total={total}")
TRACE i=1 name=Ben mark=50 total=50
TRACE i=2 name=Cleo mark=45 total=45
TRACE i=3 name=Dev mark=91 total=91
TRACE i=4 name=Eve mark=63 total=63
Students: 5
Passes : 2
Average : 12.6
Top : Cleo 91
There it is, in the fourth column. total is 50, then 45, then 91, then 63. It never grows. A running total that does not run is not accumulating, it is replacing, and the line total = marks[i] says exactly that. It should be total = total + marks[i].
That is the accumulator from Lesson 6 written wrong, and it is worth seeing why the output was 12.6 rather than obviously broken: the last mark was 63, and 63 divided by 5 is 12.6. A wrong number that came out of real arithmetic on real data always looks more plausible than it deserves to.
Note the trace format. The pass number, the data for that pass, and the accumulator. Prefixing with TRACE and some spaces makes the debugging output easy to see and easy to delete afterwards.
Bug two: four lines of trace for five students
Fix the accumulator and run again, and the trace has already handed you the next bug for free:
TRACE i=0 name=Ana mark=72 total=72
TRACE i=1 name=Ben mark=50 total=122
TRACE i=2 name=Cleo mark=45 total=167
TRACE i=3 name=Dev mark=91 total=258
TRACE i=4 name=Eve mark=63 total=321
Students: 5
Passes : 3
Average : 64.2
Top : Cleo 91
Wait. This trace has five lines and starts at i=0, because fixing the accumulator meant reading the loop line properly, and range(1, len(names)) was starting at 1. Ana was never in the loop at all. In the previous trace there were four lines for five students, which was visible the whole time if anyone had counted them.
That is worth stating plainly: count the lines of your trace. Five students should produce five lines. This is the off-by-one error from Lesson 6, and the trace makes it a counting problem rather than a reasoning problem.
With both fixed, the average is 64.2, which matches the prediction exactly. One line down, three to go, and the two remaining are still visible.
Bug three: a pass mark that does not pass
Passes now says 3 and should say 4. Which student is missing? The condition is marks[i] > 50, and Ben got exactly 50. A mark of 50 is a pass, and > excludes it.
This one needs no trace at all, because the discrepancy is exactly one and there is exactly one mark sitting on the boundary. When a count is out by one, look at the boundary case first. The fix is >=, and it is the reason Lesson 8 insisted on testing grade_for at 90, 89 and 59 rather than at 75.
Notice also that the earlier version reported 2 passes rather than 3 for a different reason as well: Ana was being skipped. Two bugs were contributing to one wrong number, which is normal and is the strongest argument for fixing them one at a time. Had you changed the comparison first, passes would have gone from 2 to 3, still wrong, and you might have concluded the comparison was not the problem.
Bug four: the right mark with somebody else's name
Top says Cleo 91. The mark is right and the name is wrong, which narrows it before any printing: the two are assigned on adjacent lines, and only one of them can be at fault. Print inside the if, where the best is replaced:
if marks[i] > best_mark:
best_name = names[i - 1]
best_mark = marks[i]
print(f" TRACE new best at i={i}: names[i]={names[i]} but stored {best_name}")
TRACE new best at i=0: names[i]=Ana but stored Eve
TRACE new best at i=3: names[i]=Dev but stored Cleo
Students: 5
Passes : 4
Average : 64.2
Top : Cleo 91
Two lines, and both are absurd. At i=3 the best mark is Dev's, and the program stored Cleo, the student before. names[i - 1] is fetching the previous name. Someone wrote the minus one out of a half-remembered idea that lists are off by one, which they are not.
The first trace line is the better lesson, though. At i=0, names[i - 1] is names[-1], and a negative index does not fail. It quietly returns the last item, Eve. A list that raises IndexError when you go one past the end says nothing at all when you go one before the start, which makes minus one a genuinely dangerous typo.
The repaired program, and the test that would have caught all of it
def summary(names, marks):
"""Report how many passed, the mean mark, and who scored highest."""
passes = 0
total = 0
best_name = ""
best_mark = 0
for name, mark in zip(names, marks):
total = total + mark
if mark >= 50:
passes = passes + 1
if mark > best_mark:
best_name = name
best_mark = mark
print("Students:", len(names))
print("Passes :", passes)
print("Average :", total / len(names))
print("Top :", best_name, best_mark)
summary(names, marks)
print()
print("One student:")
summary(["Solo"], [50])
Students: 5
Passes : 4
Average : 64.2
Top : Dev 91
One student:
Students: 1
Passes : 1
Average : 50.0
Top : Solo 50
All four numbers now match the predictions made before any debugging started. And the loop has changed shape: zip(names, marks) hands out the name and the mark together, pass by pass, so there is no index anywhere. Two of the four bugs lived in that index. Removing the index removes the whole category, which is a better kind of fix than correcting the arithmetic on it.
The second call is the other habit worth stealing. A class of one, with a mark sitting exactly on the boundary, exercises the range start, the accumulator, the comparison and the name lookup in four lines of output you can check by eye. That is a test, written as a print, and it is what the next lesson turns into something that checks itself.
Three habits that prevent the next four bugs
Explain it out loud to something that cannot help you. This has a name, rubber duck debugging, and the joke conceals something real: saying "then I add the mark to the total" while looking at a line that does not add anything is how a surprising number of bugs are found. Reading silently lets you see what you meant. Speaking makes you say what is there.
Change one thing, then run. Every time. If you have three ideas, try them one at a time and keep the ones that help.
Keep a copy that worked. Before a big change, save the file under another name. When the new version is worse and you can no longer remember what you altered, you have somewhere to go back to. Professionals do this with version control; a copy called summary_working.py will do until then.
When trace printing is not enough there is a step up: pdb, Python's built-in debugger, which stops a program mid-run and lets you inspect every variable. It is worth meeting eventually. It is not worth meeting yet, because print statements find almost everything and cost nothing to learn.
Common misconceptions
- "Debugging means reading the code until the mistake looks wrong." Rereading shows you what you intended, which is precisely the thing that is not there. Printing shows you what the machine actually has.
- "If it runs without an error, the hard part is over." A logic error produces confident, well-formatted, incorrect output. It is more expensive than a crash, because nothing announces it.
- "Fix everything you can see, then run once." Then a still-wrong result tells you nothing about which change did what. One change, one run.
- "A negative index will raise an error like a too-large one." It will not.
names[-1]is the last item, so an accidental minus one silently returns the wrong thing. - "Trace prints clutter the program." They are temporary. Prefix them, use them, delete them. The permanent version of a trace print is a test, which is the next lesson.
What to carry forward
The point: Work out the right answer first, then print, reason, isolate and fix one thing at a time, re-running after every change.
You have taken a program with four independent bugs apart without once guessing. You know to start from the wildest wrong number, because a large error usually has a simple cause; to print the loop variable, the data and the accumulator together; to count the lines of a trace against the number of items; and to look at the boundary case whenever a count is out by one. You know that two bugs can contribute to one wrong number, which is why fixing them singly matters. You know that a negative index fails silently. And you have seen a structural fix, replacing an index with zip, that removed a whole category of mistake rather than correcting one instance of it.
All the checking in this lesson was done by eye. That works for four lines and stops working somewhere around forty. The next lesson makes the program check itself.
Sources
- Sweigart, A. (2019). Automate the Boring Stuff with Python (2nd ed.), Chapter 11: Debugging. No Starch Press. automatetheboringstuff.com
- Python Software Foundation. (2026). pdb: The Python debugger. Python 3.14 documentation. docs.python.org
- Wikipedia contributors. (2026). Rubber duck debugging. Wikipedia. en.wikipedia.org
- Zeller, A. (2009). Why Programs Fail: A Guide to Systematic Debugging (2nd ed.). Morgan Kaufmann. (On isolating a defect by narrowing the difference between a working and a failing run.)
- Key terms
- Logic error
- A program that runs without complaint and produces a wrong answer, so no message points at it.
- Trace print
- A temporary print inside a loop showing the loop variable, the data and the accumulator, used to find where reality parts from expectation.
- Isolate
- To narrow a fault to one line before changing anything, so the fix can be judged.
- Boundary case
- A value sitting exactly on a threshold, such as a mark of 50 for a pass at 50, where < and <= disagree.
- zip
- A built-in that hands out items from two lists together, removing the index and the bugs that live on it.
- Rubber duck debugging
- Explaining the code aloud line by line, which forces you to say what is there rather than what you meant.
- pdb
- Python's built-in debugger, which stops a program mid-run so variables can be inspected.
Five Pence Printed as 0.5, and Nobody Noticed for a Week
- Write an assert that states what a function must produce for a given input.
- Build a small test function that reports every failure instead of stopping at the first.
- Choose test cases at the boundaries, and change a function's insides without breaking its behaviour.
A function from Lesson 8, formatting whole pence as pounds. It has been used in three programs and looked fine in all of them:
def pounds(pence):
"""Format a whole number of pence as pounds and pence."""
return f"{pence // 100}.{pence % 100}"
print("pounds(899):", pounds(899))
print("pounds(250):", pounds(250))
print("pounds(5) :", pounds(5))
pounds(899): 8.99
pounds(250): 2.50
pounds(5) : 0.5
Five pence came out as 0.5, which is fifty pence. The bug was always there and nobody hit it, because the test data in every program that used this function happened to have two digits of pence. It took a five pence item to expose it, and by then the function was in three places.
The Lesson 13 method would find this in about a minute, once somebody noticed. The point of this lesson is to make the program notice, every time it runs, without anybody being clever.
assert: a sentence the program refuses to be wrong about
The assert statement is the smallest testing tool there is. You write the word assert and something that ought to be true. In a normal run, a true condition passes quietly and a false one raises AssertionError. If no handler catches that exception, execution stops. Python omits assert statements with optimization enabled by -O, so use ordinary checks for required input validation.
assert pounds(899) == "8.99"
assert pounds(250) == "2.50"
print("The first two checks passed.")
The first two checks passed.
Silence is success. That takes some getting used to, because every other line you have written produces something. An assert is a claim about the program, sitting in the program, and a claim that is holding has nothing to say.
Now the one that fails, with a message attached after a comma:
assert pounds(899) == "8.99"
assert pounds(5) == "0.05", f"five pence formatted as {pounds(5)}"
print("never reached")
Traceback (most recent call last):
assert pounds(5) == "0.05", f"five pence formatted as {pounds(5)}"
^^^^^^^^^^^^^^^^^^^
AssertionError: five pence formatted as 0.5
An AssertionError, the failing line quoted with carets under the claim that broke, and your message saying what actually happened. Notice that the print at the end never ran: this failing assert has no exception handler, so execution stops. Notice also how much better the message is than a bare failure would be. AssertionError on its own tells you which line; the message tells you the wrong value, which is usually the thing you wanted to know next.
For nonnegative whole pence, add :02 to the remainder field: return f"{pence // 100}.{pence % 100:02}". This pads 5 to 05. Replace the return line before running the next suite.
From asserts to a test function that keeps going
A sequence of unhandled asserts stops at the first failure. To see several failures in one run, use a small test function that reports each result and carries on:
def check(description, got, expected):
"""One test. Prints a line either way and counts the failures."""
if got == expected:
print(f" ok {description}")
return 0
print(f" FAIL {description}: got {got!r}, expected {expected!r}")
return 1
It takes what the function produced and what it should have produced, prints a line either way, and returns 1 for a failure so the caller can add them up. The !r inside the f-string asks for the value's repr representation; for the strings here, that includes quotation marks, which is the difference between a confusing report saying got 0.5, expected 0.05 and a clear one saying got '0.5', expected '0.05'.
Now a function under test with a second, different bug in it, and a suite of checks aimed at its boundaries. These postage bands are invented for the exercise, not postal-service prices or rules:
def band(grams):
"""Postage band for a parcel weight in grams."""
if grams <= 0:
return "Not a parcel"
if grams < 100:
return "Letter"
if grams < 500:
return "Large letter"
if grams < 2000:
return "Small parcel"
return "Too heavy"
def run_tests():
failures = 0
print("pounds:")
failures += check("899 pence", pounds(899), "8.99")
failures += check("250 pence", pounds(250), "2.50")
failures += check("5 pence", pounds(5), "0.05")
failures += check("0 pence", pounds(0), "0.00")
failures += check("100 pence", pounds(100), "1.00")
print("band:")
failures += check("0 g", band(0), "Not a parcel")
failures += check("1 g", band(1), "Letter")
failures += check("99 g", band(99), "Letter")
failures += check("100 g", band(100), "Letter")
failures += check("101 g", band(101), "Large letter")
failures += check("500 g", band(500), "Large letter")
failures += check("2000 g", band(2000), "Small parcel")
failures += check("2001 g", band(2001), "Too heavy")
if failures == 0:
print("All tests passed.")
else:
print(f"{failures} test(s) failed.")
run_tests()
pounds:
ok 899 pence
ok 250 pence
ok 5 pence
ok 0 pence
ok 100 pence
band:
ok 0 g
ok 1 g
ok 99 g
FAIL 100 g: got 'Large letter', expected 'Letter'
ok 101 g
FAIL 500 g: got 'Small parcel', expected 'Large letter'
FAIL 2000 g: got 'Too heavy', expected 'Small parcel'
ok 2001 g
3 test(s) failed.
failures += check(...) is shorthand for failures = failures + check(...), which Python allows for every arithmetic operator and which is worth adopting now.
Three failures, one cause
Read the three failing lines together. 100 g should be a Letter and came out a Large letter. 500 g should be a Large letter and came out a Small parcel. 2000 g should be a Small parcel and came out Too heavy. In every case the boundary value fell into the band above where it belonged.
One cause: the conditions use < where they should use <=. The bands are meant to include their top weight, and a strict less-than pushes each boundary up a band. One character, in three places, and the tests named all three in one run.
Now look at the eight band tests and notice their shape. Not 250 g and 800 g and 1500 g, which are comfortable middle-of-the-band values that would all have passed. Instead: 99, 100, 101 around the first boundary, and the exact boundary at each of the others. Test where the answer changes. Boundary mistakes are easy to miss when you test only middle values. Ordinary cases can fail too. Every test above is either a boundary, one either side of a boundary, or an impossible input.
With <= in all three places, every line reads ok and the suite prints All tests passed.
The real payoff: changing the code without fear
Suppose you decide the chain of ifs is clumsy and want to rewrite band as a table:
def band(grams):
"""Same behaviour, written as a table instead of a chain of ifs."""
if grams <= 0:
return "Not a parcel"
limits = [(100, "Letter"), (500, "Large letter"), (2000, "Small parcel")]
for limit, name in limits:
if grams <= limit:
return name
return "Too heavy"
After replacing band and rerunning run_tests(), the final line is:
All tests passed.
Completely different insides, same eight tests, zero failures. Rewriting working code to make it clearer is called refactoring, and without tests it is gambling: you change something, it looks right, and you find out in three weeks. One run checks the cases in the suite; passing them does not prove every possible input works.
This is also where the payoff turns up in your own projects. The moment a program is big enough that you are nervous about touching part of it, tests can make changes easier to check. Ten lines of check calls buy the confidence back.
What to test, and what to leave alone
Not everything is worth a test, and a beginner who tries to test everything gives up by Thursday. A short list that earns its keep:
- Every boundary. The exact threshold, and one either side.
- The empty and the impossible. An empty list, a zero, a negative, a word where a number was expected.
- The case you got wrong once. When a bug turns up, write the test before the fix. The test can catch that same failure when the suite runs again.
- Anything that returns a value and has no input or output of its own. Those functions are easy to test, which is another reason Lesson 8 preferred return over print.
Start with functions that return values, where inputs and expected results are easy to compare. Functions that print and short wrappers can matter too: printed output can be captured for a test. If a function is hard to reach or test, consider separating its calculation from input and output.
Professional Python has proper tools for this, chiefly unittest, which ships with Python, and pytest, which does not. They do the counting and reporting for you and run every test file in a project with one command. They are worth learning when you have a project big enough to need them. The small check function above is the same idea with the scaffolding removed, and it will carry you a long way.
Common misconceptions
- "A passing assert should print something." Silence is success. An assert only speaks when the claim fails.
- "Tests prove the program is correct." They prove it behaves correctly on the cases you thought of. The bug at the top of this lesson survived three programs because nobody thought of five pence.
- "Test the normal cases first." Check ordinary cases and boundaries. All three band failures in this example were at boundaries.
- "Tests are extra work at the end." They are the trace prints from Lesson 13, made permanent. You were going to check those values anyway; this way the check stays.
- "If I rewrite a function, the tests have to be rewritten too." Only if you change what it does. Tests describe behaviour, which is why they survive a change of insides and catch you when you break one by accident.
The short version
What matters here: A test is a claim about what a function must produce, written down where it can be re-checked. Include boundary cases alongside ordinary inputs, since either can reveal failures.
You can write an assert and read the AssertionError it produces, with a message that names the wrong value. You can build a small check function that reports every failure rather than stopping at the first, and a run_tests function that counts them. You know to choose cases at the thresholds, one either side, plus the empty and impossible ones, and you have seen a suite point at three failures with one cause. You have refactored a function and confirmed that its results were unchanged on the tested cases. And you know that unittest and pytest exist for when a project outgrows this.
That closes Module 4. You can find errors that announce themselves, errors that stay quiet, and errors before they reach anybody else. Module 5 gives your programs something they have never had: data that is still there tomorrow.
Sources
- Python Software Foundation. (2026). Simple statements: The assert statement. Python 3.14 language reference. docs.python.org
- Python Software Foundation. (2026). unittest: Unit testing framework. Python 3.14 documentation. docs.python.org
- Sweigart, A. (2019). Automate the Boring Stuff with Python (2nd ed.), Chapter 11: Debugging (raising exceptions and using assertions). No Starch Press. automatetheboringstuff.com
- Wikipedia contributors. (2026). Software testing. Wikipedia. en.wikipedia.org
- Key terms
- assert
- In a normal run, a statement that raises AssertionError when its condition is false. An unhandled failure stops execution; -O omits the check.
- AssertionError
- The exception raised by a failing assert, carrying whatever message was written after the comma.
- Test case
- One input paired with the output the function must produce for it.
- Test function
- A function that runs many checks, reports each one, and counts the failures instead of stopping at the first.
- Boundary test
- A test at the exact value where the answer changes, plus one on each side of it.
- Refactoring
- Rewriting working code to make it clearer without changing what it does. Tests help check the behavior on selected cases.
- unittest
- The testing framework included with Python, which handles running and reporting for larger projects.
Module 5: Data That Is Still There Tomorrow
Write to a file and read it back, take a real CSV of data apart and summarise it, and draw a chart from it using nothing but the standard library.
The First Program That Remembers Yesterday
- Open a file for reading, writing or appending, and say what each mode does to what is already there.
- Read a file whole or a line at a time, and strip the newline that comes with each line.
- Use with so the file closes itself, and handle a file that is not there.
Every program in this course so far has forgotten everything the moment it finished. Here is the first one that does not:
notes = open("revision.txt", "w")
notes.write("Electricity: 45 minutes\n")
notes.write("Waves: 30 minutes\n")
notes.write("Forces: 60 minutes\n")
notes.close()
print("Written. Now reading it back:")
notes = open("revision.txt", "r")
contents = notes.read()
notes.close()
print(contents)
print("That string is", len(contents), "characters long.")
print(repr(contents[:30]))
Written. Now reading it back:
Electricity: 45 minutes
Waves: 30 minutes
Forces: 60 minutes
That string is 61 characters long.
'Electricity: 45 minutes\nWaves:'
A file called revision.txt now exists on the disk. Open it in any text editor and the three lines are there. Turn the computer off and back on and they are still there. That is the whole idea of this lesson, and it is the difference between a program and a tool.
open, the mode letter, and close
The built-in open takes a file name and a mode, and hands back a file object you can read from or write to.
| Mode | Means | If the file exists | If it does not |
|---|---|---|---|
| "r" | Read | Opens it for reading | FileNotFoundError |
| "w" | Write | Empties it completely, then writes | Creates it |
| "a" | Append | Adds to the end | Creates it |
"r" is what you get if you leave the mode out, so open("revision.txt") means read. "w" is the dangerous one. It does not ask and it does not warn: the moment the file is opened for writing, whatever was in it is gone. There is no undo and no recycle bin. More than one person has emptied a file of real work by running a program with the wrong name in it, which is a good reason to practise on files called test.txt.
close matters more than it looks. Writing is buffered, which means Python collects what you write and only actually puts it on the disk in batches. Until the file is closed, some of your text may still be in memory, and a program that crashes before closing can leave a file half written. Closing flushes the lot and releases the file.
The newline you cannot see
Look again at the last line of that output: 'Electricity: 45 minutes\nWaves:'. There is no line break in the file's contents in the way there is on the screen. There is a character, written \n, called a newline, and printing it is what moves the cursor down.
That is why every write above ends with \n. Unlike print, write adds nothing of its own; leave the newline out and all three topics run together on one line. And it is why repr is so useful with files: printed normally the newline is invisible, and printed with repr you can see exactly what is there.
It also explains the blank line in the output. The file ends with a newline, and print adds another one of its own.
with: the version that closes the file for you
Writing close every time is a promise you will eventually break, usually on the path where something went wrong. Python has a construction that closes the file for you, no matter how the block ends:
with open("revision.txt") as notes:
for line in notes:
print(repr(line))
'Electricity: 45 minutes\n'
'Waves: 30 minutes\n'
'Forces: 60 minutes\n'
with open(...) as name: opens the file, gives it the name, runs the indented block, and closes the file when the block ends. Even if an exception is raised inside the block, the file is closed on the way out; it is the finally behaviour from Lesson 12, built in. From here on, use with for every file, always.
And notice what a file gives a for loop: one line at a time, in order, each one still carrying its newline. A file behaves like a list of lines for the purposes of iterating over it, without ever loading the whole thing into memory, which is how a program reads a file larger than the computer's memory.
Pulling values out of a line
Reading a file is only half the job. The other half is turning a line of text into values, and that is Lesson 11's string methods doing the work:
total = 0
with open("revision.txt") as notes:
for line in notes:
line = line.strip()
if line == "":
continue
topic, rest = line.split(": ")
minutes = int(rest.split()[0])
total = total + minutes
print(f"{topic:12} {minutes:3} min")
print("Total:", total, "minutes")
with open("revision.txt") as notes:
lines = notes.readlines()
print("readlines gives a list of", len(lines), "items")
Electricity 45 min
Waves 30 min
Forces 60 min
Total: 135 minutes
readlines gives a list of 3 items
Four habits in that loop, all of which you will use every time you read a file.
Strip first. line.strip() removes the trailing newline along with any stray spaces. Forget it and int("45\n") still works, by luck, while topic == "Forces" quietly fails because the topic is really "Forces" followed by a newline.
Skip blank lines. Real files end with a newline, and files edited by people pick up empty lines. One continue saves a crash.
Split into names. topic, rest = line.split(": ") splits into a two-item list and unpacks it into two names in one line, which only works when there really are two pieces.
Convert what should be a number. Everything from a file is text, exactly as everything from input is text.
readlines is the third way to read: the whole file as a list of lines at once. Convenient for a small file, wasteful for a large one, and the loop above is what you should normally write.
w destroys, a adds
with open("log.txt", "w") as f:
f.write("first run\n")
with open("log.txt", "w") as f:
f.write("second run\n")
with open("log.txt") as f:
print("after two w opens:", repr(f.read()))
with open("log.txt", "a") as f:
f.write("third run\n")
with open("log.txt", "a") as f:
f.write("fourth run\n")
with open("log.txt") as f:
print("after two a opens:", repr(f.read()))
try:
with open("no_such_file.txt") as f:
print(f.read())
except FileNotFoundError as problem:
print("Caught:", problem)
after two w opens: 'second run\n'
after two a opens: 'second run\nthird run\nfourth run\n'
Caught: [Errno 2] No such file or directory: 'no_such_file.txt'
Two opens in "w" mode left one line: the first run was destroyed by the second open before a single character was written. Two opens in "a" mode left everything. If your program is keeping a log, a score table or a diary, "a" is what you want. If it is saving the current state of something, "w" is right, because you want the new version rather than a pile of old ones.
The last three lines show the other half of file handling. A missing file raises FileNotFoundError, which is an exception like any other and can be caught. That matters most the first time a program runs, when the file it wants to load has not been created yet.
A high score table that survives being closed
Here is all of it together. Load if there is something to load, add, save, and the file is the memory between runs.
FILENAME = "scores.txt"
def load_scores():
"""Read the table, or start an empty one if the file is not there yet."""
scores = []
try:
with open(FILENAME) as f:
for line in f:
line = line.strip()
if line == "":
continue
points, name = line.split(",")
scores.append((int(points), name))
except FileNotFoundError:
print("No score file yet. Starting a new one.")
return scores
def save_scores(scores):
"""Write the table back, highest first, keeping only the top five."""
scores.sort(reverse=True)
with open(FILENAME, "w") as f:
for points, name in scores[:5]:
f.write(f"{points},{name}\n")
def show(scores):
if len(scores) == 0:
print(" (no scores yet)")
return
position = 1
for points, name in scores:
print(f" {position}. {name:8} {points:6}")
position = position + 1
scores = load_scores()
print("Loaded table:")
show(scores)
scores.append((4200, "Priya"))
scores.append((3100, "Marcus"))
scores.append((5600, "Ana"))
save_scores(scores)
print("After adding three players:")
show(load_scores())
The first time it runs, with no file on the disk:
No score file yet. Starting a new one.
Loaded table:
(no scores yet)
After adding three players:
1. Ana 5600
2. Priya 4200
3. Marcus 3100
And the second time, with nothing changed in the program:
Loaded table:
1. Ana 5600
2. Priya 4200
3. Marcus 3100
After adding three players:
1. Ana 5600
2. Ana 5600
3. Priya 4200
4. Priya 4200
5. Marcus 3100
The second run loaded the table the first one left behind. That is the whole point, and the duplicated names are the proof: the program added the same three players again on top of the three it remembered. In a real game the adding would come from actual play, not from three fixed lines.
Two design details. Each score is stored as points first and name second, (5600, "Ana"), because Python sorts a list of tuples by the first item and there is nothing more to write. And the whole table is rewritten with "w" rather than appended, because the file holds the current top five rather than a history; scores[:5] throws the rest away.
Where the file actually goes
A bare name like "scores.txt" means: in whatever folder the program was run from. Not the folder the .py file is in, if those differ, which surprises people using an editor that runs from somewhere else. If the file appears somewhere you did not expect, that is why.
You can give a fuller path, and on Windows there is a trap: open("C:\Users\me\notes.txt") contains \U and \n, which Python reads as escape characters, not as folder separators. The simplest fix is forward slashes, which Python accepts on Windows too: open("C:/Users/me/notes.txt"). For anything more serious there is the pathlib module, which is worth meeting after this course.
Common misconceptions
- "Opening a file in w mode is safe until you write." The file is emptied at the moment it is opened, before any write happens. If the name was wrong, the damage is already done.
- "write adds a new line like print does." It adds nothing. Without an explicit newline character, everything you write lands on one line.
- "A line read from a file is clean." It carries its newline, so comparisons fail in ways that are invisible on screen. Strip it, and use repr when something inexplicable is happening.
- "A number in a file is a number." It is text, exactly as with input. Convert it with int or float before doing arithmetic.
- "with is just a tidier way to write open." It also guarantees the close, including when an exception is raised inside the block, which a plain close on the last line does not.
What you now know
Why this matters: A file is the only memory a program has between runs, and the two things to be careful about are the mode letter, which can destroy what is there, and the newline, which is invisible until it breaks a comparison.
You can open a file for reading, writing or appending and say what each mode does to what is already on the disk. You can read a whole file with read, a list of lines with readlines, or a line at a time with a for loop, and you know why the loop is usually right. You can strip a line, skip the blank ones, split it into parts and convert the numbers. You can write lines, remembering the newline, and you know that "w" empties and "a" adds. You use with for every file, so it closes itself. And you can handle the file that is not there yet, which is the ordinary case on a program's first run.
revision.txt had three lines in a shape you invented. The next lesson works with a file in a shape the whole world already agrees on, holding real numbers, from a real source.
Sources
- Python Software Foundation. (2026). Input and output: Reading and writing files (modes, with, read, readlines). Python 3.14 tutorial. docs.python.org
- Python Software Foundation. (2026). Built-in functions: open (the mode argument and what each letter does). docs.python.org
- Severance, C. (2016). Python for Everybody: Exploring data in Python 3, Chapter 7: Files. py4e.com
- Sweigart, A. (2019). Automate the Boring Stuff with Python (2nd ed.), Chapter 9: Reading and writing files. No Starch Press. automatetheboringstuff.com
- Key terms
- open
- The built-in that connects a program to a file on disk and hands back a file object.
- Mode
- The letter given to open: r to read, w to write from empty, a to add to the end.
- Newline
- The invisible character that ends a line in a text file, written as a backslash followed by n.
- with
- A block that closes the file when it ends, including when an exception is raised inside it.
- readlines
- A method returning the whole file as a list of lines, convenient for small files and wasteful for large ones.
- Append mode
- Opening with a, so new text is added after what is already there rather than replacing it.
- FileNotFoundError
- The exception raised when a file opened for reading does not exist, which is the normal case on a first run.
- Working directory
- The folder a program was run from, where a bare file name is looked for and created.
Eleven Real Numbers From NASA, and the Bug They Found
- Read a CSV file with a header row and turn its text into numbers.
- Skip blank, short and unreadable rows without losing the pairing between columns.
- Summarise a column of numbers and group it with a dictionary, and say what the numbers do not show.
NASA publishes a file called GLB.Ts+dSST.csv. It holds the global mean surface temperature anomaly for every year since 1880, which is how much warmer or cooler that year was than the average of the years 1951 to 1980. Here are eleven rows of it, at five-year intervals, written into a small file a first program can chew on:
year,anomaly_c
1975,-0.01
1980,0.26
1985,0.12
1990,0.45
1995,0.45
2000,0.39
2005,0.68
2010,0.72
2015,0.9
2020,1.01
2025,1.19
That is a CSV file, comma-separated values, and it is the format that most of the world's public data actually arrives in. A first line naming the columns, then one row per record, values separated by commas. It is a plain text file: open it in a text editor and that is exactly what you see. A spreadsheet will open it too, and that is what a spreadsheet mostly is.
These are real numbers from a real source. The value for 2025 is 1.19, which means 2025 averaged 1.19 degrees Celsius warmer than the 1951 to 1980 mean. By the end of this lesson your program will have read them, summarised them, grouped them by decade, and found a bug in itself.
Making the file, with the source written into it
# Build temps.csv from NASA GISTEMP v4, Land-Ocean Global Means, J-D column.
# Source: https://data.giss.nasa.gov/gistemp/tabledata_v4/GLB.Ts+dSST.csv
# Anomaly in degrees Celsius against the 1951-1980 mean.
rows = [
(1975, -0.01), (1980, 0.26), (1985, 0.12), (1990, 0.45),
(1995, 0.45), (2000, 0.39), (2005, 0.68), (2010, 0.72),
(2015, 0.90), (2020, 1.01), (2025, 1.19),
]
with open("temps.csv", "w") as f:
f.write("year,anomaly_c\n")
for year, anomaly in rows:
f.write(f"{year},{anomaly}\n")
print("Wrote temps.csv with", len(rows), "rows.")
Wrote temps.csv with 11 rows.
Those three comment lines are the most important part of the file and they are not code. A data file with no record of where it came from is worthless six months later, because you cannot check it, cannot update it and cannot answer anybody who asks. Write the source in. Every time.
One thing to notice in the written file: 0.90 came out as 0.9. Python stores the number, not the characters you typed, and 0.90 and 0.9 are the same number. The trailing zero was never there.
The three-line reader that fails on line one
The obvious way to read it, using only Lesson 15:
total = 0.0
count = 0
with open("temps.csv") as f:
for line in f:
year, anomaly = line.strip().split(",")
total = total + float(anomaly)
count = count + 1
print("Mean anomaly:", total / count)
Traceback (most recent call last):
total = total + float(anomaly)
~~~~~^^^^^^^^^
ValueError: could not convert string to float: 'anomaly_c'
The first line of a CSV is the header, and it is text. Every CSV reader you ever write has to deal with it, and the tidiest way is to read that one line before the loop starts, with f.readline(), which takes exactly one line and leaves the loop to handle the rest.
A loader that survives a real file
def load_csv(filename):
"""Read a two-column CSV of year and anomaly. Skips the header and blanks."""
years = []
values = []
with open(filename) as f:
header = f.readline().strip()
print("header:", header)
for line in f:
line = line.strip()
if line == "":
continue
parts = line.split(",")
if len(parts) != 2:
print("skipping odd line:", repr(line))
continue
try:
years.append(int(parts[0]))
values.append(float(parts[1]))
except ValueError:
print("skipping unreadable line:", repr(line))
return years, values
Four defences, and each one exists because real files are not tidy. The header is taken off the front. Blank lines are skipped, and every file that has ever been opened in a spreadsheet and saved again has at least one. A row with the wrong number of commas is reported and skipped rather than crashing the program. And a value that will not convert is caught with try, because a data file will eventually contain the word unknown, or a dash, or nothing at all between two commas.
Reporting what it skipped matters as much as skipping it. A loader that silently discards rows will one day discard half your data and tell you the mean of what was left.
The questions worth asking of any column of numbers
def summarise(years, values):
print("rows :", len(values))
print("first year:", years[0], "at", values[0])
print("last year :", years[-1], "at", values[-1])
print("lowest :", min(values), "in", years[values.index(min(values))])
print("highest :", max(values), "in", years[values.index(max(values))])
print("mean :", f"{sum(values) / len(values):.3f}")
print("change :", f"{values[-1] - values[0]:+.2f}")
rising = 0
for i in range(1, len(values)):
if values[i] > values[i - 1]:
rising = rising + 1
print("steps up :", rising, "out of", len(values) - 1)
header: year,anomaly_c
rows : 11
first year: 1975 at -0.01
last year : 2025 at 1.19
lowest : -0.01 in 1975
highest : 1.19 in 2025
mean : 0.560
change : +1.20
steps up : 7 out of 10
Count, ends, extremes, middle, change, direction. Those six questions are worth asking of any column of numbers you meet, and they take nine lines.
years[values.index(max(values))] is the parallel-list trick from Lesson 10: find where the largest value sits, then read the year at the same position. It works, and it is precisely the pattern that broke in Lesson 10, which is a warning we are about to collect on.
The + in {...:+.2f} forces a sign to be shown, so a rise prints as +1.20 rather than 1.20. For a column of changes, where some are negative, that keeps the meaning obvious.
And note the last measure. Seven steps up out of ten is not ten out of ten: 1985 was cooler than 1980 and 2000 cooler than 1995. A rising series is not a series that rises every time, and a program that only reported first and last would have hidden that.
Grouping with a dictionary
To go from a list of readings to a summary by decade, use the counting pattern from Lesson 10 with two dictionaries: one for the totals, one for how many went into each:
decade_totals = {}
decade_counts = {}
for year, value in zip(years, values):
decade = year // 10 * 10
decade_totals[decade] = decade_totals.get(decade, 0.0) + value
decade_counts[decade] = decade_counts.get(decade, 0) + 1
for decade in sorted(decade_totals):
mean = decade_totals[decade] / decade_counts[decade]
print(f" {decade}s: {mean:+.3f} from {decade_counts[decade]} reading(s)")
1970s: -0.010 from 1 reading(s)
1980s: +0.190 from 2 reading(s)
1990s: +0.450 from 2 reading(s)
2000s: +0.535 from 2 reading(s)
2010s: +0.810 from 2 reading(s)
2020s: +1.100 from 2 reading(s)
year // 10 * 10 is the floor division from Lesson 3 turning 1987 into 1980: divide by ten discarding the remainder, then multiply back. It is the standard way to bucket numbers into ranges, and it works for any size of bucket.
Printing the count beside each mean is not decoration. The 1970s figure comes from a single reading, so calling it a decade mean would be misleading, and a reader who can see the 1 can judge that for themselves. Any summary that hides how many values it is built from is hiding the thing that decides whether to believe it.
A bug the loader found when it met a bad file
Now feed the loader a file with the sort of damage real data has: a blank line, a value that is a word, and a row with an extra comma.
year,anomaly_c
1975,-0.01
1980,0.26
1985,unknown
1990,0.45,extra
1995,0.45
header: year,anomaly_c
skipping unreadable line: '1985,unknown'
skipping odd line: '1990,0.45,extra'
kept 3 rows: [(1975, -0.01), (1980, 0.26), (1985, 0.45)]
Read that last line carefully. It kept three rows, which is right: 1975, 1980 and 1995 are the three good ones. But it reports the third as 1985 with a value of 0.45. There is no such reading. 1985's value was the word unknown, and 0.45 belongs to 1995.
The fault is in two lines of the loader:
try:
years.append(int(parts[0]))
values.append(float(parts[1]))
except ValueError:
The year is appended, and then the value is converted. When the conversion fails, the year is already in its list and the value never arrives, so from that row onwards the two lists are out of step by one. This is exactly the parallel-list failure from Lesson 10, and it produced a plausible-looking wrong number rather than an error, which is the Lesson 13 warning about logic errors.
The repair is to convert both values first and only append when both have succeeded:
try:
year = int(parts[0])
value = float(parts[1])
except ValueError:
print("skipping unreadable line:", repr(line))
continue
years.append(year)
values.append(value)
header: year,anomaly_c
skipping unreadable line: '1985,unknown'
skipping odd line: '1990,0.45,extra'
kept 3 rows: [(1975, -0.01), (1980, 0.26), (1995, 0.45)]
header: year,anomaly_c
clean file still gives 11 rows
1995 with 0.45, which is correct, and the clean file still gives all eleven rows. The rule that comes out of this is worth writing down: do all the work that can fail before you change anything. Convert, validate, and only then store. A half-finished update is worse than no update, because it leaves the data quietly wrong instead of loudly missing.
And notice how it was found: by writing a deliberately broken input file and looking at what came out. That is a test, in the sense of Lesson 14, aimed at a loader rather than a calculation.
What these numbers do and do not say
Reading data carelessly is as easy as writing code carelessly, so be exact about what is in this file. Each value is an anomaly, not a temperature: the difference from the 1951 to 1980 average, in degrees Celsius. 1.19 for 2025 does not mean the world averaged 1.19 degrees. NASA's own documentation for the dataset explains why anomalies are published rather than absolute temperatures: anomalies can be combined reliably across stations that sit at different altitudes and in different climates, while absolute regional averages carry much larger uncertainties.
Two more limits on what the summary above shows. Eleven readings five years apart are a sample, not the record; the real file has one row per year since 1880 and a program that read all of it would get slightly different figures. And a year-to-year rise or fall is not by itself a trend, which is why the steps-up count was reported alongside the change from first to last rather than instead of it.
Common misconceptions
- "A CSV is a spreadsheet file." It is plain text with commas in it. A spreadsheet can open one, and opening and saving one in a spreadsheet is a common way to get blank lines and reformatted numbers into your data.
- "The header row can be treated like any other." It is text and will fail the first conversion you try. Read it off with readline before the loop.
- "Skipping bad rows quietly is tidier." A loader that says nothing can discard half a file and still report a confident mean. Print what you skip.
- "If one column fails, only that column is affected." When two lists are built in step, a failure partway through a row knocks every later pair out of alignment, and the result is wrong rather than missing.
- "An anomaly of 1.19 means a temperature of 1.19 degrees." It is a difference from a stated baseline. Always find out what the baseline is before quoting a number from a dataset.
Putting it together
Remember: Take the header off first, convert everything that might fail before storing anything, and report what you skipped.
You can read a CSV with a header row, turn its text into numbers, and defend the loader against blank lines, short rows and values that will not convert. You can summarise a column with count, ends, extremes, mean, change and direction, and you know that a rising series need not rise at every step. You can group readings with a pair of dictionaries and bucket a year into a decade with floor division, and you know to print the number of values behind every mean. You have watched a loader produce a plausible wrong pairing, traced it to the order of two appends, and fixed it with a rule that applies far beyond this lesson. And you know the difference between an anomaly and a temperature, which is the kind of thing that separates using data from quoting it.
The summary is a column of numbers, and a column of numbers is hard to feel. The next lesson turns it into a picture, twice: once out of characters in the terminal, and once as a real image file written by hand, with no library at all.
Sources
- GISTEMP Team. (2026). GISS Surface Temperature Analysis (GISTEMP), version 4, Land-Ocean Global Means table (GLB.Ts+dSST.csv). NASA Goddard Institute for Space Studies. Values read for 1975 to 2025, J-D column. data.giss.nasa.gov
- NASA Goddard Institute for Space Studies. (2026). GISTEMP frequently asked questions (why anomalies rather than absolute temperatures, and the 1951-1980 base period). data.giss.nasa.gov
- Lenssen, N. J. L., Schmidt, G. A., Hansen, J. E., Menne, M. J., Persin, A., Ruedy, R., and Zyss, D. (2019). Improvements in the GISTEMP uncertainty model. Journal of Geophysical Research: Atmospheres, 124(12), 6307-6326.
- Wikipedia contributors. (2026). Comma-separated values. Wikipedia. en.wikipedia.org
- Sweigart, A. (2019). Automate the Boring Stuff with Python (2nd ed.), Chapter 16: Working with CSV files and JSON data. No Starch Press. automatetheboringstuff.com
- Key terms
- CSV
- Comma-separated values: a plain text file with one record per line and commas between the fields.
- Header row
- The first line of a CSV, naming the columns. It is text and must be taken off before any conversion.
- readline
- A file method that reads exactly one line, used to consume the header before a loop handles the rest.
- Anomaly
- A difference from a stated baseline rather than an absolute value. GISTEMP uses the 1951-1980 average.
- Bucketing
- Grouping values into ranges, as year // 10 * 10 turns any year into the start of its decade.
- Parallel lists
- Two lists kept in step by position, which fall out of alignment if one is appended to and the other is not.
- Defensive loading
- Converting and validating every field before storing anything, so a bad row is skipped rather than half-applied.
Drawing the Chart Yourself, With No Chart Library at All
- Scale a column of values to a bar width and print a readable text chart.
- Handle negative values by putting the axis at zero rather than at the left edge.
- Write an SVG image file from Python using only open, write and close.
Eleven numbers from the last lesson, turned into something you can read at a glance:
Global temperature anomaly, degrees C against the 1951-1980 mean
Source: NASA GISTEMP v4
1975 -0.01
1980 +0.26 #########
1985 +0.12 ####
1990 +0.45 ###############
1995 +0.45 ###############
2000 +0.39 #############
2005 +0.68 #######################
2010 +0.72 ########################
2015 +0.90 ##############################
2020 +1.01 ##################################
2025 +1.19 ########################################
That is a chart. It has no pixels in it, it took twelve lines of Python, and it shows the shape of the data instantly: the dip in the mid 1980s, the flat stretch around 1990 to 2000, and the steady climb after 2005. Anyone who claims a chart needs a library has not tried this.
A note about libraries. The usual Python answer for charts is matplotlib, which is a third-party package you have to install. It is not installed on the machine these programs ran on, and it is not used anywhere in this course, because nothing here should depend on a download. Everything in this lesson uses the standard library and the built-in functions you already know. The second half writes a real image file, and a real image file, it turns out, is just text.
Scaling: turning a value into a number of characters
def text_chart(years, values, width=40):
"""For finite values, scale each magnitude to the largest magnitude."""
if len(years) != len(values):
raise ValueError("labels and values must have equal lengths")
if not values:
return
biggest = max(abs(value) for value in values) or 1.0
for year, value in zip(years, values):
blocks = round(abs(value) / biggest * width)
bar = "#" * blocks
print(f"{year} {value:+5.2f} {bar}")
The calculation divides abs(value) by the largest absolute value. For a nonempty set of finite numbers with at least one nonzero value, that ratio is between 0 and 1. The fallback scale of 1.0 lets an all-zero dataset produce empty bars. Multiply by the width you want and you get a number of characters. Round it to a whole number, because you cannot print two thirds of a hash mark.
Three details make it readable. max(abs(value) for value in values) takes the largest distance from zero in either direction, so a big negative value gets a full-width bar too. "#" * blocks is the string repetition from Lesson 8. And {value:+5.2f} lines the numbers up in a column five characters wide with the sign always shown, so the decimal points stack.
The width=40 default is there so the chart can be made narrower for a small screen without editing the function.
The bar that vanished
Look at the 1975 row. The value is -0.01 and the bar is completely empty. That is arithmetically correct: 0.01 divided by 1.19 times 40 is 0.34, and round(0.34) is 0. It is also a bad chart, because an empty bar looks identical to a missing row, and a reader cannot tell a value of zero from no data at all.
There is a second problem hiding in the same row. Every bar here is drawn to the right from the same starting column, which means -0.01 and +0.01 would look the same. For data that crosses zero, the axis has to be at zero, not at the left edge.
Putting the axis where zero is
LEFT = 34
RIGHT = 34
biggest = max((abs(value) for value in values), default=0.0) or 1.0
print(" " * (11 + LEFT) + "0")
for year, value in zip(years, values):
if value >= 0:
blocks = round(value / biggest * RIGHT)
line = " " * LEFT + "|" + "#" * blocks
else:
blocks = round(abs(value) / biggest * LEFT)
line = " " * (LEFT - blocks) + "#" * blocks + "|"
print(f"{year} {value:+5.2f} {line}")
0
1975 -0.01 |
1980 +0.26 |#######
1985 +0.12 |###
1990 +0.45 |#############
1995 +0.45 |#############
2000 +0.39 |###########
2005 +0.68 |###################
2010 +0.72 |#####################
2015 +0.90 |##########################
2020 +1.01 |#############################
2025 +1.19 |##################################
Now there is a vertical line down the chart at zero, positive bars go right from it and negative ones would go left. Both sides reserve 34 characters and use the same scale, so equal magnitudes have equal lengths. The 1975 bar is still too small to see, which is honest, because -0.01 really is almost nothing, and now the axis makes that visible rather than ambiguous: the bar starts at the line and stops there.
The trick in the negative branch is worth reading twice. The bar has to end at the axis, so it is padded with " " * (LEFT - blocks) spaces first and the hashes come after. Drawing rightwards is easy; drawing leftwards means working out where to start.
An image file is a text file
A text chart is fine in a terminal and useless in a report. For a real picture, the format to reach for is SVG, Scalable Vector Graphics, because an SVG file is plain text describing shapes. Every browser opens one. Here are the first lines of the one this lesson produces:
<svg xmlns="http://www.w3.org/2000/svg" width="640" height="360">
<rect x="0" y="0" width="640" height="360" fill="white"/>
<text x="60" y="20" font-family="sans-serif" font-size="14">Global temperature anomaly (NASA GISTEMP v4), degrees C vs 1951-1980</text>
<line x1="60" y1="317.6" x2="620" y2="317.6" stroke="black"/>
<rect x="67.6" y="317.6" width="35.6" height="2.4" fill="steelblue"/>
<text x="85.5" y="336" font-family="sans-serif" font-size="10" text-anchor="middle">1975</text>
Read it and you can see the picture. A white rectangle covering everything, which is the background. A line of text near the top, which is the title. A black line across at y of 317.6, which is the axis. Then, for each year, a coloured rectangle and a small centred label underneath it.
That is the whole format, as far as this lesson needs it. An opening svg tag saying how big the canvas is, any number of rect, line and text elements, and a closing tag. The MDN SVG reference lists the rest, and there is a great deal of rest, but four elements will draw most charts anybody needs.
Screen coordinates run downwards
There is one thing to get straight before writing any drawing code, and it catches everybody. In maths, y goes up. In SVG's initial coordinate system, y goes down: y of 0 is the top edge and y of 360 is the bottom. A rectangle is positioned by its top left corner, so a taller bar has a smaller y.
That is why the program has a function whose only job is to turn a data value into a screen position:
top = max(max(values), 0.0)
bottom = min(min(values), 0.0)
span = top - bottom
if span == 0:
top = 1.0
span = 1.0
def y_for(value):
return MARGIN_TOP + (top - value) / span * plot_height
This fragment assumes a nonempty list of finite values. Both bounds include zero, so it also works for an all-negative dataset. If every value is zero, the fallback gives a nonzero plotting span. Check the mixed-sign sample against its two ends. The largest value gives top - value of 0, so it lands at the top margin. The smallest gives the full span, so it lands at the bottom of the plot. And y_for(0.0) gives 317.6 for this data, which is where the axis line went.
bottom is min(min(values), 0.0) rather than just the smallest value, and top also includes zero, so the plotting range contains zero. A bar chart whose axis is not at zero exaggerates every difference on it, which is one of the oldest ways to mislead with a graph and is worth refusing to do even by accident.
Writing the chart out, element by element
The following is a fragment, not a complete runnable writer. First define the canvas and margins, calculate plot dimensions and bar sizes, and initialize parts with the opening svg tag, background, title and zero-axis line. The loop adds bars and labels; append the closing svg tag before writing.
for index in range(len(values)):
value = values[index]
x = MARGIN_LEFT + index * bar_slot + (bar_slot - bar_width) / 2
y = y_for(max(value, 0.0))
height = abs(y_for(value) - zero_y)
if value >= 0:
colour = "crimson"
else:
colour = "steelblue"
parts.append(f'<rect x="{x:.1f}" y="{y:.1f}" width="{bar_width:.1f}" height="{height:.1f}" fill="{colour}"/>')
label_y = HEIGHT - MARGIN_BOTTOM + 16
parts.append(f'<text x="{x + bar_width / 2:.1f}" y="{label_y}" font-family="sans-serif" font-size="10" text-anchor="middle">{years[index]}</text>')
parts.append("</svg>")
with open(filename, "w") as f:
f.write("\n".join(parts))
f.write("\n")
Expected beginning of the assembled SVG file, not terminal output:
<svg xmlns="http://www.w3.org/2000/svg" width="640" height="360">
<rect x="0" y="0" width="640" height="360" fill="white"/>
<text x="60" y="20" font-family="sans-serif" font-size="14">Global temperature anomaly (NASA GISTEMP v4), degrees C vs 1951-1980</text>
<line x1="60" y1="317.6" x2="620" y2="317.6" stroke="black"/>
<rect x="67.6" y="317.6" width="35.6" height="2.4" fill="steelblue"/>
<text x="85.5" y="336" font-family="sans-serif" font-size="10" text-anchor="middle">1975</text>
Open temps.svg in a browser and you get a picture 640 by 360: a white background, the title across the top, a horizontal black axis about seven eighths of the way down, ten red bars rising from that axis in the shape of the text chart, one small blue stub hanging below it for 1975, and the year under each bar. The bars for 2015, 2020 and 2025 are the three tallest, and 2025 reaches the top of the plot.
Four things in that loop are worth naming.
Each bar gets a slot. bar_slot is the plot width divided by the number of bars, and the bar itself is 70 percent of its slot, centred in it. That is what puts a gap between bars without any gap arithmetic.
The top of a bar is not always the value. y_for(max(value, 0.0)) gives the axis for a negative bar, because a rectangle is drawn from its top edge and a negative bar's top edge is the axis.
The height is a distance. abs(y_for(value) - zero_y) works for bars on either side of the axis. A negative height is invalid in SVG, so use a nonnegative distance.
The parts are collected in a list and joined at the end. "\n".join(parts) from Lesson 11 puts a newline between every element. Building a list and joining once is tidier than a hundred write calls, and it lets you count the elements, so it can report how many text fragments were written. A line count is not an SVG element count: closing tags can occupy lines too.
Two charts, two jobs
| Text chart | SVG file | |
|---|---|---|
| Where it appears | In the terminal, straight away | In a file, opened in a browser |
| Lines of code | About a dozen | About forty |
| Good for | Checking data while you work | Anything anyone else will see |
| Resizes | Change the width argument | Infinitely; it is vector, not pixels |
| Colour, fonts, curves | No | Yes |
| Needs installing | Nothing | Nothing |
Write the text version first, always. It takes two minutes, it shows you immediately whether the data is what you think it is, and every mistake you would have made in the SVG is cheaper to find there. Then build the picture once you know the numbers are right.
Common misconceptions
- "Drawing a chart needs a chart library." Everything in this lesson used built-in functions and open. A library saves time on a complicated chart; it is not a requirement for a simple one.
- "An SVG is an image, so Python cannot write one." It is a text file describing shapes. Python writes it exactly as it writes any other text.
- "y of 0 is the bottom of the picture." In SVG's initial coordinate system y of 0 is the top and y increases downwards, which is why a taller bar has a smaller y.
- "An empty bar means the value is missing." A value too small to round up to one character draws nothing, which is why an axis line matters: it shows where the bar starts even when it does not go anywhere.
- "Starting the axis at the smallest value shows the differences better." It exaggerates them. Including zero in both the lower and upper bounds is a decision about honesty, not about layout.
Pulling it together
The upshot: For nonnegative bars, scale each value against a positive maximum. For signed bars, use a consistent magnitude scale on both sides of zero.
You can scale a column of numbers into bars, print a text chart that lines up, and read the shape of a dataset without a single pixel. You know why a value that rounds to nothing still needs an axis to sit against, and how to draw bars in both directions from a zero line. You know that an SVG is a text file holding rect, line and text elements, that SVG's initial y coordinates increase downwards, and that one small function converting a data value into a screen position is what keeps the whole drawing consistent. You have written an image file with open and write, and nothing was installed to do it.
That closes Module 5. Your programs can now remember, read real data, summarise it and draw it. Module 6 is about making things: what else is already inside Python, a drawing you can watch happen, and two games.
Sources
- Mozilla. (2026). SVG: Scalable Vector Graphics (the svg, rect, line and text elements and their attributes). MDN Web Docs. developer.mozilla.org
- World Wide Web Consortium. (2018). Scalable Vector Graphics (SVG) 2, W3C Candidate Recommendation (the coordinate system, with y increasing downwards). w3.org
- GISTEMP Team. (2026). GISS Surface Temperature Analysis (GISTEMP), version 4. NASA Goddard Institute for Space Studies. (The data charted here.) data.giss.nasa.gov
- Tufte, E. R. (2001). The Visual Display of Quantitative Information (2nd ed.). Graphics Press. (On baselines, and on distortion caused by a truncated axis.)
- Key terms
- Scaling
- Mapping data to drawing distances. These text charts divide magnitude by the largest magnitude, then multiply by the width.
- Baseline
- The value the bars are measured from. Forcing it to zero keeps the lengths honest.
- SVG
- Scalable Vector Graphics: a plain text image format describing shapes, which any browser can open.
- Screen coordinates
- SVG's initial coordinate system, where y of 0 is the top edge and y increases downwards.
- rect
- An SVG element drawn from its top left corner, with a width and a height, which is how a bar is made.
- Bar slot
- The width available to one bar, usually the plot width divided by the number of bars, with the bar filling part of it.
- Vector graphic
- A picture stored as shapes rather than pixels, so it stays sharp at any size.
Module 6: What Was Already in the Box, and Two Games
Meet the standard library that came with Python, draw with turtle, make a program that rolls dice, and build a game you can play and then extend.
Four Lines to Count the Days Until an Exam
- Import a standard library module and call its functions through a dotted name.
- Choose between math, random, datetime and json for a given job, and read a library page to find the function you need.
- Split your own code into a module and a program that imports it, guarded by the main check.
Seven short lines, beginning with an import:
import datetime
today = datetime.date(2026, 9, 22)
exam = datetime.date(2027, 5, 17)
left = exam - today
print(left)
print(left.days)
print(today.strftime("%A %d %B %Y"))
237 days, 0:00:00
237
Tuesday 22 September 2026
Nobody worked out that 2027 is not a leap year, that September has thirty days, or which day of the week the 22nd falls on. Nothing was downloaded either. That code came with Python, sitting in the standard library, which is a few hundred modules of finished work that install themselves when Python does. This lesson is about getting at them, and about putting your own code somewhere it can be got at the same way.
What import actually does
Square roots are a fair test, because sqrt sounds like the sort of word Python would know:
print(sqrt(49))
NameError: name 'sqrt' is not defined
It does not know it. Python starts each program with a small set of built-in functions, about seventy of them, and print, len, range, int and sum are in that set because you have used them all course. sqrt is not. It lives in a module called math, and a module groups code and names that another program can import. Some are Python files; others, including math on usual CPython installations, are implemented as compiled extensions.
An import statement goes and fetches one:
import math
print(math.sqrt(49))
print(math.pi)
print(math.floor(3.7), math.ceil(3.2))
print(math.hypot(3, 4))
print(round(math.degrees(math.pi / 4), 2))
print(math.sqrt(2) ** 2)
print(math.isclose(math.sqrt(2) ** 2, 2.0))
print(math.factorial(5))
7.0
3.141592653589793
3 4
5.0
45.0
2.0000000000000004
True
120
Notice the dot. After import math, the name math exists in your program and everything the module offers hangs off it: math.sqrt, math.pi, math.floor. The module has a namespace, a mapping from its names to objects. The dot accesses a name in that namespace, and it is why you can import twenty modules without any of them quarrelling over a name. If two modules both had a function called pick, one would be random.pick and the other cards.pick, and Python would never be confused about which you meant.
Two lines of that output are old friends. math.sqrt(2) ** 2 comes out as 2.0000000000000004, exactly the floating point behaviour from Lesson 3, and math.isclose is the fix: it compares the difference with a tolerance. Its default relative tolerance is 1e-9 and its default absolute tolerance is zero. Choose tolerances for the problem; near zero you often need a nonzero abs_tol.
Spell the module name wrong and the failure comes at once, on the import line in this example:
import maths
print(maths.sqrt(49))
ModuleNotFoundError: No module named 'maths'
The module is math, singular, whatever your maths teacher calls the subject. A ModuleNotFoundError almost always means a typo, an American spelling you did not expect, or a module that really does need installing, although the examples use standard-library modules. Some installations omit optional components such as Tk.
Four ways to write an import, and the one that bites
| Written as | You then say | Use it when |
|---|---|---|
import math | math.sqrt(49) | Almost always. The dot says where the function came from. |
from math import sqrt, pi | sqrt(49) | One or two names used constantly in a short program. |
import math as m | m.sqrt(49) | A long module name you type forty times. |
from math import * | sqrt(49) | Usually avoid. It imports the module's public names and can hide your own names. |
The second form has a trap in it that is worth meeting once:
from math import pi
print(pi)
pi = 3
print(pi)
print(2 * pi * 10)
3.141592653589793
3
60
Nothing complained. pi was an ordinary variable holding 3.14159..., you assigned 3 to it, and from then on every circumference in the program was wrong by four and a half percent. Had the import been import math, the constant would have been math.pi, which this local assignment does not change, and your own pi = 3 would have sat harmlessly beside it. That is the real argument for the dotted form: not typing, but protection.
PEP 8, the style guide Python programmers actually follow, says imports go at the top of the file, one module per line, before anything else. It also says plainly that wildcard imports should be avoided.
math, and the four modules worth knowing first
Hundreds of modules is not a reading list. Four of them cover most of what a beginner wants, and the tour below is a working program each. First, what happens when you hand math something outside its domain:
import math
print(math.sqrt(-1))
ValueError: expected a nonnegative input, got -1.0
The library page says so in as many words: the implementation raises ValueError for invalid operations such as sqrt(-1.0). Complex numbers are a different module. The point is general. A library function is allowed to refuse, and it refuses with an exception you already know how to read and, since Lesson 12, how to catch.
datetime, which knows how many days February had
import datetime
start = datetime.date(2026, 9, 7)
print(start.year, start.month, start.day)
print(start.weekday())
print(start.isoformat())
half_term = start + datetime.timedelta(weeks=7)
print(half_term, half_term.strftime("%A"))
born = datetime.date(2011, 3, 14)
age_days = (start - born).days
print(age_days, "days old, which is", age_days // 365, "years")
stamp = datetime.datetime(2026, 9, 7, 8, 45, 0)
print(stamp.strftime("%d/%m/%Y at %H:%M"))
parsed = datetime.datetime.strptime("17/05/2027 09:00", "%d/%m/%Y %H:%M")
print(parsed)
print(parsed > stamp)
2026 9 7
0
2026-09-07
2026-10-26 Monday
5656 days old, which is 15 years
07/09/2026 at 08:45
2027-05-17 09:00:00
True
Four ideas are doing the work there. A date holds a year, a month and a day and knows the calendar. A timedelta is a length of time, so date + timedelta gives another date and date - date gives a timedelta whose .days you can read. strftime turns a date into a string using percent codes, and strptime reads a string back into a date using the same codes. And dates compare with > and < like numbers, which is how you sort them or ask whether a deadline has passed.
weekday() printing 0 catches everybody. The documentation is exact: Monday is 0 and Sunday is 6. So 7 September 2026 was a Monday, which the strftime("%A") on the half-term line confirms in English.
The age line is only an approximation: age_days // 365 ignores leap days and can report a year too many before a birthday. Working out somebody's age exactly is fiddlier than it looks, and the right move is to say so in a comment rather than pretend the integer division is a birthday calculator.
random, and why a seed makes a random program testable
import random
random.seed(2026)
print(random.randint(1, 6), random.randint(1, 6), random.randint(1, 6))
print(random.choice(["rock", "paper", "scissors"]))
cards = ["A", "K", "Q", "J", "10"]
random.shuffle(cards)
print(cards)
print(round(random.random(), 4))
random.seed(2026)
print(random.randint(1, 6), random.randint(1, 6), random.randint(1, 6))
1 3 5
scissors
['10', 'J', 'Q', 'K', 'A']
0.7834
1 3 5
The last line is the interesting one. Setting the same seed again produced the same three dice: 1, 3, 5. random produces pseudorandom values from deterministic state. Setting a seed initializes that state. With no explicit seed, Python uses operating-system randomness when available, otherwise system time. Reusing a seed with the same calls in the same environment lets you reproduce a run; sequences from higher-level functions can change across Python versions. Lesson 20 does nothing but this.
Note that randint(1, 6) can return 6. It is one of the few places in Python where the upper end is included, precisely because dice, and it is worth remembering next to range(1, 6), which stops at 5.
json, so a dictionary survives being written down
Lesson 15 wrote text to a file, and Lesson 16 read a CSV. You can serialize nested data into text; JSON supplies a standard format for doing so. JSON can, and the json module turns one into the other in a line each way:
import json
scores = {"Ada": 18, "Blaise": 14, "Grace": 20}
settings = {"name": "Sam", "high_scores": scores, "sound": True, "lives": 3}
with open("save.json", "w") as f:
json.dump(settings, f, indent=2)
print(open("save.json").read())
with open("save.json") as f:
loaded = json.load(f)
print(loaded["high_scores"]["Grace"])
print(loaded == settings)
print(type(loaded["sound"]))
text = json.dumps({"level": 4, "done": False})
print(text)
print(json.loads(text)["done"])
numbered = {1: "one", 2: "two"}
back = json.loads(json.dumps(numbered))
print(back)
{
"name": "Sam",
"high_scores": {
"Ada": 18,
"Blaise": 14,
"Grace": 20
},
"sound": true,
"lives": 3
}
20
True
<class 'bool'>
{"level": 4, "done": false}
False
{'1': 'one', '2': 'two'}
Read the saved file and it is almost Python, with three differences: double quotes only, true and false in lower case, and null where Python writes None. The module handles all of that. json.dump writes to an open file and json.load reads from one; dumps and loads, with the s, do the same to and from a string. indent=2 is the difference between a file you can read and one long line.
Two results in there deserve care. loaded == settings printed True, so the whole nested structure came back intact, booleans still booleans. But the last line shows a dictionary with keys 1 and 2 going out and coming back with keys '1' and '2', as strings. That is not a bug in Python; it is JSON, whose documentation states that keys in key and value pairs are always of type str. If your keys are numbers, JSON will quietly stringify them, and the lookup that worked before saving will fail after loading.
And not everything fits:
import json
print(json.dumps({"topics": {"algebra", "trig"}}))
TypeError: Object of type set is not JSON serializable
when serializing dict item 'topics'
Dictionaries, lists, strings, numbers, booleans and None convert. A set does not, because JSON has no set. Turn it into a list first with list(topics) and it saves fine.
Which of the four, for which job
| You want to | Reach for | The call |
|---|---|---|
| Take a square root, round down, get pi | math | math.sqrt, math.floor, math.pi |
| Compare two floats sensibly | math | math.isclose(a, b) |
| Roll a die, shuffle, pick one at random | random | randint, shuffle, choice |
| Repeat a random run exactly | random | random.seed(7) |
| Count days between two dates | datetime | date2 - date1 |
| Print a date as people write it | datetime | strftime("%d/%m/%Y") |
| Save a nested dictionary and get it back | json | json.dump, json.load |
The phrase people use for this is that Python comes with batteries included. There is also statistics for means and medians, os.path for asking whether a file exists, csv for the job you did by hand in Lesson 16, time for pausing, and a great deal more. The habit worth forming is not memorising them. It is checking the library index before you write forty lines, to see whether it already provides the operation you need.
Your own file is a module too
Here is grades.py, an ordinary file with two functions in it and one line that is about to cause trouble:
"""grades.py: the two mark functions this course keeps needing."""
def average(numbers):
return sum(numbers) / len(numbers)
def letter(mark):
if mark >= 70:
return "A"
if mark >= 60:
return "B"
if mark >= 50:
return "C"
return "D"
print("grades.py is being read")
if __name__ == "__main__":
print("Self test:", average([10, 20, 30]), letter(64))
And a second file, report.py, in the same folder:
import grades
results = [71, 64, 48, 55]
print("Average:", round(grades.average(results), 1))
for mark in results:
print(mark, grades.letter(mark))
grades.py is being read
Average: 59.5
71 A
64 B
48 D
55 C
Two things happened. Your own functions arrived through a dotted name exactly as math's did, because both modules expose names. Python first checks its module cache; for a new import it uses its import machinery, including built-in modules and the module search path. In this ordinary script example, the script's directory is on that path. And the line grades.py is being read printed, although nothing in report.py asked for it. The first import initializes the module by executing its top-level code. Function definitions create functions without calling their bodies. Conditional blocks run only if their conditions hold, and later ordinary imports normally reuse the cached module. The two def statements ran, which is how the functions came to exist; the print ran too.
The self test did not run, though. Now run grades.py on its own:
grades.py is being read
Self test: 20.0 B
That is what if __name__ == "__main__": is for. Python sets the variable __name__ inside every module: to the string "__main__" when that file is the one you launched, and to the module's own name when it was imported by something else. So the guarded block runs when you run the file directly and stays quiet when another program imports it. The modules chapter of the tutorial puts it as making the file usable as a script as well as an importable module.
That guard is where your tests from Lesson 14 belong. Put the asserts under it and a module carries its own checks: run it normally, it checks the chosen cases; import it, the guarded checks stay quiet. Remember that -O omits asserts. The bare print("grades.py is being read") above the guard is the mistake to avoid, and the only reason it is in the example is to show you what an unguarded line does to every program that imports the file.
Common misconceptions
- "Importing a module downloads it." Import does not install a package. These examples use installed standard-library modules or files you write. A module's initialization code can itself perform network operations, so do not infer that imports can never use the network.
- "from math import sqrt is the modern way, and import math is old fashioned." Neither is newer. The dotted form is preferred because the dot records where a name came from. Assigning pi locally does not change math.pi, although assigning math.pi directly can change that module attribute.
- "random.seed makes the numbers less random." It makes them repeatable, which is a different property. Without an explicit seed, initialization uses operating-system randomness when available, otherwise time. Record the seed and calls when reproducibility matters.
- "json.dump gives back exactly what I put in." It gives back dictionaries, lists, strings, numbers, booleans and None. Numeric dictionary keys come back as strings, tuples come back as lists, and a set does not go out at all.
- "A module has to be something famous that other people wrote." Any .py file in the folder is importable by name. Splitting a long program into two files and importing one from the other is the ordinary thing to do.
- "__name__ is something I have to set." Python sets it. You only ever read it, and almost always in that one comparison.
Looking back
In short: Before you write forty lines, spend two minutes on the library index. Someone has probably already written them, tested them, and fixed the leap year bug.
You can import a module and call its functions through a dotted name, and you know why that dot is worth the extra characters. You have used math for roots and for comparing floats properly, datetime for arithmetic on real dates, random for dice and for the seed that makes dice repeatable, and json for saving a nested dictionary and getting it back with its booleans intact. You know the three shapes of import and which one to refuse. You know that a first import executes a module's top-level code, while later ordinary imports normally use its cached object, and that if __name__ == "__main__": is the switch that separates a file's own self test from what it offers other programs.
Next lesson the output stops being text. One import, a few instructions, and a window opens with a line being drawn across it.
Sources
- Python Software Foundation. (2026). The Python Tutorial, 6: Modules (the import statement, the module search path, and executing modules as scripts). docs.python.org
- Python Software Foundation. (2026). math: Mathematical functions (sqrt, floor, ceil, hypot, degrees, factorial, isclose with a default relative tolerance of 1e-09, and ValueError for invalid operations such as sqrt of -1.0). docs.python.org
- Python Software Foundation. (2026). datetime: Basic date and time types (date, timedelta, weekday with Monday as 0, isoformat, strftime and strptime format codes). docs.python.org
- Python Software Foundation. (2026). json: JSON encoder and decoder (dump, load, dumps, loads, the indent argument, the Python to JSON conversion table, and the rule that JSON keys are always of type str). docs.python.org
- van Rossum, G., Warsaw, B., and Coghlan, N. (2001, revised). PEP 8: Style Guide for Python Code (imports at the top of the file, one per line, and wildcard imports to be avoided). Python Software Foundation. peps.python.org
- Key terms
- Module
- An importable unit with its own names, often a .py file but also possibly a built-in or extension module.
- Standard library
- The few hundred modules that install with Python itself, so no download is needed to use them.
- import
- A statement that locates a module, initializes it if needed, and binds its name or selected names in your program.
- Namespace
- A mapping from names to objects. A module's namespace keeps its names separate from those in other modules.
- Seed
- A value used to initialize pseudorandom state. The same seed and calls reproduce a run in the same environment.
- timedelta
- A length of time. Adding one to a date gives another date; subtracting two dates gives one of these.
- JSON
- A text format for nested data. Python dictionaries and lists convert both ways; sets do not.
- __name__
- A variable Python sets in every module: the string __main__ when the file was launched, otherwise the module's name.
Six Lines That Open a Window and Draw a Square
- Draw with the turtle module using forward, right, penup, goto and the fill pair.
- Explain why the exterior angles of any closed path add to 360 degrees, and use it to draw any polygon.
- Predict where a turtle program will draw by tracking position and heading on paper.
Every program in this course so far has answered in text. This one answers in a picture, so read the next paragraph before you run it.
The output of a turtle program is a drawing in a window, not printed text. Nothing appears in the terminal. A separate window opens, titled Python Turtle Graphics, with a small arrowhead in the middle of it, and the drawing happens in front of you at a speed you can watch. Because this page cannot show you that window, every turtle program below is followed by a careful description of what it draws, and later in the lesson there is a way to check a drawing with arithmetic instead of eyes. Close the window when you have finished looking; the program is waiting for you to do exactly that.
import turtle
t = turtle.Turtle()
for side in range(4):
t.forward(120)
t.right(90)
turtle.done()
What the window shows: a white area with an arrowhead at the centre pointing right. The arrowhead slides 120 pixels to the right, leaving a thin black line behind it, then pivots a quarter turn clockwise, then slides down 120 pixels, and so on, until a square 120 pixels on each side sits in the window with its top left corner at the centre. The turtle finishes where it started, pointing right again. The window then stays open, doing nothing, until you close it.
Six lines. One of them is a loop you have written since Lesson 7, and the other five are the whole of the drawing.
Three names in those six lines
import turtle fetches the turtle module, which is part of the standard library, so nothing was installed. It is an implementation of the drawing tools from Logo, a language Wally Feurzeig, Seymour Papert and Cynthia Solomon built in 1967 to teach programming to children. The documentation still describes it as an educational tool, and it earns that description: the reason a beginner can draw a square in six lines is that the turtle's instructions are the instructions you would give a person walking on a floor.
turtle.Turtle() makes a turtle and the window it lives in. The name t is then a thing with methods, so every drawing instruction is t.something(...). That dotted shape is exactly the namespace idea from the last lesson.
turtle.done() hands control to the window. The documentation says it must be the last statement in a turtle graphics program, and the script waits, without exiting, until the window is closed. Leave it out and on most systems the window flashes up and vanishes before you can see anything, which is the single most common first turtle problem.
If the import fails with No module named '_tkinter', turtle is there but the graphics toolkit underneath it, tkinter, has not been installed with your Python. On Windows and macOS installers from python.org it is included; on Linux it is usually a separate package.
Why the square went downwards
You said right, and the square appeared below the starting point. Both facts come from one rule: the turtle starts at (0, 0) in the middle of the window, facing east, and right turns it clockwise by however many degrees you ask for. Degrees are the default unit.
So the four corners are: start at the centre, walk east to (120, 0), turn clockwise to face south, walk to (120, -120), turn to face west, walk to (0, -120), turn to face north, walk home.
Turtle's y axis points up, like a maths graph and unlike the SVG coordinates of Lesson 17, where y of 0 was the top edge and y increased downwards. Two drawing systems, two opposite conventions, and the only defence is to check which one you are in before you calculate anything. A quick test settles it: t.goto(0, 100) draws upwards in turtle and downwards in SVG.
Checking a drawing without opening a window
You do not need a window to know where a turtle ends up. A position and a heading are two numbers and an angle, and trigonometry moves them. This program is ordinary text-printing Python, and its output below is real:
import math
def step(x, y, heading, distance):
x = x + distance * math.cos(math.radians(heading))
y = y + distance * math.sin(math.radians(heading))
return round(x, 1), round(y, 1)
def draw_polygon(sides, length):
x, y, heading = 0.0, 0.0, 0.0
corners = [(x, y)]
for corner in range(sides):
x, y = step(x, y, heading, length)
heading = (heading - 360 / sides) % 360
corners.append((x, y))
return corners
print("square:", draw_polygon(4, 120))
square: [(0.0, 0.0), (120.0, 0.0), (120.0, -120.0), (0.0, -120.0), (0.0, 0.0)]
Those are the four corners of the square the window drew, produced without any window at all. Three things in that code are worth pointing at. math.radians converts because math.cos and math.sin want radians while turtle talks in degrees. Subtracting from the heading is what makes it a right turn, since clockwise is the negative direction. And % 360 keeps the heading in the range 0 to 359 no matter how many times you go round, which is the same remainder operator you met dividing seconds into hours in Lesson 3.
Keep this program. When a turtle drawing comes out wrong, printing the corners is the fastest way to find out whether the geometry is wrong or only the colours.
The 360 rule, and a polygon function
Why 90 degrees for a square? Because the turtle went all the way round once and finished facing its original direction, so the turns must add up to a full circle: four turns, 360 degrees, 90 each. That generalises to a rule with no exceptions for any closed path drawn by turning always the same way:
| Shape | Sides | Turn each time | Sum of turns |
|---|---|---|---|
| Triangle | 3 | 120 | 360 |
| Square | 4 | 90 | 360 |
| Pentagon | 5 | 72 | 360 |
| Octagon | 8 | 45 | 360 |
| Sixty-sided figure | 60 | 6 | 360 |
The turn is 360 divided by the number of sides, and it is the exterior angle, not the interior angle you may have learned in geometry. For a triangle the interior angle is 60 and the turtle turns 120. Confusing the two is the most common reason a beginner's polygon fails to close.
One function covers all of them, and it takes the turtle as a parameter so the same function can drive any turtle:
import turtle
def polygon(t, sides, length):
for step in range(sides):
t.forward(length)
t.right(360 / sides)
def move_to(t, x, y):
t.penup()
t.goto(x, y)
t.pendown()
pen = turtle.Turtle()
pen.speed(0)
pen.hideturtle()
move_to(pen, -220, 0)
polygon(pen, 3, 90)
move_to(pen, -80, 0)
polygon(pen, 5, 60)
move_to(pen, 60, 0)
polygon(pen, 8, 40)
move_to(pen, 200, 0)
polygon(pen, 60, 6)
turtle.done()
What the window shows: four black outlines in a row across the middle, all of them hanging below the horizontal line through the centre, because every turn is a right turn. On the left, a triangle about 90 wide and 78 tall. Next, a pentagon about 97 wide and 92 tall. Next, an octagon about 97 across. On the right, a figure with sixty tiny sides that your eye reads as a circle, about 115 across. No arrowhead is visible anywhere, because hideturtle removed it, and the whole thing appears almost instantly because speed(0) means draw as fast as possible.
Those measurements are not guesses: run the polygon corners through the checking program above and the triangle's lowest corner comes out at y of -77.9, the pentagon's at -92.3.
move_to is the useful trick. penup lifts the pen so movement leaves no line, goto jumps to an absolute position, and pendown puts the pen back. Without the pen lift, four thin lines would join the four shapes together across the window.
One number changed, a different picture
Five sides, five equal turns, one forward distance. Here are two programs differing in a single number:
import turtle
pen = turtle.Turtle()
pen.pensize(3)
pen.color("crimson", "gold")
pen.begin_fill()
for point in range(5):
pen.forward(200)
pen.right(144)
pen.end_fill()
turtle.done()
What the window shows: a five-pointed star, the kind you draw without lifting your pencil, outlined in thick crimson and filled with gold. It is 200 pixels wide. One point sits at the top, two stick out sideways level with the centre of the window, and two hang down below. The starting position, the centre of the window, is the left-hand point of the star.
Now change 144 to 72 and run it again. The window shows a plain gold pentagon with a crimson edge, roughly 324 pixels wide, sitting below and to the right of where it started. No points, no crossings, nothing star-like.
The arithmetic explains both. Five turns of 72 add to 360: one lap, one convex pentagon. Five turns of 144 add to 720: two laps, and a path that crosses itself three times on the way, which is what a five-pointed star is. Feed both through the checking program and the numbers say the same thing:
star, turning 144:
point at (200.0, 0.0) now facing 216
point at (38.2, -117.6) now facing 72
point at (100.0, 72.6) now facing 288
point at (161.8, -117.6) now facing 144
point at (-0.0, -0.0) now facing 0
pentagon, turning 72: [(0.0, 0.0), (200.0, 0.0), (261.8, -190.2), (100.0, -307.8), (-61.8, -190.2), (0.0, 0.0)]
Read the star's five points as coordinates and the shape is unmistakable: one at y of +72.6 (the top point), two at y of 0, two at y of -117.6. The pentagon's corners, by contrast, march steadily away and back, never rising above the start. Both paths return exactly to (0, 0), which is the 360 rule and the 720 rule both keeping their promise.
color("crimson", "gold") sets two colours at once: the pen colour first, the fill colour second. begin_fill and end_fill bracket the shape to be filled, and turtle fills whatever region the path enclosed between them, which for the star means the middle pentagon fills too. Named colours such as crimson, gold and steelblue work, as do hex strings like "#ff8800".
A spiral out of three numbers
import turtle
pen = turtle.Turtle()
pen.speed(0)
length = 5
for step in range(60):
pen.forward(length)
pen.right(91)
length = length + 4
turtle.done()
What the window shows: a square spiral that slowly rotates. Each side is four pixels longer than the last, from 5 up to 241, and each turn is 91 degrees rather than 90, so every group of four sides sits one degree round from the group before. The result is a widening square coil about 320 pixels across, filling a region roughly from -163 to +162 in both directions, with a dense knot at the centre and long straight strokes at the outside. The turtle finishes at the top of the picture, facing up and to the left.
Change 91 to 90 and the rotation disappears: you get a plain rectangular spiral with everything parallel. Change it to 121 and the coil becomes triangular. Sixty turns of 91 degrees is 5460 degrees, just over fifteen full laps, and that one spare degree per corner is the whole effect.
This is the pattern worth stealing: a variable that changes a little on every pass of the loop. It is the same idea as the running total in Lesson 6, used for a length instead of a sum.
Common misconceptions
- "turtle prints something in the terminal too." It does not. If you want text as well, use
printfor the terminal orpen.write("text")to put words in the window. - "The angle to turn is the angle inside the shape." It is the exterior angle, 360 divided by the number of sides. A triangle needs turns of 120, not 60.
- "right(90) turns anticlockwise because y goes up." Right is always clockwise from the turtle's point of view. The y axis pointing up is a separate fact, and it is the opposite of the SVG convention from Lesson 17.
- "The window closed instantly, so the program crashed." Far more likely,
turtle.done()is missing, so the program finished and the window went with it. - "goto draws a line, so it cannot be used for moving." It draws only while the pen is down. Bracket it with
penupandpendownand it becomes a jump. - "A drawing cannot be checked, only looked at." Position and heading are numbers. Twenty lines of trigonometry will tell you every corner before you open a window, which is how the measurements on this page were produced.
The takeaway
The core of it: A turtle program is a loop and an angle. Get the angle from 360 divided by the number of sides, and the rest is decoration.
You can open a window, draw with forward, right, penup, goto and the fill pair, and you know why turtle.done() has to be the last line. You know the turtle starts at the centre facing east with y increasing upwards, and that this is the reverse of the SVG system you used a lesson ago. You can work out the turn for any polygon, you know it is the exterior angle, and you have seen one number, 72 against 144, decide between a pentagon and a star. And you can check a drawing with arithmetic rather than eyes, which is the only way to debug a picture.
The next two lessons make something you can play. The first ingredient arrives in the one you have already met: random, and a number the program refuses to tell you.
Sources
- Python Software Foundation. (2026). turtle: Turtle graphics (the turtle starting at (0, 0) facing east, degrees as the default angle unit, forward, right, penup, goto, pensize, color, begin_fill, end_fill, speed, hideturtle, write, and done() as the required last statement). docs.python.org
- Python Software Foundation. (2026). tkinter: Interface to Tcl/Tk for graphical user interfaces (the toolkit turtle draws with, and the source of the No module named '_tkinter' error). docs.python.org
- Wikipedia. (2026). Logo (programming language) (Feurzeig, Papert and Solomon, 1967, and turtle graphics as a teaching device). en.wikipedia.org
- Papert, S. (1980). Mindstorms: Children, Computers, and Powerful Ideas. Basic Books. (The argument for learning geometry by walking a path, which is what the turtle is.)
- Key terms
- turtle
- A standard library module that draws in a window by moving a pen you steer with forward and right.
- Heading
- The direction the turtle faces, in degrees, starting at 0 for east and decreasing as you turn right.
- Exterior angle
- The amount the turtle turns at a corner, equal to 360 divided by the number of sides for a regular polygon.
- penup and pendown
- The pair that lets the turtle move without drawing, so shapes can be placed apart from each other.
- begin_fill and end_fill
- The pair that brackets a path so turtle colours in the region it enclosed.
- turtle.done()
- The last statement of a turtle program. It keeps the window open until you close it.
- tkinter
- The graphics toolkit turtle is built on. If it is missing, the import fails with No module named '_tkinter'.
The Game That Said Not That One to the Right Answer
- Use randint, choice, sample and shuffle, and explain what seeding does to a random program.
- Build a guessing game with a loop, a comparison, a guess counter and a limit on attempts.
- Explain why halving the range finds any number from 1 to 100 in at most seven guesses.
Here is a guessing game that looks finished, and a run of it where the player guessed correctly and was told they had not:
import random
random.seed(4)
secret = random.randint(1, 20)
print("I am thinking of a number between 1 and 20.")
guess = input("Your guess: ")
while guess != secret:
print("Not that one.")
guess = input("Your guess: ")
print("Correct!")
Fed the guesses 7, then 12, then 8, it produced the output below. The traceback has been shortened to its final error line. The guesses do not appear after the prompts because they were fed in from a file rather than typed; when you type them, each one shows up where the cursor is:
I am thinking of a number between 1 and 20.
Your guess: Not that one.
Your guess: Not that one.
Your guess: Not that one.
Your guess: EOFError: EOF when reading a line
In this Python 3.14 example, the first randint(1, 20) call after seed 4 produces 8, so the third guess was right. The program said Not that one, asked again, ran out of input and stopped with an error. No typed guess can make the loop's condition false, although an input error or interruption can still stop the program.
A string is not a number, and == does not care
These six lines expose the mismatch:
guess = input("Your guess: ")
secret = 8
print(repr(guess), repr(secret))
print(type(guess), type(secret))
print(guess == secret)
print(int(guess) == secret)
Your guess: '8' 8
<class 'str'> <class 'int'>
False
True
There it is. When input returns successfully, it returns a string, as it has since Lesson 2, and '8' is not 8. The equality comparison is False without raising an error, so the original loop keeps asking until input fails or the program is interrupted. repr is the tool that made it visible: it shows a string with its quotes, so '8' and 8 stop looking identical.
Worth holding on to: when a comparison is always false and you cannot see why, print repr and type of both sides. It is two lines and it ends the argument.
The fix is int(input(...)), and with a counter added the game works:
import random
random.seed(4)
secret = random.randint(1, 20)
guesses = 0
print("I am thinking of a number between 1 and 20.")
while True:
guess = int(input("Your guess: "))
guesses = guesses + 1
if guess < secret:
print("Too low.")
elif guess > secret:
print("Too high.")
else:
print(f"Correct, in {guesses} guesses.")
break
I am thinking of a number between 1 and 20.
Your guess: Too high.
Your guess: Too low.
Your guess: Too low.
Your guess: Correct, in 4 guesses.
That was fed 10, 5, 7, 8. Two things changed besides the conversion. The loop is now while True with a break, because the test that ends the game is in the middle of the work rather than at the top, and telling the player too low or too high turns guessing into searching.
What random actually gives you
The default generator used by the random module is the Mersenne Twister. It is pseudo-random: it calculates a sequence from an internal state, and a seed initializes that state. This default generator is deterministic and unsuitable for cryptographic purposes. Computers can also obtain randomness from the operating system; the default generator used in these examples is the one we are discussing.
A seed helps you reproduce a bug while building a game. With the same Python version and the same sequence of calls in a single thread, resetting the seed replays the results. Some random algorithms can change between Python versions, so use Python 3.14 for these examples:
import random
random.seed(99)
print("randint(1, 6) ", random.randint(1, 6))
print("randrange(0, 10, 2)", random.randrange(0, 10, 2))
print("choice of a list ", random.choice(["north", "south", "east", "west"]))
print("sample of 6 from 49", sorted(random.sample(range(1, 50), 6)))
print("choices, weighted ", random.choices(["hit", "miss"], weights=[1, 3], k=12))
print("random() ", round(random.random(), 4))
print("uniform(1.5, 2.5) ", round(random.uniform(1.5, 2.5), 4))
random.seed(99)
rolls = []
for throw in range(20):
rolls.append(random.randint(1, 6))
print("twenty dice ", rolls)
counts = {}
random.seed(1)
for throw in range(6000):
face = random.randint(1, 6)
counts[face] = counts.get(face, 0) + 1
print("6000 dice, counts ", dict(sorted(counts.items())))
randint(1, 6) 4
randrange(0, 10, 2) 6
choice of a list south
sample of 6 from 49 [6, 9, 12, 15, 16, 39]
choices, weighted ['miss', 'miss', 'miss', 'miss', 'miss', 'miss', 'miss', 'miss', 'hit', 'miss', 'miss', 'miss']
random() 0.6824
uniform(1.5, 2.5) 1.6523
twenty dice [4, 4, 2, 5, 2, 2, 2, 2, 1, 3, 6, 4, 5, 6, 6, 5, 1, 5, 4, 2]
6000 dice, counts {1: 979, 2: 1002, 3: 1007, 4: 995, 5: 1045, 6: 972}
Eight functions, and each answers a different question:
| Call | Gives | Boundary or use |
|---|---|---|
randint(1, 6) | A whole number, a die roll | Yes: 1 and 6 can both come up |
randrange(0, 10, 2) | One of 0, 2, 4, 6, 8 | No: stops before 10, like range |
choice(list) | One item from a list | Any item |
sample(range, 6) | Six different items, no repeats | Lottery numbers |
choices(list, weights=weights, k=k) | k items, repeats allowed, with the supplied weights | Loaded dice; k must be named |
random() | A float from 0.0 up to but not including 1.0 | Probabilities |
uniform(1.5, 2.5) | A float between 1.5 and 2.5 | Rounding determines whether 2.5 can occur |
shuffle(list) | Nothing: it reorders the list in place | Dealing cards |
The line beginning twenty dice restarts with the same roll, 4, because it resets seed 99 and makes the same first call. After that, this loop makes different calls from the earlier example. A seed alone does not fix every output independently of what the program asks the generator to do.
sample draws without replacement: it cannot select the same position twice. The range 1 to 49 has unique values, so its sampled values are different too. If a list contains two copies of the same value, both can appear in a sample. To shuffle a list, write random.shuffle(cards) and then use cards; assigning the return value to cards would replace the list with None.
With weights 1 and 3, hit has probability 1/4 on each draw. Three hits in twelve is the expected count, not a quota. The one hit shown here does not by itself establish a bug or bias. In the 6000-roll example, the six counts are all within five percent of the expected 1000. That describes this run; another seed need not give those counts or stay inside that band.
For passwords or security tokens, the documentation sends you to secrets, which uses randomness supplied by the operating system. The default random generator is suitable for this guessing game, but not for protecting secrets.
The finished game, in three functions
One long block of code is hard to test, so the game gets split the way Lesson 8 taught: each function does one thing and is checkable on its own.
"""guess.py: the finished number guessing game."""
import random
LOW = 1
HIGH = 100
ALLOWED = 7
def ask_guess(prompt):
while True:
answer = input(prompt)
try:
return int(answer)
except ValueError:
print(f"'{answer}' is not a whole number.")
def play(secret, allowed):
for attempt in range(1, allowed + 1):
guess = ask_guess(f"Guess {attempt} of {allowed}: ")
if guess == secret:
return attempt
if guess < secret:
print("Too low.")
else:
print("Too high.")
return 0
def report(attempts, secret):
if attempts == 0:
print(f"Out of guesses. It was {secret}.")
elif attempts == 1:
print("First time. Nobody is that lucky.")
else:
print(f"Got it in {attempts} guesses.")
if __name__ == "__main__":
random.seed(11)
number = random.randint(LOW, HIGH)
print(f"I have a number from {LOW} to {HIGH}. You have {ALLOWED} guesses.")
report(play(number, ALLOWED), number)
Fed 50, then the word seventy, then 75, 63, 56, 60 and 58:
I have a number from 1 to 100. You have 7 guesses.
Guess 1 of 7: Too low.
Guess 2 of 7: 'seventy' is not a whole number.
Guess 2 of 7: Too high.
Guess 3 of 7: Too high.
Guess 4 of 7: Too low.
Guess 5 of 7: Too high.
Guess 6 of 7: Got it in 6 guesses.
Look at the two lines labelled Guess 2. The word seventy did not crash the program and did not cost the player an attempt, because ask_guess loops until it gets a number and only then returns, and the attempt counter lives in play, one level up. That division is the whole reason for having two functions instead of one.
And fed seven wrong guesses, 1 to 7:
I have a number from 1 to 100. You have 7 guesses.
Guess 1 of 7: Too low.
Guess 2 of 7: Too low.
Guess 3 of 7: Too low.
Guess 4 of 7: Too low.
Guess 5 of 7: Too low.
Guess 6 of 7: Too low.
Guess 7 of 7: Too low.
Out of guesses. It was 58.
play returns the attempt number on a win and 0 on a loss, and 0 is a safe signal here because no win can ever happen on attempt zero. report does nothing but turn that number into a sentence. Splitting the game this way means you can test report(0, 58) and report(1, 58) without playing anything.
Three capital-letter names at the top, LOW, HIGH and ALLOWED, are intended as constants. PEP 8 recommends capitals for these names, but Python still allows you to assign new values to them. Putting them in one place lets you change the game to 1 to 1000 with fifteen guesses by editing two lines instead of hunting through the file.
Seven guesses, and why not more
Is seven attempts for a hundred numbers generous or mean? Guessing 1, 2, 3 in order needs up to a hundred tries. Halving needs seven, and this program proves it by playing every possible game:
def guesses_needed(secret, low, high):
count = 0
while True:
middle = (low + high) // 2
count = count + 1
if middle == secret:
return count
if middle < secret:
low = middle + 1
else:
high = middle - 1
worst = 0
for secret in range(1, 101):
needed = guesses_needed(secret, 1, 100)
if needed > worst:
worst = needed
print("worst case over 1 to 100:", worst, "guesses")
print("guesses for 1, 50, 73, 100:", guesses_needed(1, 1, 100),
guesses_needed(50, 1, 100), guesses_needed(73, 1, 100),
guesses_needed(100, 1, 100))
total = 0
for secret in range(1, 101):
total = total + guesses_needed(secret, 1, 100)
print("average:", round(total / 100, 2))
for size in [10, 20, 100, 1000, 1000000]:
worst = 0
step = 1
while step <= size:
step = step * 2
worst = worst + 1
print(f"a range of {size:>7} needs at most {worst:>2} halving guesses")
worst case over 1 to 100: 7 guesses
guesses for 1, 50, 73, 100: 6 1 6 7
average: 5.8
a range of 10 needs at most 4 halving guesses
a range of 20 needs at most 5 halving guesses
a range of 100 needs at most 7 halving guesses
a range of 1000 needs at most 10 halving guesses
a range of 1000000 needs at most 20 halving guesses
Seven is exactly enough when you use the midpoint strategy and the clues are correct. That strategy is binary search. One guess can cover one possible secret. Two guesses can cover three: the first midpoint, plus one value on either side. Each additional guess changes the capacity from c to 2c + 1. The capacities are 1, 3, 7, 15, 31, 63 and 127. Six guesses cannot cover 100 numbers; seven can.
The counting loop doubles step until it is strictly greater than the range size, because the capacity is step - 1. The equals sign in step <= size matters: a range of 8 needs up to 4 guesses, not 3, and a range of 1 still needs 1 guess. Doubling a positive range size adds one to this worst-case count. A million numbers needs twenty. The reported average of 5.8 weights all 100 possible secrets equally; it is not a prediction for every player's average.
Notice the middle is (low + high) // 2 with integer division, from Lesson 3, because there is no such thing as guessing 50.5. And low = middle + 1 rather than low = middle is the off-by-one care from Lesson 6: the middle has just been ruled out, so leaving it in the range would let the program guess it twice and, in the worst case, loop forever.
Turning the game round
Let the computer guess and the same strategy becomes code:
low = 1
high = 100
tries = 0
print("Think of a number from 1 to 100. Answer h, l or c.")
while True:
middle = (low + high) // 2
tries = tries + 1
answer = input(f"Is it {middle}? ")
if answer == "c":
print(f"Found it in {tries} tries.")
break
elif answer == "h":
low = middle + 1
elif answer == "l":
high = middle - 1
else:
print("Answer h for higher, l for lower, c for correct.")
tries = tries - 1
if low > high:
print("Those answers contradict each other.")
break
Fed h, l, l, h, h, c, the whole conversation came out on one line, because the prompts have no newline of their own and the fed answers are not echoed:
Think of a number from 1 to 100. Answer h, l or c.
Is it 50? Is it 75? Is it 62? Is it 56? Is it 59? Is it 60? Found it in 6 tries.
Follow the arithmetic: 50, told higher, so the range becomes 51 to 100 and the middle is 75. Told lower, so 51 to 74 and the middle is 62. Lower again, 51 to 61, middle 56. Higher, 57 to 61, middle 59. Higher, 60 to 61, middle 60. Correct. Six questions for a hundred numbers.
The final check catches an empty interval. If the player's answers contradict each other, low passes high. Without that check the program could keep asking about an impossible interval, although c, an input error or an interruption could still stop it. tries = tries - 1 in the else branch keeps an invalid letter from adding to the count.
Common misconceptions
- "input gives a number when the person types a number." A successful call returns a string.
'8' == 8is False and no error is raised, which is what makes this bug so quiet. - "The default random generator is suitable for passwords." Its sequence is deterministic. Use secrets for passwords and security tokens.
- "Removing the seed gives every player a different game." Removing the fixed seed lets Python initialize the generator from operating-system randomness, or system time if that is unavailable. It does not guarantee different secret numbers. With only 100 possible secrets, repeats are allowed and expected across enough rounds.
- "randint(1, 6) behaves like range(1, 6)." It does not. randint includes both ends, range stops one short. Mixing them up is a classic way to make a die that never rolls six.
- "Seven guesses for a hundred numbers is unfair." Halving finds any of them in at most seven, and the average is 5.8. The limit is a puzzle, not a punishment.
- "while True is a sign of bad code." It can fit a loop whose exit test belongs in the middle of the work. The original bug was that no typed guess could make its loop condition false. An error could stop it, but a correct answer could not.
What to remember
Bottom line: Convert the input, then loop, then count, then limit. A game is a loop with a condition and a little bookkeeping.
You can draw numbers, select items and reorder a list using the random module. You know how the default generator uses a seed to reproduce the same sequence of calls in the same Python version, and why removing a fixed seed permits variation without guaranteeing unique games. You have built a guessing game in three functions where a typed word costs nothing, an attempt limit ends the game, and the report is a separate testable function. And you can explain why midpoint guesses find any secret from 1 to 100 in at most seven guesses and any secret from 1 to a million in twenty.
The next lesson keeps the loop and the input and throws away the number. Instead of a secret, the program holds a map, and instead of too high or too low, it says you are in a corridor and there are two doors.
Sources
- Python Software Foundation. (2026). random: Generate pseudo-random numbers. Python 3.14 documentation. See seed, choices, sample, uniform and Notes on Reproducibility for the qualifications used here. docs.python.org
- Python Software Foundation. (2026). secrets: Generate secure random numbers for managing secrets. Python 3.14 documentation. The introduction distinguishes operating-system randomness from the default random generator. docs.python.org
- Black, P. E. (2022, April 21). Binary search. NIST Dictionary of Algorithms and Data Structures. The definition describes midpoint narrowing until the value is found or the interval is empty; the exact guess counts here are derived and checked by the lesson's code. NIST
- Python Software Foundation. (2026). Built-in Functions. Python 3.14 documentation. See input, int and repr for string input, conversion and printable representations. docs.python.org
- van Rossum, G., Warsaw, B., and Coghlan, A. (2001, updated guidance). PEP 8: Style Guide for Python Code, Constants. Capitalization is a naming convention. peps.python.org
- Key terms
- Pseudo-random
- Produced by a deterministic algorithm. For the default generator, the seed and subsequent calls determine the sequence within a given Python version.
- randint(a, b)
- A whole number from a to b with both ends included, unlike range, which stops one short.
- sample
- Items drawn without replacement: each population position can be selected at most once. Repeated values in the population can still appear more than once.
- choices
- Items drawn with repeats allowed and optionally unequal weights, as in loaded dice.
- while True with break
- The loop shape to use when the test that ends the work sits in the middle of it rather than at the top.
- Binary search
- Halving the range of possibilities at every step, which finds one value in a hundred in at most seven tries.
- Constant
- A value intended to stay unchanged, conventionally assigned to a capitalized name at module level. Python does not enforce that intention.
- repr
- A built-in that shows a value as Python would write it, with quotes on strings, which is how a string 8 is told from a number 8.
A House Made of Dictionaries
- Store a game world as a dictionary of rooms, each holding its own dictionary of exits.
- Read a typed command, split it into words, and act on the verb.
- Carry an inventory in a list, gate a room behind an item, and save the game state to JSON.
In 1975 and 1976 Will Crowther, a programmer and caver, wrote a program that described a cave in words and let you type where to go next. Don Woods expanded it in 1977, and Colossal Cave Adventure became the first well known adventure game: no pictures, one or two words per command. You are about to build one, and the surprising part is how little new you need. A map is a dictionary. A command is a string you split. An inventory is a list. You have had all three since Module 3.
The map is a dictionary whose values are dictionaries
ROOMS = {
"hall": {
"text": "A cold stone hall. Dust on the floor, and one shoe.",
"exits": {"north": "library", "east": "kitchen"},
},
"library": {
"text": "Shelves to the ceiling. Something metal glints on a high shelf.",
"exits": {"south": "hall"},
},
"kitchen": {
"text": "A kitchen that has not cooked anything in years.",
"exits": {"west": "hall", "down": "cellar"},
},
}
def describe(room_name):
room = ROOMS[room_name]
print(room["text"])
print("Exits:", ", ".join(sorted(room["exits"])))
describe("hall")
print()
describe("library")
print()
print(ROOMS["kitchen"]["exits"]["down"])
print(len(ROOMS), "rooms")
A cold stone hall. Dust on the floor, and one shoe.
Exits: east, north
Shelves to the ceiling. Something metal glints on a high shelf.
Exits: south
cellar
3 rooms
Read the structure from the outside in. ROOMS is a dictionary whose keys are room names. Each value is another dictionary, with a text and an exits. And exits is a third dictionary, mapping a direction to the name of the room it leads to. So ROOMS["kitchen"]["exits"]["down"] reads as: the kitchen, its exits, the one called down, which is the cellar. Three square brackets, three lookups, left to right.
This is the shape that makes the whole game small. A dictionary answers the question what is at this key in one step, which is exactly the question a game asks constantly: what does this room look like, where does this direction lead. A list of rooms would force you to search it every move.
", ".join(sorted(room["exits"])) is worth unpacking, because three ideas from earlier lessons are stacked in it. Looping over a dictionary gives its keys, so sorted receives the direction names; sorted puts them in alphabetical order so the output does not wander about between runs; and join from Lesson 11 glues them with a comma and a space.
The loop that reads what you type
Add one variable for where the player is, and one loop:
here = "hall"
describe(here)
while True:
command = input("> ").strip().lower()
if command == "quit":
print("Suit yourself.")
break
exits = ROOMS[here]["exits"]
if command in exits:
here = exits[command]
describe(here)
else:
print("You cannot go that way.")
Fed north, north, south, east and quit, with the kitchen's down exit removed for the moment:
A cold stone hall. Dust on the floor, and one shoe.
Exits: east, north
> Shelves to the ceiling. Something metal glints on a high shelf.
Exits: south
> You cannot go that way.
> A cold stone hall. Dust on the floor, and one shoe.
Exits: east, north
> A kitchen that has not cooked anything in years.
Exits: west
> Suit yourself.
That is a working game in twelve lines. The second north failed because the library has only a south exit, and if command in exits caught it; without that test the program would have raised a KeyError and stopped, which is the difference between a game and a crash.
here is the only thing that changes as the player moves, and the whole map stays constant. A program with one variable holding which of several named situations it is in is a finite-state machine, and recognising that shape is useful well beyond games: menus, traffic lights and vending machines are all written this way.
.strip().lower() does a lot of quiet work. strip removes a stray space before or after what was typed, and lower means North, NORTH and north are all the same command. Two method calls, and a whole category of player complaints disappears.
Two words, not one
One-word commands run out fast, because take and go north are different shapes. Splitting solves it:
command = input("> ").strip().lower()
words = command.split()
if words == []:
continue
verb = words[0]
split with no argument breaks on any run of whitespace, so go north with three spaces still gives two words. words[0] is the verb and words[1], when there is one, is what the verb acts on. The if words == [] check handles the player pressing Enter on an empty line: continue jumps back to the top of the loop, and without it words[0] would raise an IndexError on nothing at all.
The finished game
Here it is whole, with items to pick up, a locked door, a dark room, a win and a save file.
"""adventure.py: a text adventure in one file."""
import json
import os
ROOMS = {
"hall": {
"text": "A cold stone hall. Dust on the floor, and one shoe.",
"exits": {"north": "library", "east": "kitchen"},
"item": None,
},
"library": {
"text": "Shelves to the ceiling. Something metal glints on a high shelf.",
"exits": {"south": "hall"},
"item": "brass key",
},
"kitchen": {
"text": "A kitchen that has not cooked anything in years.",
"exits": {"west": "hall", "down": "cellar"},
"item": "candle",
},
"cellar": {
"text": "Black as anything, and colder than the hall.",
"exits": {"up": "kitchen", "north": "garden"},
"item": None,
"needs": "brass key",
"dark": True,
},
"garden": {
"text": "Daylight, wet grass, and a gate standing open.",
"exits": {"south": "cellar"},
"item": None,
},
}
SAVE = "adventure.json"
def describe(here, carrying):
room = ROOMS[here]
if room.get("dark") and "candle" not in carrying:
print("It is too dark to see anything at all.")
print("Exits: up")
return
print(room["text"])
if room["item"] is not None:
print("You can see:", room["item"])
print("Exits:", ", ".join(sorted(room["exits"])))
def try_move(here, direction, carrying):
exits = ROOMS[here]["exits"]
if direction not in exits:
print("You cannot go that way.")
return here
target = exits[direction]
needed = ROOMS[target].get("needs")
if needed is not None and needed not in carrying:
print(f"The way is locked. You would need the {needed}.")
return here
if ROOMS[here].get("dark") and "candle" not in carrying and direction != "up":
print("Not in the dark. You would walk into something.")
return here
return target
def take(here, carrying):
item = ROOMS[here]["item"]
if item is None:
print("There is nothing here to take.")
elif item in carrying:
print("You already have it.")
else:
carrying.append(item)
ROOMS[here]["item"] = None
print(f"Taken: {item}")
return carrying
def save_game(here, carrying):
with open(SAVE, "w") as f:
json.dump({"here": here, "carrying": carrying}, f, indent=2)
print(f"Saved to {SAVE}.")
def load_game():
if not os.path.exists(SAVE):
print("No saved game found. Starting in the hall.")
return "hall", []
with open(SAVE) as f:
state = json.load(f)
print("Game loaded.")
return state["here"], state["carrying"]
def main():
here = "hall"
carrying = []
print("THE HOUSE. Type go north, take, look, inventory, save, load or quit.")
describe(here, carrying)
while True:
command = input("> ").strip().lower()
words = command.split()
if words == []:
continue
verb = words[0]
if verb == "quit":
print("You give up and go home.")
return
elif verb == "look":
describe(here, carrying)
elif verb == "inventory":
if carrying == []:
print("You are carrying nothing.")
else:
print("You are carrying:", ", ".join(carrying))
elif verb == "take":
carrying = take(here, carrying)
elif verb == "save":
save_game(here, carrying)
elif verb == "load":
here, carrying = load_game()
describe(here, carrying)
elif verb == "go" and len(words) > 1:
moved_to = try_move(here, words[1], carrying)
if moved_to != here:
here = moved_to
describe(here, carrying)
if here == "garden":
print("You are out. Well done.")
return
else:
print(f"I do not know how to {command}.")
if __name__ == "__main__":
main()
A whole game played, fed from a file: go east, go down, dance, go west, go north, take, inventory, go south, go east, go down, go north, go up, take, inventory, go down, go north.
THE HOUSE. Type go north, take, look, inventory, save, load or quit.
A cold stone hall. Dust on the floor, and one shoe.
Exits: east, north
> A kitchen that has not cooked anything in years.
You can see: candle
Exits: down, west
> The way is locked. You would need the brass key.
> I do not know how to dance.
> A cold stone hall. Dust on the floor, and one shoe.
Exits: east, north
> Shelves to the ceiling. Something metal glints on a high shelf.
You can see: brass key
Exits: south
> Taken: brass key
> You are carrying: brass key
> A cold stone hall. Dust on the floor, and one shoe.
Exits: east, north
> A kitchen that has not cooked anything in years.
You can see: candle
Exits: down, west
> It is too dark to see anything at all.
Exits: up
> Not in the dark. You would walk into something.
> A kitchen that has not cooked anything in years.
You can see: candle
Exits: down, west
> Taken: candle
> You are carrying: brass key, candle
> Black as anything, and colder than the hall.
Exits: north, up
> Daylight, wet grass, and a gate standing open.
Exits: south
You are out. Well done.
Five moments in that transcript are the design of the game, and each is two or three lines of code.
The locked door. ROOMS["cellar"]["needs"] is the string "brass key", and try_move looks it up on the target room, not the current one, then checks the inventory. The player who walked straight to the cellar was told which item to look for, which is far kinder than a flat refusal and costs one f-string.
room.get("needs") rather than room["needs"]. Only the cellar has that key. Square brackets would raise a KeyError in every other room, so get is used, which returns None when the key is absent. That is the Lesson 10 lesson doing real work: get is for keys that might not be there.
The dark room. Two separate places check it, and they do different jobs. In describe, darkness hides the room's text. In try_move, darkness blocks every direction except up, which is the way you came. Split like that, a player can always escape a dark room, which is a rule worth keeping: never let a game trap somebody with no move at all.
The item moves. take appends to carrying and then sets ROOMS[here]["item"] = None. Forget the second line and the brass key can be picked up forever, which is the commonest bug in a first adventure game. An item exists in exactly one place: a room or your hands.
The unknown command. The final else answers anything unrecognised. Every branch of that chain ends somewhere, which is why typing dance produced a sentence instead of silence.
Saving, and the bug the save file revealed
Two lines of json from Lesson 18 save the game. Here is a short run that took the key and saved:
> Shelves to the ceiling. Something metal glints on a high shelf.
You can see: brass key
Exits: south
> Taken: brass key
> Saved to adventure.json.
> You give up and go home.
And the file it wrote:
{
"here": "library",
"carrying": [
"brass key"
]
}
Then a fresh run of the program, typing load and inventory:
THE HOUSE. Type go north, take, look, inventory, save, load or quit.
A cold stone hall. Dust on the floor, and one shoe.
Exits: east, north
> Game loaded.
Shelves to the ceiling. Something metal glints on a high shelf.
You can see: brass key
Exits: south
> You are carrying: brass key
> You give up and go home.
Read those last lines again. The player is carrying the brass key, and the library says the brass key is still lying there. Take it again and there would be two.
This is a real bug, and it is instructive about what a save file is for. ROOMS is rebuilt from the source code every time the program starts, so the emptied library fills back up. The save recorded where the player was and what they held, and forgot that the world had changed. A save file has to hold everything that changed during play, not just the parts you were thinking about. The fix is to save the items too:
json.dump({"here": here, "carrying": carrying,
"items": {name: ROOMS[name]["item"] for name in ROOMS}}, f, indent=2)
and on loading, write those values back into ROOMS. The general rule worth taking from this: the state of a program is every variable that can change, and a saved game that stores some of them will contradict itself.
Common misconceptions
- "A bigger game needs more code." Adding a sixth room to this game is four lines of data and no new code at all. That is what putting the world in a dictionary buys you.
- "The room descriptions belong in the print statements." Keeping text in the data and printing it from one function means you can add rooms without touching the loop. Mixing them is why beginner games become impossible to extend.
- "take only needs to add the item to the inventory." It must also remove it from the room, or the item exists twice over.
- "ROOMS[room]['needs'] is fine, since the cellar has that key." Every other room does not, and the lookup runs for every room. Use get and test for None.
- "Saving the position is saving the game." The save has to cover every changed variable. This one missed the room items and contradicted itself the moment it was loaded.
- "A player typing nonsense will crash the game." Only if no branch catches it. A final else that answers anything unknown is one line and makes the program unbreakable by typing.
Where this leaves us
Worth holding on to: Put the world in data and the rules in code. Then a new room is data, and the code never grows.
You can hold a map as a dictionary of rooms whose exits are themselves dictionaries, and reach any part of it with a chain of square brackets. You can read a command, strip and lower it, split it into a verb and a target, and route it through a chain of branches that always ends in an else. You can carry an inventory in a list, gate a room behind an item with get, hide a dark room's text while still letting the player leave, and finish a game with a win. You know that a save file must record every changed part of the state, because you have seen a save that did not, and read the contradiction it produced.
That closes Module 6. The last module turns the direction round: instead of writing programs, you read one somebody else wrote, and then you plan and finish something of your own.
Sources
- Python Software Foundation. (2026). The Python Tutorial, 5: Data Structures (nested dictionaries, dictionary membership with in, list append, and looping over dictionary keys). docs.python.org
- Python Software Foundation. (2026). Built-in Types: str.split, str.strip, str.lower and str.join; dict.get. docs.python.org
- Wikipedia. (2026). Colossal Cave Adventure (Crowther's 1975 to 1976 original, Woods's 1977 expansion, and one and two word commands). en.wikipedia.org
- Severance, C. (2016). Python for Everybody: Exploring Data in Python 3, chapter 9: Dictionaries (dictionaries as lookup tables, and the get method with a default). py4e.com
- Key terms
- Nested dictionary
- A dictionary whose values are themselves dictionaries, reached with a chain of square brackets.
- State
- Every variable that can change while a program runs. Here it is the room you are in and what you carry.
- Finite-state machine
- A program with one variable naming which of several situations it is in, and rules for moving between them.
- Parser
- The part of a program that turns typed text into a verb and its target, usually with split.
- Inventory
- A list of what the player carries. An item belongs either to a room or to the inventory, never both.
- dict.get
- A lookup that returns None, or a default you give it, instead of raising KeyError when the key is absent.
- continue
- A statement that abandons the rest of this pass of the loop and goes back to the top, used here for an empty line.
- Save file
- A record of the whole changed state. Leave part of it out and the loaded game contradicts itself.
Module 7: Reading What Others Wrote, and Finishing Something of Your Own
Work out what an unfamiliar program does, then plan, build, test and finish a project you choose yourself against a stated rubric.
Nggnpx ng qnja!
- Work out what an unfamiliar program does by running it, tracing one value, and renaming as you go.
- Name the things that make code readable: docstrings, meaningful names, constants and small functions.
- Read a standard library function's own source and documentation to answer a question about behaviour.
Somebody hands you this and says it works. What does it do?
def m(s):
t = ""
for c in s:
if c.isalpha():
b = 65
if not c.isupper():
b = 97
t = t + chr((ord(c) - b + 13) % 26 + b)
else:
t = t + c
return t
print(m("Attack at dawn!"))
print(m(m("Attack at dawn!")))
print(m("Why did the chicken cross the road?"))
print(m("Jul qvq gur puvpxra pebff gur ebnq?"))
Nggnpx ng qnja!
Attack at dawn!
Jul qvq gur puvpxra pebff gur ebnq?
Why did the chicken cross the road?
Line two is the clue. Running m twice gives back what you started with, and line four decodes what line three encoded. You have found ROT13 without being told: a Caesar cipher with a shift of thirteen, its own inverse because 26 is twice 13. In the 1980s Usenet readers used it to hide punchlines and spoilers, which is why the chicken's answer needed decoding.
Notice how little reading that took. You ran it and looked at the output. Most code you meet in your life will be code you did not write, and this lesson is the method for meeting it.
Five questions, in this order
What does it print? Run it. An unfamiliar program is a black box until you see one input become one output, and the fastest path to understanding is almost never reading from line 1.
Where does it start? Find the bottom of the file, or the if __name__ == "__main__": block. Function definitions do not run; they wait. The calls are where the program actually begins.
What are the names telling me? Good names shorten this step to nothing. Bad names, like m, t, c and b above, mean you have to earn the meaning from the code, which is the next question.
What happens to one value? Pick one small input and follow it. Not the whole string: one character. This is the step beginners skip and experienced programmers never do.
What changes if I change one thing? Edit a number, add a print, run it again. The program is on your machine and cannot be broken in a way that matters.
Following one character through
Question four, done in code rather than in your head:
for c in "At!":
if c.isalpha():
b = 65
if not c.isupper():
b = 97
print(c, "ord", ord(c), "base", b, "offset", ord(c) - b,
"plus 13 mod 26", (ord(c) - b + 13) % 26,
"becomes", chr((ord(c) - b + 13) % 26 + b))
else:
print(c, "not a letter, kept as it is")
A ord 65 base 65 offset 0 plus 13 mod 26 13 becomes N
t ord 116 base 97 offset 19 plus 13 mod 26 6 becomes g
! not a letter, kept as it is
Now the mystery numbers have names. 65 is ord("A") and 97 is ord("a"), from Lesson 11, so b is the number the alphabet starts at for this letter's case. Subtracting it turns a character code into a position in the alphabet, 0 for A and 19 for t. Adding 13 and taking the remainder on division by 26 moves thirteen places and wraps round the end, which is why t, position 19, becomes position 6, which is g. Adding the base back turns the position into a character again.
Three of those steps are a pattern worth naming: subtract the base, do the arithmetic, add the base back. Any time you see that shape you are looking at arithmetic on positions rather than on character codes.
And the self-inverse property falls out of the arithmetic. Thirteen places twice is twenty-six places, and twenty-six modulo twenty-six is zero. One function, no separate decoder.
Renaming is how you prove you understood
Here is the same algorithm written so nobody has to do any of that work:
"""rot13.py: shift every letter thirteen places, leaving everything else alone."""
ALPHABET_SIZE = 26
SHIFT = 13
def rot13(text):
"""Return text with each letter moved 13 places round the alphabet."""
result = ""
for character in text:
if character.isalpha():
base = ord("A")
if character.islower():
base = ord("a")
offset = (ord(character) - base + SHIFT) % ALPHABET_SIZE
result = result + chr(base + offset)
else:
result = result + character
return result
if __name__ == "__main__":
secret = rot13("Attack at dawn!")
print(secret)
print(rot13(secret))
assert rot13(rot13("anything at all")) == "anything at all"
print("Self test passed.")
Nggnpx ng qnja!
Attack at dawn!
Self test passed.
Identical behaviour, and the second version needs no detective work. Four changes did it.
Names that say what the thing is. text, character, result, base, offset. None of them is clever and all of them are correct.
The magic numbers became named constants. SHIFT = 13 and ALPHABET_SIZE = 26 at the top, in capitals as PEP 8 asks. A reader can now see that this is a general Caesar cipher with the shift set to 13, and change it to 5 in one place.
65 and 97 became ord("A") and ord("a"). Exactly the same values, computed instead of remembered, and the reader needs no ASCII table.
Docstrings. PEP 257 defines a docstring as a string literal that is the first statement in a module or function, and it becomes that object's __doc__ attribute. Its advice on wording is precise and easy to follow: a one-line docstring is a phrase ending in a period that prescribes the effect as a command, so Return text with each letter moved, not Returns the text.
One thing to notice about the assert: it is the Lesson 14 habit, and it tests the property rather than a particular output. Double application returns the original, for any input.
A program you have inherited
Reading is not only for admiring. This one arrived with a complaint attached: a student says their mark of 70 was recorded as a B.
D = [("Ada", 71), ("Blaise", 70), ("Cleo", 69), ("Dai", 50), ("Eve", 49)]
def g(m):
if m > 70:
return "A"
elif m > 60:
return "B"
elif m > 50:
return "C"
else:
return "D"
def r(d):
t = 0
for p in d:
t = t + p[1]
print(f"{p[0]:<8}{p[1]:>4} {g(p[1])}")
print(f"{'mean':<8}{t / len(d):>4.1f}")
r(D)
Ada 71 A
Blaise 70 B
Cleo 69 B
Dai 50 D
Eve 49 D
mean 61.8
Apply the questions. It prints a table of names, marks and letters, plus a mean, so g grades and r reports. D is a list of tuples, and p[0] and p[1] are a name and a mark, which is what the f-strings confirm. The program starts at the last line.
The complaint is correct, and it is visible in the output: Blaise on 70 got a B, and Dai on exactly 50 got a D. Every boundary is wrong by one mark, because > excludes the boundary value that a grade table means to include. Here are both versions side by side at every boundary:
mark old new
49 D D
50 D C
51 C C
59 C C
60 C B
69 B B
70 B A
71 A A
Only the boundary marks differ, which is exactly the signature of a wrong comparison operator, and exactly why Lesson 14 insisted that a test suite include the boundary values rather than a comfortable 65.
Two repairs are needed and they are different in kind. The first is one character, four times: > becomes >=. The second is that the program said nothing about which it meant. Nowhere does it state that 70 and above is an A, so the bug could not be spotted by reading, only by testing. A docstring saying Return the letter grade for a mark, where 70 and above is an A turns an invisible bug into an obvious one.
Key idea: a comment or docstring that states the intention is what makes a bug findable. Code can only tell you what it does; it cannot tell you what it was supposed to do.
What readable code gives you, item by item
| Feature | The version without it | The version with it |
|---|---|---|
| Names | g(m), p[1], t | letter_for(mark), mark, total |
| Docstring | You infer the rule from the code | The rule is stated, so a wrong rule shows |
| Named constants | 65, 97, 13, 26 in the middle of a line | SHIFT, ALPHABET_SIZE, changed in one place |
| Small functions | One block doing four jobs | Four testable functions |
| A main guard | Demo code runs on import | Self test runs only when launched |
| An assert | You trust it | It proves itself every run |
A comment earns its place by saying why, not what. i = i + 1 # add one to i is noise. i = i + 1 # the header row is not data is the reason somebody wrote the line, and it is the thing you cannot recover from the code.
Reading Python's own code
The standard library is Python, written by people, and you can read it. The inspect module will hand you the source of most functions:
import inspect
import statistics
print(inspect.getsource(statistics.mean))
That printed the real function. Three of its docstring examples are replaced by dots below, to keep the block short:
def mean(data):
"""Return the sample arithmetic mean of data.
>>> mean([1, 2, 3, 4, 4])
2.8
...
If data is empty, StatisticsError will be raised.
"""
T, total, n = _sum(data)
if n < 1:
raise StatisticsError('mean requires at least one data point')
return _convert(total / n, T)
Several things are worth noticing in eight lines of professional code. The docstring comes first and follows PEP 257: Return the sample arithmetic mean of data, prescribing rather than describing. The examples inside it begin with three angle brackets, which is the format the doctest module runs as tests, so those examples are checked rather than decorative. The empty-data case is documented and then handled explicitly with a raise. And two names start with an underscore, _sum and _convert, which is the convention for a helper that is not part of the public interface: a signal to readers rather than a rule Python enforces.
The same habit applies to documentation. A library page is not an essay to read from the top. Find the function, read its signature for what it takes and gives, read the one-line summary, then read the examples, and only then the prose. Four steps, usually under a minute, and it answers far more questions than searching for someone who has had your problem.
Common misconceptions
- "Read a program from line 1 to the end." Definitions only wait to be called. Find what actually runs, usually at the bottom or under the main guard, and read outwards from there.
- "Short names make code faster." They make it shorter to type and slower to understand. Python does not care either way.
- "If I cannot see the bug, I need to read harder." Trace one value, or print one. The ROT13 arithmetic became obvious the moment three characters were printed with their intermediate numbers.
- "A comment should explain what the line does." The line already does that. A comment should say why it is there, which is the part that cannot be recovered from the code.
- "The standard library is too advanced to read." statistics.mean is eight lines. Reading real code is how you learn what good code looks like, and inspect.getsource puts it in front of you.
- "A program with no comments and clear names is badly documented." Clear names and a docstring stating the rule are worth more than paragraphs. The inherited grader failed because it stated no rule, not because it had few comments.
Summing up
Key idea: Run it, find the entry point, trace one value, rename as you understand. Reading code is an activity, not a stare.
You can take an unfamiliar program, work out what it does from its output, follow a single character through its arithmetic, and rewrite it so the next reader needs none of that effort. You know what makes the difference: names that say what a thing is, magic numbers turned into named constants, a docstring that prescribes the effect, small functions, a main guard, and an assert that proves the property. You have found a real boundary bug in inherited code by testing every boundary rather than by reading, and you know why the missing docstring was as much the problem as the wrong operator. And you can open the standard library's own source and read it.
One lesson remains, and it is yours. You choose what to build, and the page gives you a rubric to build it against and one finished project to measure yours by.
Sources
- van Rossum, G., Warsaw, B., and Coghlan, N. (2001, revised). PEP 8: Style Guide for Python Code (naming conventions, constants in capitals, and comments that explain intent). Python Software Foundation. peps.python.org
- Goodger, D., and van Rossum, G. (2001). PEP 257: Docstring Conventions (a docstring as the first statement, becoming __doc__, and the rule that a one-line docstring prescribes the effect as a command rather than describing it). Python Software Foundation. peps.python.org
- Python Software Foundation. (2026). inspect: Inspect live objects (getsource, used here to read statistics.mean). docs.python.org
- Wikipedia. (2026). ROT13 (a Caesar cipher of shift 13, its own inverse since 26 is twice 13, and its use on Usenet to hide spoilers). en.wikipedia.org
- Key terms
- Entry point
- The place a program actually starts doing work: the calls at the bottom of the file or under the main guard.
- Tracing
- Following one value through a program step by step, usually by printing the intermediate results.
- Magic number
- A bare number in the middle of an expression whose meaning is not stated, such as 65 or 26.
- Docstring
- A string as the first statement of a module or function. PEP 257 asks it to prescribe the effect as a command.
- ROT13
- A Caesar cipher with a shift of thirteen, which decodes itself because thirteen twice is a full alphabet.
- Boundary value
- An input exactly on a decision line, such as a mark of 70, where a wrong comparison operator shows up.
- Leading underscore
- A naming convention marking a helper as internal to a module. It is a signal to readers, not a rule Python enforces.
- inspect.getsource
- A standard library call that prints the source code of a function, which makes the library itself readable.
Eight Cards, Three Right, and a File That Remembers Which One You Missed
- Choose a project you can describe in one sentence and finish in a week.
- Build it in working slices, testing each before adding the next.
- Score your own program against a rubric, and say honestly where it loses marks.
Two runs of a finished program, a few minutes apart. Nothing was typed between them except the answers:
Self test passed.
8 cards in the deck, 0 marked hard.
What does open(name; "a") do to a file? Right.
Which module gives you randint? Right.
What is 7 // 2? No, it is 3.
What type does input() always return? Right.
Score: 3 of 4 (75 percent)
Still to learn, worst first:
1x What is 7 // 2
Self test passed.
8 cards in the deck, 1 marked hard.
What is 7 // 2? Right.
Which statement runs a module's self test only when launched? Right.
What does a dictionary lookup use instead of a position? Right.
What does len("abc") give? Right.
Score: 4 of 4 (100 percent)
Nothing outstanding. Deck clear.
The card missed in the first run came back first in the second, because the program wrote it down. That is 110 lines of Python using nothing you have not met: a file, a list of dictionaries, a loop, some functions, random, json and three asserts. This lesson builds it in front of you, scores it against a rubric, and then hands both the rubric and the method to you for a project of your own.
Choosing something you will actually finish
The commonest way a first project dies is being too big to start. Three tests, all of which your idea has to pass.
One sentence. If you cannot say what the program does in one sentence with no and then in it, it is two projects. The example above is: it asks me questions from a file and remembers the ones I get wrong.
A version 1 that runs in an hour. Not the finished thing: the smallest version that does something visible. For the flashcards, version 1 was eleven lines that read the file and printed how many cards it found.
You want the output. Not the mark, the output. A program you would use is a program you will debug at ten at night; one you were assigned is a program you will abandon at the first traceback.
| Project | One sentence | Uses from this course |
|---|---|---|
| Flashcard trainer | Quizzes me from a file and tracks what I miss | Files, dictionaries, random, json, functions |
| Homework tracker | Lists what is due and how many days are left | datetime, json, sorting, f-strings |
| Text adventure of your own | Walks the player through a map you invented | Nested dictionaries, a parser, a save file |
| Data report | Summarises a CSV I found and charts it | Files, CSV, dictionaries, a text chart, SVG |
| Cipher workbench | Encodes, decodes and cracks a Caesar cipher | Strings, ord and chr, loops, letter counting |
| Turtle drawing generator | Draws a pattern from numbers I choose | turtle, loops, functions, angles |
Pick one of those or invent your own. The rubric below does not care which.
The rubric you are building against
Read this before you write any code, because half of it is about decisions you make in the first twenty minutes.
| Criterion | Marks | What full marks looks like |
|---|---|---|
| It runs | 10 | Runs from a clean start with no crash, on input a stranger would try. |
| It does what it says | 15 | The one-sentence description is true of the finished program. |
| Functions | 15 | Four or more functions, each doing one nameable job, most returning a value. |
| Data structures | 10 | A list or dictionary chosen because it fits, and you can say why. |
| Input handling | 10 | A typed word where a number belongs does not stop the program. |
| Persistence or real data | 10 | Reads or writes a file, or works on data you did not invent. |
| Tests | 10 | At least three asserts, including one boundary case, run by the file itself. |
| Readability | 10 | Names that explain themselves, a module docstring, constants in capitals. |
| Honest limitations | 10 | A short list of what it cannot do, written by you, and correct. |
The last row is the one people leave out and the one that separates a student program from a professional one. A program with three known limitations written down is more trustworthy than one claiming none.
Planning, on paper, in four minutes
Lesson 4 gave you decomposition and pseudocode. Here they are on this project. The jobs, each a line:
read the deck from a file
remember which questions were missed before
choose a few cards, hardest first
ask one card and say whether it was right
count the score and update the misses
print a report
save the misses
Seven lines, and every one of them became a function with almost the same name. That is not a coincidence; it is what decomposition is for. Then the pseudocode for the one job with real thinking in it, choosing the round:
split the cards into hard (missed before) and easy
sort hard by how many times it was missed, worst first
shuffle easy so the order changes between runs
join them, hard first, and take the first few
Notice what the plan does not contain: any Python. Deciding that the hard cards come first and the easy ones are shuffled is a decision about teaching, not about syntax, and it is much cheaper to change in a list of four sentences than in twenty lines of code.
Slice 1: get the data in
The deck is a CSV, question and answer per line, with a header:
question,answer
What type does input() always return,str
What is 7 // 2,3
What does len("abc") give,3
Which module gives you randint,random
What does a dictionary lookup use instead of a position,key
What keyword ends a loop early,break
What does open(name; "a") do to a file,append
Which statement runs a module's self test only when launched,main
def load_cards(filename):
cards = []
with open(filename) as f:
header = f.readline()
for line in f:
line = line.strip()
if line == "":
continue
parts = line.split(",")
cards.append({"question": parts[0], "answer": parts[1]})
return cards
deck = load_cards("cards.csv")
print(len(deck), "cards loaded")
print(deck[0])
print(deck[-1]["answer"])
8 cards loaded
{'question': 'What type does input() always return', 'answer': 'str'}
main
That is a working program after eleven lines, and it is worth stopping to notice what it proves: the file is where you think, the header is being skipped, the split is working, and a card is a dictionary with two keys. Four assumptions checked in one run. Every further slice is added to something known to work, which is the difference between debugging one new idea and debugging seven at once.
A list of dictionaries is the structure here, rather than a dictionary keyed by question, because the cards have an order that matters and the same question could in principle appear twice. Being able to say that sentence is the Data structures row of the rubric.
The finished program
"""flashcards.py: quiz yourself from a CSV deck, and remember what you got wrong.
Usage: put cards.csv beside this file, one card per line, question,answer.
Progress is kept in progress.json so a later run can drill the hard cards first.
"""
import json
import os
import random
DECK = "cards.csv"
PROGRESS = "progress.json"
ROUND_SIZE = 4
def load_cards(filename):
"""Return a list of card dictionaries read from a two-column CSV."""
cards = []
with open(filename) as f:
f.readline()
for line in f:
line = line.strip()
if line == "":
continue
parts = line.split(",")
if len(parts) != 2:
print(f"Skipping a malformed line: {line}")
continue
cards.append({"question": parts[0], "answer": parts[1]})
return cards
def load_progress(filename):
"""Return a dictionary of question to number of times it was missed."""
if not os.path.exists(filename):
return {}
with open(filename) as f:
return json.load(f)
def save_progress(filename, misses):
with open(filename, "w") as f:
json.dump(misses, f, indent=2)
def is_right(given, expected):
"""Return True when the answer matches, ignoring case and outer spaces."""
return given.strip().lower() == expected.strip().lower()
def pick_round(cards, misses, size, seed):
"""Return up to size cards, hardest first, then a random selection."""
random.seed(seed)
hard = []
easy = []
for card in cards:
if misses.get(card["question"], 0) > 0:
hard.append(card)
else:
easy.append(card)
hard.sort(key=lambda card: misses[card["question"]], reverse=True)
random.shuffle(easy)
return (hard + easy)[:size]
def ask(card):
"""Ask one card and return True if it was answered correctly."""
given = input(card["question"] + "? ")
if is_right(given, card["answer"]):
print("Right.")
return True
print(f"No, it is {card['answer']}.")
return False
def run_round(cards, misses):
score = 0
for card in cards:
if ask(card):
score = score + 1
if card["question"] in misses:
misses[card["question"]] = misses[card["question"]] - 1
if misses[card["question"]] <= 0:
del misses[card["question"]]
else:
misses[card["question"]] = misses.get(card["question"], 0) + 1
return score, misses
def report(score, total, misses):
percent = round(100 * score / total)
print(f"Score: {score} of {total} ({percent} percent)")
if misses == {}:
print("Nothing outstanding. Deck clear.")
return
print("Still to learn, worst first:")
for question in sorted(misses, key=lambda q: misses[q], reverse=True):
print(f" {misses[question]}x {question}")
def self_test():
assert is_right(" STR ", "str")
assert not is_right("int", "str")
cards = [{"question": "a", "answer": "1"}, {"question": "b", "answer": "2"}]
chosen = pick_round(cards, {"b": 3}, 2, seed=1)
assert chosen[0]["question"] == "b"
assert len(pick_round(cards, {}, 1, seed=1)) == 1
print("Self test passed.")
def main():
cards = load_cards(DECK)
misses = load_progress(PROGRESS)
print(f"{len(cards)} cards in the deck, {len(misses)} marked hard.")
chosen = pick_round(cards, misses, ROUND_SIZE, seed=5)
score, misses = run_round(chosen, misses)
report(score, len(chosen), misses)
save_progress(PROGRESS, misses)
if __name__ == "__main__":
self_test()
main()
The first run, fed the answers append, RANDOM, 4 and str with spaces round it, produced the transcript at the top of this lesson, and wrote this file:
{
"What is 7 // 2": 1
}
The second run read that file and drilled the missed card first, and getting it right emptied the record:
{}
Five decisions inside that program are worth stating, because each is the kind of thing a rubric is really asking about.
is_right forgives the player. given.strip().lower() == expected.strip().lower() accepted RANDOM and accepted str with a space either side. A trainer that marks you wrong for a capital letter does not get used twice.
A right answer reduces the miss count rather than clearing it. One lucky recall is not learning. Miss a card twice and you need two correct answers before it leaves the list, which is why the count is a number and not a set.
The seed is an argument, not a fixed value. pick_round(cards, misses, size, seed) takes the seed, so the self test can ask for a predictable round while the real game passes something that changes. Hard-coding random.seed(5) inside the function would make the test easy and the program boring.
The malformed line is skipped, not fatal. Run load_cards on a deck with a broken line and a blank line in it, and importing the module ran no quiz at all, thanks to the main guard:
Skipping a malformed line: this line has no comma at all
Skipping a malformed line: What does open(name, "a") do,append
2 cards survived
What is 7 // 2 -> 3
Which module gives randint -> random
The second skipped line is a real limitation, not a win. That card was perfectly good; it was thrown out because the question contains a comma, so split(",") found three parts. It is why the deck above writes open(name; "a") with a semicolon, which is a workaround rather than a fix. The real fix is the csv module, whose whole purpose is handling quoting and embedded delimiters that a plain split cannot.
Marking the example honestly
| Criterion | Awarded | Why |
|---|---|---|
| It runs | 10 of 10 | Runs with no progress file, with one, and with a broken deck. |
| It does what it says | 15 of 15 | Asks from a file, remembers misses, drills them first. |
| Functions | 15 of 15 | Nine functions, seven returning values, each with one job. |
| Data structures | 10 of 10 | A list of dictionaries for ordered cards, a dictionary for counts. |
| Input handling | 6 of 10 | Any typed answer is accepted, but a malformed CSV line is silently dropped rather than reported at the end. |
| Persistence | 10 of 10 | Reads a CSV, reads and writes JSON. |
| Tests | 8 of 10 | Four asserts including the case and space boundary, but nothing tests run_round. |
| Readability | 9 of 10 | Module docstring, constants, clear names; save_progress has no docstring. |
| Honest limitations | 10 of 10 | Listed below, including the comma problem. |
That is 93 out of 100, and the four marks lost on input handling are the most useful part of the table. Writing down where your own program is weak is a skill, and it is much easier than pretending.
The limitations, in the program's own words: a question containing a comma cannot be stored, because the loader splits on commas; answers must match exactly once case and outer spaces are ignored, so there is no allowance for a spelling mistake; the round size is fixed at four; and progress is keyed on the question text, so editing a question in the deck loses its history.
Finishing, which is a separate job from building
A program that works on your machine is about eighty percent of a project. The rest:
Run it from clean. Delete the progress file, the saved game, whatever your program creates, and run it as a stranger would. This is how the flashcard program's missing-file branch came to exist at all.
Try the input a stranger would try. Empty input, a word where a number goes, a very large number, a file that is not there. Each one you survive is a mark.
Read your own code once, slowly. Apply Lesson 22 to yourself: are the names right, is any number magic, does each function do one thing. Rename freely. It is the cheapest improvement available.
Write six lines of usage. What it does, how to run it, what files it needs, what it makes, what it cannot do. In the flashcard program that is the module docstring at the top, which is where a reader looks first.
Stop. There is always one more feature. A finished project with three limitations written down beats an unfinished one with none, and the whole point of the rubric is that it tells you when you are done.
Common misconceptions
- "A good project is a big project." A small program that runs, is readable and admits its limits scores far higher than an ambitious one that crashes. The rubric has 10 marks for running and none for ambition.
- "Plan the whole thing, then write the whole thing." Plan the jobs, then build one slice at a time and run it. Eleven working lines beat a hundred untested ones.
- "Tests are for after it works." The asserts here shaped the code: needing to test pick_round is why the seed became a parameter instead of a fixed number inside the function.
- "Handling bad input means never crashing." It means failing usefully. Skipping a malformed line and saying so is good; skipping it silently, as this program does, costs marks.
- "Listing what my program cannot do makes it look worse." It makes it look finished. An unstated limitation is a bug waiting to be found by somebody else.
- "If it runs, it is readable enough." The inherited grader in Lesson 22 ran perfectly and was wrong at every boundary. Running and being understandable are separate properties.
What to carry forward
Why this matters: One sentence, one hour to version 1, one slice at a time, and a written list of what it cannot do. That is how a program gets finished rather than abandoned.
You have a rubric with nine criteria, a worked example built in slices with every output on this page produced by running it, and an honest score of 93 with the reasons for the seven missing marks. You know how to choose a project that survives contact with a traceback: one sentence, a version 1 within the hour, and output you actually want. You know that finishing is its own job, made of running from clean, trying a stranger's input, rereading your own names and writing six lines of usage.
Twenty-three lessons ago you printed one line. Since then you have built a cipher, repaired a program with four bugs in it, read a CSV of NASA temperature data and drawn it twice, opened a window and drawn a star, and written a game with a locked door and a save file. Everything you used is installed on your machine already, and the standard library index is the part of the documentation worth reading for pleasure. Pick your project. Version 1 within the hour.
Sources
- Python Software Foundation. (2026). The Python Tutorial (the whole tutorial, as the reference to return to while building a project). docs.python.org
- Python Software Foundation. (2026). csv: CSV File Reading and Writing (reader and DictReader, and the quoting rules a plain split on commas cannot handle). docs.python.org
- van Rossum, G., Warsaw, B., and Coghlan, N. (2001, revised). PEP 8: Style Guide for Python Code (constants in capitals, naming, and layout, used for the Readability row of the rubric). Python Software Foundation. peps.python.org
- Sweigart, A. (2019). Automate the Boring Stuff with Python: Practical Programming for Total Beginners (2nd ed.). No Starch Press. (On choosing small automation projects a beginner will finish.) automatetheboringstuff.com
- Key terms
- Slice
- A version of the program that runs and does something visible, added to only once it works.
- Rubric
- The stated criteria a project is marked against, read before writing code rather than after.
- One sentence test
- If the project cannot be described in one sentence without the words and then, it is two projects.
- Seed as a parameter
- Passing the random seed into a function so a test can ask for a predictable result and the real program cannot.
- Failing usefully
- Reporting bad input and carrying on, rather than crashing or silently discarding it.
- Known limitation
- Something the program cannot do, written down by its author. Unstated, it is a bug waiting to be found.
- Running from clean
- Deleting every file the program creates and running it as a stranger would, to test the first-run path.
- Usage docstring
- The few lines at the top of a file saying what it does, how to run it, and what files it needs.