💻 Computer Science · Undergraduate · INF 301

Information Science & Digital Literacy

You already spend hours a day finding, judging, and passing along information. Almost none of that was ever taught to you. This course fixes that. You will start with the field that has thought hardest about the problem, information science, and the long history of recorded information from clay tablets through the codex, print, the card catalog, and the web. Then you will do the practical work:…

Start the interactive course (quizzes, progress, videos) →

Free forever. No sign-up, no ads. 16 lessons. The full lesson text is below so you can read it right here.

Module 1: Foundations of Information

Before you can organize, find, or judge information, you need to know what the word means and where the machinery came from. This module introduces information science as a field, walks the five thousand year history of recorded information from clay to the web, and pins down the difference between data, information, and knowledge.

What Information Science Is, and Where It Came From

  • Define information science and name the professions that practice it.
  • Trace how recorded information moved from clay tablets to the codex, print, the card catalog, and the web.
  • Explain why each format change altered what people could actually find.

The big picture

You have a question. Somewhere, someone has already written down the answer. The distance between those two facts is the entire subject of this course. Information science is the field that studies how recorded information gets created, described, stored, found, judged, and kept, and the people who do that work professionally are librarians, archivists, records managers, data curators, indexers, and search engineers. This first lesson introduces the field by walking its history, because every tool you will use later in the course, from a subject heading to a search ranking, is a solution somebody invented for a problem that had become unbearable.

Key idea: information science studies the gap between a question and the recorded answer, and every technique in this course is an attempt to shrink that gap.

A field defined by a problem, not by a subject

Most academic fields are defined by what they study: biology studies life, geology studies rocks. Information science is different. It is defined by a problem that shows up everywhere: there is more recorded information than any person can read, so how do you get to the part you need? That framing explains why the field looks so scattered from the outside. All of these are information science:

  • A public librarian deciding which of two hundred books on diabetes to put on the display table for a patient who reads at a sixth grade level.
  • An archivist appraising forty boxes of a dead senator's papers and deciding, permanently, which ten boxes survive.
  • A hospital records manager writing the retention schedule that says how long a chart is kept and when it is destroyed.
  • A data curator at a research university writing the documentation that lets a stranger reuse a dataset in ten years.
  • A search engineer at a company deciding whether recency should outrank relevance for news queries.
  • A user experience researcher watching six people fail to find the returns policy on a website, and rewriting the navigation labels.

Those six people rarely meet. They use different vocabularies and read different journals. But they are all solving one problem with the same handful of moves: describe the thing so it can be found later, arrange things so browsing works, provide a way to search, and decide what to keep. Since the 1990s many universities have gathered these threads into what are informally called iSchools, information schools that grew out of library schools and now teach librarianship alongside human computer interaction, data curation, and information policy. The name change was an argument: the object of study is information, and the library is one important place it lives, not the whole of it.

Key idea: information science is one problem, too much recorded information for any reader, attacked by many professions using the same four moves: describe, arrange, search, and select what to keep.

Clay, scrolls, and the first catalogs

Start about five thousand years ago in Mesopotamia. The earliest surviving writing is not poetry or scripture. It is accounting: how much grain, from whom, to whom. That is worth sitting with, because it tells you that recorded information was invented as an administrative technology, a way to make a memory that outlives the person who has it and can be checked by someone else. Clay tablets are durable, which is why we have so many, but they are heavy and one tablet holds very little.

The scroll, made from papyrus and later parchment, held far more text in far less weight. But a scroll has a design flaw you can feel in your hands: to reach a passage in the middle you must physically unroll everything before it. Scholars call this sequential access. Finding a specific line means winding through the whole work. There is no page number to turn to, because there are no pages.

Even so, the ancients understood the finding problem. The library at Alexandria, founded in the third century BCE, is famous for its size, but its more interesting achievement was organizational: the scholar Callimachus compiled the Pinakes, a set of lists that recorded authors, their works, and identifying details so a reader could learn what the library held without walking the shelves. That is a catalog, and it is the ancestor of every catalog since. Note what it already separates: the description of a work lives apart from the work itself, so you can search the descriptions quickly and only then go fetch the object.

Key idea: the catalog was invented as soon as a collection outgrew memory, and its core trick has never changed: search short descriptions instead of searching the objects.

The codex: the invention that made looking things up possible

Somewhere in the first centuries of the common era, the scroll gave way to the codex, which is simply the book as you know it: leaves stacked and bound along one edge. We tend to treat this as a packaging change. It was much more than that. The codex introduced random access, the ability to open directly to any point without traversing what comes before.

Once you can open to any point, a family of finding tools becomes possible, and every one of them was invented in the centuries that followed:

  • Page numbers, so a location can be named.
  • Tables of contents, which describe a work in its own order.
  • Indexes, which describe a work in an order the reader chooses, usually alphabetical, and which are impossible in a scroll because there is no stable address to point at.
  • Running heads and chapter divisions, so a browsing eye can land in roughly the right place.

Alphabetical order itself had to be invented as a habit, and it took a surprisingly long time to feel natural. Arranging entries by their letters is arbitrary and slightly insulting to a medieval sense of order, in which things ought to be arranged by their importance or their place in creation. Alphabetization won because it has one enormous virtue: it requires no knowledge of the subject. Anyone who knows the alphabet can find the entry, without knowing whether angels outrank animals.

Key idea: random access, the ability to open straight to a point, is what made indexes and page references possible, and arbitrary orders like the alphabet won because they let a stranger find things without expertise.

Print, and the first information overload

Movable type printing in Europe from the 1450s changed the arithmetic. A manuscript was copied one at a time by hand, at enormous cost, with errors creeping in at every copy. Print made copies cheap, numerous, and identical. Identical matters more than cheap for our purposes: if every copy of an edition has the same words on the same page, then a citation to page 112 works for every reader, and scholarship can become a conversation between texts rather than between people who each hold a unique object.

Cheapness brought the complaint that has never stopped since. By the sixteenth century scholars were already writing that there were more books than anyone could read, that the good was buried under the bad, and that something had to be done. Their solutions look familiar: bibliographies, which are lists of what exists on a subject; commonplace books, in which a reader copied passages under topic headings for later retrieval, essentially a personal tagged database; and encyclopedias, which promised a survey so you could skip the originals. If you have ever kept a notes app organized by topic, you have made a commonplace book.

Key idea: information overload is not a modern condition; every jump in the cheapness of copying has produced the same complaint and the same family of fixes, which are all forms of description and summary.

The card catalog: the search engine of the paper era

The nineteenth century industrialized the catalog. The key insight was to stop writing catalogs in bound volumes, which are impossible to keep current because every new book has to be squeezed in somewhere, and instead to put one record on one card. Cards can be interfiled forever. A collection can grow without the catalog having to be rewritten.

The people who built this system set out rules that still govern description today. Anthony Panizzi at the British Museum fought through ninety-one cataloging rules in the 1840s, insisting that a catalog needed consistent principles rather than a cataloger's taste. Charles Ammi Cutter, in 1876, wrote objectives for a catalog that read as a requirements document for a modern search system: a reader should be able to find a book when they know the author, or the title, or the subject; the catalog should show what the library has by a given author, on a given subject, and in a given kind of literature; and it should help the reader choose among them by edition and character. Melvil Dewey, the same year, published the Decimal Classification that put every subject at a numbered spot on a shelf so that related books stand together.

Look closely at the card catalog and you will see a database with no computer in it. Each card is a record. Each record has fields in fixed positions: author, title, publisher, date, and at the bottom, the tracings, a list of every other card filed for this same book. The drawers are indexes. Filing rules are the query language. When you later meet metadata standards and database fields in this course, you are meeting the card in a new medium, and you should feel the continuity rather than the novelty.

Key idea: the card catalog was a working database of records, fields, and indexes, and its design goals, find a known item and gather everything on a subject, are still the two jobs every search system tries to do.

From documentation to the web

Two threads run from the card catalog to your phone. The first is the documentation movement in Europe, whose central figure, Paul Otlet, is worth knowing. With Henri La Fontaine he built the Mundaneum in Brussels, an attempt to index the world's published knowledge on standard cards, millions of them, with a scheme for linking related entries and a proposal for consulting them remotely through a screen and a telephone line. It was hand-built and it failed as an institution, but it was not naive. Otlet was arguing that the unit of information should be the fact or the passage, not the book, and that the links between units matter as much as the units.

The second thread is American and wartime. In 1945 Vannevar Bush, who had directed United States scientific research during the Second World War, published an essay called As We May Think, describing a hypothetical desk called the memex in which a researcher could store documents on microfilm and, crucially, build associative trails between them, then follow those trails later and share them. Bush's problem was the one this whole lesson circles: the specialist could no longer keep up with the literature of their own field. His answer, linking rather than only classifying, is the intellectual ancestor of the hyperlink. Meanwhile Claude Shannon's 1948 mathematical theory of communication gave the word information a precise technical meaning for engineers, measured in bits, which is a different sense from the one librarians use, and confusing the two causes real trouble. We will separate them carefully in the next lesson.

The web, from 1989 onward, delivered the linking part at global scale and delivered it fast. What it did not deliver, and this is the point of the rest of the course, is the description part. A library catalog contains records made deliberately by trained people according to shared rules. The web contains documents that describe themselves however they like, or lie about themselves on purpose, with no shared vocabulary and no guarantee that anything stays where you left it. Search engines are an extraordinary engineering response to that mess. They are not a replacement for description; they are what you get when description is unavailable and you must infer quality from links, text, and behavior instead.

Key idea: the web solved linking and scale but abandoned shared description, which is exactly why evaluating sources became the reader's job rather than the cataloger's.

Why the history is practical

None of this is antiquarian. Three working habits fall directly out of it. First, whenever you cannot find something, ask which of the four moves failed: was it never described in terms you would use, is the arrangement hiding it, is your query wrong, or was it never kept? That question resolves most search failures faster than trying more keywords. Second, remember that description is always made by someone with a viewpoint, which is the subject of Module 2. Third, remember that nothing survives by default. The clay tablets survived because clay is stubborn. Your favorite webpage from 2011 probably did not survive, and the reason is that nobody made it their job to keep it, a problem we will take up in the last lesson of the course.

Key idea: when a search fails, diagnose which of the four moves broke, description, arrangement, retrieval, or preservation, instead of just typing more keywords.

Common misconceptions

  • "Information science is just library science with a new name." Librarianship is one practice inside it. The field also covers archives, records management, data curation, information retrieval, human computer interaction, and information policy. The rebranding to iSchools in the 1990s was a claim about scope, and it stuck because the work really did spread past library walls.
  • "Information overload is a modern, internet problem." Scholars complained about too many books within a century of movable type, and the fixes they invented, bibliographies, indexes, encyclopedias, and topic-sorted notebooks, are the same fixes we use now in software form.
  • "The card catalog was primitive." It was a functioning database with records, fields, controlled headings, and multiple indexes, and its stated objectives from 1876 describe what modern search still tries and often fails to do.
  • "Search engines made cataloging obsolete." Search engines infer meaning from text, links, and user behavior because deliberate description is missing. Where description does exist, in library databases, medical literature, and legal research, professionals still use it, because it is far more precise.
  • "Everything is online now." A large share of published books, most archival material, most local government records, and most historical newspapers are not online in searchable form, and much that was online has vanished. Absence from a search result is not evidence of absence.

Recap

  • Information science studies how recorded information is described, arranged, found, judged, and kept, across libraries, archives, records, data, and search.
  • Catalogs appeared as soon as collections outgrew memory; the Pinakes at Alexandria separated descriptions of works from the works themselves.
  • The codex introduced random access, which made page numbers, indexes, and precise citation possible.
  • Print made copies cheap and identical, enabling shared citation and producing the first documented information overload.
  • The card catalog was a paper database; Panizzi, Cutter, and Dewey set rules and objectives still in use.
  • Otlet's Mundaneum and Bush's memex proposed linking as the answer to overload, anticipating the hyperlink.
  • The web delivered linking and scale without shared description, which is why source evaluation is now the reader's job.

Sources

  1. Library of Congress. (n.d.). About the Library. Library of Congress. loc.gov
  2. International Federation of Library Associations and Institutions. (n.d.). About IFLA. IFLA. ifla.org
  3. Encyclopaedia Britannica. (n.d.). Library. Britannica. britannica.com
  4. Bush, V. (1945). As We May Think. The Atlantic. theatlantic.com
  5. Wikipedia contributors. (n.d.). Information science. Wikipedia. en.wikipedia.org
  6. American Library Association. (n.d.). About ALA. American Library Association. ala.org
Key terms
Information science
The study of how recorded information is created, described, organized, retrieved, evaluated, and preserved.
Catalog
A set of descriptions of the items in a collection, searchable separately from the items themselves.
Codex
The bound book format that replaced the scroll and introduced random access to any page.
Random access
The ability to go directly to any point in a document without passing through what comes before it.
Index
A finding tool that lists terms with pointers to where they appear, arranged in an order chosen by the reader rather than the author.
Tracings
The list at the bottom of a catalog card naming every other heading under which that same item is filed.
Memex
Vannevar Bush's 1945 hypothetical machine for storing documents and linking them into associative trails, an ancestor of the hyperlink.
iSchool
An information school, typically grown from a library school, teaching librarianship alongside data curation, human computer interaction, and information policy.

What Information Actually Means

  • Distinguish data, information, and knowledge, and state the honest limits of that hierarchy.
  • Separate information as a thing, as a process, and as knowledge, and say which sense a given statement uses.
  • Describe an information need and explain why the query a person types is rarely the question they have.

The big picture

The word information is doing at least four different jobs in ordinary English, and people argue past each other constantly because they are each using a different one. An engineer means bits on a wire. A librarian means documents on a shelf. A teacher means what ended up in a student's head. A manager means the report that arrived this morning. Getting these apart is not word games. It changes what you do when a search fails, and it changes what you promise when you say a system delivers information.

Key idea: information is not one concept; sorting out which sense is in play is the first move in most arguments about information.

The data, information, knowledge ladder

The most common teaching model is a ladder, sometimes drawn as a pyramid, running from data at the bottom through information and knowledge, occasionally with wisdom on top. Take it as a rough sorting device rather than a law of nature. Here is the version worth remembering, with a single running example.

LevelWhat it isExample
DataRaw recorded observations or symbols, without interpretation37.9
InformationData placed in a context that answers a questionPatient temperature at 08:00 was 37.9 degrees Celsius
KnowledgeInformation integrated with experience and other information into something you can act onThat reading is within normal range for this patient, so the fever has broken and the antibiotic is working

The number 37.9 by itself answers nothing. You do not know what was measured, in what units, of what, or when. Add context and it becomes a statement about the world. Add clinical experience, a baseline, and a treatment history and it becomes a decision.

Now the honest part, because textbooks usually skip it. The ladder is criticized for good reasons. There is no such thing as truly raw data: someone chose to measure temperature rather than something else, chose Celsius, chose 08:00, and chose one decimal place. Every one of those choices is interpretation baked in before the number exists. The phrase raw data has been called an oxymoron for exactly this reason. The ladder also implies a tidy upward flow that does not match how people work, since knowledge is usually what tells you which data to collect in the first place. Use the ladder to remember that context does the heavy lifting; do not use it to claim that information is just data with sauce on it.

Key idea: data becomes information when context answers a question, but no data is ever raw, because choosing what and how to measure is already interpretation.

Three senses of the word: thing, process, knowledge

A cleaner cut, and one worth internalizing, comes from the information scientist Michael Buckland, who pointed out that we use information in three distinct ways.

  • Information as thing. Objects and documents that can be stored and retrieved: a book, a file, a photograph, a dataset, a rock in a museum drawer. This sense is the only one that systems can actually handle. A database stores things, not understanding.
  • Information as process. The act of becoming informed: reading, being told, working something out. This sense is an event that happens in time, to a person.
  • Information as knowledge. What a person ends up believing or understanding. This sense lives in a head and is intangible.

Watch how sloppiness here produces bad claims. A university announces an information system that will make its research available to the public. Which sense? It can deliver information as thing, files that can be downloaded. It cannot deliver information as knowledge, because whether anyone becomes informed depends on reading level, background, motivation, and time. Systems traffic in things; the process and the knowledge are the user's, and they are the part that usually fails.

This is also why the phrase information is power is half true and often misleading. Possession of documents is not understanding, and a document you cannot read, in a language you do not have, behind a paywall, or written for specialists, is information as thing with no path to information as knowledge. Much of the practical work in this course is about building that path.

Key idea: systems can only store and move information as thing; whether that becomes knowledge in a reader depends on the person and the context, which is where most information projects actually fail.

The engineer's sense: information as surprise

There is a fourth sense you will meet in computing, and it is precise and narrow. In Claude Shannon's 1948 theory, information is measured by how much uncertainty a message removes. A message that tells you something you already knew carries no information at all in this sense. A coin flip result carries one bit, because it resolves one of two equally likely possibilities. This is enormously useful for designing communication systems, compression, and error correction, and it is the reason your files are measured in bytes.

Notice what Shannon's definition deliberately excludes: meaning, truth, and usefulness. In the engineering sense, a page of accurate history and a page of random letters of the same length can carry the same number of bits, and the random page may carry more, because it is less predictable. Shannon said so explicitly; he was solving a transmission problem, not a meaning problem.

So when someone argues that we live in an age of unprecedented information, ask which sense. Bits transmitted, certainly, by many orders of magnitude. Documents available, yes. People becoming informed and knowing more true things, that is a separate empirical question with a much less obvious answer, and it is the question the rest of this course cares about.

Key idea: the engineering measure of information counts uncertainty removed and deliberately ignores meaning and truth, so a growth in bits proves nothing about a growth in understanding.

Documents, and the strange edges of the category

If systems handle information as thing, what counts as a thing? The usual answer is the document, and its boundaries are stranger than you expect. The documentation theorist Suzanne Briet posed the classic puzzle: an antelope running wild on a plain is not a document. But an antelope captured, brought to a zoo, catalogued, labeled, and studied has become one, because it has been made into evidence and placed into a system of description. The object did not change; its position in a descriptive system did.

That test, is it treated as evidence within a system of description, is more useful than any list of formats. It correctly includes a museum specimen, a soil sample with a field label, a screenshot in a court filing, and a dataset. It correctly excludes an unlabeled rock in your pocket. And it tells you why so much digital material is fragile: a photograph with no filename, no date, and no context is barely a document, because nothing about it supports retrieval or evidence.

Key idea: something becomes a document when it is treated as evidence inside a system of description, which is why context and labeling, not format, decide whether a thing can be found and used later.

Information need: the question behind the query

Here is the concept that will do the most practical work for you. What a person types into a search box is not their question. Information science distinguishes several layers, following Robert Taylor's classic account:

  1. The visceral need: a vague dissatisfaction, an itch. Something is wrong with my knee.
  2. The conscious need: the itch put into words, badly. I want to know if this knee thing is serious.
  3. The formalized need: a real question stated clearly. What are the common causes of pain on the inside of the knee in a thirty year old runner, and which ones need a doctor?
  4. The compromised need: what you actually type, shaped by what you think the system will accept. knee pain running inside

Almost every frustrating search failure lives in the gap between levels three and four. You compromised your question into keywords, the system answered the keywords faithfully, and the answer was useless. Librarians run a reference interview precisely to walk a person back up that ladder, which is why a good librarian's first response to a question is another question. You can do this for yourself: before searching, write the formalized question as a full sentence, then build the query from it deliberately. In Module 3 you will do exactly that, several times, as a drill.

One more concept belongs here: relevance, which is not a property of a document. It is a relation between a document and a person's need at a moment. The same article is relevant to you today and irrelevant tomorrow. Search systems can only estimate relevance from proxies, words in common, links, popularity, and behavior, which is both why they work at all and why they fail in the specific ways you will study next module.

Key idea: the query is a compromised version of a real question, and relevance is a relation to a person at a moment, not a fixed property of a document.

Common misconceptions

  • "Data is objective and information is interpreted." Choosing what to measure, in what units, at what moment, and to what precision is already interpretation. Raw data is a useful shorthand, not a real category.
  • "More information means better decisions." Beyond a point, additional documents raise cost and confusion without raising accuracy, and can raise confidence without raising accuracy, which is worse. What helps is better selection, not more volume.
  • "Information is inherently true." In everyday use we say misinformation to mark falsehood, which reveals that the plain word carries a whiff of truth it does not deserve. Nothing in any of the four senses requires a statement to be accurate.
  • "If a system holds the answer, the user has the answer." Holding a document is information as thing. Becoming informed is a process that can fail for reasons of language, level, access, time, or motivation.
  • "A good search is about better keywords." More often it is about a better question. Keywords chosen before the question is stated clearly will faithfully retrieve the wrong thing.

Recap

  • Data becomes information when context answers a question; knowledge adds experience and lets you act. The ladder is a rough tool, not a law.
  • Buckland's three senses: information as thing, as process, and as knowledge. Systems only handle things.
  • Shannon's engineering sense measures uncertainty removed in bits and ignores meaning and truth entirely.
  • A document is anything treated as evidence inside a system of description, which is why labeling and context matter more than format.
  • Taylor's four levels show that a typed query is a compromised version of a formalized question.
  • Relevance is a relation between a document and a person's need at a moment, which systems can only estimate.

Sources

  1. Encyclopaedia Britannica. (n.d.). Information theory. Britannica. britannica.com
  2. Wikipedia contributors. (n.d.). DIKW pyramid. Wikipedia. en.wikipedia.org
  3. Wikipedia contributors. (n.d.). A Mathematical Theory of Communication. Wikipedia. en.wikipedia.org
  4. Association of College and Research Libraries. (2016). Framework for Information Literacy for Higher Education. American Library Association. ala.org
  5. Wikipedia contributors. (n.d.). Relevance (information retrieval). Wikipedia. en.wikipedia.org
Key terms
Data
Recorded observations or symbols before interpretation, though the choice of what and how to record is already interpretive.
Information as thing
Buckland's sense covering documents and objects that can be stored, described, and retrieved by a system.
Information as process
The act of becoming informed, which happens to a person over time and cannot be delivered by a system.
Shannon information
A measure of how much uncertainty a message removes, expressed in bits, deliberately indifferent to meaning and truth.
Document
Anything treated as evidence within a system of description, from a book to a labeled museum specimen.
Information need
The underlying question a person is trying to answer, which the query they type only partially expresses.
Compromised need
Taylor's term for the query as actually submitted, reshaped by what the searcher believes the system will accept.
Relevance
A relation between a document and a particular person's need at a particular moment, not a fixed property of the document.

Module 2: Organizing Information

Somebody decided what your search results are called before you searched for them. This module takes apart classification schemes and the politics inside their categories, teaches you to write real metadata for a set of items, and covers controlled vocabularies, tagging, and the identifiers that keep things findable when everything else moves.

Classification and Its Politics

  • Explain how Dewey Decimal and Library of Congress classification arrange knowledge, and why a shelf order is a claim about the world.
  • Identify specific documented biases in major classification schemes and how they were challenged and changed.
  • Apply the tradeoffs of any classification, warrant, hospitality, and literary versus user warrant, to a scheme you encounter.

The big picture

Walk into any library and the shelves are making an argument. Putting all the books about religion together, and putting eight tenths of that space under Christianity, is a claim about how the world is shaped. Putting books about the history of women in a separate spot from the history of the countries those women lived in is another claim. Classification schemes look like neutral plumbing. They are not, they never were, and learning to see the argument in a category list is one of the most transferable skills in this course, because you will meet the same problem in a company's product taxonomy, a government's crime statistics, and a hospital's diagnostic codes.

Key idea: a classification scheme is an argument about how the world divides up, frozen into an order that then shapes what people find.

What a classification actually does

A classification does two jobs at once, and confusing them causes most beginner mistakes. First, it groups like with like, so that browsing works: stand in one spot and everything near you is on your topic. Second, it assigns each item a unique address so it can be shelved and retrieved. The address is called a call number, and it is a location as much as a label.

The tension is immediate. A physical book can be in exactly one place. A book about the economics of climate policy in Brazil belongs, honestly, with economics, with climate science, with policy, and with Brazil. Classification forces a choice, which is why librarians speak of the primary subject and why the catalog record, which can carry many subject headings, matters so much: the shelf gives you one arrangement, and the catalog gives you all the others. Once collections went digital this constraint relaxed, but classification survived, partly because browsing a coherent shelf remains genuinely useful, and partly because thousands of libraries and millions of records already use it.

Key idea: classification gives an item one place for browsing, while the catalog record gives it many subject headings for searching, and you need both.

Dewey: ten classes and their long shadow

Melvil Dewey published his Decimal Classification in 1876, and it is still the most widely used scheme in public and school libraries worldwide. Its structure is genuinely elegant: ten main classes, each divided into ten divisions, each divided into ten sections, with decimals extending as far as needed. The number is meaningful, so 973 is United States history, and 973.7 is the Civil War period. You can read the topic out of the number.

ClassSubject
000Computer science, information, general works
100Philosophy and psychology
200Religion
300Social sciences
400Language
500Science
600Technology
700Arts and recreation
800Literature
900History and geography

Now look at the politics, which are not hidden and not disputed. In the original scheme, the 200s for religion allotted roughly 200 through 289 to Christianity and squeezed every other religion of the world into 290 through 299. Judaism, Islam, Hinduism, Buddhism, and every indigenous tradition shared the space that Christian denominations enjoyed several times over. That is not a slur against Dewey personally so much as a description of a nineteenth century American library serving a nineteenth century American town: the scheme mirrored its users. But once frozen into a standard, that mirror became a mold, and libraries in Mumbai and Cairo inherited an arrangement built around a different world.

The 300s and 900s carried similar assumptions, with a large Eurocentric weighting in history and geography and, for a long time, homosexuality classed under abnormal psychology and later under social problems before it moved. Dewey has been revised continuously, now in its twenty-third print edition and maintained as a live product by OCLC, and many of these placements have changed. Some libraries have abandoned it entirely: a number of public libraries have reorganized into bookstore-style subject neighborhoods, and several First Nations and tribal libraries in North America have adopted schemes such as the Brian Deer Classification, built from the community's own categories rather than an imported one.

Key idea: Dewey's proportions encoded the worldview of its time and place, and because a standard outlives its context, those proportions kept shaping collections long after the context was gone.

Library of Congress Classification: built from a collection, not a theory

Most academic libraries in the United States use Library of Congress Classification instead. Its logic is different in an instructive way. Dewey starts from a theory of how knowledge divides and then fits books into it. LCC was built outward from an actual collection, the Library of Congress itself, in the early twentieth century, with each class developed more or less independently by people working on that literature. That makes it far more hospitable to new topics and far less tidy.

Its notation is alphanumeric: a letter or letters for the class, then numbers, then a Cutter number derived from the author or subject. So a book on American literature might sit at PS 3537, and a book on plant physiology at QK 711. Twenty-one main letter classes cover the whole scheme, from A for general works to Z for library science and bibliography.

The useful comparison for you is this:

AspectDewey DecimalLibrary of Congress
OriginA theory of knowledge, 1876One library's actual holdings, from 1897 onward
NotationPure numbers with decimalsLetters plus numbers plus a Cutter number
Typical settingPublic and school librariesAcademic and research libraries
StrengthReadable, teachable, consistent structureRoom to expand; handles very large specialized collections
WeaknessFixed proportions built into a decimal frameInconsistent between classes; harder to learn

Key idea: Dewey imposes a theory of knowledge on a collection while LCC grew out of one collection's shelves, which is why LCC expands more easily and coheres less.

Warrant: where do categories come from?

Classification theory has a precise word for the justification of a category: warrant. It comes in several flavors, and naming them lets you criticize any scheme, including a company's website menu, without hand-waving.

  • Literary warrant: a category exists because enough documents exist on that topic. This is the dominant principle in library classification, and it is conservative by construction. A topic with little published on it gets no category, so emerging fields and marginalized subjects are systematically slower to appear.
  • User warrant: a category exists because users ask for that topic in those words. This is the principle behind most website navigation and behind modern subject heading revision.
  • Scientific or philosophical warrant: a category exists because a discipline says it is real.
  • Cultural warrant: the scheme should fit the worldview of the community it serves, which is the explicit basis for community-built schemes.

A second technical virtue is hospitality: can the scheme accept a new topic without breaking? Dewey's decimals can always be extended rightward, but they cannot be reproportioned. If your religion class is full at 290, a new tradition cannot get equal room without renumbering the world, and renumbering means relabeling and reshelving every affected book in every library on earth. That practical inertia, not stubbornness, is the main reason bad categories persist. Changing a standard has a physical cost.

Key idea: most classification is justified by literary warrant, which lags behind the world, and hospitality is limited by the enormous physical cost of renumbering, so schemes change slowly even when everyone agrees they are wrong.

The record of repair: Sanford Berman and after

The politics of classification is not only a critique from outside; it has a long internal reform history. The best known figure is Sanford Berman, a cataloger at Hennepin County Library in Minnesota, who from the 1970s campaigned against specific offensive and inaccurate Library of Congress subject headings and published detailed proposals for replacements. Many were adopted, some after decades of argument. His method is a model worth copying: name the exact heading, show the specific harm or inaccuracy, and propose the specific replacement.

A recent and well documented case is the heading Illegal aliens. Following a campaign begun by students at Dartmouth College in 2014, the Library of Congress approved replacing it with Noncitizens and Illegal immigration. The change was blocked by Congress in 2016, an unusual direct intervention in cataloging practice, and the revised headings were finally implemented in 2021. Study that timeline: seven years, from a student petition to an implemented standard, with a legislative fight in the middle. It tells you both that these systems can change and that they change on the scale of years, through committees, with politics in the room.

Similar work continues on headings and classifications for Indigenous peoples, disability, gender and sexuality, and enslavement, and it is worth knowing that the change is never only symbolic. A heading determines what a search retrieves. If material about a community is filed under a term nobody in that community uses, the community cannot find its own record.

Key idea: categories are revisable, and the successful method is specific: name the heading, document the harm, propose the replacement, and expect the process to take years.

Reading any classification critically

You now have a repeatable procedure. Given any category system, from a library scheme to the department list on a job site to the crime categories in a police report, ask five questions:

  1. What is the top level, and what does its size imply? Proportion is argument. Which topic got a whole branch and which got a footnote?
  2. Where is the residual category? Every scheme has an Other, a Miscellaneous, or a 290. Whatever lands there is being treated as a deviation from a norm defined elsewhere.
  3. Whose words are the labels? Insider terms or outsider terms. Current or historical.
  4. What cannot be expressed? Find a real item that does not fit and see what the scheme forces you to do with it.
  5. Who can change it, and how long does that take? A scheme with no visible revision process is a scheme that will be wrong permanently.

Key idea: to critique any scheme, look at the proportions, the residual category, the source of the labels, what cannot be expressed, and the revision process.

Common misconceptions

  • "Classification is a neutral technical activity." Every scheme encodes choices about what is central and what is marginal, and those choices came from somewhere and served someone.
  • "Bias in classification means the cataloger was prejudiced." Usually it means the scheme reflected the assumptions of its time and place and then hardened into a standard. The mechanism is inertia at least as much as intent.
  • "Digital collections make classification obsolete." Full-text search finds words, not concepts, and it cannot show you what exists nearby that you did not know to name. Amazon, Netflix, and every online store all run large internal taxonomies.
  • "Nothing ever changes." Headings and classes do change, through documented proposals and committees. The Illegal aliens case took seven years and a fight with Congress, and it changed.
  • "An item belongs in one true place." Only shelving forces that. The catalog record can and should carry several subject headings, so search reaches an item by many routes.

Recap

  • Classification groups like with like for browsing and assigns a unique shelf address; the catalog supplies the other arrangements.
  • Dewey divides knowledge into ten classes with meaningful decimal numbers, and its proportions, notably in religion, encoded its era.
  • Library of Congress Classification grew from an actual collection, uses letter plus number notation, and is more hospitable but less consistent.
  • Warrant explains where categories come from: literary, user, scientific, and cultural warrant each justify categories differently.
  • Hospitality is limited by the physical cost of renumbering, which is the main reason bad categories persist.
  • Reform works when it is specific, as with Sanford Berman's proposals and the Illegal aliens to Noncitizens change of 2021.

Sources

  1. Library of Congress. (n.d.). Library of Congress Classification Outline. Library of Congress. loc.gov
  2. OCLC. (n.d.). Dewey Decimal Classification. OCLC. oclc.org
  3. American Library Association. (n.d.). Association for Library Collections and Technical Services. ALA. ala.org
  4. Wikipedia contributors. (n.d.). Dewey Decimal Classification. Wikipedia. en.wikipedia.org
  5. Wikipedia contributors. (n.d.). Sanford Berman. Wikipedia. en.wikipedia.org
  6. Library of Congress. (n.d.). Library of Congress Subject Headings. LC Linked Data Service. id.loc.gov
Key terms
Classification
A system that groups items by subject and assigns each a unique address, usually a call number, for shelving and retrieval.
Call number
The code that gives an item its exact place on the shelf and encodes its subject.
Dewey Decimal Classification
An 1876 scheme dividing knowledge into ten numbered classes, each subdividable by decimals, used mainly in public and school libraries.
Library of Congress Classification
An alphanumeric scheme built outward from the Library of Congress collection, used mainly in academic libraries.
Literary warrant
The principle that a category exists because enough documents exist on that topic, which makes classification lag behind emerging subjects.
User warrant
The principle that a category exists because users ask for that topic in those words.
Hospitality
A scheme's ability to accommodate new topics without restructuring, limited in practice by the cost of renumbering existing items.
Residual category
The Other or Miscellaneous slot in any scheme, which reveals what the scheme treats as a deviation from its norm.

Controlled Vocabularies, Subject Headings, and Metadata

  • Explain the vocabulary problem and how a controlled vocabulary with authority control solves it.
  • Read a Library of Congress subject heading, including its subdivisions, and use broader, narrower, and related terms.
  • Write descriptive metadata for a set of items using Dublin Core elements and consistent value formats.

The big picture

Here is a problem you can feel immediately. Search a catalog for heart attack and miss everything filed under myocardial infarction. Search for film and miss movie, motion picture, and cinema. Now flip it: search for mercury and get the planet, the metal, the god, and a car. Natural language gives you many words for one thing and one word for many things, and search engines paper over this with statistics and guesswork. Libraries and specialized databases solved it a different way, by agreeing in advance on which word to use. This lesson teaches you that machinery, and then puts you to work writing real descriptions of real items, which is the fastest way to understand why metadata standards exist.

Key idea: natural language has synonyms and homographs, and a controlled vocabulary fixes both by designating one preferred term for each concept and pointing every alternative at it.

The vocabulary problem, precisely

Two failure modes, and you should be able to name them:

  • Synonymy costs you recall. Several words mean the same thing, so a search on one of them misses documents that used another. You do not see what you missed, which makes this failure invisible and therefore dangerous.
  • Homography or polysemy costs you precision. One word means several things, so a search returns a pile of irrelevant results. This failure is visible and annoying, which is why beginners worry about it more, even though the invisible one usually hurts more.

Recall and precision are the two measures every retrieval system trades off. Recall is the share of the relevant material that you actually got. Precision is the share of what you got that is actually relevant. Push one up and the other tends to fall: a broad search catches more of what exists and buries it in noise, while a narrow search returns a clean list that quietly omits things. Knowing which one your task needs is a real decision. A systematic review of medical evidence needs recall above all, because a missed trial can change the conclusion. A quick check of a date needs precision, because you want one reliable answer now.

Key idea: synonymy silently costs recall and homography noisily costs precision, and every search is a deliberate choice about which of those two you can afford to lose.

Controlled vocabulary and authority control

A controlled vocabulary is a fixed list of terms in which each concept gets exactly one preferred form. The practice of enforcing it is authority control, and the file that records the decisions is an authority file. Three relationship types do all the work:

  • Equivalence: the vocabulary says use this term, not that one. In thesaurus notation you see USE and UF, meaning used for. Heart attack UF; use Myocardial infarction. This is what lets you type the wrong word and still land in the right place, because the system silently redirects you.
  • Hierarchy: broader term and narrower term, written BT and NT. Cardiovascular diseases is the BT of Heart diseases, which is the BT of Myocardial infarction. Hierarchy is what lets a database explode a search: ask for the broad term and automatically get every narrower one underneath it.
  • Association: related term, written RT, for concepts that are neighbors without being parent and child. Rheumatic fever RT Heart diseases.

Authority control also handles names, and this is the part people underrate. Samuel Clemens and Mark Twain are one author; Cambridge is several places; a woman who published under three surnames across a career is one researcher. An authority record establishes the preferred form, records the variants, and adds enough detail, dates, fields, affiliations, to distinguish the seven different people named Robert Smith who publish in chemistry. The international VIAF service links national authority files so that the same person resolves across countries and languages.

Key idea: USE, BT, NT, and RT are the four relationships that turn a word list into a vocabulary, and authority control extends the same discipline to names so that one person or place resolves to one record.

Library of Congress Subject Headings, read closely

Library of Congress Subject Headings, universally abbreviated LCSH, is the largest general subject vocabulary in English, with hundreds of thousands of headings, and you have used it without knowing. Its distinctive feature is pre-coordination: instead of leaving you to combine concepts at search time, the cataloger builds a compound heading in advance, joined by double hyphens. Read this one part by part:

World War, 1939-1945 -- Women -- United States -- Historiography

PartTypeWhat it does
World War, 1939-1945Main headingThe core topic
WomenTopical subdivisionNarrows the aspect
United StatesGeographic subdivisionNarrows the place
HistoriographyForm or genre subdivisionSays what kind of document this is

Four kinds of subdivision exist: topical, geographic, chronological, and form. The last one is genuinely useful and often forgotten. Adding Bibliography, Statistics, Periodicals, or Juvenile literature to a heading tells you what sort of thing you are getting, not what it is about. If you want data rather than argument, the form subdivision Statistics is a precision instrument.

Two practical moves follow. First, when you find one good item in a catalog, click its subject headings rather than searching again. You are handing the system a cataloger's exact controlled string, which is far better than any keywords you would have guessed. Second, browse the heading upward and downward using the broader and narrower links to calibrate how specific your search should be.

Key idea: LCSH pre-coordinates topic, place, period, and form into one string, and the fastest catalog technique is to find one good record and then follow its subject headings.

A specialist contrast: MeSH and query expansion

Medical Subject Headings, or MeSH, is the vocabulary the United States National Library of Medicine uses to index the biomedical literature in PubMed, and it works differently in a way worth seeing. MeSH is post-coordinated: rather than building long compound strings, indexers attach several separate descriptors to an article, and you combine them at search time with Boolean logic, which you will practice in Module 3.

The payoff is automatic explosion. Search PubMed for a concept and it will, by default, map your text to the matching MeSH descriptor and include every narrower descriptor beneath it in the tree. Type heart attack and the system understands Myocardial Infarction and picks up the narrower terms too. This is why a trained searcher gets systematically better recall in PubMed than a keyword typist, and it is the single strongest argument for learning that a controlled vocabulary exists behind a database you use.

Key idea: pre-coordinated vocabularies build the compound heading in advance, post-coordinated ones let you combine descriptors at search time and explode a term to include everything below it in the hierarchy.

Metadata: what it is and its three kinds

Metadata is usually glossed as data about data, which is true and nearly useless. A better working definition: metadata is structured description that makes a thing findable, usable, and manageable by someone who is not holding it. The word structured is doing the work. A paragraph about a photograph is description. The same facts in labeled fields, with consistent value formats, are metadata, because a machine can sort, filter, and match on them.

Three kinds, and you should be able to sort any field into one:

  • Descriptive: what it is and what it is about. Title, creator, date, subject, description, language. This is what supports discovery.
  • Structural: how the parts fit together. Which scan is page 12, which file is the high resolution master, which track belongs to which album, what order the chapters go in.
  • Administrative: what you are allowed and required to do with it. Rights and licence, provenance, file format and checksum, retention date, who deposited it and when. Two subtypes matter enough to have names: rights metadata and preservation metadata.

Most beginners write only descriptive metadata and are then astonished, two years later, when nobody can tell whether the file may be republished, which of the four versions is authoritative, or whether the file has silently corrupted.

Key idea: descriptive metadata gets a thing found, structural metadata keeps its parts in order, and administrative metadata records rights and provenance, which is the kind everyone forgets and later needs most.

Standards, and why they are not bureaucracy

Any group can invent field names. The point of a standard is that a description written here can be read there, without a human translating. Four you should recognize:

StandardHomeTypical use
Dublin CoreDublin Core Metadata InitiativeFifteen simple elements, a lowest common denominator for exchange between very different systems
MARC 21Library of CongressThe detailed machine-readable format behind library catalog records worldwide
Schema.orgSearch engine consortiumMarkup embedded in web pages so search engines can recognize an event, recipe, product, or article
EXIFCamera industryTechnical metadata written into a photo file by the device: date, exposure, often GPS coordinates

The fifteen Dublin Core elements are worth memorizing because they cover most needs and are easy to apply: Title, Creator, Subject, Description, Publisher, Contributor, Date, Type, Format, Identifier, Source, Language, Relation, Coverage, Rights. Notice that Rights is in the core fifteen. That was not an accident.

Standards also fail in an instructive way. A field named Date is only useful if everyone writes dates the same way. If half your records say 3/4/2019 and half say 2019-04-03, sorting is broken and international readers cannot tell March from April. This is why standards specify not only which fields exist but which value formats and which vocabularies fill them: dates as year-month-day, languages as codes, subjects from a named list. The rule to carry away is that a field without a stated value format is a field that will be inconsistent within a year.

Key idea: a metadata standard specifies field names, value formats, and which vocabulary fills each field, because field names alone do not prevent the inconsistency that makes data unusable.

Common misconceptions

  • "Metadata is just extra paperwork." It is the only reason retrieval works at scale. Every filter, sort, and facet you have ever clicked runs on metadata, and everything that is hard to find is usually hard to find because its description was thin.
  • "Full-text search makes subject headings unnecessary." Full text finds the words an author happened to use. A subject heading records what the work is about even when that phrase never appears, and it groups items that use different words for the same idea.
  • "Metadata is neutral." Subject vocabularies contain the same politics as classification, which is why LCSH heading revisions matter and why they are argued over.
  • "Tagging everything with many terms is best." Indiscriminate terms destroy precision and make the vocabulary meaningless. Assign the terms that describe what the work is substantially about, not everything it mentions.
  • "Metadata does not travel with the file." Some does, such as EXIF inside a photo, and it can travel further than you intend: uploaded images have leaked home addresses through embedded GPS coordinates. Knowing the metadata is there is also a privacy skill.

Recap

  • Synonymy costs recall invisibly; homography costs precision visibly; every search trades the two off.
  • A controlled vocabulary designates one preferred term per concept and links alternatives with USE, BT, NT, and RT.
  • Authority control does the same for names and places, and VIAF links national authority files internationally.
  • LCSH pre-coordinates main heading plus topical, geographic, chronological, and form subdivisions into one string.
  • MeSH is post-coordinated and explodes a term to include narrower descriptors, which is why trained PubMed searching finds more.
  • Metadata comes in descriptive, structural, and administrative kinds; the administrative kind carries rights and provenance.
  • Dublin Core, MARC 21, schema.org, and EXIF are the standards to recognize, and value formats matter as much as field names.

Sources

  1. Library of Congress. (n.d.). MARC Standards. Library of Congress. loc.gov
  2. Library of Congress. (n.d.). Library of Congress Subject Headings. LC Linked Data Service. id.loc.gov
  3. Dublin Core Metadata Initiative. (n.d.). DCMI Metadata Terms. DCMI. dublincore.org
  4. National Library of Medicine. (n.d.). Medical Subject Headings (MeSH). NLM. nlm.nih.gov
  5. Schema.org. (n.d.). Getting started with schema.org. Schema.org. schema.org
  6. OCLC. (n.d.). VIAF: The Virtual International Authority File. OCLC. viaf.org
Key terms
Controlled vocabulary
A fixed list of terms in which each concept has exactly one preferred form, with alternatives pointing to it.
Authority control
The practice of enforcing one preferred form for a name, place, or subject, and recording its variants in an authority record.
Recall
The share of all relevant material that a search actually retrieved.
Precision
The share of retrieved results that are actually relevant.
Pre-coordination
Building a compound subject string in advance, as LCSH does with double-hyphen subdivisions.
Post-coordination
Assigning separate descriptors and letting the searcher combine them at search time, as in MeSH and PubMed.
Explosion
Automatically including every narrower term beneath a chosen descriptor in a hierarchy, which raises recall.
Metadata
Structured description in labeled fields that makes an item findable, usable, and manageable by someone not holding it.
Dublin Core
A fifteen-element metadata standard designed as a lowest common denominator for exchange between different systems.
EXIF
Technical metadata a camera writes inside a photo file, including date, exposure settings, and sometimes GPS coordinates.

Taxonomies, Folksonomies, and Identifiers

  • Contrast a designed taxonomy with a user-generated folksonomy and state what each does well and badly.
  • Explain faceted classification and recognize it in the filter panels you already use.
  • Read an ISBN, ISSN, DOI, and ORCID, and explain why persistent identifiers exist.

The big picture

There are exactly two directions a vocabulary can come from. It can be designed from above by people whose job it is, which gives you a taxonomy. Or it can accumulate from below out of what users actually type, which gives you a folksonomy. Both are everywhere, both work, and both fail in predictable ways. Then there is a third thing that is not a vocabulary at all but solves a related problem: the persistent identifier, a name that is guaranteed not to move. This lesson covers all three, and by the end you should be able to look at any filter panel, tag cloud, or citation and say what kind of system you are looking at.

Key idea: taxonomies are designed from above and are consistent but slow, folksonomies grow from below and are current but messy, and identifiers sidestep naming entirely by assigning something that never changes.

Taxonomy: designed hierarchy

A taxonomy is a controlled vocabulary arranged as a hierarchy of broader and narrower terms. The classification schemes from the last two lessons are taxonomies. So is the department tree of a retail site, the ticket category list in a help desk, and the org chart of subject areas in a company intranet. The defining features are that somebody designed it, somebody maintains it, and each term has a defined place relative to the others.

Its strengths are consistency and predictability. If the taxonomy says Outerwear contains Coats, then every coat is under Outerwear, and a person who learns the structure once can navigate forever. Its weaknesses are cost and lag. Someone must be paid to maintain it, and it will always trail the vocabulary of the people using it, because adding a term requires a decision and decisions take time.

Key idea: a taxonomy buys consistency and predictable navigation at the cost of maintenance effort and a permanent lag behind live language.

Facets: the idea you use every day without a name

Strict hierarchy has a flaw that the Indian librarian S. R. Ranganathan attacked in the 1930s. Real things have several independent dimensions, and forcing them into one tree means choosing which dimension wins. His answer was faceted classification: describe an item along several independent axes, then combine the axes as needed. His scheme used five fundamental categories, abbreviated PMEST for Personality, Matter, Energy, Space, and Time.

You do not need his scheme, but you use his idea constantly. Every filter panel on a shopping or streaming site is faceted classification:

FacetExample values on a shoe site
TypeBoot, sandal, sneaker
Size7, 8, 9
ColorBlack, brown, white
PriceUnder 50, 50 to 100
BrandIndependent list

The power is combinatorial. Five facets with five values each describe over three thousand combinations from twenty-five terms, and the user assembles the one they want by clicking rather than by finding a pre-built category. This is why faceted navigation beat deep menu trees on the web almost completely. It is also the right design when you are organizing anything of your own: instead of asking which single folder this belongs in, ask what its independent attributes are.

Key idea: faceted classification describes items on several independent axes and lets the user combine them, which produces enormous coverage from a small vocabulary and is why filter panels replaced menu trees.

Folksonomy: vocabulary from below

A folksonomy is what emerges when users apply their own free-text tags to items and the system aggregates them. The bookmarking service Delicious and the photo site Flickr made this famous in the mid 2000s; hashtags are the same mechanism with different manners. Nobody approves the terms, there is no hierarchy, and any user can add any label.

The advantages are real and worth respecting:

  • Cost. The labor is distributed and free, so enormous collections get described at all.
  • Currency. New vocabulary appears the day people start using it, with no committee.
  • The long tail. Tags capture aspects no taxonomy would bother with, including feelings, uses, and to-read style personal workflow labels.
  • User words. The vocabulary is by construction the one users actually type.

The failures are equally real and are exactly the problems controlled vocabularies exist to solve:

  • No synonym control. nyc, newyork, new-york, and NewYorkCity split one concept across four tags.
  • Plurals and case. cat and cats are different tags to a naive system.
  • No hierarchy. Tagging something terrier does not connect it to dog, so you cannot explode a search.
  • Ambiguity. apple stays ambiguous forever.
  • Basic level variation. One person tags a photo animal, another dog, another Jack Russell. All three are correct and they do not retrieve one another.
  • Spam and gaming. When tags affect visibility, people add popular tags to unrelated items.

One genuinely interesting empirical finding: tag distributions stabilize. Across many items, the proportions of the top tags settle into a consistent power-law shape after enough taggers, meaning that a few tags dominate and a long tail trails off. Aggregate folksonomies are therefore noisier than a taxonomy at the item level but surprisingly stable at the collection level, which is why they are useful for recommendation and trending even when any single tag is unreliable.

Key idea: folksonomies are cheap, current, and in the user's own words, and they fail at synonyms, hierarchy, and ambiguity, which is precisely the work a controlled vocabulary does.

Hybrids, which is what serious systems actually build

The mature answer is not to choose. Working systems combine the two:

  • Suggest as you type. Offering existing tags while a user types quietly collapses most synonym and plural variation, at almost no cost to the user. This one intervention does more than any rule.
  • Map tags to controlled terms. Keep the user's tag for display, store a mapping to the official vocabulary for retrieval.
  • Mine tags for vocabulary maintenance. A tag that keeps appearing and has no controlled equivalent is a proposal for a new heading, backed by user warrant.
  • Split the roles. Let the institution supply subject terms and let users supply the aspects the institution would never record.

Wikipedia is worth studying here as a working example rather than a punchline, and this course will keep returning to it. Its category system is a user-built taxonomy with a formal structure, its templates carry structured metadata, its Wikidata companion assigns identifiers, and its content policies, verifiability rather than truth, no original research, and neutral point of view, are an explicit information policy that non-specialists debate in public on every article's talk page. Whatever you think of any individual article, the process is unusually visible, which makes it an ideal thing to study.

Key idea: the practical design is a hybrid, and the cheapest high-value intervention in any tagging system is suggesting existing tags as the user types.

Persistent identifiers: names that do not move

Now the third thing. Titles change, authors change names, websites reorganize, publishers get bought. If your only handle on an item is its title and a URL, you will lose it. A persistent identifier is a permanently assigned code, plus an organization that promises to keep it resolving. Four to know:

IdentifierIdentifiesShape
ISBNA specific edition and format of a book13 digits, ending in a check digit
ISSNA serial title, such as a journal or newspaper8 digits in two groups of four
DOIA document, most often a journal article or datasetA prefix starting 10., a slash, then a publisher-chosen suffix
ORCIDA researcher16 digits in four groups of four

Read an ISBN carefully, because the detail matters in practice: the ISBN identifies a specific edition and format, so the hardcover, the paperback, and the ebook of the same book have different ISBNs. That is a feature when you are buying and a nuisance when you are searching, because looking up one ISBN will not find the others. It is also why library catalogs group editions under a work-level record rather than relying on ISBNs alone.

A DOI looks like 10.1000/182. The part before the slash is the prefix assigned to a registrant, and everything after is that registrant's own suffix, which may be any string. The crucial property is indirection: the DOI itself carries no location. To resolve it you hand it to a resolver, which looks up the current address and redirects you. When a publisher moves an article, they update the record and every citation in the world keeps working. That is why journals ask you to cite the DOI rather than the URL you happened to land on, and why you should write the resolvable form, doi.org followed by a slash and the DOI, into your citations.

ORCID solves the person problem. Three researchers named Wei Zhang publish in the same field; one of them changes surname mid-career; a third publishes as both J. Smith and Jane Smith. An ORCID is claimed by the researcher, travels with them across institutions, and is now requested by most major publishers and funders. If you intend to publish anything, register one before your first submission, because retrofitting attribution across a career is genuinely painful.

Key idea: persistent identifiers work by indirection, since the identifier holds no location and a resolver supplies the current one, which is why a DOI survives a publisher moving a file and a URL does not.

Common misconceptions

  • "Tags are just a worse version of subject headings." They fail at different things and succeed at different things. Tags capture live vocabulary and personal use aspects that no cataloger would record, and they get collections described that would otherwise have no description at all.
  • "A folksonomy is unstructured, so it is useless in aggregate." Individual tags are unreliable, but tag distributions stabilize across many taggers, which makes aggregate folksonomy data genuinely useful for recommendation and trend detection.
  • "A DOI is just a URL." A DOI carries no location. It is resolved to a current location by a registry, which is exactly what lets it survive reorganizations that break URLs.
  • "Two copies of the same book share an ISBN." The ISBN identifies an edition and format. Hardcover, paperback, and ebook each get their own, which is why ISBN searching alone misses versions.
  • "Faceted navigation is a shopping site feature." It is a general classification technique from the 1930s. Use it whenever you are organizing your own material: ask what the independent attributes are instead of which one folder something belongs in.

Recap

  • Taxonomies are designed hierarchies: consistent and navigable, but costly to maintain and always behind live language.
  • Faceted classification, from Ranganathan, describes items on independent axes and combines them, producing huge coverage from few terms.
  • Folksonomies emerge from user tags: cheap, current, in user words, but broken on synonyms, plurals, hierarchy, and ambiguity.
  • Tag distributions stabilize in aggregate even though individual tags are unreliable.
  • Real systems build hybrids; suggesting existing tags as a user types is the highest-value single fix.
  • ISBN identifies an edition and format, ISSN a serial title, DOI a document, ORCID a researcher.
  • Persistent identifiers work by indirection through a resolver, which is why they outlive URLs.

Sources

  1. International DOI Foundation. (n.d.). What is a DOI?. DOI.org. doi.org
  2. ORCID. (n.d.). What is ORCID?. ORCID. orcid.org
  3. International ISBN Agency. (n.d.). About the ISBN standard. International ISBN Agency. isbn-international.org
  4. Wikipedia contributors. (n.d.). Folksonomy. Wikipedia. en.wikipedia.org
  5. Wikipedia contributors. (n.d.). Faceted classification. Wikipedia. en.wikipedia.org
  6. Encyclopaedia Britannica. (n.d.). S. R. Ranganathan. Britannica. britannica.com
Key terms
Taxonomy
A designed and maintained hierarchy of broader and narrower terms in which every term has a defined place.
Faceted classification
Describing items along several independent axes and combining those axes at search time, the principle behind filter panels.
Folksonomy
A vocabulary that emerges from free-text tags applied by users, with no approval process or hierarchy.
Basic level variation
The problem that different taggers correctly label the same item at different levels of generality, such as animal, dog, and Jack Russell.
Persistent identifier
A permanently assigned code plus an organization that keeps it resolving to the item's current location.
DOI
A Digital Object Identifier: a prefix beginning 10., a slash, and a registrant-chosen suffix, resolved through a central registry.
ORCID
A sixteen-digit identifier claimed by an individual researcher that travels with them across institutions and name changes.
ISBN
A thirteen-digit identifier for one specific edition and format of a book, so hardcover and ebook differ.

Module 3: Finding Information

Search looks like magic and is not. This module opens up crawling, indexing, and ranking so you can predict what a search engine will and will not find, drills Boolean and field searching as real exercises, and moves you into library databases, Google Scholar, and systematic searching you can document and repeat.

How Search Engines Actually Work

  • Trace the three stages of web search: crawling, indexing, and ranking.
  • Explain an inverted index and describe conceptually how link analysis such as PageRank scores a page.
  • Assess the evidence on search personalization and filter bubbles honestly.

The big picture

You type four words and get ten results in under a second, out of a corpus of hundreds of billions of pages. That performance is only possible because almost none of the work happens when you press Enter. The engine did the expensive part months ago. Understanding the three stages, crawling, indexing, and ranking, lets you predict what a search engine can and cannot find, which turns a lot of frustrating searching into deliberate searching. It also puts you in a position to judge the loudest claim made about search, that it traps each of us in a personalized bubble, against what the evidence actually shows.

Key idea: search is fast because the engine crawled and indexed the web in advance, so your query is matched against a prepared index rather than against the web itself.

Stage one: crawling

A crawler, also called a spider or bot, is a program that fetches a page, extracts the links on it, adds those links to a queue, and repeats. The queue is called the frontier. Starting from a set of seed addresses, a crawler can in principle walk the entire linked web, and in practice runs continuously, revisiting pages at rates it estimates from how often they change.

Three consequences fall out immediately, and they explain most of what search misses:

  • If nothing links to it, it may never be found. A page with no inbound links is invisible to a link-following crawler unless it is submitted directly or listed in a sitemap.
  • Crawling is rationed. No engine can fetch everything as often as it would like, so it allocates a crawl budget, favoring sites that change often and appear important. A small site may be revisited rarely, which is why a new page there can take days or weeks to appear in results.
  • Sites can ask to be excluded. A file named robots.txt at the root of a site tells cooperating crawlers which paths not to fetch. It is a request honored by convention, not an access control, a distinction that matters and that CS 305 Cybersecurity Fundamentals treats properly.

Now the part students consistently underestimate. An enormous share of recorded information is not crawlable at all. Content behind a login, behind a paywall, or generated only in response to a form submission cannot be reached by following links. Library databases, most government records systems, court filings, and the majority of scholarly literature live there. This is the deep web, and it is not sinister; it is the ordinary condition of anything that sits behind a query interface. When people say everything is on the internet, they generally mean everything is on the crawlable web, and that is a much smaller claim.

Key idea: crawlers follow links, so unlinked pages, and everything behind a login, paywall, or search form, is simply absent from the index no matter how good your query is.

Stage two: indexing

Having fetched a page, the engine does not store it and scan it later. Scanning hundreds of billions of documents per query is impossible. Instead it builds an inverted index, which flips the relationship around. A normal index maps documents to their words. An inverted index maps each word to the documents containing it. Here is the whole idea in miniature, with three tiny documents.

DocumentText
D1the cat sat on the mat
D2the dog sat on the log
D3a cat and a dog

The inverted index built from those three documents looks like this:

TermAppears in
catD1, D3
dogD2, D3
satD1, D2
matD1
logD2

Now search for cat and dog. Instead of reading three documents, look up two rows and intersect the lists: D1 and D3 intersected with D2 and D3 gives D3. That intersection is the mechanical meaning of the Boolean AND you will use in the next lesson, and understanding it is why Boolean feels obvious once you have seen the index. Scale the same operation to billions of documents and it is still just list lookups and intersections.

Before indexing, text is processed. Tokenization splits text into terms and decides what to do with hyphens, apostrophes, and scripts without spaces. Normalization lowercases and strips accents. Stemming or lemmatization reduces running, ran, and runs toward a common form so they match. Stop words, extremely common terms like the and of, may be dropped or handled specially. Positions are usually stored too, which is what makes a phrase search possible: to match an exact phrase, the engine checks that the terms appear in adjacent positions rather than merely in the same document. If you want to know how indexes work inside a database rather than a search engine, CS 340 Databases and SQL covers that machinery directly.

Key idea: an inverted index maps terms to the documents containing them, so a query becomes a set of list lookups and intersections, which is literally what Boolean operators do.

Stage three: ranking

Matching gives you millions of documents. Ranking decides the order, and the order is what people mean when they judge a search engine. Three families of signal, conceptually:

Text signals. How often does the term appear in the document, and how rare is that term across the whole collection? A word that appears in almost every document tells you nothing; a rare word that appears many times in one document is a strong hint. That intuition is formalized as term frequency weighted by inverse document frequency, and while modern rankers are far more elaborate, the intuition still holds: rare matching words matter more than common ones.

Link signals. This is the innovation that made Google. The idea, called PageRank, is that a link is a vote, but not all votes count equally. A link from a page that is itself heavily linked counts for more than a link from an obscure page, and a page that links to a thousand things spreads its vote thin. The definition is recursive, since a page's score depends on the scores of the pages linking to it, and the whole thing is solved by iterating until the numbers settle. A useful mental picture is a random surfer clicking links forever: the score of a page is roughly how often that wanderer lands there. What made the idea powerful is that it uses the judgment of millions of independent page authors, which is much harder to fake than the words on your own page.

Everything else. Freshness, because a query about an election wants today and a query about photosynthesis does not. Location and language. Device, since a page that fails on a phone should rank lower for a phone user. Page quality signals and the enormous ongoing effort against spam. And, increasingly, learned models that match meaning rather than words, so that a query about a stiff neck after sleeping can retrieve a page that never uses the word stiff.

Two things to hold on to. First, ranking is adversarial. Search engine optimization is an entire industry devoted to influencing that order, which is why ranking methods change constantly and why the top results are the outcome of a contest, not a neutral readout. Second, ranking is not truth. It estimates likely usefulness for a typical person issuing that query, which is a different thing from accuracy, and Module 4 exists because of that gap.

Key idea: ranking blends text rarity, link authority, freshness, context, and learned meaning, and it estimates likely usefulness rather than truth in an environment where many parties are actively trying to move the order.

Personalization and the filter bubble, with the evidence

Now the claim you have certainly heard. The filter bubble hypothesis, popularized by Eli Pariser in 2011, holds that personalized ranking shows each of us more of what we already agree with, narrowing our information diet without our noticing and hardening political division.

Start with what is definitely true. Personalization is real. Search results vary by location, by language, by device, by whether you are signed in, and by a limited amount of prior activity. Social feeds are ranked far more aggressively than web search and by very different signals. And the mechanism is plausible.

Now what the research has found, which is more complicated and less alarming than the headline. Studies that ran identical queries from different accounts have generally found meaningful but modest differences in web search results, with the largest variation driven by location and by the query itself rather than by inferred ideology, and with substantial result overlap even between very different profiles. Large-scale studies of social feeds have repeatedly found that the strongest driver of a narrow information diet is not the algorithm but self-selection: whom you chose to follow, and which links you choose to click. Several studies have found that people encounter more ideologically varied material online than they do through their offline social circles, precisely because online networks include weak ties who differ from you. And a set of large experiments that actually changed people's feeds for months found effects on exposure but, in the shorter run, little detectable movement in political attitudes.

How to hold this honestly, which is the real lesson:

  • Personalization exists and location effects are substantial, so two people genuinely do see different results.
  • The measured ideological narrowing from search personalization is smaller than the popular account implies.
  • Human choice, whom you follow and what you click, does more narrowing than ranking does.
  • Platforms differ enormously. Findings about one feed do not transfer to another, and the systems change faster than the studies.
  • The more defensible worry is not personalization but engagement-driven ranking, which optimizes for attention rather than accuracy. That is Module 5.

Practically, you can test your own exposure. Run the same query signed in and in a private window, compare the first ten results, and note what changed. Do it for a political query and a non-political one. Most people find the differences smaller than expected on web search and larger than expected on social feeds, which is itself the finding.

Key idea: personalization is real but the measured filter bubble effect in search is modest, since your own choices about whom to follow and what to click narrow your diet more than ranking does.

Common misconceptions

  • "Searching the web searches the web." It searches an index built earlier from what a crawler could reach. Everything behind a login, paywall, or query form is absent regardless of your query.
  • "If it is not in the results, it does not exist." It may be uncrawled, deliberately excluded, deep web, offline, or simply ranked below where you stopped looking.
  • "PageRank counts links." It weights them. A link from a highly linked page counts more, and a page linking to everything dilutes its own vote. The definition is recursive and solved by iteration.
  • "Top ranking means most accurate." Ranking estimates likely usefulness for a typical searcher, in an adversarial environment where an industry works to move it. Accuracy is a separate question you must answer yourself.
  • "Everyone is trapped in a filter bubble." Personalization is real, but research finds the ideological narrowing from search ranking is modest, and that self-selection matters more. Overstating the algorithm's power conveniently removes our own responsibility for what we choose to follow.

Recap

  • Crawling follows links from a frontier, is rationed by crawl budget, and honors robots.txt by convention.
  • The deep web, everything behind logins, paywalls, and query forms, is uncrawlable and includes most scholarly and government material.
  • An inverted index maps terms to documents, so queries become list lookups and intersections; positions enable phrase search.
  • Ranking blends term rarity, link authority such as PageRank, freshness, context, and learned semantic matching.
  • Ranking is adversarial and estimates usefulness, not truth.
  • Personalization is real, driven mostly by location and language; measured filter bubble effects in search are modest, and self-selection matters more.

Sources

  1. Google. (n.d.). How Search Works. Google. google.com
  2. Google. (n.d.). In-depth guide to how Google Search works. Google Search Central. developers.google.com
  3. Wikipedia contributors. (n.d.). PageRank. Wikipedia. en.wikipedia.org
  4. Wikipedia contributors. (n.d.). Inverted index. Wikipedia. en.wikipedia.org
  5. Pew Research Center. (n.d.). Internet and Technology research. Pew Research Center. pewresearch.org
  6. Encyclopaedia Britannica. (n.d.). Search engine. Britannica. britannica.com
Key terms
Crawler
A program that fetches pages, extracts their links, and queues those links to fetch next, building the engine's view of the web.
Crawl budget
The rationed rate at which an engine will fetch pages from a given site, which is why new pages on small sites appear slowly.
robots.txt
A file at a site's root asking cooperating crawlers not to fetch certain paths; a convention, not an access control.
Deep web
Material reachable only through a login, paywall, or query form, and therefore absent from crawler-built indexes.
Inverted index
A structure mapping each term to the list of documents containing it, turning search into list lookups and intersections.
Stemming
Reducing word forms such as running and ran toward a common root so that they match one another.
PageRank
A link analysis method that treats a link as a vote weighted by the linking page's own score, defined recursively and solved by iteration.
Filter bubble
The hypothesis that personalized ranking narrows each person's information diet toward what they already agree with.
Self-selection
The narrowing of an information diet caused by a person's own choices about whom to follow and what to click.

Constructing Queries That Work

  • Build Boolean queries with AND, OR, NOT and correct nesting, and predict what each returns.
  • Use phrase searching, truncation, proximity, and field limits to control precision and recall.
  • Diagnose a search that returned too much or too little and apply the specific fix.

The big picture

This is the most immediately practical lesson in the course, and it is a drill, not a lecture. You are going to learn a small set of operators, and then you are going to build a real search block by block, the way a professional searcher does, and diagnose it when it misbehaves. The operators take twenty minutes to learn and will save you hours every term for the rest of your life. The underlying machinery is the inverted index from the previous lesson: AND is a list intersection, OR is a list union, NOT is a subtraction. Once you see it that way, nothing about Boolean is mysterious.

Key idea: Boolean operators are set operations on lists of documents, so AND narrows by intersecting, OR widens by uniting, and NOT subtracts.

The three operators, worked

Suppose a small database where the term diet appears in 400 records, exercise in 600, and both appear together in 120.

QueryOperationRecords returnedEffect
diet AND exerciseIntersection120Narrows
diet OR exerciseUnion880Widens
diet NOT exerciseSubtraction280Narrows, and discards

Check the union arithmetic, because it is the one people get wrong: 400 plus 600 is 1000, minus the 120 counted twice, gives 880. Now the two rules that matter more than the definitions.

AND narrows more than beginners expect. Every additional AND term cuts the set again. Three or four ANDed concepts is usually the practical maximum; beyond that you are asking for a document that mentions everything you thought of, which few real documents do.

NOT is dangerous and should be your last resort. If you search depression NOT anxiety, you have thrown away every study that examined depression thoroughly and mentioned anxiety once in the discussion. Those are often the best studies. Use NOT only for a genuine unrelated homograph, such as excluding the fruit sense of a word, and even then check what you lost by running the search both ways.

Key idea: use AND sparingly because each one cuts the set, and treat NOT as a last resort because it discards documents for merely mentioning a term.

Nesting: the parentheses rule

Mixing AND and OR without parentheses produces results you did not ask for, because the system applies a default order of operations that is probably not yours. Compare:

  • teen OR adolescent AND depression. Many systems read this as teen OR (adolescent AND depression), returning every document mentioning teenagers about anything at all. Your set explodes and you blame the database.
  • (teen OR adolescent) AND depression. This returns documents about depression that use either word for young people. This is what you meant.

The rule to memorize: OR goes inside parentheses; AND joins the parentheses. Synonyms are gathered with OR inside a bracket, and distinct concepts are joined with AND between brackets. Every well-formed professional search has this shape.

Key idea: group synonyms with OR inside parentheses and join separate concepts with AND between them, because unparenthesized mixtures are evaluated in an order you did not choose.

The other four tools

  • Phrase search. Double quotes require adjacent words in order. Searching climate change without quotes finds documents about climate and about change; with quotes it finds the phrase. This is the single highest-value keystroke in searching, and it uses the position data stored in the index.
  • Truncation. A symbol, usually an asterisk, matches any ending. adolescen* catches adolescent, adolescents, and adolescence in one term. Truncate carefully: cat* also catches catalog, cathedral, and catastrophe. The fix is to truncate at the longest safe stem, so adolescen* is safe and ad* is not.
  • Wildcards inside a word. Many databases use a question mark for one character, so wom?n matches woman and women, and behavio?r handles the British and American spellings. Symbols vary between systems; check the help page once and note them.
  • Proximity. Some databases let you require that two terms appear within a set distance, written variously as N3, W3, or NEAR/3. Searching for teacher within three words of burnout is far more precise than ANDing them across a whole article and far less brittle than demanding an exact phrase. When a database offers proximity, it is usually the best tool it has.

Key idea: quotes force adjacency, truncation captures word endings, wildcards handle spelling variants, and proximity operators sit usefully between a loose AND and a rigid phrase.

Field searching

A database record has fields: title, author, abstract, subject, publication, year, document type. Searching everything is the default and is usually too broad. Restricting to a field is how you get precision without losing the right documents. Typical field limits, though the exact codes differ by system:

FieldUse it whenEffect
TitleYou want documents substantially about the topicVery high precision, some lost recall
AbstractTitle alone is too narrowA good middle setting
Subject or descriptorYou know the controlled termHighest quality results, since a human assigned it
AuthorFollowing a researcher's workCombine with affiliation or ORCID to disambiguate
PublicationYou know the journal or paperUseful for tracking a specific venue

The move that will improve your work most: run a loose search, find one genuinely good record, open it, and read its subject or descriptor field. Then run a new search using those exact controlled terms in the subject field. You have replaced your guesses with the vocabulary a professional indexer assigned. This is the technique from Module 2 applied at the keyboard, and it is what separates a five-minute search from a thirty-minute one.

Key idea: find one good record, steal its subject headings, and search those in the subject field, because a human already decided what that document is about.

A full worked search, block by block

Question: does social media use affect depression in teenagers? Here is the whole professional workflow.

Step 1, write the formalized question. Among adolescents, is time spent on social media associated with depressive symptoms? Note the three concepts: the population, the exposure, the outcome. Most research questions decompose into exactly this shape.

Step 2, build one block per concept, synonyms joined by OR.

  • Block A, population: (adolescen* OR teen* OR youth OR "young people")
  • Block B, exposure: ("social media" OR "social networking" OR Instagram OR TikTok OR Facebook OR "screen time")
  • Block C, outcome: (depress* OR "depressive symptoms" OR "mental health")

Step 3, join the blocks with AND. The full query is Block A AND Block B AND Block C. Written out: (adolescen* OR teen* OR youth OR "young people") AND ("social media" OR "social networking" OR Instagram OR TikTok OR Facebook OR "screen time") AND (depress* OR "depressive symptoms" OR "mental health").

Step 4, run it and inspect, do not just read. Look at the number of results and at the first ten titles. Then diagnose.

Step 5, iterate deliberately. Too many results and the first ten are off topic? Move the blocks into the title or abstract field, or drop the broadest synonym, which here is mental health. Too few? Add synonyms to the thinnest block, remove the least essential block entirely, or truncate more aggressively. Getting adult studies? That is the population block failing; check whether the database has an age limiter, which is far more reliable than keywords for age.

Step 6, write the search down. Record the database, the exact query string, the limits applied, the date, and the number of results. You will need it, either to resume next week or because a marker or a co-author asks how you found things. This is not bureaucracy; a search you cannot reproduce is a search you cannot defend.

Key idea: one block per concept, synonyms with OR inside, blocks joined by AND, then diagnose and iterate, and write down the exact query with its limits and date.

Diagnosing a bad result set

SymptomLikely causeFix
Millions of results, mostly irrelevantTerms are too general, or a phrase was not quotedQuote phrases, limit to title or abstract, add a concept block
Almost nothingToo many ANDs, or a term nobody usesDrop a block, add synonyms, truncate, check spelling and controlled terms
Right topic, wrong populationPopulation block relies on keywordsUse the database's age, species, or geography limiter
Results about the wrong sense of a wordHomographAdd a disambiguating AND term, or as a last resort a narrow NOT
Everything is thirty years oldNo date limit, and old items have accumulated citationsApply a date range, and sort by date to check
Only reviews and opinion piecesNo document type limitLimit publication type to the study kind you want

Key idea: every bad result set has a diagnosis, so read the symptom and apply the matching fix instead of retyping keywords and hoping.

Why the same query behaves differently on the web

Web search engines and library databases are built for different jobs, and the operators behave accordingly.

Library databaseWeb search engine
ContentCurated, described recordsWhatever was crawlable
BooleanStrictly honoredSoftened; the engine may ignore or reinterpret operators
VocabularyControlled subject terms availableFree text only
Result orderOften by date or relevance you can setRanked by many hidden signals
ReproducibleYes, the same query returns the same setNo, results vary by time, place, and person

That last row is the one to remember for any serious work: a web search is not reproducible, so it cannot be the documented backbone of a literature search. Web engines do offer their own useful operators, and these four are worth having in your fingers: quotation marks for a phrase, a minus sign immediately before a word to exclude it, site: followed by a domain to search within one site, and filetype: followed by an extension to find documents of one kind. Searching a topic with site: restricted to a government domain, for example, is a fast way to reach official material without wading through commercial pages.

Key idea: databases honor Boolean strictly and return reproducible sets, while web engines soften operators and personalize results, so only the database search can be documented and repeated.

Common misconceptions

  • "More search terms means better results." Each ANDed term shrinks the set. More terms usually means fewer and often means none. Add synonyms with OR to widen; add concepts with AND only when you truly need all of them.
  • "OR narrows the search." The word sounds exclusive in English but OR is a union and always widens. This is the most common beginner error.
  • "NOT is a good way to clean up results." It discards any document that merely mentions the excluded term, which frequently means throwing out the best sources.
  • "Truncation always helps." Truncating too early is a classic disaster: ad* is not a search, it is a lottery. Truncate at the longest safe stem.
  • "Operators work the same everywhere." Symbols and defaults differ by system, and web engines soften Boolean entirely. Read the help page of a database once and write the symbols down.

Recap

  • AND intersects and narrows, OR unions and widens, NOT subtracts and discards.
  • Group synonyms with OR inside parentheses, join concepts with AND between them.
  • Quotes force phrases, truncation catches endings, wildcards handle spelling variants, proximity sits between AND and phrase.
  • Field searching, especially the subject field with a controlled term copied from a good record, gives precision cheaply.
  • Build one block per concept, run, diagnose from the symptom table, iterate, and record the query, limits, and date.
  • Databases honor Boolean and are reproducible; web engines soften operators and personalize, so they cannot anchor a documented search.

Sources

  1. National Library of Medicine. (n.d.). PubMed User Guide. NLM. pubmed.ncbi.nlm.nih.gov
  2. Library of Congress. (n.d.). Search the Library of Congress Online Catalog. Library of Congress. catalog.loc.gov
  3. Association of College and Research Libraries. (2016). Framework for Information Literacy for Higher Education. American Library Association. ala.org
  4. Wikipedia contributors. (n.d.). Boolean operations on sets. Wikipedia. en.wikipedia.org
  5. Google. (n.d.). Refine web searches. Google Search Help. support.google.com
Key terms
Boolean AND
An intersection that returns only records containing every joined term, narrowing the result set.
Boolean OR
A union that returns records containing any of the joined terms, widening the result set; used for synonyms.
Nesting
Using parentheses to control evaluation order, conventionally OR inside brackets and AND between them.
Phrase search
Quotation marks requiring terms to appear adjacent and in order, enabled by position data in the index.
Truncation
A symbol, usually an asterisk, that matches any word ending, applied at the longest safe stem.
Proximity operator
A database operator requiring two terms to appear within a stated number of words of each other.
Field searching
Restricting a term to one part of a record, such as title, abstract, or subject, to raise precision.
Search block
One concept expressed as a bracketed set of synonyms joined by OR, later joined to other blocks with AND.

Databases, Google Scholar, and Systematic Searching

  • Choose between a discovery layer, a subject database, and a web search tool for a given information need.
  • Use Google Scholar effectively while naming its documented blind spots.
  • Run citation chaining in both directions and document a search well enough to reproduce it.

The big picture

Most of the material that would actually settle your question is not on the open web. It is in subscription databases your library already pays for, in repositories you have never opened, and in books nobody digitized. The gap between an average student search and a good one is almost never cleverness. It is knowing that four or five other places exist and being willing to spend ten minutes in them. This lesson is a tour of those places, an honest account of Google Scholar's strengths and its blind spots, and a method for searching systematically enough that you could hand your process to someone else and they could repeat it.

Key idea: the difference between a weak search and a strong one is usually which tool you opened, not how clever your keywords were.

What a library database is, and why it costs money

A library database is a curated collection of records, usually with controlled subject vocabulary applied by indexers, searchable with strict Boolean and real field limits. Some are published by scholarly societies, some by commercial aggregators who license content from many publishers and resell access as a bundle. Libraries pay substantial subscription fees, often the largest line in a collections budget, and that is why access runs through your login rather than through the open web.

Two practical consequences. First, your library card or student login is a key to a large amount of material that is otherwise paywalled, including newspaper archives, industry reports, historical collections, and streaming media. Most people never use more than one of the databases they already have. Second, when you leave an institution you lose that access, which is one of the strongest arguments for the open access movement you will meet in Module 6, and a reason to note down what you found while you still can.

Key idea: library databases are curated, indexed, and licensed, which is why they are better searchable than the web and why they sit behind a login you already have.

Discovery layers versus subject databases

Most library home pages now offer one big search box that searches across many databases at once. That is a discovery layer, and it is genuinely useful for starting out and for known-item lookups. It has a real cost: it searches many sources shallowly, cannot apply any one source's controlled vocabulary properly, and its relevance ranking is opaque. For a serious search you want the subject database itself, where the subject headings, limiters, and proximity operators actually work.

A rough map of where to go, by field:

FieldWhere to lookFree?
Medicine, health, biologyPubMed and MEDLINEPubMed is free to search
EducationERICFree
PsychologyPsycInfoSubscription
Humanities and older journalsJSTOR, Project MUSESubscription, some open
Engineering and computingIEEE Xplore, ACM Digital LibrarySubscription
Cross-disciplinary citation trackingWeb of Science, ScopusSubscription
PreprintsarXiv, bioRxiv, medRxiv, SSRNFree
Open access journalsDOAJ, CORE, PubMed CentralFree
Digitized booksHathiTrust, Internet Archive, Google BooksFree to search

Key idea: a discovery layer is a broad shallow start, while the subject database is where controlled vocabulary, limiters, and proximity operators actually function.

Interlibrary loan, the most underused service in education

Here is the single most valuable paragraph in this lesson for most students. If your library does not have an article or a book, you can request it from another library through interlibrary loan, usually abbreviated ILL. At most academic institutions and many public libraries it is free to the user. Articles commonly arrive as a scan by email within one to five days. Books take longer, typically one to three weeks, and arrive as a physical loan.

Almost nobody uses it, and the reason is that people assume a paywall is a wall. It is a queue. Practically: when you hit a paywall, note the full citation, open your library's ILL form, paste it in, and move on with your work. Requesting five things at once costs you ten minutes and dramatically changes what you can write about. Build the habit now, because the moment you stop treating paywalls as final is the moment your research stops being limited by what happens to be free.

Key idea: a paywall is a queue, not a wall, and interlibrary loan is usually free and takes days, so the correct response to a paywall is to request the item and keep working.

Google Scholar: what it does well

Google Scholar is genuinely useful and you should use it. Its strengths are real:

  • Breadth across disciplines. One box covers journals, conference papers, theses, preprints, and books from many fields at once.
  • Cited by. Every record shows how many later works cite it, and lets you search inside those citing works. This is the fastest forward-chaining tool that is free.
  • Versions. It clusters preprints, repository copies, and publisher versions, which is often how you get a legal free full text of a paywalled article.
  • Library links. In its settings you can add your institution, after which results show a link straight into your library's licensed copy. Set this up once; it takes two minutes and pays off forever.

Key idea: use Google Scholar for breadth, for cited-by chaining, and for finding legal free versions, and configure library links the first time you use it.

Google Scholar: the blind spots, stated plainly

  • Undocumented coverage. Nobody outside Google knows exactly what is indexed or how completely. You therefore cannot say what your search covered, which disqualifies it as the sole basis for any systematic review.
  • No controlled vocabulary. There are no subject headings and no explosion, so you carry the entire synonym problem yourself.
  • Weak Boolean and few limiters. Operators are softened, there is a cap on the number of terms it will honor, and limiting is largely restricted to date and a crude title option. There is no publication type limiter, so you cannot ask for randomized trials only.
  • No quality filter. It indexes predatory journals, working papers, marketing white papers, and student theses alongside peer-reviewed research, with no signal telling you which is which. That is the blind spot that catches students hardest.
  • Citation counts are noisy and gameable. Counts include self-citations and citations from non-peer-reviewed material, duplicate records split counts, and the numbers have been shown to be manipulable.
  • Uneven coverage and language skew. Coverage is stronger in the sciences than in the humanities, stronger for recent material than for older, and skewed toward English.
  • Not reproducible. Results shift with time and context, so the same query today and next month gives a different set.

The practical rule: use Scholar to find things and to chain citations, then verify what you found in a database that documents its coverage and tells you the publication type. If a source matters to your argument, check the journal, check whether the piece was peer reviewed, and check the version you are reading against the version of record.

Key idea: Google Scholar's coverage is undocumented and unfiltered by quality, so treat it as a discovery tool whose findings you verify elsewhere rather than as an authority on what exists.

Citation chaining, in both directions

Once you have one genuinely good paper, you have two ready-made search strategies that require no keywords at all.

  1. Backward chaining. Read its reference list. Those are works the author judged relevant, already filtered by an expert. Backward chaining walks you into the past and toward the foundational sources of a conversation.
  2. Forward chaining. Find who has cited it since, using Google Scholar's cited by, or Web of Science or Scopus if you have them. This walks you toward the present and, crucially, is how you discover that a finding was later contradicted, refined, or retracted.

A concrete routine that works: find one strong recent review article on your topic; mine its reference list for the five most-cited older works; then forward-chain from each of those to see what the last three years did with them. In about forty minutes you will have mapped the shape of a literature, including its disagreements, which is more than a week of undirected keyword searching typically produces.

Key idea: one good paper gives you two keyword-free strategies, backward into its references and forward into what cited it, and forward chaining is how you learn a finding was later overturned.

Searching systematically

At the far end of rigor sits the systematic review, which is a whole methodology. Its defining features are worth knowing even if you never conduct one, because they define what a defensible search looks like:

  • A protocol written and often registered before searching, stating the question and the inclusion and exclusion criteria in advance.
  • A search strategy designed for each database separately, using that database's controlled vocabulary, and reported in full so it can be rerun.
  • Multiple databases plus grey literature, meaning material outside commercial publishing such as government reports, theses, conference abstracts, and trial registries, because relying on published journals alone tilts the evidence toward positive findings.
  • Screening by two independent people, with disagreements resolved by a third.
  • A flow record of counts at each stage, records found, duplicates removed, screened, excluded with reasons, and included, which the PRISMA reporting guideline standardizes.

You can scale this down honestly for a term paper, and doing so makes your work markedly better. Write your question first. Search at least two databases plus one open source. Record every query string with its database, limits, date, and result count in a plain table. Note your inclusion criteria before you start, such as English language, last ten years, empirical studies only. Keep a list of what you excluded and why. That is perhaps thirty minutes of extra work and it converts your search from a story you tell into a process you can defend.

Key idea: a defensible search states criteria in advance, covers several sources including grey literature, and records every query with its date and counts so someone else could rerun it.

Common misconceptions

  • "If I cannot get the full text, I cannot use it." Interlibrary loan is usually free and fast, and many paywalled articles have a legal author copy in a repository that Scholar's versions link will find.
  • "Google Scholar is a peer-reviewed database." It indexes predatory journals, white papers, and theses alongside peer-reviewed work with no marker distinguishing them.
  • "A high citation count means a paper is good." Counts measure attention, including critical attention. Papers are heavily cited for being wrong, and counts vary enormously by field, age, and gaming.
  • "The library's single search box searches everything." A discovery layer searches many sources shallowly and cannot apply any one source's controlled vocabulary or specialist limiters.
  • "Grey literature is low quality." Government reports, trial registries, and theses are often the most detailed sources available, and excluding them systematically biases a review toward published positive findings.

Recap

  • Library databases are curated and indexed, which is why they support strict Boolean, controlled vocabulary, and real limiters.
  • Discovery layers are broad and shallow; subject databases are where the specialist tools work.
  • Interlibrary loan is usually free and fast, so treat a paywall as a queue.
  • Google Scholar is excellent for breadth, cited-by chaining, versions, and library links.
  • Its blind spots are undocumented coverage, no controlled vocabulary, weak limiters, no quality filter, noisy citation counts, and no reproducibility.
  • Backward and forward citation chaining from one good paper maps a literature fast and reveals later contradictions.
  • A defensible search sets criteria first, uses several sources including grey literature, and records queries, limits, dates, and counts.

Sources

  1. National Library of Medicine. (n.d.). PubMed. NLM. pubmed.ncbi.nlm.nih.gov
  2. Institute of Education Sciences. (n.d.). ERIC: Education Resources Information Center. U.S. Department of Education. eric.ed.gov
  3. Directory of Open Access Journals. (n.d.). About DOAJ. DOAJ. doaj.org
  4. Google. (n.d.). About Google Scholar. Google Scholar. scholar.google.com
  5. PRISMA. (n.d.). PRISMA statement. PRISMA. prisma-statement.org
  6. HathiTrust. (n.d.). About HathiTrust Digital Library. HathiTrust. hathitrust.org
Key terms
Discovery layer
A single library search box covering many databases at once, broad and shallow, without any one source's controlled vocabulary.
Interlibrary loan
A service that obtains an item your library does not hold from another library, usually free to the user and delivered in days.
Backward chaining
Following a paper's own reference list to find the earlier work it relied on.
Forward chaining
Finding later works that cite a paper, which reveals refinements, contradictions, and retractions.
Grey literature
Material outside commercial publishing, such as government reports, theses, and trial registries, whose exclusion biases reviews toward positive findings.
Systematic review
A review using a pre-registered protocol, documented per-database strategies, multiple sources, and independent screening.
PRISMA
A reporting guideline standardizing the flow of record counts through a review, from records found to records included.
Predatory journal
A publication that charges fees while providing little or no genuine peer review, and which appears in undocumented indexes without a warning label.

Module 4: Evaluating Information

Checklists feel rigorous and perform badly. This module teaches the method professional fact-checkers actually use, then takes you inside research papers, statistics, images, and generative AI output so you can judge a claim by how it was produced rather than by how the page looks.

Lateral Reading: How Fact-Checkers Evaluate Sources

  • Explain the Stanford research comparing historians, fact-checkers, and students, and what it found about vertical versus lateral reading.
  • Apply the four moves and click restraint to evaluate an unfamiliar site in about ninety seconds.
  • Distinguish primary, secondary, and tertiary sources and trace a claim back to its origin.

The big picture

In a set of studies at Stanford University, researchers gave the same web evaluation tasks to three groups: professional fact-checkers, PhD historians, and Stanford undergraduates. The historians and the students behaved similarly, and both performed badly. They stayed on the unfamiliar page, read it carefully, examined its design and its About section, and reasoned about whether it seemed credible. The fact-checkers did something that looks almost lazy: within seconds of arriving, they left. They opened new tabs and searched for what other sources said about the organization behind the page. They reached correct conclusions faster and far more often. That contrast is the most important practical finding in modern digital literacy, and this lesson is about turning it into a habit you can execute in ninety seconds.

Key idea: reading down a page tells you what a source says about itself, and the only reliable way to evaluate an unfamiliar source is to leave it and find out what others say about it.

Vertical reading and why expertise did not save the historians

Reading down a page, weighing its arguments, checking its internal consistency, and judging its presentation is vertical reading. It is what school trains you to do, and it is the right tool for a text you have already decided to take seriously. It is the wrong first tool for an unknown website, and the historians in the study are the proof: enormous subject expertise and excellent close-reading skills did not protect them, because the question was not about the text at hand. The question was who made this and what is their record, and the answer to that is never on the page.

Consider what a page's self-presentation actually costs. Professional design is a template and an afternoon. A .org domain is available to anyone for a few dollars a year and signals nothing about nonprofit status or honesty. An About page is written by the organization about itself. A long reference list can cite sources that do not say what the page claims, and almost nobody checks. Citations, credentials, and a serious tone are all cheap to fabricate and expensive to verify from the inside. Everything a vertical reader examines is under the control of the party being evaluated.

Key idea: every signal available on the page itself, design, domain, About text, tone, and even the reference list, is controlled by the party you are trying to evaluate, which is why on-page evaluation fails.

Lateral reading: the four moves

Lateral reading means opening new tabs and investigating the source from outside. The most widely taught version organizes it into four moves, easy to remember because the first one is doing nothing.

  1. Stop. Notice your reaction. Strong emotion, especially outrage or delight, is the reliable signal that you are about to share something without checking. Also stop to ask how much verification this claim deserves; a restaurant address needs less than a health decision.
  2. Investigate the source. Not what it says about itself. Open a new tab, search the organization's name, and read what independent sources say. You are looking for who funds it, who runs it, and what its track record is. Thirty seconds is often enough to reclassify a page entirely.
  3. Find better coverage. Often you do not care about this particular page; you care about the claim. Search the claim itself and see whether established sources report it, and how they characterize it. If a dramatic claim appears nowhere else, that absence is your answer.
  4. Trace claims, quotes, and media to the original context. Most of what you see online has been re-reported. Follow it back. Statements get compressed, statistics lose their qualifiers, and images get recaptioned in ways that reverse their meaning.

To this add click restraint, which the Stanford researchers documented as a distinguishing fact-checker behavior. Novices click the first result. Fact-checkers scan the whole results page first, reading titles, snippets, and domains, and then choose which result to open. Ranking is not a credibility ordering, as you know from Module 3, and the useful result is frequently the fourth one.

Key idea: stop, investigate the source from outside, find better coverage of the claim, trace it to the original, and read the whole results page before clicking anything.

Why checklists underperform

You have probably been taught a checklist, most likely CRAAP: Currency, Relevance, Authority, Accuracy, Purpose. Checklists are not worthless. They name real dimensions and they are a reasonable way to introduce the vocabulary. But as a working method they have documented weaknesses, and it is worth being precise about why rather than just declaring them out of fashion.

  • They evaluate the page using the page. Authority is judged from stated credentials, purpose from the About page, accuracy from the reference list. All of these are supplied by the source.
  • They reward effort rather than accuracy. Working through fifteen questions feels rigorous and produces a confident conclusion that can be entirely wrong. That combination, high confidence with no external check, is worse than no method.
  • They are slow, so people skip them. A method that takes fifteen minutes will not be applied to the forty things you read today. Lateral reading takes ninety seconds, which is why it survives contact with real life.
  • They handle the modern failure mode badly. The characteristic problem now is not a shoddy amateur page. It is a professionally produced site with real citations, a plausible name, and undisclosed funding or an undisclosed agenda. Such a site passes CRAAP comfortably.

The honest summary: use checklist dimensions as a vocabulary for describing what you found, and use lateral reading as the method for finding it. If you were taught CRAAP, you were not taught something false. You were taught the labels without the procedure.

Key idea: checklists judge a source using evidence the source supplies, are slow enough to be skipped, and pass exactly the professionally produced advocacy sites that cause the most trouble today.

Ninety seconds, worked

You land on an unfamiliar site with a striking claim about a medication. Here is the whole procedure, with times.

  1. 0 to 10 seconds. Stop. Note that you feel alarmed, and note that this is a health claim, which raises the verification you owe it.
  2. 10 to 40 seconds. Open a new tab. Search the organization's name plus a neutral word such as funding or who runs. Read the results page before clicking. Look for an independent encyclopedia entry, news coverage, or a nonprofit registry record. You are answering: who are these people and what is their record?
  3. 40 to 70 seconds. New tab. Search the claim itself in your own words, not the site's phrasing, since their phrasing may be designed to retrieve only their allies. Check whether health agencies or established outlets report anything similar, and how they frame it.
  4. 70 to 90 seconds. If the page cites a study, find the study. Read its abstract. Check whether it says what the page says it says, what population it studied, and whether the effect described is the one claimed.

Three outcomes are typical. The claim is broadly supported but overstated, which is by far the most common. The claim comes from an organization with an undisclosed interest, in which case you now know how to read everything else on the site. Or the claim traces to nothing at all, and the absence of any independent coverage is itself decisive.

Key idea: ninety seconds of tabs beats fifteen minutes of close reading, because you spend it on the one question the page cannot answer: who made this, and what does everyone else say?

Primary, secondary, and tertiary sources

Tracing a claim requires knowing what you are looking at. The three categories are relational, not fixed, and that is the part students find genuinely confusing.

TypeWhat it isExamples
PrimaryDirect evidence created at the time by a participant or observer, or an original report of new researchA diary, a treaty, a photograph, a dataset, a journal article reporting the authors' own experiment
SecondaryInterpretation, analysis, or synthesis of primary materialA history book, a review article, a news analysis of a study
TertiaryCompilations that organize the other twoEncyclopedias, textbooks, bibliographies, most Wikipedia articles

The relational part: a 1935 newspaper article is a secondary source about the event it reports, and a primary source about how the press covered that event in 1935. A textbook is tertiary for its subject, and primary evidence for what was taught in the year it was published. So never ask what type a document is; ask what type it is for the question you are asking.

In the sciences the vocabulary shifts slightly. A journal article reporting original research is primary. A systematic review is secondary, and for most practical purposes a good review is what you actually want, since a single study is weak evidence on its own. Knowing which one you are reading is the first thing to establish before you cite it.

Key idea: primary, secondary, and tertiary are relations to your question rather than properties of a document, so the same newspaper article can be either depending on what you are asking.

The chain: how a claim degrades

Trace almost any striking statistic you see and you will find a chain like this:

  1. A study finds a modest effect in a specific population, with stated limits.
  2. The university press office writes a release emphasizing the strongest framing.
  3. A news outlet writes from the release, sometimes without reading the study, and drops the population and the limits.
  4. An aggregator rewrites the news article, sharpening the headline.
  5. A social post states it as a general fact, with no link.
  6. Later posts cite the social post, and the origin disappears.

Nobody in that chain necessarily lied. Each step made a small, defensible compression, and the result is a claim the original study does not support. This is why the fourth move exists, and why it is worth doing even when you trust every party involved. The practical version: when a number matters, do not stop at the article that told you. Find the study, and read what it says about who was studied and how large the effect was.

Key idea: claims degrade through ordinary compression at each retelling, so tracing to the original is necessary even when nobody in the chain is dishonest.

Common misconceptions

  • "A professional-looking site is more trustworthy." Design is a template and an afternoon. Presentation quality carries almost no information about reliability.
  • "A .org domain means nonprofit or trustworthy." Anyone can register one. Domain suffixes with real gatekeeping are the exception, such as .gov and .edu in the United States, and even those tell you about the registrant, not about a given page's accuracy.
  • "Lots of citations means the claims are supported." Only if the citations say what the page claims. Checking two of them at random is more informative than counting all of them.
  • "Reading carefully is the best defense." Careful reading of a page you cannot verify is exactly what failed for the PhD historians in the Stanford studies. Careful reading is for after you have established what the source is.
  • "Primary sources are always better." Primary sources are evidence, not conclusions. A single primary study can be an outlier; a good secondary review weighs many. For most questions you want the review.

Recap

  • Stanford research found fact-checkers outperformed both historians and students by leaving the page immediately to check the source from outside.
  • Vertical reading judges a page on evidence the page itself supplies, all of which is cheap to fake.
  • The four moves are stop, investigate the source, find better coverage, and trace to the original; add click restraint before opening any result.
  • Checklists such as CRAAP name real dimensions but evaluate a source using the source, take too long to be used, and pass professional advocacy sites.
  • Primary, secondary, and tertiary are relations to your question, not fixed properties of a document.
  • Claims degrade through ordinary compression at each retelling, which is why tracing to the original is always worth it.

Sources

  1. Stanford History Education Group. (n.d.). Civic Online Reasoning. Stanford University. cor.stanford.edu
  2. Stanford History Education Group. (n.d.). About SHEG. Stanford University. sheg.stanford.edu
  3. Wikipedia contributors. (n.d.). Lateral reading. Wikipedia. en.wikipedia.org
  4. Poynter Institute. (n.d.). MediaWise. Poynter. poynter.org
  5. American Library Association. (n.d.). About ALA. American Library Association. ala.org
Key terms
Vertical reading
Evaluating a source by reading down the page itself, judging its design, tone, credentials, and citations.
Lateral reading
Evaluating a source by leaving it and searching what independent sources say about the organization behind it.
Click restraint
Scanning the entire search results page, including titles, snippets, and domains, before deciding which result to open.
The four moves
Stop, investigate the source, find better coverage, and trace claims to the original context.
Primary source
Direct evidence created by a participant or observer, or an original report of new research.
Secondary source
An interpretation, analysis, or synthesis of primary material, such as a history book or a review article.
Tertiary source
A compilation that organizes primary and secondary material, such as an encyclopedia or textbook.
Claim degradation
The loss of qualifiers, population limits, and effect size as a finding is retold through press releases, news, aggregators, and social posts.

Reading Research Critically and Reading Numbers

  • Interrogate a study by its sample, design, effect size, and funding rather than by its abstract.
  • State what peer review catches and what it does not, and handle preprints and retractions correctly.
  • Convert relative risk to absolute risk and apply base rates to a positive test result.

The big picture

Lateral reading tells you whether a source is worth your attention. This lesson tells you what to do once it is. Almost every consequential claim you meet, about a drug, a policy, a diet, a technology, ultimately rests on a study and a number, and both can be entirely honest and still mislead you badly. You do not need statistics coursework to handle this. You need about six questions and three arithmetic moves, and you can carry all of them in your head.

Key idea: judge a research claim by how it was produced, who was studied, and how large the effect was, because a true finding stated without those three things is not yet usable.

Read a paper in the right order

Research papers follow a standard shape, often abbreviated IMRaD: Introduction, Methods, Results, and Discussion, with an abstract on the front. Most people read the abstract and stop, which is the worst possible strategy, because the abstract is the authors' summary of their own work written to attract readers. Read in this order instead:

  1. Methods first. Who was studied, how many, how were they chosen, what was done to them, and how long were they followed? If the methods do not support the headline, nothing later will fix that.
  2. Results, especially the tables. Look for the actual numbers and the confidence intervals, not the adjectives.
  3. The limitations paragraph near the end of the discussion. Authors are usually honest here, and this paragraph frequently contains the sentence that undoes the press release.
  4. Funding and conflict of interest statements, usually at the very end or the very front.
  5. Abstract last, as a summary to check against what you found.

Key idea: read methods, results tables, limitations, and funding before the abstract, because the abstract is the authors' advertisement for their own work.

The six questions

  • Who was studied? Thirty undergraduates at one university is not the general population. Mice are not people. A study of men aged 40 to 65 tells you little about women aged 20. Population is the qualifier most often lost in retelling.
  • How many? Small samples produce large, unstable effects that shrink or vanish on repetition. This is not a flaw in any one study; it is arithmetic.
  • What was the design? This is the big one, and it deserves its own table.
  • How large was the effect? Not whether it was statistically significant. Significance is about whether an effect is distinguishable from noise; size is about whether it matters.
  • Who paid, and who benefits? Funding does not make a study wrong, but industry-funded studies of a sponsor's product return favorable results more often than independent ones, which is a reason for extra scrutiny rather than dismissal.
  • Has anyone repeated it? A single study is a hypothesis with evidence attached. Replication is what turns it into knowledge.
DesignWhat it can showMain weakness
Randomized controlled trialCausation, because assignment is randomCostly, often short, sometimes unrepresentative participants
Cohort studyAssociation over time in real populationsConfounding: groups differ in ways you did not measure
Case-control studyAssociations for rare outcomes, cheaplyRecall and selection bias
Cross-sectional surveyA snapshot of prevalence and correlationCannot establish which came first
Case reportThat something can happen onceNo comparison group at all
Systematic review and meta-analysisThe weight of evidence across studiesOnly as good as the studies included, and inherits their biases

Key idea: the design sets the ceiling on what a study can claim, so an observational study reporting a causal headline has been over-read by someone in the chain.

Peer review: what it does and does not do

Peer review means that before publication, an editor sends a manuscript to several researchers in the field who comment on it and recommend acceptance, revision, or rejection. It is a genuine quality filter and it is also routinely oversold.

What it usually catches: unclear reporting, missing detail, obviously inappropriate statistics, claims that overreach the data, missing prior literature. What it usually does not catch: fabricated data, since reviewers rarely see raw data; subtle analysis errors; whether the result is actually true, since reviewers do not repeat the experiment; and questionable practices such as trying many analyses and reporting the one that worked. Review quality also varies enormously between journals, and reviewers are unpaid volunteers working in their spare time.

Two related things to know. Predatory journals charge publication fees while providing little or no genuine review, and they can look convincing. Check whether a journal appears in the Directory of Open Access Journals, whether its editorial board is real and reachable, and whether it is indexed in a database that documents its criteria. Preprints are manuscripts posted publicly before review, on servers such as arXiv, bioRxiv, medRxiv, and SSRN. They are genuinely valuable, since they make results available immediately and are openly criticized in public, and they are genuinely risky, since nothing has been filtered. The correct handling is simple: use them, and say in your writing that they are preprints.

Key idea: peer review catches unclear reporting and overreach but does not detect fraud or verify truth, so it raises the floor rather than certifying a result.

Retraction, and the claims that outlive it

When a paper is found to be seriously flawed or fraudulent it can be retracted, which means the publication record formally withdraws it. Two facts about retraction matter for you. First, it is easy to check: search the title or DOI along with the word retracted, look at the publisher's page for a notice, and consult a retraction database. Second, retracted papers keep being cited for years afterward, often by authors who never learned. The most famous case is the 1998 paper linking the MMR vaccine to autism, fully retracted in 2010 after the findings were found to be false and the author lost his medical licence; the claim continues to circulate globally, more than a decade after the record was corrected.

The general lesson is one you should keep: correcting the record and correcting belief are different projects, and the first does not accomplish the second. Module 5 is about why.

Key idea: always check whether a key source has been retracted, because withdrawal from the literature does not remove a claim from circulation.

Three arithmetic moves that protect you

Move one: convert relative risk to absolute risk. A headline says a drug increases the risk of a condition by 50 percent. That is a relative risk. Ask: fifty percent of what? Suppose the baseline risk was 2 in 10,000 per year. A 50 percent increase takes it to 3 in 10,000. The absolute increase is 1 in 10,000, which most people would weigh very differently from the headline. Neither number is a lie; the relative figure is simply useless without the baseline. The rule: whenever you see a percentage change in risk, find the baseline before you react.

Move two: apply base rates. This one surprises nearly everyone, so work it slowly. A disease affects 1 percent of a population. A test detects 99 percent of true cases and gives a false positive in 5 percent of healthy people. You test positive. What is the chance you have the disease? Take 10,000 people:

GroupNumberTest positive
Have the disease10099
Do not have the disease9,900495
Total positives594

Of 594 positive results, 99 are correct. That is about 17 percent, not 99 percent. The reason is that the rare condition is swamped by false positives drawn from a much larger healthy group. This single calculation explains why screening rare conditions in whole populations produces so many false alarms, and it generalizes far beyond medicine to fraud detection, security alerts, and any test for anything uncommon.

Move three: separate correlation from causation, and name the confounder. Everyone can recite that correlation is not causation, and almost nobody applies it, because doing so requires proposing a specific alternative. Practice the move: ice cream sales correlate with drowning deaths, and the confounder is hot weather. Students who use a tutoring service get better grades, and the confounder is that motivated students seek tutoring. When you meet a correlation, ask three questions in order. Could the causation run backwards? Could a third factor drive both? Could the pattern be a selection artifact, because of who ended up in the data?

Key idea: find the baseline behind any percentage change, work the base rate before trusting a positive test, and always name a specific confounder rather than reciting the slogan.

Reading a graph adversarially

Charts persuade faster than sentences, which makes them worth checking harder. Five things to look at every time, before you look at the shape:

  • The vertical axis baseline. A truncated axis that starts at 95 instead of 0 turns a trivial change into a cliff. Sometimes truncation is legitimate, for a temperature series for example, but you must know it is there.
  • The time range. Cherry-picking a start date is the most common trick in the world. Any noisy series contains a window that shows any trend you like.
  • Dual axes. Two series on two differently scaled vertical axes can be made to appear to move together by choosing the scales. Treat any dual-axis chart as an assertion, not evidence.
  • Area versus length. If a quantity doubles and the graphic doubles both the width and the height of a shape, the area quadruples and the eye reads a fourfold increase.
  • What is not plotted. Missing categories, excluded years, and unstated denominators do more damage than any distortion of what is shown. A raw count map of anything mostly shows where the people are.

Two related traps worth naming. Survivorship bias: analyzing only what remained, such as studying successful companies and concluding their habits cause success, when the failed companies had the same habits and are absent from the data. Regression to the mean: an extreme measurement tends to be followed by a less extreme one for purely statistical reasons, which is why a punishment after a terrible performance and a reward after a spectacular one both appear to work.

Key idea: check the axis baseline, the time range, dual scales, area encoding, and what was left out before you look at the shape of a chart.

Common misconceptions

  • "Statistically significant means important." It means the effect is unlikely to be pure noise given the model and the sample size. With a large enough sample, a trivially small effect becomes significant. Size and significance are different questions.
  • "Peer reviewed means true." It means several specialists thought it worth publishing. Fraud, subtle errors, and non-replicable results routinely pass, and journals vary enormously in rigor.
  • "A preprint is not real research." It is research that has not yet been reviewed. Use it, label it, and treat it with the caution you would give any unreviewed claim.
  • "Industry funding means the study is fake." It is a reason for extra scrutiny, not automatic dismissal. Check whether the design was preregistered, whether the data are available, and whether independent work agrees.
  • "A big percentage change is a big deal." Not without the baseline. Fifty percent of a very small number is a very small number.

Recap

  • Read methods, results tables, limitations, and funding before the abstract.
  • Ask who was studied, how many, what design, how large the effect, who paid, and whether it replicated.
  • Study design caps what can be claimed; only randomized assignment supports causal language directly.
  • Peer review catches unclear reporting and overreach, not fraud or falsity, and predatory journals provide none of it.
  • Preprints are usable if labeled; always check whether a key paper has been retracted.
  • Convert relative risk to absolute, work base rates before trusting a positive test, and name a specific confounder.
  • Check a chart's axis baseline, time range, dual scales, area encoding, and omissions before reading its shape.

Sources

  1. National Library of Medicine. (n.d.). PubMed. National Institutes of Health. pubmed.ncbi.nlm.nih.gov
  2. Retraction Watch. (n.d.). About Retraction Watch. The Center for Scientific Integrity. retractionwatch.com
  3. Directory of Open Access Journals. (n.d.). About DOAJ. DOAJ. doaj.org
  4. Encyclopaedia Britannica. (n.d.). Statistics. Britannica. britannica.com
  5. Wikipedia contributors. (n.d.). Base rate fallacy. Wikipedia. en.wikipedia.org
  6. Wikipedia contributors. (n.d.). Peer review. Wikipedia. en.wikipedia.org
Key terms
IMRaD
The standard research paper structure of Introduction, Methods, Results, and Discussion, best read out of order starting with Methods.
Effect size
How large a difference or association is, as distinct from whether it is statistically distinguishable from noise.
Confounder
An unmeasured third factor that influences both variables in a correlation and can create the appearance of causation.
Relative risk
A change in risk expressed as a percentage of the baseline, which is uninterpretable until the baseline is stated.
Absolute risk
A change in risk expressed in actual cases per population, such as 1 additional case in 10,000.
Base rate
How common a condition is in the population being tested, which determines how much a positive result actually means.
Preprint
A manuscript posted publicly before peer review, usable if explicitly labeled as unreviewed.
Retraction
Formal withdrawal of a published paper from the literature, which does not stop the claim from circulating.
Survivorship bias
Drawing conclusions from only the cases that remained, ignoring the ones that failed and are absent from the data.
Regression to the mean
The statistical tendency of an extreme measurement to be followed by a less extreme one, which makes interventions look effective.

Verifying Images, Video, and AI Output

  • Run a reverse image search and use it to find an image's earliest appearance and original caption.
  • Apply a four-question verification routine to a photograph or video circulating online.
  • Explain why generative AI systems produce confident false citations and apply a rule for using them safely.

The big picture

Here is the finding that should reorganize how you think about visual misinformation: the dominant technique is not fabrication. It is a real, unaltered photograph with a false caption. Verification services and fact-checking organizations consistently report that miscontextualized authentic media outnumbers manipulated media by a wide margin, and it is easy to see why. Faking an image takes skill and leaves traces. Taking a genuine photograph of a genuine crowd from a different country five years ago and captioning it as today's protest takes ten seconds and survives every forensic test, because there is nothing wrong with the pixels. This lesson teaches the workflow that catches that, then turns to the newest source that produces confident falsehoods at scale: generative AI.

Key idea: most visual misinformation is a real image with a false caption, so the primary verification question is not is this image edited but where and when did this image first appear.

Reverse image search: your first move

A reverse image search takes a picture as the query instead of words. The engine computes a compact fingerprint of the image's visual features and finds other copies and near-copies in its index, including ones that were cropped, resized, or lightly recolored. You can use it from a phone or a laptop: save or copy the image, then upload or paste it into a reverse image tool. Several exist, they index different things, and running two is worth the extra thirty seconds.

What you are looking for, in priority order:

  1. The earliest appearance. Sort or scan for the oldest instance. If a photograph presented as today's news appeared on a stock site in 2016, you are finished.
  2. Other captions. The same image described three different ways in three different conflicts is a strong signal on its own.
  3. A higher resolution or uncropped version. Crops remove context deliberately. Finding the full frame frequently changes the meaning entirely, revealing a film crew, a sign, or a much smaller crowd.
  4. A fact-check that already exists. Widely circulated images have usually been checked by someone. Searching a distinctive phrase from the caption plus the word fact check often takes five seconds.

Key idea: reverse image search answers the question that matters most, which is when and where this picture first appeared, and finding an uncropped version often reverses its apparent meaning.

The four questions

Professional verification, as practiced in newsrooms, comes down to four questions asked in order about any piece of media.

  • Provenance: is this the original? Reverse search it. Find the earliest posting and the account that posted it. A screenshot of a post is not a source; find the post.
  • Source: who captured it, and what do we know about them? Look at the account's history. An account created last week that posts only about one conflict is a different thing from a photographer with a decade of local work.
  • Date: when was it captured, not when was it posted? These differ constantly. Look for evidence inside the frame: seasonal vegetation, snow, clothing, construction stages, visible dates, and events in the background.
  • Location: where was it captured? Read the frame for language on signage, script, license plate formats, road markings, architectural style, vehicle models, and terrain. Then confirm against satellite and street-level map imagery. Matching a distinctive roofline or a row of pylons to a map is slower than the other steps and is often decisive.

Weather is an underrated cross-check. Historical weather records are freely available, and a photograph supposedly taken during a downpour on a day that was dry in that city is finished as evidence. Shadows are similar: their direction and length constrain the time of day and, roughly, the season.

A caution about metadata. A photo file can carry EXIF data including timestamp and GPS coordinates, which sounds like it settles everything. In practice most social platforms strip EXIF on upload, so its absence proves nothing, and its presence proves less than you think because it can be edited. Treat metadata as a lead, never as proof.

Key idea: ask provenance, source, date, and location in that order, read the frame for language, vehicles, vegetation, and shadows, and treat EXIF as a lead rather than evidence.

Video, and why it is only slightly harder

Video verification follows the same routine with one extra step: take screenshots of several distinct keyframes and reverse search each one, because the whole video may not be indexed while a frame from it is. Then check the specifics. Does the audio match the visuals, including whether the ambient sound suits the environment? Do the visuals cut in a way that hides a transition? Is this a re-upload of a longer video with the informative part removed?

Cropping in time is the video equivalent of cropping in space and it is just as effective. A thirty-second clip beginning after the provocation and ending before the resolution is not a lie in any single frame, and it is a lie. When a clip is doing heavy rhetorical work, the right question is always what happened in the ninety seconds before it started.

Key idea: reverse search several keyframes rather than the whole video, and remember that trimming a clip in time misleads exactly as effectively as cropping an image in space.

Synthetic media, honestly

Generated and manipulated media is now cheap and rapidly improving. It is worth being precise about what has changed and what advice has expired.

Expired advice: count the fingers, look for the mismatched earrings, check whether the subject blinks. These artifacts were real in earlier generations of image and video synthesis and are largely gone. Teaching people to look for them is now actively harmful, because it produces false confidence in both directions, clearing convincing fakes and condemning real photographs with an odd shadow.

What still helps:

  • Context beats forensics. The same four questions apply. A dramatic video from an event where no other footage exists, posted by a new account, is suspicious regardless of how it looks.
  • Provenance infrastructure. The durable answer is cryptographic provenance rather than detection: standards such as C2PA Content Credentials attach a signed record of where an image came from and how it was edited, supported by camera makers, software vendors, and platforms. Detection is an arms race that defenders lose; provenance changes the question from does this look fake to can this prove where it came from.
  • Expect the reverse problem. The existence of convincing fakes lets people dismiss genuine evidence as fabricated, a dynamic called the liar's dividend. In practice this is already the more common harm.

Key idea: visual artifact checklists have expired, so rely on context and provenance, and expect the bigger harm to be genuine evidence dismissed as fake.

Generative AI as an information source

A large language model produces text by predicting what plausibly comes next, trained on an enormous body of writing. That description explains both its usefulness and its characteristic failure. It is genuinely good at producing fluent, well-organized, plausible prose. It has no separate store of verified facts to consult and no internal signal that distinguishes a remembered fact from a well-formed guess. When it does not know, it does not produce a gap; it produces something shaped like an answer.

The technical name usually used is hallucination, though confabulation is the better word, because the phenomenon resembles a person confidently filling a memory gap with plausible material rather than seeing something that is not there. The most dangerous version for students is the fabricated citation, and it is worth understanding exactly why it happens. A citation has a highly regular form: author names, a year, a title, a journal, volume, pages, a DOI. Generating something in that shape is easy. Generating one that exists is a completely different task. So you get references with real authors who work in the field, in real journals, with plausible titles and well-formed DOIs, that were never written. They are far more convincing than an obvious error would be, and this has produced real professional consequences, including sanctions against lawyers who filed court documents citing cases that did not exist.

How to use these tools well, since refusing to use them is not the lesson:

Good useWhy it works
Building search vocabularyAsk for synonyms and technical terms for a concept, then take those words to a real database
Explaining a concept you will verifyAn explanation you check against a textbook costs you nothing if it is wrong
Drafting and restructuring your own textThe content is yours, so there is nothing to fabricate
Naming what you do not knowAsking what questions should I be asking generates leads, not claims
Bad useWhy it fails
Asking for sources or citationsReference format is easy to generate and existence is not checked
Asking for specific numbers or datesPrecise values are exactly what gets confabulated, and they look authoritative
Asking about recent eventsTraining data has a cutoff, and a confident answer will still be produced
Asking about niche topicsThe thinner the training material, the more the output is reconstruction

The rule, and it is not negotiable: every citation must be verified independently before you use it. Search the exact title in a library database or Google Scholar. Resolve the DOI. If you cannot find it in a source that documents what it indexes, it does not exist, and pasting it into your paper is your error, not the tool's. Note also that a system which cites live web sources is doing something different and better, but the same verification rule applies, because a real link can still be attached to a claim the linked page does not make. Open the link and read the passage.

Key idea: a language model predicts plausible text with no internal check on existence, so it produces convincing fake citations, and the non-negotiable rule is to verify every reference in a source that documents its coverage.

Common misconceptions

  • "Most visual misinformation is faked images." Miscontextualized authentic media dominates. A real photo with a false caption passes every forensic test because nothing is wrong with it.
  • "You can spot AI images by the hands." Those artifacts are largely gone. Artifact checklists now create false confidence in both directions.
  • "EXIF data proves when and where a photo was taken." Platforms usually strip it, and it can be edited. It is a lead, not evidence.
  • "If a video is unedited, it is honest." Trimming a clip in time removes the provocation or the resolution without altering a single frame.
  • "AI made up a citation because it was trying to be helpful." It generated text in the shape of a citation because that shape is highly regular. There is no separate step that checks whether a reference exists.

Recap

  • The dominant technique is a genuine image with a false caption, so start with when and where it first appeared.
  • Reverse image search finds the earliest appearance, other captions, uncropped versions, and existing fact-checks.
  • Ask provenance, source, date, and location, reading the frame for language, vehicles, vegetation, and shadows, and cross-check weather.
  • For video, reverse search several keyframes and ask what happened before the clip started.
  • Artifact checklists for synthetic media have expired; context and cryptographic provenance such as C2PA are the durable answers.
  • Language models predict plausible text and therefore produce well-formed citations that do not exist.
  • Verify every reference independently in a source that documents its coverage, and open every link to check the passage.

Sources

  1. TinEye. (n.d.). How reverse image search works. TinEye. tineye.com
  2. Coalition for Content Provenance and Authenticity. (n.d.). C2PA overview. C2PA. c2pa.org
  3. Poynter Institute. (n.d.). International Fact-Checking Network. Poynter. poynter.org
  4. Wikipedia contributors. (n.d.). Hallucination (artificial intelligence). Wikipedia. en.wikipedia.org
  5. Wikipedia contributors. (n.d.). Reverse image search. Wikipedia. en.wikipedia.org
Key terms
Miscontextualization
Presenting genuine, unaltered media with a false caption about its date, place, or subject; the most common form of visual misinformation.
Reverse image search
Searching with a picture as the query to find copies and near-copies, revealing earliest appearance and other captions.
Provenance
The documented origin and chain of handling of a piece of media.
Keyframe search
Reverse searching several distinct still frames from a video, since frames may be indexed when the full video is not.
C2PA Content Credentials
A standard attaching a signed, verifiable record of an image's origin and edit history, shifting verification from detection to provenance.
Liar's dividend
The benefit accruing to bad actors from the mere existence of convincing fakes, which lets genuine evidence be dismissed as fabricated.
Confabulation
A language model's production of fluent, plausible, false content that fills a gap, since it has no internal check on whether a claim exists.
Fabricated citation
A reference generated in correct bibliographic form, with plausible authors, journal, and DOI, for a work that was never written.

Module 5: Misinformation and the Information Environment

Why false claims travel, why corrections only partly work, and what the attention economy rewards. This module reports the evidence honestly, including where the popular story is stronger than the research supports, and covers the interventions that actually have measured effects.

How False Information Spreads, and Why Corrections Only Partly Work

  • Distinguish misinformation, disinformation, and malinformation, and describe how novelty and emotion drive spread.
  • Explain the continued influence effect and state honestly what the evidence shows about correction and backfire.
  • Identify the features of conspiracy thinking and the tactics of organized influence operations.

The big picture

A false claim and a true one are not competing on equal terms. The false one can be built to be interesting, because it is not constrained by what happened. That single asymmetry explains a surprising amount, and it is the foundation of this lesson. We will look at what the research actually found about spread, then at the harder question of why corrections work less well than they should, and finish with organized influence and conspiracy thinking. Throughout, the discipline is to report effect sizes honestly, including the places where the popular story has outrun the evidence.

Key idea: false claims can be optimized for interest because they are not constrained by what happened, which gives them a structural advantage over true ones.

Three words that are not synonyms

TermFalse?Intent to harm?Example
MisinformationYesNoSharing a wrong health tip you believed
DisinformationYesYesA fabricated document planted to discredit someone
MalinformationNoYesPublishing someone's genuine private records to harm them

The distinction matters practically because the responses differ. Misinformation responds to better information. Disinformation is a strategic act by an actor with goals, so the response is exposure and attribution. Malinformation is true, so accuracy checking does nothing at all; the issue is context and harm. Notice also that the same content can move categories: a deliberately planted falsehood becomes misinformation the moment a sincere person shares it.

Key idea: misinformation is false without intent, disinformation is false with intent, and malinformation is true but weaponized, and only the first responds to fact-checking alone.

What the research found about spread

The most cited empirical study of this question analyzed roughly 126,000 rumor cascades on Twitter over about eleven years, classified by six independent fact-checking organizations, and was published in Science in 2018. Its headline findings:

  • False stories spread significantly farther, faster, deeper, and more broadly than true ones, and the gap was largest for political stories.
  • True stories rarely reached more than about a thousand people, while the top one percent of false ones routinely reached between a thousand and a hundred thousand.
  • The likely mechanism was novelty and emotional response: false stories were measurably more novel relative to a user's recent exposure, and replies to them expressed more surprise and disgust, while replies to true stories expressed more sadness, anticipation, and trust.
  • Bots accelerated true and false stories at roughly equal rates. Humans did the differential spreading. This is the finding people most often get backwards.

Now the honest caveats, because a course that only reports headline findings has taught you nothing about reading evidence. This is one platform, in one period, with a specific definition of contested news based on what fact-checkers happened to examine, which is a sample skewed toward the disputed. It does not establish that most people see mostly falsehoods; other research consistently finds that misinformation is a small fraction of the average person's media diet and is heavily concentrated among a small, distinctive minority of users. Both things are true: false content has a structural spread advantage, and it is still a minority of what most people encounter.

Add two more robust structural findings. Superspreaders: a small number of accounts generate a large share of misinformation exposure, which means the problem is more concentrated, and more tractable, than a picture of universal viral chaos suggests. And the illusory truth effect: repeated exposure to a statement increases the feeling that it is true, even when you knew better initially, and even for implausible claims. Repetition is a mechanism of belief independent of evidence, which is why amplifying a false claim in order to mock it is a genuinely counterproductive act.

Key idea: false stories spread farther and faster mainly because humans find novel and emotionally arousing content more shareable, and repetition alone raises the feeling of truth.

Why corrections only partly work

Suppose you correct someone with clear evidence. Research finds two things at once, and holding both is the skill.

First, the continued influence effect is well established: after a correction, people often continue to use the retracted information in their reasoning. Told that a warehouse fire was caused by paint cans stored improperly, then told there were no paint cans, people still refer to something flammable when explaining the fire. Belief updates partially, and the discredited story keeps doing work in the mental model.

Second, and this is the important correction to what you may have been taught: the backfire effect, the claim that corrections make people believe falsehoods more strongly, was reported in influential early studies but has largely failed to replicate at scale. Larger and better-powered studies have generally found that corrections do move beliefs in the right direction on average, including among people predisposed to disagree. The effect is smaller than we would like, it fades over time, and it does not always change downstream attitudes or behavior. But the widespread advice that correcting people is counterproductive is not supported.

This matters practically. If you believed corrections backfire, you would stay silent. The evidence says: correct, and correct well. What better means is reasonably well established:

  1. Lead with the truth, not the myth. The first thing you say is the thing most likely to be remembered.
  2. Provide an alternative explanation, not just a negation. A gap in a causal story gets filled by the old explanation. Replace it rather than deleting it. This is the single most effective element.
  3. Keep it simple and concrete. A simple myth beats a complicated truth. If your correction is harder to process, it loses.
  4. Repeat the correct version. The illusory truth effect works for accurate statements too.
  5. Do not lead with mockery. A correction that also attacks identity invites defense of the identity rather than reconsideration of the claim.

Key idea: corrections work partially rather than not at all, the backfire effect largely failed to replicate, and the strongest single technique is replacing the false explanation with an alternative rather than merely negating it.

Propaganda and organized influence operations

Propaganda is communication designed to shape belief and action in the service of an agenda, and it is much older than the internet. Its modern networked form has documented features worth recognizing:

  • Amplifying existing divisions rather than inventing new ones. Operations typically find genuine grievances and inflame them from multiple sides simultaneously, which is why the content often looks locally authentic.
  • Astroturfing: manufacturing the appearance of grassroots support with coordinated accounts, so that a fringe position appears widely held. Perceived consensus is persuasive on its own.
  • Narrative laundering: planting a claim in a low-credibility outlet, amplifying it, then citing the resulting coverage until a mainstream outlet reports the controversy. Each step adds apparent legitimacy without adding evidence.
  • The firehose model: high volume, rapid, repetitive, and openly inconsistent messaging. Analysts have noted that contradicting yourself is not a bug in this model. When the aim is to make people believe nothing rather than something, inconsistency does the job.

That last point is the one to carry: the goal of much organized disinformation is not to convince you of a specific falsehood. It is to exhaust you into treating all sources as equally unreliable, because a public that believes nothing cannot be organized around anything. Cynicism is the product. Which means that your refusal to slide from healthy skepticism into blanket dismissal is not naivety; it is the actual defense.

Key idea: organized influence usually amplifies real divisions, manufactures apparent consensus, and launders claims into credible outlets, and its frequent goal is generalized cynicism rather than belief in any specific claim.

Conspiracy thinking

Conspiracies genuinely happen, are sometimes uncovered, and are prosecuted. So the useful question is never whether a conspiracy is possible but what distinguishes conspiracy thinking as a style of reasoning. Its markers:

  • Unfalsifiability. Contrary evidence is absorbed as further proof, since anyone producing it must be in on it. A claim no evidence could dent is not a claim about the world.
  • Implausible coordination. The theory requires very many people to keep a secret perfectly and indefinitely, against every incentive to defect.
  • Pattern-seeking over evidence. Coincidences are treated as connections and randomness is read as design.
  • Nothing is an accident, and nothing is as it seems. Errors, incompetence, and chance are excluded as explanations, though they explain most of the world.

Why people hold them is better understood than it used to be, and it is not stupidity. Conspiracy belief correlates with feelings of powerlessness, anxiety and uncertainty, a desire to feel that one possesses uncommon knowledge, and membership in a community that supplies belonging. Each of those needs is real, and a debunking that satisfies none of them competes badly against a belief that satisfies all of them. That is why arguing facts at someone rarely works, and it is not a reason for despair; it is a reason to change tactics. What practitioners report working: maintaining the relationship, asking genuine questions about how the person came to a view, and inviting them to specify what evidence would change their mind, all without ridicule. Public ridicule reliably hardens positions, and it is aimed at an audience rather than the person.

Key idea: conspiracy thinking is marked by unfalsifiability and implausible coordination, it meets real needs for control and belonging, and facts alone compete badly against a belief that supplies community.

Common misconceptions

  • "Bots are why false news spreads." In the largest study of the question, bots amplified true and false stories at similar rates. Humans did the differential spreading.
  • "Correcting people backfires, so do not bother." The backfire effect largely failed to replicate. Corrections generally help on average, modestly and imperfectly.
  • "Most of what people see online is misinformation." Research consistently finds it is a small share of the average diet, heavily concentrated among a small minority of users. It has a spread advantage without dominating exposure.
  • "Debunking a myth means repeating it clearly first." Leading with the myth risks reinforcing it through repetition. Lead with the truth and mention the myth once.
  • "Anyone who believes a conspiracy theory is unintelligent." Belief tracks powerlessness, uncertainty, and community more than reasoning ability, and contempt reliably hardens positions.

Recap

  • Misinformation is false without intent, disinformation false with intent, malinformation true but weaponized.
  • False stories spread farther and faster, driven by novelty and by surprise and disgust, and humans rather than bots did the differential spreading.
  • Misinformation is still a minority of most people's exposure and is concentrated among superspreaders and a small minority of users.
  • Repetition raises perceived truth through the illusory truth effect, so amplifying a claim to mock it helps it.
  • The continued influence effect is real, but the backfire effect largely failed to replicate; corrections work partially.
  • Good corrections lead with the truth, supply an alternative explanation, stay simple, repeat, and avoid ridicule.
  • Influence operations amplify real divisions, astroturf consensus, launder narratives, and often aim at cynicism rather than belief.

Sources

  1. Massachusetts Institute of Technology. (2018). Study: On Twitter, false news travels faster than true stories. MIT News. news.mit.edu
  2. RAND Corporation. (n.d.). Truth Decay. RAND. rand.org
  3. Pew Research Center. (n.d.). Journalism and Media research. Pew Research Center. pewresearch.org
  4. Encyclopaedia Britannica. (n.d.). Propaganda. Britannica. britannica.com
  5. Wikipedia contributors. (n.d.). Illusory truth effect. Wikipedia. en.wikipedia.org
  6. Wikipedia contributors. (n.d.). Conspiracy theory. Wikipedia. en.wikipedia.org
Key terms
Misinformation
False information shared without intent to deceive, which is the kind that responds to better information.
Disinformation
False information created and spread deliberately to deceive or harm, requiring exposure and attribution rather than correction alone.
Malinformation
Genuine information deployed to cause harm, such as publishing private records, where accuracy checking is beside the point.
Illusory truth effect
The increase in perceived truth of a statement caused by repeated exposure, independent of evidence.
Superspreader
One of the small number of accounts responsible for a large share of misinformation exposure.
Continued influence effect
The tendency to keep using retracted information in reasoning after a correction has been accepted.
Backfire effect
The claimed strengthening of a false belief after correction; reported in early studies but largely not replicated at scale.
Astroturfing
Manufacturing the appearance of grassroots support through coordinated accounts, exploiting the persuasive power of perceived consensus.
Narrative laundering
Planting a claim in a low-credibility outlet and amplifying it until credible outlets report the resulting controversy.

Attention, Ranking, and What Actually Helps

  • Explain how engagement-based ranking follows from an advertising business model and what it systematically rewards.
  • Describe how newsrooms actually produce and correct reporting, and read the difference between news, analysis, opinion, and sponsored content.
  • Compare interventions such as prebunking, accuracy prompts, and fact-check labels with honest effect sizes.

The big picture

Module 3 showed you how ranking works mechanically. This lesson asks the question underneath it: ranked toward what? A ranking algorithm optimizes something, and what it optimizes is set by how the service makes money. Once you can trace that line, from business model to metric to what appears at the top of your screen, a great deal of otherwise mysterious online behavior becomes predictable. Then we will look at how journalism actually produces information, since it is the institution most people rely on and least understand, and finish with the interventions that have been tested, reported at their real size rather than their advertised one.

Key idea: a ranking system optimizes a metric, the metric is chosen by the business model, and what reaches your screen follows from that chain rather than from any judgment about what is true.

The attention economy

The economist Herbert Simon put the core observation plainly in 1971: a wealth of information creates a poverty of attention. Information became abundant; attention did not. Anything abundant becomes cheap, and the scarce complementary good becomes valuable. So attention is what gets sold.

In an advertising-funded service, revenue is roughly the number of advertisements shown multiplied by what advertisers pay for each. More time on the service means more advertisements, so time and interaction become the target. Nobody needs to be cynical for this to happen; it follows from the accounting. The problem is that nothing in that equation refers to whether what you saw was true, useful, or good for you.

This produces a proxy problem worth naming, because it generalizes far beyond social media. Platforms cannot measure value directly, so they measure proxies: time spent, clicks, shares, comments, replays. Proxies begin as reasonable correlates of value and stop being so the moment they become targets, because everyone starts optimizing the proxy. That is Goodhart's law: when a measure becomes a target, it ceases to be a good measure. A comment count cannot tell the difference between a thoughtful discussion and a furious argument, and the furious argument produces more comments.

Key idea: platforms optimize measurable proxies for value, and once a proxy becomes the target it stops tracking value, which is why engagement metrics reward argument over understanding.

What engagement optimization rewards

If the metric is interaction, then content that reliably produces interaction rises. Research on what spreads points repeatedly at the same features:

  • Emotional arousal, especially high-arousal negative emotion. Anger and outrage produce action; sadness produces scrolling.
  • Moral and emotional language. Studies of political messages have found that each additional moral-emotional word measurably raises sharing, and that the effect is strongest within a group rather than across groups.
  • Identity conflict. Content framing an out-group as threatening reliably outperforms content about policy.
  • Curiosity gaps. A headline that withholds the answer forces a click, which is exactly what the metric rewards.
  • Novelty. As in the previous lesson, unusual claims travel further, and unusual claims are disproportionately false.

Now the honest limits, because this argument is often stated far more strongly than the evidence supports. Platforms also actively demote some categories of content, and their systems change constantly. Large field experiments that actually altered people's feeds for months found that changing the ranking changed what people were exposed to, but produced small and often undetectable movement in political attitudes over those months. Attributing polarization primarily to feed ranking is not well supported. The defensible claim is narrower and still important: engagement optimization systematically advantages emotionally arousing and conflict-framed content over careful content, and it does so without any component that evaluates truth.

Key idea: engagement ranking demonstrably advantages arousing and conflictual content, but the evidence that it is the main cause of political polarization is weak, and overstating it obscures the mechanism that is real.

How journalism actually works

Most criticism of the press is made by people who have never seen how a story gets made. Here is the process at a functioning outlet.

A reporter covering a beat, a defined subject area, develops sources over years. A story starts from a tip, a document, a data set, or an observed pattern. The reporter gathers evidence, seeks documents, and interviews people, including the subject of any accusation, who is given a genuine opportunity to respond. An editor challenges the framing and demands support for each claim. At larger outlets a copy editor checks details and, for legally risky stories, a lawyer reviews. Only then does it publish. Wire services such as the Associated Press and Reuters supply much of the raw reporting that smaller outlets run, which is why the same story appears in many places with near-identical wording; that is shared sourcing, not collusion.

Four things follow that you can use immediately:

  • Distinguish the genres. News reports events. Analysis explains and contextualizes. Opinion argues. Sponsored content is advertising in editorial clothing and is required to be labeled, usually in the smallest permissible type. Confusing an opinion column for reporting and then concluding the outlet is biased is the single most common reading error.
  • Headlines are not written by the reporter. They are written by editors under space and traffic pressure, which is why a headline sometimes overstates a careful story. Judge the article.
  • Anonymous sources are governed by rules at serious outlets: the identity is known to the editor, the reason for anonymity is stated, and corroboration is required. An unexplained anonymous source is a legitimate criticism; anonymity as such is not.
  • The correction policy is the best single credibility signal available. Does the outlet publish corrections, are they findable, and do they say plainly what was wrong? Outlets that never correct are not more accurate; they are less accountable. This one check does more work than any list of trusted sources.

The economics matter more than most bias arguments. Advertising revenue that once funded newsrooms moved to platforms, and local newspapers have closed in large numbers across the United States, leaving many counties with little or no local coverage, a situation researchers call a news desert. The consequences documented in that literature are concrete: less scrutiny of local government, lower voter turnout, and higher municipal borrowing costs, because nobody is watching. When people debate national media bias, the more consequential story is usually that nobody at all is covering their county commission.

Key idea: reporting is a documented process with editing, response, and correction, and the most useful credibility signal is whether an outlet corrects itself publicly and findably.

Interventions, at their real size

What actually reduces the spread and acceptance of false information? Several things have been tested. All of them work somewhat, and none of them is large. Reporting that accurately is itself a piece of information literacy.

InterventionWhat it isHonest assessment
Prebunking or inoculationExposing people in advance to a weakened form of a manipulation technique, then explaining itConsistently produces small to moderate improvements in spotting manipulation, in games and in short videos, at scale. Effects decay within weeks without boosters, and most studies measure recognition rather than behavior
Accuracy promptsA brief nudge asking people to consider accuracy before sharingImproves the quality of what people share by a modest amount. The finding is real and replicated by several groups, effect sizes are small, and the size and durability are actively debated
Fact-check labelsAttaching a verified-false marker to specific postsReduces belief and sharing of labeled items. Coverage is the limit: labeling is slow and partial, and the implied truth effect means unlabeled false items can gain credibility by comparison
Source credibility indicatorsShowing outlet-level trust ratings next to linksSmall effects; many users ignore them, and they shift the argument to who rates the raters
FrictionPrompting people to open an article before resharing, or limiting forward countsMeasurably reduces thoughtless resharing. Cheap, unglamorous, and among the better-supported interventions
Lateral reading instructionTeaching the fact-checker method directlyThe best-supported educational intervention: short courses produce measurable, transferable improvements in evaluation, which is why Module 4 taught it
Generic media literacy coursesBroad instruction about mediaMixed evidence. Specific, practiced procedures beat general awareness, and some approaches risk producing indiscriminate distrust

Two conclusions follow. First, treat any claim of a decisive fix with suspicion. This is a portfolio problem: many small interventions, applied together, at design level and personal level. Second, notice that the intervention with the strongest evidence is a procedure you can perform yourself, for free, in ninety seconds, without anyone's permission. That is the part of the problem you fully control.

Key idea: every tested intervention produces small effects, prebunking and accuracy prompts decay and are modest, and the best-supported one is teaching people to read laterally.

Common misconceptions

  • "Platforms want you to see false content." They optimize interaction, and false content happens to produce interaction. The absence of a truth term in the objective is the problem, not a preference for falsehood.
  • "The algorithm caused polarization." Field experiments that changed feeds for months found exposure effects but small attitude effects. Engagement ranking advantages conflictual content, which is a narrower and better-supported claim.
  • "If two outlets print the same words, they are colluding." Both are probably running the same wire copy, which is shared sourcing rather than coordination.
  • "An outlet that publishes corrections is unreliable." The opposite. A findable, specific correction record is the strongest available credibility signal, and outlets that never correct are simply less accountable.
  • "Media literacy education fixes this." Generic awareness courses show mixed results and can produce indiscriminate distrust. Specific practiced procedures such as lateral reading are what the evidence supports.

Recap

  • Attention is the scarce good in an information-abundant environment, and advertising models sell it.
  • Platforms optimize proxies for value, and Goodhart's law means a proxy stops tracking value once it becomes the target.
  • Engagement ranking advantages arousal, moral-emotional language, identity conflict, curiosity gaps, and novelty.
  • Evidence that feed ranking is the main driver of polarization is weak; the narrower claim about advantaged content is well supported.
  • Reporting involves beats, documents, response from subjects, editing, legal review, and wire sourcing; headlines are written by editors.
  • A public, findable, specific correction policy is the best single credibility signal.
  • Local news collapse and news deserts have documented civic consequences, including reduced scrutiny and higher municipal borrowing costs.
  • Prebunking, accuracy prompts, labels, indicators, and friction all help modestly; lateral reading instruction has the strongest support.

Sources

  1. Pew Research Center. (n.d.). Journalism and Media research. Pew Research Center. pewresearch.org
  2. Poynter Institute. (n.d.). International Fact-Checking Network. Poynter. poynter.org
  3. Google Jigsaw. (n.d.). Jigsaw. Google. jigsaw.google.com
  4. Wikipedia contributors. (n.d.). Inoculation theory. Wikipedia. en.wikipedia.org
  5. Wikipedia contributors. (n.d.). Attention economy. Wikipedia. en.wikipedia.org
  6. Encyclopaedia Britannica. (n.d.). Journalism. Britannica. britannica.com
Key terms
Attention economy
The condition in which information is abundant and attention is the scarce good, so attention becomes the thing that is sold.
Engagement metric
A measurable proxy for value such as time spent, clicks, shares, or comments, used to rank content.
Goodhart's law
The principle that when a measure becomes a target it ceases to be a good measure, because everyone optimizes the proxy.
Beat
A defined subject area a reporter covers continuously, building sources and background over years.
Wire service
An organization such as the Associated Press or Reuters supplying reporting that many outlets run, explaining near-identical wording across sites.
News desert
A community with little or no local news coverage, associated with reduced government scrutiny, lower turnout, and higher municipal borrowing costs.
Prebunking
Inoculation against manipulation by exposing people in advance to a weakened example of a technique and explaining it.
Accuracy prompt
A brief nudge to consider accuracy before sharing, which modestly improves the quality of what people share.
Implied truth effect
The tendency for unlabeled false items to gain credibility when other items carry fact-check labels.

Module 6: Information and Society

Who collects your information, who owns the information you use, who can reach it, and who keeps it. This closing module covers privacy and surveillance, copyright and citation, open access, the digital divide, algorithmic accountability, digital preservation, and the careers built on all of it.

Privacy, Surveillance, and What You Can Actually Do

  • Describe what is collected about you, by whom, and how inference turns ordinary data into a profile.
  • Read a privacy policy quickly for the four things that matter, and explain why notice-and-consent fails.
  • Build a personal privacy threat model and apply practical steps in order of benefit per unit of effort.

The big picture

Privacy is not secrecy. Almost nothing you want to protect is shameful; you close the bathroom door anyway. Privacy is control over the flow of information about you: who gets it, in what context, and for what purpose. That definition does more work than any list of hiding techniques, because it explains why a fact being technically public does not make its aggregation harmless. Your address is public. Your pharmacy visits are semi-public. Your movements are observable. Assembled into one profile by a party you never met, those become something you never agreed to. This lesson covers what is actually collected, why the consent system fails, and what to do about it in an order that respects your limited time.

Key idea: privacy is control over how information about you flows between contexts, which is why aggregating individually harmless facts produces a genuine harm.

What is actually collected

KindWho collects itExample
First-party activityThe service you are usingWhat you searched, watched, bought, how long you paused
Third-party trackingAdvertising and analytics companies embedded in pages and appsCookies, tracking pixels, and code kits inside apps that report your activity elsewhere
Device fingerprintingTrackers, when cookies are blockedScreen size, fonts, time zone, and dozens of other settings combined into a near-unique signature
LocationApps, and advertising kits inside themPrecise coordinates over time, which are sold onward and are extremely identifying
Purchases and loyaltyRetailers and payment intermediariesItemized histories linked to a phone number or card
Public records and broker filesData brokersAddresses, property, court records, licences, assembled and resold
InferenceEveryone aboveEstimated age, income band, health interests, political leaning, life events

That last row is the one that matters most and gets the least attention. The valuable product is not the raw log; it is the inferred profile built from it. You never disclosed that you are probably pregnant, job-hunting, or in financial trouble. It was inferred from timing, sequence, and correlation with millions of other people. This is why the phrase I have nothing to hide misses the mechanism: you are not being read, you are being modeled, and the model asserts things about you that you never said and cannot see.

Location deserves a specific warning. A few days of precise location data identifies almost anyone, because the pair of places where you spend your nights and your weekdays is nearly unique. Location histories collected through advertising kits inside ordinary apps have repeatedly been shown to be resold with weak controls. If you take one technical action from this lesson, make it your location permissions.

Key idea: the product is the inferred profile rather than the raw log, and location data is the most identifying category because your home and work pair is nearly unique.

Why privacy policies do not protect you

The governing model in most places is notice-and-consent: a company discloses its practices, you agree, and consent is presumed. It fails for reasons that are structural rather than fixable by trying harder.

  • Length. A frequently cited estimate found that reading every privacy policy an average person encounters in a year would take several weeks of full-time work. Nobody has that time, and everyone knows it.
  • Readability. Policies are typically written at a level requiring years of higher education, in language chosen by lawyers to preserve options.
  • Vagueness by design. Phrases such as we may share information with our partners and affiliates for business purposes authorize almost anything while promising almost nothing.
  • No real alternative. Consent obtained under a take-it-or-leave-it condition, for a service required for school or work, is not meaningfully voluntary.
  • Downstream invisibility. Even a policy you read fully cannot tell you what the fourth company down the chain does, because it does not know either.

Still, reading one is worth ten minutes as a skill, and there is a fast method. Do not read it linearly. Search the document for four things: what categories are collected, who it is shared with or sold to, how long it is retained, and how to delete your account and data. Those four answers, which usually take about ninety seconds to locate with a text search, tell you more than the whole document read front to back.

Key idea: notice-and-consent fails because of length, readability, deliberate vagueness, and lack of alternatives, so read policies by searching for collection, sharing, retention, and deletion.

The legal landscape, briefly

Rules differ enormously by jurisdiction, and it is worth knowing which regime you live under.

  • European Union, the GDPR. Establishes rights that travel with the person: access what is held about you, correct it, erase it in many circumstances, port it elsewhere, and object to certain processing. It requires a lawful basis for processing rather than merely a disclosure, which is the important structural difference.
  • United States, sectoral and state-based. There is no single national consumer privacy law. Specific sectors are covered, such as health information under HIPAA and student education records under FERPA, and a growing number of states, beginning with California, grant rights to know, delete, and opt out of sale.
  • Everywhere. Rights that exist on paper require someone to exercise them. Where you have a right to request deletion, exercising it against the largest brokers is a real, if tedious, action with real effect.

Note the boundary of this course: how systems defend data technically, threat modeling for attackers, encryption, and authentication belong to CS 305 Cybersecurity Fundamentals. Here we are concerned with who collects information, what they are permitted to do with it, and what you can decide.

Key idea: the GDPR requires a lawful basis and grants portable individual rights, while United States protection is sectoral and state-based, so your options depend heavily on where you are.

Build a threat model before you buy tools

The most common privacy mistake is installing tools without asking what you are protecting and from whom. Different answers imply completely different actions.

ConcernWhat actually helps
Advertisers and brokers building a profileTracker blocking, location permissions, ad identifier reset, broker opt-outs
A stalker or abusive former partnerAccount security, location sharing audit, removing metadata from photos, people-search site removals; a specialist support organization, not a browser extension
An employer or schoolDo not use their device or network for private matters; policy, not technology, governs this
A future admissions officer or hiring managerAudit what is publicly attached to your name; nothing technical helps after publication
Casual snooping by people nearbyDevice lock, screen privacy, notification settings

Notice that the most serious case in that table, an abusive person with prior physical access to your devices and knowledge of your accounts, is the one where generic privacy advice is least adequate and where specialist help matters most. Naming the threat first prevents you from spending your effort in the wrong place.

Key idea: decide who you are protecting what from before choosing any tool, because the correct action for advertisers and the correct action for an abusive individual have almost nothing in common.

Practical steps, in order of benefit per unit of effort

  1. Audit phone permissions, starting with location. Go through your installed apps and set location to never or only while using for everything that does not obviously need it. Do the same for microphone, camera, contacts, and photos. This takes fifteen minutes and is the highest-value action available.
  2. Reset and limit your advertising identifier. Both major mobile platforms provide a setting to reset or disable the identifier used to link your activity across apps.
  3. Install a reputable content blocker and turn off third-party cookies in your browser. This removes a large share of routine cross-site tracking for a one-time cost of a few minutes.
  4. Separate your identities. Use a distinct email address for shopping and signups, and reserve one for anything consequential. This limits how easily separate profiles get joined.
  5. Reduce what you volunteer. Loyalty programs, quizzes, and forms requesting a birth date and phone number for no reason are all collection points. A partial answer is usually accepted.
  6. Exercise your legal rights. If your jurisdiction offers deletion or opt-out rights, use them against the largest data brokers and people-search sites. Expect the process to be deliberately tedious.
  7. Practice metadata hygiene. Remember from Module 2 that images can carry GPS coordinates. Check what your camera writes and strip it before publishing photos of anywhere you live or work.
  8. Search yourself. Once a term, search your own name and see what is publicly attached to it. You cannot manage what you have not looked at.

One honest caveat to close on. None of this makes you invisible, and it should not have to be your job. Individual action is genuinely worth doing and is also an inadequate substitute for regulation, because the burden of reading thousands of policies and opting out of hundreds of brokers cannot reasonably fall on each person. Do the eight things above, and support the policy conversation, without pretending the first substitutes for the second.

Key idea: permission auditing and tracker blocking give the largest return for the least effort, and individual action is worthwhile without being a substitute for regulation.

Common misconceptions

  • "I have nothing to hide." The issue is aggregation and inference, not secrets. A model built from your behavior asserts things about you that you never said and cannot inspect or correct.
  • "Private browsing makes me anonymous." It stops your browser from keeping local history. Your network provider, the sites you visit, and any embedded trackers see the same things they always did.
  • "If it is free, it is not worth much." The service is genuinely valuable; the exchange is simply attention and data rather than money, and the terms of that exchange are not visible to you.
  • "Anonymized data is anonymous." Location traces and combinations of ordinary attributes re-identify individuals with high accuracy. Removing a name is not the same as removing identity.
  • "Accepting cookies is required to browse." Most consent banners must offer a reject option for non-essential cookies, and choosing it rarely affects the site's function.

Recap

  • Privacy is control over how information flows between contexts, not secrecy.
  • Collection spans first-party activity, third-party trackers, fingerprinting, location, purchases, broker files, and inference.
  • The inferred profile is the product, and location is the most identifying single category.
  • Notice-and-consent fails on length, readability, vagueness, and absence of alternatives; read policies by searching for collection, sharing, retention, and deletion.
  • The GDPR requires a lawful basis and grants portable rights; United States coverage is sectoral and state-based.
  • Build a threat model first, because different adversaries require completely different responses.
  • Permission auditing, ad identifier limits, tracker blocking, identity separation, and rights requests give the best return.

Sources

  1. Electronic Frontier Foundation. (n.d.). Surveillance Self-Defense. EFF. ssd.eff.org
  2. Electronic Frontier Foundation. (n.d.). About EFF. EFF. eff.org
  3. Federal Trade Commission. (n.d.). Privacy, Identity and Online Security. FTC Consumer Advice. consumer.ftc.gov
  4. Pew Research Center. (n.d.). Privacy and Information Sharing research. Pew Research Center. pewresearch.org
  5. Wikipedia contributors. (n.d.). General Data Protection Regulation. Wikipedia. en.wikipedia.org
Key terms
Contextual integrity
The idea that privacy means information flowing appropriately within its original context, which aggregation across contexts violates.
Third-party tracking
Collection by advertising and analytics companies embedded in pages and apps you did not choose to visit.
Device fingerprinting
Identifying a device from a combination of settings such as screen size, fonts, and time zone, which works even when cookies are blocked.
Data broker
A company that assembles and resells personal profiles built from public records, purchases, and app-derived data.
Inference
Deriving attributes such as health interests, income band, or life events from behavior, producing claims you never disclosed.
Notice-and-consent
The regulatory model in which disclosure plus agreement is treated as consent, undermined by length, readability, and lack of alternatives.
GDPR
European Union regulation requiring a lawful basis for processing and granting rights of access, correction, erasure, and portability.
Privacy threat model
An explicit statement of what you are protecting and from whom, which determines which measures are actually relevant.
Re-identification
Linking supposedly anonymized records back to individuals using location traces or combinations of ordinary attributes.

Divides, Algorithms, Archives, and Careers in Information

  • Describe the three levels of the digital divide and why access alone does not close it.
  • Identify the sources of algorithmic bias and the mechanisms proposed for accountability.
  • Explain link rot and digital preservation, and name the professions built on information work.

The big picture

The last lesson widens the frame. Everything you have learned assumes a person with a device, a connection, the skills to use them, and something left to find. Each of those assumptions fails for large numbers of people, and the failures are patterned rather than random. This lesson covers who is excluded, how automated decisions encode old patterns, why so much of the record disappears, and finally what it looks like to do this work for a living. It also closes the course, so the last section asks what you should now actually do differently.

Key idea: access, skill, and survival of the record are three separate preconditions for information to reach anyone, and each fails in patterned ways.

The digital divide has three levels

The phrase digital divide originally meant who has a connection and who does not. Researchers now describe three distinct levels, and confusing them produces policies that fail.

  1. Access. Whether a person has a device and a connection adequate for the task. In the United States a substantial minority of adults still lack home broadband, and a meaningful share are smartphone-dependent, meaning their only internet access is a phone. Adoption is consistently lower among lower-income households, older adults, and rural residents, and the gaps have narrowed without closing.
  2. Skills and usage. Whether a person can do consequential things with the connection they have. Someone using a phone for messaging and video is online in the counting sense and cannot easily complete a job application, write a paper, or fill in a benefits form. This is why access-only programs disappoint.
  3. Outcomes. Whether being online translates into tangible benefit: better employment, health information used correctly, civic participation. This is where the divide compounds existing advantage rather than reducing it.

Two consequences worth carrying. The homework gap names students expected to complete assigned work online without adequate home connectivity, which turns an academic requirement into a test of household income. And moving a service online is never neutral: every migration to digital-only reduces cost for the provider and transfers cost to the least connected users. Public libraries absorb an enormous amount of this, functioning as the connectivity and skills provider of last resort, which is a considerable part of what a library does now and almost none of what people picture.

Key idea: access, skills, and outcomes are three separate divides, so counting connections measures the easiest level and misses the two that determine whether being online helps.

Algorithmic bias and what accountability would mean

Automated systems now rank, filter, score, and sort people. When those systems produce patterned disadvantage, it is rarely because someone encoded prejudice deliberately. The mechanisms are more mundane and therefore harder to remove:

  • Training data reflects history. A model learns from past decisions. If past hiring favored a group, a model trained to predict past hiring outcomes reproduces that pattern faithfully, and its consistency reads as objectivity.
  • Proxy variables. Removing a protected attribute does not remove it, because postal code, school attended, and purchasing patterns correlate with it. The model reconstructs what you deleted.
  • Label choice. The thing being predicted is a human choice, and it is often a convenient stand-in for what you actually care about. A widely reported healthcare example used prior spending as a proxy for medical need, and because less money had historically been spent on Black patients with the same conditions, the system understated their need.
  • Feedback loops. Predicting where problems will be found, then sending resources there, generates more records there, which confirms the prediction. The system measures its own attention.
  • Deployment context. A tool validated on one population and applied to another fails in ways nobody tested for.

What accountability looks like in practice, and where each mechanism runs out:

MechanismWhat it doesLimit
AuditingIndependent testing for disparate outcomes across groupsRequires access that owners often refuse on trade secrecy grounds
DocumentationStandardized descriptions of a dataset or model, its intended use, and its known limitsVoluntary in most places, and easy to make thin
Impact assessmentStructured evaluation before deployment, as encouraged by risk management frameworksCan become a compliance ritual if nothing hinges on the finding
ContestabilityAn affected person can see the basis of a decision and appeal it to a humanThe most useful mechanism and the least commonly implemented
TransparencyPublishing how a system worksFull disclosure invites gaming, and technical explanations rarely help the affected person

Notice that this is a documentation and description problem, which is to say an information science problem. Datasets need provenance, models need metadata, decisions need records that survive. The vocabulary you learned in Module 2 is exactly the vocabulary this field has reached for.

Key idea: algorithmic bias usually comes from historical training data, proxy variables, and the choice of what to predict, and the most useful accountability mechanism is a person's ability to see and contest a decision.

Archives, and why they are not libraries

A library collects published material, usually in many copies held by many institutions, and organizes it by subject so that any copy substitutes for any other. An archive holds unique, mostly unpublished material, the records of a person or organization, and organizes it on entirely different principles:

  • Provenance: records from one creator stay together and are not intermixed with another's, because who created a record and why is evidence in itself.
  • Original order: the creator's own arrangement is preserved, because it reveals how they worked and what they considered related.
  • Description at the collection level: archivists describe boxes and series through a finding aid rather than cataloging every item, since item-level description of forty boxes is impossible.
  • Appraisal: deciding what to keep permanently and what to destroy. This is the most consequential and least visible act in the whole information world. Most records are not kept, and someone chooses.

Key idea: archives keep unique records organized by provenance and original order, described at collection level, and appraisal decides permanently what survives.

Digital preservation and link rot

Digital material is fragile in ways paper is not. Paper degrades slowly and remains readable. Digital material fails abruptly and in several distinct ways:

  • Media decay. Storage fails, and silent corruption alters bits without announcing itself, which is why serious preservation stores checksums and verifies them on a schedule.
  • Format obsolescence. A file whose format has no current reader is a rectangle of bytes.
  • Dependency. Interactive work needs a running environment, so preserving the file preserves only part of the object.
  • Link rot and content drift. The web address you cited stops resolving, or worse, resolves to different content than it did when you cited it. Studies of scholarly and legal citations have repeatedly found that a large share of cited web addresses no longer work after a few years, including significant proportions of links in law review articles and in United States Supreme Court opinions. Content drift is the quieter problem, because a working link to changed content produces no error at all.

Against this stand deliberate efforts: the Internet Archive, whose Wayback Machine has captured hundreds of billions of web pages and which also lends digitized books; national library web archiving programs; Perma.cc, which creates permanent citation links for legal and scholarly writing; and distributed schemes that keep multiple independent copies precisely because a single copy in a single institution is a single point of failure.

What you should personally do is short and worth adopting today. Cite DOIs when they exist. When you cite a web page, save it to a web archive first and cite the archived snapshot alongside the live link. Download and keep a copy of anything your own work truly depends on. And remember that absence of a page today tells you nothing about whether it existed, which is why checking an archive is a research technique and not just a curiosity.

Key idea: digital material fails through media decay, format obsolescence, dependency, and link rot, and content drift is worse than link rot because a working link to changed content raises no error.

Working in information

These skills are a profession, and the profession is larger and stranger than most people assume.

RoleWhat the work actually isUsual entry route
LibrarianPublic, academic, school, medical, law, or corporate; instruction, reference, collections, systems, and increasingly data servicesA master's degree in library and information science, commonly the MLIS
ArchivistAppraisal, arrangement, description, and access for unique recordsMLIS with an archives concentration, or a history graduate degree plus certification
Records managerRetention schedules, compliance, and disposition across an organizationInformation or business degrees, with professional certification
Data curator or data librarianDocumentation, metadata, and repository work that makes datasets reusableMLIS or a domain degree plus data skills
Metadata or taxonomy specialistDesigning and maintaining controlled vocabularies for organizations and productsMLIS, linguistics, or applied experience
UX researcher or information architectStudying how people look for things and structuring products so they find themHuman computer interaction, psychology, design, or an iSchool degree
Search relevance or knowledge engineerImproving retrieval quality inside a company's search or knowledge systemsComputing plus retrieval knowledge
Digital preservation specialistFormat migration, integrity checking, and repository operationsMLIS or archives training plus technical skills

Two honest notes. First, the traditional library job market is competitive and salaries vary enormously by sector, so anyone considering the MLIS should look at real postings and real salary data before enrolling rather than after. Second, and more useful for most readers: you are unlikely to become a cataloger, and you will absolutely find yourself deciding what fields a system should store, what a category should be called, what gets kept, and who can find it. Every organization has these problems and most handle them badly. Being the person in the room who knows that vocabulary is a genuine and portable advantage.

Key idea: information work spans libraries, archives, records, data curation, taxonomy, user experience, and search, and its core skills apply in any organization that stores anything.

Where this leaves you

Six habits carry most of this course's value, and they cost almost nothing:

  1. Write the formalized question before you search, and build the query from it in blocks.
  2. Find one good record, then steal its subject headings and its reference list.
  3. Treat a paywall as a queue and file the interlibrary loan request.
  4. Read laterally, in ninety seconds, before you believe or share anything unfamiliar.
  5. Find the baseline behind any percentage, and verify every citation you did not personally open.
  6. Archive what you cite, and keep a copy of what you depend on.

Key idea: the whole course reduces to six repeatable habits, and repeating them is what distinguishes someone who can find and judge information from someone who cannot.

Common misconceptions

  • "The digital divide is basically solved." Access gaps persist, and the skills and outcomes divides are larger and less measured than the access divide.
  • "Removing race or gender from a model removes bias." Proxy variables such as postal code and school reconstruct the attribute you deleted.
  • "Archives are old libraries." They hold unique unpublished records, organize by provenance and original order, describe at collection level, and appraise what to destroy.
  • "The Internet Archive means nothing is lost." It captures an enormous but partial slice, misses material behind logins and paywalls, and honors exclusion requests. It is a magnificent safety net with holes.
  • "A working link means the source still says what I cited." Content drift means a live link can silently point to different content, which produces no error and is therefore easy to miss.

Recap

  • The digital divide has access, skills, and outcomes levels, and counting connections measures only the first.
  • Libraries function as the connectivity and skills provider of last resort, and every digital-only migration transfers cost to the least connected.
  • Algorithmic bias arises from historical training data, proxy variables, label choice, feedback loops, and deployment context.
  • Accountability mechanisms include auditing, documentation, impact assessment, contestability, and transparency, each with real limits.
  • Archives keep unique records by provenance and original order, and appraisal permanently decides what survives.
  • Digital material fails through media decay, format obsolescence, dependency, link rot, and content drift.
  • Information careers span libraries, archives, records, data curation, taxonomy, user experience, search, and preservation.

Sources

  1. Pew Research Center. (n.d.). Internet and Technology research. Pew Research Center. pewresearch.org
  2. Internet Archive. (n.d.). About the Internet Archive. Internet Archive. archive.org
  3. National Archives and Records Administration. (n.d.). About the National Archives. NARA. archives.gov
  4. Society of American Archivists. (n.d.). So You Want to Be an Archivist. SAA. archivists.org
  5. Bureau of Labor Statistics. (n.d.). Education, Training, and Library Occupations. Occupational Outlook Handbook, U.S. Department of Labor. bls.gov
  6. National Institute of Standards and Technology. (n.d.). AI Risk Management Framework. NIST. nist.gov
Key terms
Digital divide
Patterned inequality in access to devices and connectivity, in the skills to use them, and in the outcomes people obtain from being online.
Smartphone-dependent
Having a phone as one's only internet access, which permits messaging and video but makes forms, applications, and documents difficult.
Homework gap
The gap affecting students expected to complete assigned work online without adequate home connectivity.
Proxy variable
A feature such as postal code or school attended that correlates with a protected attribute, reconstructing it after it has been removed.
Contestability
The ability of a person affected by an automated decision to see its basis and appeal it to a human.
Provenance principle
The archival rule that records of one creator are kept together and not intermixed, because origin is itself evidence.
Appraisal
The archival decision about what records are kept permanently and what is destroyed.
Link rot
The failure of previously working web addresses over time, affecting large shares of citations in scholarly and legal writing.
Content drift
A live link that resolves to different content than when it was cited, producing no error and therefore going unnoticed.
MLIS
The Master of Library and Information Science, the usual professional entry credential for librarianship and many archives roles.

Open the interactive version with quizzes and progress →