Module 1: The Field, the Scene, and the Courtroom Door
What forensic science is and what a crime laboratory's caseload really contains, how television reshaped what jurors expect, how a scene is searched, documented and collected without destroying what it holds, and the legal standards that decide whether an examiner ever gets to speak to a jury at all.
What Forensic Science Actually Is
- Define forensic science and distinguish criminalistics from the wider set of forensic disciplines.
- Describe what a publicly funded crime laboratory's caseload actually contains and how requests move through it.
- Distinguish class from individual characteristics, and identification from individualization, using correct terminology.
- Evaluate claims about the CSI effect against the survey evidence rather than the anecdotes.
Two attic rooms in Lyon
In 1910 the Lyon police gave a young physician named Edmond Locard two attic rooms above the courthouse, a microscope, and almost no money. He had been asking for a laboratory attached to a police department, an idea nobody had funded before. What he got was a garret. Out of it came a working principle that still organizes the whole field: every contact between two things leaves a trace. Fibers move from a coat to a car seat. Soil moves from a driveway to a shoe. Skin cells move from a hand to a doorknob. Locard's claim was not that these traces are always found. It was that they are always made.
That is the optimistic half of forensic science. The pessimistic half is that traces degrade, get walked on, get rained on, get vacuumed up by a well-meaning property owner, and get contaminated by the people trying to collect them. Everything in this course happens between those two facts.
Forensic science is the application of scientific methods and techniques to questions that arise in legal proceedings. The word forensic comes from the Latin forum, the public place where Roman cases were argued, and the etymology carries a warning worth taking seriously from the first page: this is science with an audience, performed under adversarial conditions, for a decision maker who is usually not a scientist. The chemistry is the same chemistry. The consequences are not.
Key idea: Forensic science is ordinary science applied to legal questions, and the legal setting, not the science, is what makes it distinctive and difficult.
The disciplines, and the one word that gets misused
People often use forensic science and criminalistics as synonyms. They are not. Criminalistics is the subset that deals with physical evidence from crime scenes: drugs, DNA, fingerprints, firearms, trace, fire debris. The wider field also includes forensic pathology, forensic anthropology, forensic odontology, forensic psychiatry, forensic engineering, forensic accounting, and forensic toxicology, most of which are practiced by people who trained in a parent discipline first and applied it to legal questions second. A forensic pathologist is a physician. A forensic anthropologist is an anthropologist. A latent print examiner, by contrast, is usually trained entirely inside the forensic world, and that difference will matter enormously by the end of this course.
It is also worth naming who is not a forensic scientist. The detective who works the case is an investigator. The evidence technician who photographs and collects at the scene may or may not be a scientist, depending on the agency. The medical examiner's investigator who responds to a death scene is a third kind of person again. Television collapses all of these into one character who dusts for prints, runs the DNA, interrogates the suspect, and makes the arrest. No real system works that way, and the reasons it does not are mostly good ones.
What actually arrives at the lab
Here is a number that reorganizes most students' picture of the field. The Bureau of Justice Statistics counted roughly 400 publicly funded forensic crime laboratories in the United States and estimated that they received on the order of 3.8 million requests for analysis in a single year. The largest single category of those requests, by a wide margin, was not DNA and not fingerprints. It was controlled substances: identifying the white powder, the pills, the plant material seized in an arrest.
A working crime lab is, first and foremost, an analytical chemistry operation. A drug chemist may run dozens of samples in a week, each one a gas chromatograph and mass spectrometer run whose output is an unambiguous chemical identification. The DNA unit down the hall may take weeks per case. The latent print unit may have a backlog measured in months. Meanwhile most of the building is weighing powder.
| Unit | Typical work | What the answer looks like |
|---|---|---|
| Controlled substances | Identify seized material by chemical analysis | A definite chemical identity and a weight |
| Toxicology | Detect and quantify drugs and alcohol in blood, urine, tissue | A concentration, plus interpretation limits |
| Forensic biology and DNA | Screen for body fluids, extract, amplify, type STR loci | A profile, plus a statistic |
| Latent prints | Develop, photograph, compare friction ridge impressions | An examiner's conclusion, plus a rationale |
| Firearms and toolmarks | Compare fired components under a comparison microscope | An examiner's conclusion, plus a rationale |
| Trace | Fibers, glass, paint, hair, fire debris, residues | Usually a class association, rarely more |
| Digital | Image devices, recover files, reconstruct activity | Artifacts and timelines, with attribution problems |
Look down the right-hand column and you will see the fault line this whole course runs along. Some units produce measurements. Others produce examiner conclusions. Both go to court in the same envelope, wearing the same institutional authority, and juries have no obvious way to tell them apart.
The point: Most crime laboratory work is analytical chemistry with objective answers, but the disciplines that generate the most courtroom drama are the ones that end in a human judgment.
Class, individual, and the words examiners are allowed to use
Two terms do a great deal of work in this field. A class characteristic is a feature shared by a group of items: the tread pattern molded into every shoe of a given model and size, the number of lands and grooves cut by a particular make of barrel, the fact that a fiber is nylon rather than cotton. A individual characteristic is a feature said to arise from random events during manufacture, use, or damage, which is therefore claimed to distinguish one item from all others of its class: a nick in a knife edge, a stone cut into a shoe sole, a scratch pattern on a firing pin.
From that pair comes the field's most contested distinction. Identification in the chemical sense means determining what a substance is: this powder is cocaine hydrochloride. Individualization means asserting that a questioned item came from one specific source in the world and no other: this latent print was made by this finger. Chemical identification rests on physical properties that can be measured and reproduced. Individualization rests on a premise about uniqueness plus an examiner's judgment about how much agreement is enough. For a century, examiners in several disciplines testified to individualization in absolute terms. Much of Module 6 is about what happened when scientists outside those disciplines finally asked for the studies supporting it.
You are going to be corrected on vocabulary in this course more than in most, because in forensic science the vocabulary is the claim. The difference between saying two hairs are microscopically similar and saying they match is the difference between an association and an accusation.
Worth holding on to: Class characteristics narrow a pool; individual characteristics are claimed to pick out a single source, and that claim always needs evidence behind it.
The path a piece of evidence takes
Follow a single item through. A knife is found in a storm drain two blocks from a stabbing. A patrol officer secures the location. A crime scene technician photographs it in place with a scale, records the position, and packages it in a rigid container so nobody is cut and no trace is rubbed off. A chain of custody form is opened, and every transfer from that moment forward is signed. The knife goes to an evidence room, is logged, and sits in a temperature-controlled locker.
Later a request is submitted. In the lab, a biology screener swabs the handle for touch DNA and the blade for blood, and those swabs go one way while the knife itself goes to trace and then to a pathologist who will compare the blade dimensions to the wound tracks. Each analyst writes a report. Each report enters discovery, where the defense receives it. Months or years later, an analyst is subpoenaed, qualified as an expert, examined, and cross-examined about work she did on a Tuesday she does not remember, in a case she may have handled alongside two hundred others.
Two structural features of that path deserve early attention. First, most publicly funded crime laboratories in the United States are administratively part of a law enforcement agency: a police department, a sheriff's office, a state department of public safety. The analyst's employer is the same institution as the investigator's. Second, the analyst's report is written for a legal audience and will be read by people looking for a sentence they can use. Both facts are ordinary, both are defensible, and both create pressures that the 2009 National Research Council review took seriously enough to build a recommendation around.
The CSI effect, measured instead of asserted
Since the early 2000s, prosecutors have complained that television has taught jurors to expect DNA in every case and to acquit without it. Defense lawyers have complained about the mirror image: that television has taught jurors to treat forensic conclusions as infallible. The phrase covering both complaints is the CSI effect, and it is repeated so often that it is easy to assume somebody proved it.
Somebody tried. Judge Donald Shelton and colleagues surveyed more than a thousand people summoned for jury duty in Ann Arbor, Michigan, asking what evidence they expected to see and how they would vote in various scenarios. Expectations were indeed high: a large share of respondents said they expected some kind of scientific evidence in a criminal case. But when Shelton compared those expectations to how respondents said they would decide, the predicted acquittal effect did not appear cleanly. Watching the shows was not a good predictor of demanding scientific evidence before convicting. Shelton's own summary in the National Institute of Justice Journal was that the CSI effect, as prosecutors describe it, was not supported by his data, though jurors' expectations about technology had plainly risen, probably from living in a world full of technology rather than from any one program.
That is a useful first lesson in reading forensic claims. A vivid, widely repeated, professionally endorsed proposition turned out to be partly right, partly wrong, and mostly untested. Hold that shape in mind. You will see it again with hair, with bite marks, and with fire.
Why this matters: The best-known claim about forensic evidence and juries was asserted for years before anyone measured it, and when someone did, the data were messier than either side's version.
Common misconceptions
- Crime labs mostly do DNA. The largest category of requests at publicly funded laboratories is controlled substance analysis; DNA is a minority of the work and a large share of the delay.
- Results come back in an hour. Turnaround is measured in weeks to months, and backlogs in some disciplines are counted in the thousands of cases.
- The same person works the scene, runs the tests, and interrogates the suspect. These are separate roles in separate units, deliberately, because combining them concentrates both error and bias.
- A forensic conclusion is a measurement. Some are. Many are an examiner's judgment about whether two things agree closely enough, which is a different kind of claim and carries a different kind of error.
- Forensic scientists work for the defense as often as the prosecution. The overwhelming majority of forensic analysis in the United States is performed by laboratories inside law enforcement agencies, and defense access to independent testing is uneven and often unfunded.
Where this leaves us
- Forensic science applies scientific methods to legal questions; criminalistics is its physical-evidence core.
- Locard's exchange principle says traces are always created, not that they are always recovered.
- A real crime laboratory is mostly an analytical chemistry operation, with drug identification the largest category of requests.
- Class characteristics narrow a pool; individualization claims a single source, and that claim requires validation.
- Evidence travels scene, evidence room, laboratory, report, discovery, testimony, and its integrity can fail at any step.
- The CSI effect is real as a change in expectations and poorly supported as a claim about verdicts.
Sources
- National Institute of Standards and Technology. (n.d.). Forensic science. nist.gov
- Britannica. (2024). Forensic science. britannica.com
- National Institute of Justice. (n.d.). Forensic sciences. Office of Justice Programs. nij.ojp.gov
- Bureau of Justice Statistics. (2016). Publicly funded forensic crime laboratories: Resources and services, 2014. U.S. Department of Justice. bjs.ojp.gov
- Shelton, D. E. (2008). The CSI effect: Does it really exist? NIJ Journal, 259. National Institute of Justice.
- Key terms
- Forensic science
- The application of scientific methods and techniques to questions that arise in legal proceedings.
- Criminalistics
- The branch of forensic science dealing with physical evidence recovered from crime scenes, including drugs, DNA, prints, firearms, and trace.
- Locard's exchange principle
- The principle that contact between two objects transfers material in both directions, so traces are always created even when they are not recovered.
- Class characteristic
- A feature shared by all members of a group of items, such as a shoe model's tread pattern or a fiber's polymer type.
- Individual characteristic
- A feature claimed to arise from random events in manufacture, use, or damage, and therefore to distinguish one item from others of its class.
- Identification
- Determining what a substance or object is, such as establishing that a powder is cocaine hydrochloride.
- Individualization
- The assertion that a questioned item originated from one specific source and no other; a much stronger claim than identification.
- CSI effect
- The proposed influence of forensic television on juror expectations and verdicts, widely asserted and only weakly supported by survey research.
Working a Scene: Documentation, Collection, and the Chain of Custody
- Sequence the tasks of scene processing from securing a perimeter through release, and explain why the order matters.
- Choose an appropriate search pattern and packaging method for a given item and scene type.
- Explain what a chain of custody records, how it fails, and what a break in it does legally.
- Identify realistic contamination pathways and the controls that guard against them.
Nine days about a plastic bag
In April 1995, a Los Angeles Police Department criminalist named Dennis Fung spent parts of nine days on the witness stand. Barry Scheck's cross-examination in the O. J. Simpson trial was not about whether the DNA typing had been done correctly. It was about what happened to the blood before it ever reached a laboratory. Wet cotton swatches had been placed in plastic bags and left in a warm vehicle for hours. A reference vial of the defendant's own blood had been carried around by a detective for most of a day instead of being booked. Much of the collection had been done by a trainee on one of her earliest scenes.
None of this proved anyone tampered with anything. What it did was hand the defense a story about carelessness that the jury could follow without understanding a single thing about polymerase chain reaction. The prosecution had the science. It lost the argument about the bag.
Here is the underlying chemistry, which is not complicated and which is exactly why the mistake was a mistake. Blood on a cotton swatch is a warm, moist, nutrient-rich surface. Sealed in plastic, it stays moist. Bacteria multiply, and bacterial and endogenous nucleases degrade DNA into fragments too short to type reliably. Paper packaging breathes, the swatch dries, and degradation slows to a crawl. That is the whole reason for the rule: biological evidence is air-dried and packaged in paper. The rule is a sentence long. Nine days of testimony grew out of ignoring it.
Key idea: Scene work sets a permanent ceiling on what any laboratory can later do, and no amount of instrumentation recovers information that was destroyed before collection.
The first fifteen minutes
Scene processing has a sequence, and the sequence is not aesthetic. Life safety comes first: render aid, and if that means a paramedic walks through a blood pattern, the paramedic walks through it. Second comes scene security, which in practice means establishing a perimeter far larger than the obvious scene, because perimeters are easy to shrink later and impossible to expand once fifty people have walked through. Third comes a scene log: a written record of every person who enters and leaves, with times. That log is the single cheapest piece of paperwork in forensic science and the one most often skipped.
Only then does anything get touched. The single most common cause of lost evidence is not sabotage or exotic contamination. It is people. Officers, supervisors, detectives, and curious visitors who have no task at the scene walk through it, and each of them adds fibers, hairs, shoeprints, and skin cells while subtracting what was there. A scene that has been visited by twenty people is not a scene; it is a scene plus twenty people.
Documenting before disturbing
Everything gets recorded before anything moves, in four overlapping ways, because each one catches what the others miss.
Notes are contemporaneous and dull on purpose: time of arrival, weather, lighting, door and window positions, appliance states, who was present. A detail that seems pointless at 3 a.m. can become the case two years later, and nobody remembers whether the porch light was on.
Photography works in three ranges. Overall shots establish the scene and its context. Midrange shots tie an item to a fixed reference such as a doorway, so a jury can locate it. Close-ups capture the item itself, and each close-up is taken twice, once without a scale and once with an ABFO scale or ruler in the same plane as the evidence, so that later measurement is possible without the accusation that the scale was placed to distort. Photograph before, during, and after, and photograph things you have not yet decided are evidence.
Sketching supplies what photographs cannot: measured relationships. A rough sketch is drawn at the scene with real measurements written on it, and a finished diagram is prepared later. Measurements are taken by a repeatable method, usually rectangular coordinates from two fixed walls, baseline measurements along a stretched tape, or triangulation from two permanent points. Increasingly, agencies capture the whole room with a 3D laser scanner or photogrammetry, which produces a navigable point cloud, but the scan is a supplement, not a substitute, because a scanner records geometry and not judgment.
Video narrates the walk-through and preserves spatial relationships that still images fragment. Silence is preferable to commentary; recorded speculation at a scene has an unfortunate habit of resurfacing in cross-examination.
Searching on purpose
You do not wander a scene. You choose a pattern, cover the area systematically, and document what you covered so a later reader knows where you did not look.
| Pattern | How it runs | Best suited to |
|---|---|---|
| Line or strip | Searchers walk parallel lanes across the area | Large open outdoor areas, fields, roadsides |
| Grid | Two line searches at right angles, doubling coverage | High-value outdoor scenes where a miss is unacceptable |
| Spiral | Inward or outward circular path from a center point | Scenes with a single focal item, such as a body in open ground |
| Zone or quadrant | Area divided into sectors, each searched independently | Structures, vehicles, rooms with defined boundaries |
| Wheel or ray | Searchers move outward along spokes from a center | Small circular scenes; coverage thins with distance |
Notice the weakness built into the last row. Spokes diverge, so a ray search leaves widening gaps as it moves outward. Knowing a method's failure mode is more useful than memorizing its name.
Packaging: the rules and the reasons
Each item is packaged separately, sealed, and labeled across the seal so that opening the package breaks the writing. Beyond that, packaging follows the physics and biology of the material.
| Evidence | Package in | Because |
|---|---|---|
| Wet or damp biological stains | Air-dry, then paper bag or envelope | Moisture drives bacterial growth and nuclease activity that degrade DNA |
| Fire debris | Clean unused metal cans or specialized nylon bags | Ignitable liquid residues are volatile and escape ordinary plastic |
| Trace: fibers, hairs, glass | Druggist fold inside a labeled envelope | Small particles migrate through the corners of ordinary paper bags |
| Firearms | Unloaded, secured in a box, ammunition packaged separately | Safety, plus preservation of any residue or biological material |
| Digital devices | Shielded bag or airplane mode, power documented | A phone with a signal can be wiped remotely while in custody |
| Liquids and blood tubes | Leakproof secondary container, refrigerated | Preservatives fail and samples putrefy at ambient temperature |
Two habits separate careful collection from routine collection. First, controls: when you swab a surface, you also swab a nearby unstained area with a clean swab and submit it, so the laboratory can tell whether anything you found was on the whole surface rather than in the stain. Second, elimination samples: DNA and prints from the responders, the technicians, the residents, and the person who found the body, so that a profile turning up later can be attributed rather than pursued for two years.
The point: Every packaging rule encodes a specific mechanism of loss, and if you can state the mechanism you will never need to memorize the rule.
Chain of custody
A chain of custody is a continuous documented record of who had an item, from collection to court, with every transfer signed and dated. It is not a formality invented by lawyers. It answers a question the state has to answer in every case: is the thing you tested the thing you seized, and is it in the same condition?
What happens when the chain has a gap? In most American jurisdictions, less than students expect. Courts generally treat gaps as going to the weight of the evidence rather than its admissibility, meaning the jury hears about the gap and decides what to make of it. Admissibility is usually denied only where the gap is severe enough that the item cannot reasonably be identified as the one collected. This is precisely why a defense attorney invests days in the subject: the goal is not usually exclusion, it is doubt.
A modern evidence system is barcoded and electronic, which is a real improvement, and it does not solve the underlying human problem. Items still get signed out for court and returned late, transferred between agencies, consumed in testing, and, at the far end of the process, purged. Sexual assault kits sat unanalyzed in police storage in city after city, a failure of custody in the ordinary sense rather than the legal one, and you will meet the consequences in the DNA module.
What scene work cannot settle
Scene processing shades into scene reconstruction, and the boundary deserves care. Collection asks what is here. Reconstruction asks what happened, and it is a far more demanding claim. Bloodstain pattern analysis is the clearest example. Some of it is straightforward physics: a drop striking a surface at a shallow angle makes a longer ellipse than one striking perpendicular, and the ratio of width to length gives an impact angle. Some of it, particularly claims about the mechanism that produced a pattern or the position of a person in a room, is far less secure. The 2009 National Research Council review singled the discipline out as one where the opinions were more subjective than they appeared, and a large study published in 2021 found that experienced analysts frequently disagreed with one another, and sometimes with their own earlier conclusions on the same patterns.
The honest thing to say about this lesson is that reading it does not make you competent to work a scene. Scene processing is a physical skill learned by doing it under supervision, and the details that matter, how a swab is rolled, how a fold is made, how a photograph is framed, do not transfer through prose. What does transfer is the reasoning: preserve first, document before disturbing, package according to the mechanism of loss, and record every hand the item passes through.
Common misconceptions
- A break in the chain of custody gets evidence thrown out. Usually it goes to weight, not admissibility; the jury hears about it and the defense argues it.
- Plastic bags protect evidence. They trap moisture and destroy DNA in biological stains, and they let volatile ignitable liquid residues escape in fire debris.
- Luminol proves blood. Luminol is a presumptive test that also reacts with bleach, some metals, and certain plant materials; it directs attention and confirms nothing.
- Scene technicians solve the case. They preserve and record; interpretation belongs to analysts, pathologists, and investigators, and merging those roles is how contextual bias enters.
- A 3D scan removes human judgment. A scanner records geometry precisely and decides nothing about what matters, what to collect, or what a pattern means.
Putting it together
- Sequence is life safety, perimeter, log, documentation, then collection; each step protects the next.
- Notes, photographs at three ranges, measured sketches, and video each capture what the others miss.
- Search patterns are chosen for the terrain, and every pattern has a known failure mode.
- Packaging rules encode mechanisms of loss: moisture destroys DNA, plastic leaks volatiles, corners leak particles.
- Controls and elimination samples cost minutes at the scene and save years of misdirected investigation.
- Chain of custody answers whether the tested item is the seized item; gaps usually create doubt rather than exclusion.
Sources
- National Institute of Justice. (n.d.). Forensic sciences. Office of Justice Programs. nij.ojp.gov
- National Institute of Standards and Technology. (n.d.). Organization of Scientific Area Committees for Forensic Science. nist.gov
- Wikipedia contributors. (n.d.). O. J. Simpson murder case. en.wikipedia.org
- Fisher, B. A. J., and Fisher, D. R. (2012). Techniques of Crime Scene Investigation (8th ed.). CRC Press.
- Technical Working Group on Crime Scene Investigation. (2000). Crime scene investigation: A guide for law enforcement. National Institute of Justice, U.S. Department of Justice.
- Key terms
- Scene log
- A written record of every person entering and leaving a secured scene, with times, used to account for contamination and to identify elimination samples.
- Midrange photograph
- An image linking an item of evidence to a fixed reference point such as a doorway, bridging overall views and close-ups.
- Triangulation
- A sketching method that fixes an item's position by measuring its distance from two permanent reference points.
- Druggist fold
- A folded paper packet used to contain small trace particles such as fibers, hairs, or glass fragments before placing them in a labeled envelope.
- Substrate control
- A swab or sample of an unstained area adjacent to a stain, submitted so the laboratory can distinguish the stain from the surface it sits on.
- Elimination sample
- A reference sample from a responder, technician, resident, or finder, collected so that their DNA or prints can be recognized and set aside.
- Chain of custody
- The continuous documented record of possession and transfer of an item from collection to court, establishing that the item tested is the item seized.
- Bloodstain pattern analysis
- The interpretation of the size, shape, and distribution of bloodstains to infer events; parts rest on measurable physics, parts on contested examiner judgment.
Getting Through the Door: Frye, Daubert, and Rule 702
- Contrast the Frye general acceptance test with the Daubert reliability inquiry and identify which standard governs where.
- Apply the Daubert factors and the current text of Federal Rule of Evidence 702 to a proffered forensic method.
- Explain the Confrontation Clause requirement that the analyst who did the work testify to it.
- Assess the evidence that Daubert screening operates differently in criminal cases than in civil ones.
A blood pressure cuff in 1923
James Alphonso Frye was charged with murder in Washington, D.C. His lawyer wanted the jury to hear from a Harvard-trained psychologist named William Moulton Marston, who claimed that systolic blood pressure rose when a person lied and that his cuff could detect it. Marston is remembered today for something else entirely: a few years later he created Wonder Woman, whose lasso compels the truth. The trial judge refused the testimony. In 1923 the Court of Appeals of the District of Columbia agreed in an opinion barely two pages long, and in doing so produced the sentence that governed scientific evidence in American courts for the next seventy years. A scientific principle, the court said, must have gained general acceptance in the particular field in which it belongs.
That is the Frye standard, and its logic is a kind of outsourcing. A judge is not a scientist, so rather than evaluate the science, the judge counts heads in the relevant scientific community. The appeal is obvious. So are the two failures. A genuinely new and valid method is inadmissible until the field catches up, and a method that an entire field has accepted for a century without ever testing it sails through, because the field accepts it. Hold that second failure. It is the story of the rest of this course.
Key idea: Frye asks whether a method is accepted by its own community, which lets an untested but long-practiced discipline qualify simply because its practitioners agree with each other.
1993: the judge becomes a gatekeeper
In 1975 Congress enacted the Federal Rules of Evidence, and Rule 702 said, in substance, that a qualified expert could testify if scientific or technical knowledge would help the jury. It did not mention Frye. For nearly two decades courts argued about whether the Rules had quietly killed the general acceptance test.
The answer arrived through a morning sickness drug. Jason Daubert and Eric Schuller were born with serious birth defects, and their families sued Merrell Dow Pharmaceuticals, maker of Bendectin. Merrell Dow's expert cited the published epidemiology, which had not found an association. The plaintiffs' experts had reanalyzed the data and run animal and chemical studies, work that had not been published or accepted. Applying Frye, the lower courts excluded it. In 1993 the Supreme Court, in an opinion by Justice Blackmun, held that the Federal Rules had superseded Frye, and that trial judges must serve as gatekeepers who assess whether reasoning and methodology are scientifically valid.
The opinion offered factors, explicitly non-exclusive and not a checklist:
- Can the theory or technique be tested, and has it been?
- Has it been subjected to peer review and publication?
- Is there a known or potential error rate?
- Do standards exist that control the technique's operation?
- Is it generally accepted in the relevant community?
Notice what happened to Frye. It did not die; it was demoted from the whole test to the fifth item on a list. And notice the third factor, because it is the one this course keeps returning to. Daubert told judges to ask a discipline for its error rate. Several disciplines that had been testifying for decades discovered they had never measured one.
Two follow-on cases completed the framework. In General Electric Co. v. Joiner (1997), the Court held that appellate courts review exclusion decisions only for abuse of discretion, and added that nothing requires a judge to admit opinion evidence connected to existing data only by the expert's say-so. That phrase, the analytical gap between the data and the opinion, is now standard vocabulary. In Kumho Tire Co. v. Carmichael (1999), involving a tire failure analyst, the Court held that gatekeeping applies to all expert testimony, not only to what looks like laboratory science. Experience-based expertise is not exempt from being asked how it knows.
Why this matters: Daubert shifted the question from who agrees with you to whether the method has been tested and how often it is wrong, and Kumho Tire closed the escape hatch of calling a method technical rather than scientific.
Rule 702 today
Rule 702 was amended in 2000 to codify the trilogy, and amended again effective December 1, 2023, because courts had drifted. The current rule permits a qualified expert to testify if the proponent demonstrates to the court that it is more likely than not that the expert's knowledge will help the trier of fact, that the testimony rests on sufficient facts or data, that it is the product of reliable principles and methods, and that the expert's opinion reflects a reliable application of those principles and methods to the facts of the case.
Two changes in that sentence matter for forensic science. The 2023 amendment makes explicit that the burden is the proponent's, by a preponderance, and that these are questions for the judge rather than issues to be waved through as matters of weight. And the final clause targets the gap between a valid method and an overstated conclusion. A discipline can be perfectly reliable in principle and still produce testimony that outruns it: an examiner using a validated method who then tells the jury the error rate is zero has failed 702(d) even though the method itself is sound.
Not every court uses this framework. A minority of states, including California, Illinois, Pennsylvania, New York, and Washington, retain a general acceptance test in some form, which means the admissibility of the same technique can differ across a state line. California's version, from People v. Kelly, is often called the Kelly or Kelly-Frye rule. If you work in this field, the first question about any admissibility problem is not what the science says. It is what jurisdiction you are in.
The analyst has to show up
Admissibility is only half of getting evidence to a jury. The Sixth Amendment's Confrontation Clause supplies the other half, and a trio of cases changed daily practice in every crime laboratory in the country.
In Melendez-Diaz v. Massachusetts (2009), the state proved that a seized substance was cocaine by filing sworn certificates from analysts who never testified. The Court held the certificates were testimonial, so the defendant had a right to cross-examine the people who made them. In Bullcoming v. New Mexico (2011), the state tried a workaround: a different analyst, familiar with the procedures, testified in place of the one who had performed the blood alcohol test. The Court rejected the substitution. A surrogate cannot be cross-examined about what the actual analyst did, saw, or got wrong. Williams v. Illinois (2012) fractured the Court and left real uncertainty about expert reliance on reports by others, and lower courts have been sorting through it since.
The practical consequence is a scheduling problem with constitutional roots. Analysts spend days in courthouse hallways instead of at the bench, which is one visible cause of laboratory backlogs, and it is a cost the Court considered and accepted. Cross-examination of the person who did the work is not a formality. It is the only mechanism the system has for finding out that the calibration was overdue or that the analyst was signing results she had not run.
Does the gate actually close?
Here is the uncomfortable finding. Daubert transformed civil litigation, where exclusion motions against plaintiffs' experts became routine and often decisive. In criminal cases it changed remarkably little. Peter Neufeld, writing in 2005, examined how criminal courts applied Daubert to prosecution forensic evidence and concluded the effect was close to nonexistent. Challenges to fingerprint, firearms, bite mark, and handwriting testimony were filed and, with rare exceptions, denied, frequently on the ground that the technique had long been admitted, which is Frye reasoning wearing a Daubert caption.
Why the asymmetry? Several reasons compound. Judges hearing that a method has been accepted since 1911 are reluctant to be the first to exclude it, and stare decisis pushes hard in that direction. Public defenders rarely have funds for the competing expert a serious challenge requires. And a ruling excluding fingerprint evidence would unsettle an enormous number of closed cases, a consideration no judge states aloud and every judge understands.
Movement has come mostly since 2016, and mostly at the level of language rather than exclusion. Judges increasingly permit the examiner to testify but restrict how strongly the conclusion may be phrased, forbidding claims of certainty or of exclusion of all other sources. That is a real change, and it is a smaller change than the 2009 and 2016 scientific reviews called for.
Common misconceptions
- Daubert replaced Frye everywhere. It governs federal courts and most states, but a minority of states, including California, Illinois, Pennsylvania, and New York, still apply general acceptance in some form.
- General acceptance no longer matters after Daubert. It survives as one of the listed factors; it simply is no longer the whole inquiry.
- Daubert applies only to laboratory science. Kumho Tire extended gatekeeping to technical and experience-based expertise, including most pattern comparison work.
- If a method is admissible, the expert may say anything about it. Rule 702(d) requires reliable application to the facts, and courts increasingly limit the strength of the conclusion rather than excluding the witness.
- A supervisor can testify to another analyst's results. Bullcoming rejected surrogate testimony; the analyst who did the work is generally the one who must face cross-examination.
The short version
- Frye (1923) asked whether a method was generally accepted in its field, which admitted the untested along with the tested.
- Daubert (1993) made judges gatekeepers and asked about testing, peer review, error rate, standards, and acceptance.
- Joiner (1997) set abuse-of-discretion review and named the analytical gap; Kumho Tire (1999) extended gatekeeping to all expertise.
- Rule 702, as amended in 2023, puts the burden on the proponent by a preponderance and requires reliable application, not just a reliable method.
- Melendez-Diaz and Bullcoming require the analyst who did the work to testify and be cross-examined.
- In practice, Daubert screening bites hard in civil cases and rarely excludes prosecution forensic evidence in criminal ones.
Sources
- Legal Information Institute. (n.d.). Federal Rule of Evidence 702: Testimony by expert witnesses. Cornell Law School. law.cornell.edu
- Legal Information Institute. (n.d.). Daubert standard. Cornell Law School. law.cornell.edu
- Daubert v. Merrell Dow Pharmaceuticals, Inc., 509 U.S. 579 (1993). law.cornell.edu
- Neufeld, P. J. (2005). The (near) irrelevance of Daubert to criminal justice and some suggestions for reform. American Journal of Public Health, 95(S1), S107-S113.
- Key terms
- Frye standard
- The 1923 rule admitting scientific evidence only if the underlying principle has gained general acceptance in its particular field.
- Daubert standard
- The 1993 federal approach requiring judges to assess the scientific validity of an expert's reasoning and methodology before admitting the testimony.
- Gatekeeper
- The trial judge's role under Daubert of screening expert testimony for reliability and fit before the jury hears it.
- Analytical gap
- The distance between the underlying data and an expert's conclusion, named in Joiner; a judge need not admit an opinion bridged only by the expert's assertion.
- Rule 702
- The federal evidence rule governing expert testimony, amended in 2023 to require the proponent to show reliability by a preponderance and to demand reliable application to the facts.
- Confrontation Clause
- The Sixth Amendment guarantee that a defendant may confront witnesses, which requires the analyst who performed a forensic test to testify about it.
- Voir dire of an expert
- Questioning at trial about a witness's qualifications and methods, conducted before the court decides whether the witness may give opinion testimony.
Module 2: Prints and Biology
Friction ridge examination from ridge formation in the womb through development chemistry, database searching, and the ACE-V process, then DNA from a stain on a shirt to an electropherogram, and finally the genuinely hard part: mixtures, probabilistic genotyping, familial and genealogical searching, and the question of how a person's DNA got where it was found.
Fingerprints: Ridge Detail, ACE-V, and the Mayfield Error
- Explain how friction ridge skin forms, why it persists, and why identical twins have different ridge detail.
- Select appropriate development techniques for porous and nonporous surfaces and explain the underlying chemistry.
- Describe the ACE-V process and identify where subjectivity and contextual information enter it.
- State what published black box studies show about latent print error rates and what testimony those rates support.
Latent Fingerprint 17
On 11 March 2004, coordinated bombs on Madrid commuter trains killed 193 people. Spanish police recovered a blue plastic bag containing detonators, and on it a partial latent print, designated Latent Fingerprint 17. Digital images went to the FBI. A search of the automated system returned twenty candidates. A senior FBI examiner compared the latent to the fourth candidate and identified it. Two more FBI examiners verified. An experienced examiner appointed by the court to assist the defense, working independently, agreed as well.
The print belonged, they concluded, to Brandon Mayfield, a lawyer in Portland, Oregon, whose prints were in the system from military service. He had not left the United States in years and had no passport. On 6 May 2004 the FBI arrested him as a material witness.
The Spanish National Police had already told the FBI they did not agree. On 19 May they matched the print to Ouhnane Daoud, an Algerian national living in Spain. Mayfield was released the next day and the case against him dismissed. In 2006 the government paid two million dollars and issued a formal apology, and the Department of Justice Inspector General published a long post-mortem. Its findings are the reason this case opens the lesson rather than closing it. The examiners had not been careless. They were among the best in the world, and they had reasoned in a circle: having found agreement in some features, they explained away the disagreements, and each verifier knew what the previous examiner had concluded before looking.
Key idea: The most consequential fingerprint error on record was produced by expert examiners following normal procedure, which tells you the vulnerability is in the procedure and not in the people.
Why the ridges are there at all
Friction ridge skin covers the fingers, palms, toes, and soles. Its evolutionary function is grip and tactile sensing; the ridges channel moisture and improve the skin's response to texture. What matters forensically is how they form. Between roughly the tenth and sixteenth weeks of gestation, ridges develop in the basal layer of the epidermis over transient swellings called volar pads. The overall pattern responds to the size, shape, and timing of regression of those pads, which is under some genetic influence. The fine detail responds to local mechanical stresses in a growing sheet of tissue, which is not genetically specified at all.
That distinction produces the field's most useful teaching fact. Identical twins share a genome and often share pattern type, and their minutiae are different. Compare that with DNA, where identical twins are indistinguishable by standard typing. The two techniques fail in opposite places, which is one reason a case with both is stronger than a case with either.
Once formed, the pattern persists. Superficial abrasions regenerate exactly, because the template lies below the surface. Damage deep enough to destroy the basal layer leaves a permanent scar, and scars become identifying features themselves. Ridges grow with the hand but do not rearrange.
Patterns, minutiae, and pores
Examiners describe ridge detail at three levels.
Level 1 is overall ridge flow and pattern class. Roughly six in ten fingers are loops, about three in ten are whorls, and about one in twenty are arches. Level 1 detail alone can exclude, and it can never individualize, since millions of people share a right loop.
Level 2 is the minutiae: points where a ridge ends, splits into two (a bifurcation), or appears as an isolated dot. Their type, position, and relationship to one another carry most of the comparison's weight.
Level 3 is finer still: the positions of sweat pores along a ridge, the shapes of ridge edges, incipient ridges between the main ones. It is real detail and it is the least reliably reproduced from impression to impression, so its use is contested.
How many matching minutiae are enough? For decades various jurisdictions used point standards, twelve in some countries, sixteen in Britain until 2001. In 1973 the International Association for Identification adopted a resolution stating that no valid basis exists for requiring a predetermined minimum number of points, and American practice moved to a qualitative and quantitative assessment by the examiner. That was scientifically defensible: a point standard is arbitrary. It also removed the only objective threshold the discipline had, and replaced it with the phrase sufficient agreement, which means whatever the examiner's training tells her it means.
What matters here: There is no number of matching points that makes an identification; the standard is an examiner's judgment that agreement is sufficient, which is exactly what makes validation studies indispensable.
Getting the print off the surface
Three kinds of impression turn up. A patent print is visible without treatment, left in blood, grease, or paint. A plastic print is a three-dimensional impression in a soft material such as putty or wax. A latent print is invisible, made of eccrine sweat, sebaceous oils picked up from the face and hair, and cellular debris, and it has to be developed.
The development method follows the surface, because the residue behaves differently depending on whether it can soak in.
| Surface | Method | What it reacts with |
|---|---|---|
| Nonporous: glass, metal, plastic | Powder and lift, or cyanoacrylate fuming | Powder adheres to residue; superglue vapor polymerizes on it, forming a white ridge deposit |
| Porous: paper, cardboard, untreated wood | Ninhydrin, then physical developer | Ninhydrin reacts with amino acids to give a purple product; physical developer targets lipids and works on wetted paper |
| Porous, higher sensitivity | DFO or indanedione before ninhydrin | Amino acids, producing fluorescent product viewed under an alternate light source |
| Wet nonporous | Small particle reagent | Suspended particles adhere to the lipid fraction, which survives water better than amino acids |
| Adhesive tape | Gentian violet or specialized adhesive-side reagents | Sebaceous material on the sticky face |
| Bloody prints | Amido black, leucocrystal violet | Protein in the blood, not the sweat residue |
Sequence matters as much as choice. On a porous item you use the fluorescent amino acid reagent before ninhydrin, because ninhydrin consumes the same target. On a nonporous item you examine visually and with an alternate light source, then fume, then apply dye stain, then powder, because each step can destroy what a later step would have found. And on any item that might also yield DNA, the swabbing decision comes first, since several development reagents interfere with typing.
ACE-V, and the place the bias enters
Comparison follows a four-stage protocol known as ACE-V.
Analysis examines the unknown impression on its own: how much ridge detail is present, how clear it is, whether distortion from pressure or movement explains apparent differences, and whether it has value for comparison at all. Ideally this happens before the examiner has seen any exemplar.
Comparison places the latent beside the known print and works through corresponding features.
Evaluation reaches one of three conclusions: identification (source identification), exclusion, or inconclusive.
Verification has a second qualified examiner repeat the work.
Read that again with the Madrid case in mind. Nothing in the protocol requires the analysis stage to be completed and documented before the exemplar is seen, so an examiner can drift into deciding what the smudge contains after learning what it is supposed to contain. Nothing in the protocol requires verification to be blind, so in ordinary practice the verifier knows the first examiner's conclusion and is being asked, in effect, to disagree with a colleague. The Inspector General identified both problems in the Mayfield review. Documented analysis before comparison, and blind verification, are the standard fixes, and adoption has been partial.
How often are examiners wrong?
Before 2011, the honest answer was that nobody knew, because nobody had run the study. Then Bradley Ulery and colleagues, working with the FBI and Noblis, published a black box study in the Proceedings of the National Academy of Sciences. They gave 169 practicing examiners thousands of latent and exemplar pairs, some from the same source and some not, and simply counted outcomes.
The false positive rate, calling two prints from different sources an identification, was 0.1 percent. The false negative rate, missing a true identification, was 7.5 percent. A companion study the following year had examiners re-examine some of their own earlier comparisons and found that they did not always repeat their own conclusions.
Both numbers deserve a moment. A one in a thousand false positive rate is low, and it is emphatically not zero, which means testimony asserting zero error rate or one hundred percent certainty is false as a matter of arithmetic. The Department of Justice's uniform language rules now forbid examiners from claiming a zero error rate or absolute certainty. The much higher false negative rate is a different kind of finding: examiners are cautious, and caution buys misses. PCAST in 2016 accepted latent print comparison as foundationally valid, while insisting that examiners disclose error rates rather than testify to certainty, and noting that a second published study had produced a higher false positive rate than the FBI study did.
The upshot: Fingerprint comparison is among the better validated of the pattern disciplines, its measured false positive rate is small but real, and the correct courtroom claim is a strong opinion with a stated error rate rather than certainty.
Common misconceptions
- Fingerprints have been scientifically proven unique. Uniqueness is a well-supported working assumption from a century of experience, not a proven theorem, and in any case the courtroom question is whether an examiner can reliably tell.
- A computer identifies the person. Automated systems return a ranked candidate list; every identification is a human conclusion, and Mayfield was the fourth name on such a list.
- Twelve matching points means a match. American practice abandoned point standards after the 1973 IAI resolution; the criterion is an examiner's judgment of sufficient agreement.
- Identical twins have identical fingerprints. They share a genome and often a pattern class, but the fine ridge detail forms from local mechanical stress and differs.
- Every surface yields a usable print. Most touches leave nothing usable; texture, moisture, time, heat, and simple bad luck destroy far more impressions than are ever recovered.
What to carry forward
- Ridge detail forms in the womb from local stresses, persists for life, and differs between identical twins.
- Level 1 pattern can exclude but never individualize; Level 2 minutiae carry most of the comparison.
- Development chemistry follows the surface, and the order of treatments determines what survives to be found.
- ACE-V is a sound outline with two known weak points: unblinded analysis and unblinded verification.
- The 2011 black box study measured a false positive rate of about 0.1 percent and a false negative rate of about 7.5 percent.
- Those numbers make certainty testimony indefensible and make a stated error rate the honest alternative.
Sources
- Britannica. (2024). Fingerprint. britannica.com
- National Institute of Standards and Technology. (n.d.). Forensic science. nist.gov
- Wikipedia contributors. (n.d.). Brandon Mayfield. en.wikipedia.org
- Ulery, B. T., Hicklin, R. A., Buscaglia, J., and Roberts, M. A. (2011). Accuracy and reliability of forensic latent fingerprint decisions. Proceedings of the National Academy of Sciences, 108(19), 7733-7738.
- Office of the Inspector General. (2006). A review of the FBI's handling of the Brandon Mayfield case. U.S. Department of Justice.
- Key terms
- Friction ridge skin
- The ridged skin of the fingers, palms, toes, and soles, whose pattern forms in the womb and persists for life.
- Minutiae
- Level 2 ridge features, chiefly ridge endings, bifurcations, and dots, whose types and spatial relationships carry most of the weight in a comparison.
- Latent print
- An invisible impression composed of sweat, sebaceous oil, and cellular debris that must be developed chemically or physically to be seen.
- Cyanoacrylate fuming
- Development of prints on nonporous surfaces by exposing them to superglue vapor, which polymerizes on the residue to form a white ridge deposit.
- Ninhydrin
- A reagent that reacts with amino acids in print residue on porous surfaces such as paper to produce a purple product.
- ACE-V
- The four-stage examination protocol of Analysis, Comparison, Evaluation, and Verification used in friction ridge and other pattern disciplines.
- Blind verification
- Independent re-examination by a second examiner who does not know the first examiner's conclusion, a recommended safeguard that is not universally adopted.
- Black box study
- A test that presents practitioners with samples of known ground truth and counts their conclusions, measuring a discipline's error rate without opening up how examiners decide.
DNA I: From a Stain on a Shirt to an STR Profile
- Trace the laboratory workflow from body fluid screening through extraction, quantitation, amplification, and capillary electrophoresis.
- Read an electropherogram, name alleles by repeat number, and explain what the amelogenin locus reports.
- Calculate a random match probability across several loci using the product rule and state the assumptions it depends on.
- Explain what a single-source DNA profile establishes and what it leaves open.
Nine minutes past nine, Leicester, 1984
Alec Jeffreys has told the story often enough that he can give the time. At about 9:05 on the morning of 10 September 1984, in a laboratory at the University of Leicester, he pulled an X-ray film out of a developing tank. He had been looking at a minisatellite region of DNA, expecting a technical result about gene evolution. What he saw was a smear of bands, one column per person, and the columns were different from each other while family members shared bands in inheritable patterns. Within about half an hour, by his account, he and his team had worked out that this could settle paternity, immigration disputes, and criminal identity. He has called it a moment of pure eureka, which is a phrase scientists rarely earn.
The first criminal use came fast, and it did not go the way anyone expected. In Leicestershire villages, two schoolgirls had been raped and murdered, Lynda Mann in 1983 and Dawn Ashworth in 1986. A local kitchen porter named Richard Buckland confessed to the second killing. Police asked Jeffreys to confirm the confession with the new technique. Instead the DNA showed that the semen from both crimes came from one man, and that man was not Buckland. The first use of DNA in a criminal case exonerated the suspect. Police then asked roughly five thousand local men for blood samples. Colin Pitchfork persuaded a colleague to give a sample in his name, the substitution was mentioned in a pub and reported, and Pitchfork was convicted in January 1988.
Worth holding on to: The technique that made DNA the strongest identification evidence in forensic science began by clearing an innocent man who had already confessed, and that pattern has repeated ever since.
Where the DNA is, and how you find it
Nuclear DNA lives in nucleated cells, which rules out mature red blood cells; the DNA in a bloodstain comes from white cells. Practical sources are blood, semen, saliva, vaginal and other epithelial cells, hair with the root sheath attached, bone, teeth, and the small quantities of skin cells transferred by handling an object, usually called touch or trace DNA.
Before extraction, an analyst usually wants to know what the stain is, and the tests come in two grades. Presumptive tests are fast, sensitive, and not specific. The Kastle-Meyer test turns phenolphthalein pink in the presence of the peroxidase-like activity of hemoglobin, and it also responds to some plant peroxidases and oxidizing agents. Acid phosphatase screening flags possible semen. Amylase testing flags possible saliva. Confirmatory tests cost more and mean more: the microscopic identification of spermatozoa is definitive for semen, an immunochromatographic test for human hemoglobin or for prostate specific antigen identifies the species and fluid, and a crystal test confirms blood.
One extraction technique deserves its own paragraph, because it is the reason sexual assault evidence works at all. In a vaginal swab, the victim's epithelial cells vastly outnumber the perpetrator's sperm cells. Differential extraction exploits a physical difference: sperm heads are held together by disulfide bonds and resist ordinary lysis, so a first, gentle lysis releases the epithelial DNA, which is pipetted off, and a second lysis containing a reducing agent then opens the sperm. The result is two fractions, one dominated by the victim and one by the semen donor. Without it, nearly every such sample would be an uninterpretable mixture.
Copying the target: PCR
Extracted DNA is quantified, usually by real-time PCR, which reports not only how much human DNA is present but how degraded it is and how much of it is male. That number drives the next decision, because too little template produces unreliable results and too much overloads the instrument.
Amplification is the polymerase chain reaction, and its logic is worth holding in your head as three temperatures repeated about thirty times. Heat to roughly 95 degrees Celsius and the double helix separates. Cool to roughly 55 to 60 degrees and short synthetic primers bind to the sequences flanking each target region. Warm to roughly 72 degrees and a heat-stable polymerase extends from each primer, copying the region. Each cycle doubles the number of copies of the targeted regions and only those regions, so thirty cycles turn a few dozen molecules into something an instrument can see.
That sensitivity is the technique's power and its danger. PCR does not know whether the template came from the crime or from an analyst's sneeze in 2019, and it amplifies both with equal enthusiasm. This is why laboratories keep pre-amplification and post-amplification work in physically separate rooms, use dedicated equipment, and maintain elimination databases of staff profiles.
What gets copied: short tandem repeats
The targets are short tandem repeats, or STRs: stretches of DNA where a short unit, usually four bases, repeats a variable number of times. At a locus called TH01, for instance, the unit AATG may repeat six times on one chromosome and nine times on the other. Alleles are named by repeat number, so that person is a 6,9 at TH01. Someone with the same repeat count on both chromosomes is homozygous, and shows one peak instead of two.
Repeat counts are not always whole numbers. A partial repeat gives an allele such as 9.3, meaning nine full repeats plus three bases, and TH01 9.3 is common in European populations. The notation looks odd the first time and is standard.
The United States national database, CODIS, originally used thirteen core loci. Since 1 January 2017 the core set has been twenty, expanded to reduce the chance of adventitious matches in a database that has grown to well over ten million offender profiles and to improve compatibility with European systems. Commercial amplification kits copy all twenty at once, plus amelogenin, a locus on the sex chromosomes whose X and Y versions differ in length, giving a fast sex indication as a byproduct.
Reading the result
Amplified fragments are separated by capillary electrophoresis. The mixture is injected into a thin capillary filled with a polymer, a voltage is applied, and shorter fragments travel faster. Each primer set carries one of several fluorescent dyes, so loci overlapping in size can be told apart by color. A laser excites the dyes as they pass a detector, and software converts the signal into an electropherogram: a plot of fluorescence in relative fluorescence units against fragment size.
What you actually see is a row of peaks. Two peaks of roughly equal height at a locus indicate a heterozygote. One tall peak indicates a homozygote. Size is converted to repeat number by comparison with an allelic ladder run alongside the samples. A well-behaved single-source profile from a good sample is one of the cleanest results in forensic science: twenty loci, at most two peaks each, balanced heights, no ambiguity about what it says.
Key idea: An electropherogram is a length measurement, not a picture of a person, and everything difficult about DNA interpretation comes from peaks that are too small, too many, or unbalanced.
The number attached to a match
A profile alone means nothing until you can say how rare it is. For a single-source profile the standard statistic is the random match probability: the chance that a randomly chosen unrelated person from a reference population would have this profile.
Work a small example with three loci. Suppose the evidence profile is:
- Locus 1: genotype 14, 16, with allele frequencies 0.10 and 0.08. A heterozygote frequency is 2pq, so 2 times 0.10 times 0.08 equals 0.016.
- Locus 2: genotype 11, 11, with allele frequency 0.20. A homozygote frequency is p squared, so 0.04.
- Locus 3: genotype 8, 9.3, with frequencies 0.12 and 0.30. So 2 times 0.12 times 0.30 equals 0.072.
Because the loci are on different chromosomes or far enough apart to inherit independently, the frequencies multiply: 0.016 times 0.04 times 0.072 equals about 0.000046, or roughly one in twenty-two thousand. Three loci is a weak result. Now extend the same multiplication to twenty loci and the exponents pile up fast, which is how laboratories arrive at figures in the range of one in many quadrillions for a complete single-source profile.
Those figures rest on assumptions you should be able to name. Allele frequencies come from population databases of a few hundred people per group, so the estimate is a sample estimate. The product rule assumes independence within and across loci, which is approximately true and not exactly true, because real populations have internal structure. Laboratories apply a correction, conventionally called theta, that makes the estimate more conservative. And none of it addresses relatives: a full sibling is far more likely to share a profile than a random stranger, which is the door into familial searching in the next lesson.
Backlog, speed, and what DNA does not say
In 2009 a prosecutor's investigator in Wayne County, Michigan walked into a Detroit police storage warehouse and found 11,341 sexual assault kits sitting on shelves, most never submitted for testing. The subsequent research project, funded by the National Institute of Justice, tested thousands of them, identified hundreds of suspected serial offenders, and produced convictions in cases where the evidence had been in a box for two decades. Similar counts followed in other cities. The scientific problem was zero; the failure was organizational.
The technology has kept accelerating. Rapid DNA instruments now produce an STR profile from a buccal swab in around ninety minutes without a human analyst, and federal law permits their use for booking samples under specified conditions. Sensitivity has improved to the point where profiles come from a few dozen cells.
Which brings the necessary caution. A DNA profile answers exactly one question: whose cells are these. It does not say what body fluid they came from unless a separate test established that, and it does not say when they were deposited or how. A man's DNA on a knife handle is consistent with him stabbing someone, and equally consistent with him having made a sandwich with that knife on Tuesday. Every hard case in the next lesson turns on that gap.
Common misconceptions
- DNA evidence proves guilt. It addresses source, not conduct; how and when the material arrived is a separate question that DNA typing cannot answer.
- A one in a quadrillion figure means only one person on Earth could have left it. It is an estimate of how rare the profile is among unrelated people, built on database frequencies and independence assumptions, and it does not apply to close relatives.
- Red blood cells carry the DNA in a bloodstain. Mature red cells have no nucleus; the DNA comes from white cells.
- All the DNA in a hair can be typed. A shed hair without a root sheath usually yields only mitochondrial DNA, which is shared by everyone in a maternal line and cannot individualize.
- Testing every sexual assault kit was technically impossible. The kits sat untested because of resources, policy, and priorities; when Detroit funded the work, the laboratory science functioned normally.
Pulling it together
- Jeffreys identified variable minisatellite patterns in 1984, and the first criminal application exonerated a suspect who had confessed.
- Presumptive tests indicate, confirmatory tests establish, and differential extraction separates sperm from epithelial cells in sexual assault samples.
- PCR doubles targeted regions each cycle, which delivers both extraordinary sensitivity and a permanent contamination risk.
- STR alleles are named by repeat number; the CODIS core set expanded from thirteen to twenty loci in 2017, with amelogenin indicating sex.
- Random match probability multiplies locus frequencies under stated assumptions, and it does not apply to relatives.
- A profile identifies whose cells are present, never how or when they got there.
Sources
- Britannica. (2024). DNA fingerprinting. britannica.com
- National Institute of Standards and Technology. (n.d.). Forensic science. nist.gov
- Wikipedia contributors. (n.d.). Colin Pitchfork. en.wikipedia.org
- National Institute of Justice. (n.d.). Forensic sciences. Office of Justice Programs. nij.ojp.gov
- Butler, J. M. (2011). Advanced Topics in Forensic DNA Typing: Methodology. Academic Press.
- Key terms
- Short tandem repeat
- A DNA region where a short unit, usually four bases, repeats a variable number of times; alleles are named by repeat count.
- Polymerase chain reaction
- A cycling process of denaturation, primer annealing, and extension that doubles targeted DNA regions each cycle.
- Differential extraction
- A two-step lysis that separates resistant sperm cells from epithelial cells, producing distinguishable fractions from sexual assault samples.
- Electropherogram
- The plot of fluorescence against fragment size produced by capillary electrophoresis, on which alleles appear as peaks.
- Amelogenin
- A sex-chromosome locus whose X and Y versions differ in length, used as a sex indicator alongside the STR loci.
- Random match probability
- The estimated chance that a randomly chosen unrelated person would share the observed profile, calculated by multiplying locus frequencies.
- CODIS
- The Combined DNA Index System, the United States national DNA database, which since 2017 uses a core set of twenty STR loci.
- Presumptive test
- A rapid, sensitive screening test that indicates a substance may be present without establishing its identity or species.
DNA II: Mixtures, Probabilistic Genotyping, Familial Searching, and Transfer
- Explain stochastic effects in low-template DNA and how they complicate mixture interpretation.
- Interpret a likelihood ratio and contrast it with the older combined probability of inclusion approach.
- Describe probabilistic genotyping software, its validated range, and the disputes over access to its source code.
- Distinguish familial searching from investigative genetic genealogy, and separate source questions from activity questions.
The man who was in a hospital bed
In November 2012, a Silicon Valley businessman named Raveesh Kumra was killed during a home invasion in Monte Sereno, California. Under his fingernails, analysts found DNA belonging to Lukis Anderson, a homeless man with a record of petty offenses. Anderson was arrested and held on a charge that could have carried the death penalty.
He had an alibi that was close to perfect. At the time of the killing he was unconscious in a hospital bed, so drunk that staff had recorded his condition through the night. The explanation, when investigators reconstructed it, was that the same paramedics who had picked Anderson up off a sidewalk earlier that evening responded to the Kumra scene hours later. Something on their hands, equipment, or clothing carried his cells into the house. Anderson spent about five months in jail before the alibi was confirmed and the charges were dropped.
Nothing went wrong in the laboratory. The typing was correct, the profile was Anderson's, the statistic was overwhelming. What failed was the inference from a source conclusion to an event, and that is the subject of this lesson.
The core of it: Modern DNA typing is sensitive enough that the hard question is almost never whose DNA it is, but how it got there.
Transfer, and the ladder of questions
DNA moves. Primary transfer is direct: you touch a glass and leave cells. Secondary transfer is indirect: you shake a hand, and that person touches the glass, leaving your cells on it. Tertiary transfer adds another step. Experimental work has repeatedly produced secondary transfer in ordinary conditions, and shedder status varies between people, so one person's cells travel more readily than another's.
Forensic scientists in Europe formalized the resulting distinction, and it is the single most useful idea in this lesson. Evidence can be evaluated at different levels of proposition:
- Source level: whose DNA is in this sample? Laboratories are good at this.
- Activity level: did this person stab the victim, or did the DNA arrive by transfer? This requires data on transfer, persistence, and background levels for that material and surface, and most laboratories do not address it at all.
- Offense level: is the defendant guilty? Not a scientific question.
When an analyst testifies to a source-level conclusion and a jury hears an activity-level claim, nobody has lied and the record is still wrong. That gap put Lukis Anderson in jail.
What makes a mixture hard
A mixture is a sample containing DNA from two or more people, and it announces itself by showing three or more alleles at one or more loci, or by peak heights that do not balance the way a single source should. Sexual assault evidence, handled objects, weapon grips, and fingernail scrapings are routinely mixtures.
Interpretation runs into a set of problems that all come from having too few molecules of one contributor.
| Effect | What you see | Why it happens |
|---|---|---|
| Allele drop-out | A true allele is missing | Too few starting copies; sampling failure during early PCR cycles |
| Allele drop-in | An extra peak with no contributor | Sporadic contamination of a single amplifiable molecule |
| Heterozygote imbalance | One of a pair much shorter than the other | Unequal early amplification of two alleles from a tiny template |
| Stutter | A small peak one repeat below a real allele | Polymerase slippage during copying of a repeat region |
| Degradation slope | Larger loci lower than smaller ones | Fragmented template, so long targets are less often intact |
Every one of these makes a minor contributor's profile ambiguous, and worse, they make the number of contributors itself uncertain. Studies in which analysts were asked to state how many people were in a known mixture found frequent underestimation, especially with four or more contributors. If you cannot confidently say how many people are in a sample, every downstream statistic inherits that uncertainty.
From inclusion to likelihood ratios
For years the common statistic was the combined probability of inclusion, which asks what fraction of the population could not be excluded as a possible contributor to the mixture. It is easy to compute and easy to explain, and it throws away information: it ignores peak heights entirely, and it treats a person who fits the data poorly the same as one who fits it well. Applied to a complex mixture with drop-out, it can be badly misleading, and after a 2015 notice from the FBI about errors in the population data used for such calculations, the Texas Forensic Science Commission ordered a statewide review that changed the reported statistics in a large number of old cases.
The replacement is the likelihood ratio. Instead of asking who is excluded, it compares two explanations of the same data:
LR equals the probability of observing this evidence if the prosecution's proposition is true, divided by the probability of observing it if the defense's proposition is true.
Suppose the propositions are, first, that the mixture contains the victim and the defendant, and second, that it contains the victim and one unknown unrelated person. If the software reports a likelihood ratio of 2 million, the correct statement is that the data are 2 million times more probable under the first proposition than under the second. It is not a statement that the defendant is 2 million times more likely to be guilty; that inversion is the prosecutor's fallacy, and it is the single most common error in courtroom discussion of DNA. A likelihood ratio near 1 means the evidence does not distinguish the propositions. Below 1, it favors the defense, and reporting such values is good practice that not all laboratories follow.
Remember: A likelihood ratio measures how well the data are explained by two competing stories, and it depends entirely on which two stories were compared.
Software that does what analysts could not
Probabilistic genotyping systems, of which STRmix and TrueAllele are the best known, model the whole electropherogram: peak heights, stutter, drop-out and drop-in probabilities, degradation, and the number of contributors. They use Markov chain Monte Carlo sampling to explore the space of genotype combinations consistent with the data, then compute a likelihood ratio. Fully continuous models use peak heights; semi-continuous models use only which alleles are present.
The gain is real. Mixtures once reported as uninterpretable now yield usable numbers, and the interpretation no longer depends on where a particular analyst set a threshold. The problems are equally real, and there are three.
First, validated range. PCAST examined the published validation studies in 2016 and concluded that foundational validity had been established only within a limited range, roughly three-person mixtures in which the person of interest contributed at least about 20 percent of the DNA. Outside that range, PCAST said, validity had not been established, which is not the same as saying the software is wrong; it is saying nobody had shown it was right. Laboratories have since published validations extending further, and the boundary remains a live question that any competent cross-examination should probe.
Second, the software is proprietary. Defendants have argued that they cannot meaningfully challenge a conclusion whose reasoning is a trade secret. In 2021 a New Jersey appellate court held in State v. Pickett that a defendant was entitled to access TrueAllele's source code under a protective order, and courts elsewhere have split.
Third, different programs and different parameter settings can give different answers on the same data. That is expected of any model-based statistic and it is unsettling to a jury, which has been told that DNA gives certainty.
NIST's own scientific foundation review of DNA mixture interpretation, released in draft in 2021, made the underlying point plainly: the reliability of these methods is well supported for simple mixtures, and the publicly available data thin out quickly as mixtures get more complex.
Searching for relatives, and searching family trees
Two techniques exploit the fact that relatives share alleles. They are frequently confused and they are not the same thing.
Familial searching runs a crime scene profile against a law enforcement database and, instead of reporting only exact matches, deliberately looks for profiles similar enough to suggest a close biological relative. Candidates are then filtered, often with Y-chromosome STR testing, which passes down the male line unchanged. The best known result is the case of Lonnie David Franklin Jr., the Los Angeles serial murderer known as the Grim Sleeper, whose son's profile entered the state database after a felony conviction and produced a familial lead in 2010. Policies vary sharply: California authorized the practice in 2008 with a formal protocol, and Maryland prohibits it by statute.
Investigative genetic genealogy works differently. Instead of STR loci, the laboratory generates a dense set of single nucleotide polymorphisms from crime scene DNA and uploads a file to a consumer genealogy database that permits law enforcement use. The database returns distant relatives, often third or fourth cousins, and genealogists then build family trees from public records to find the person who fits the age, sex, and geography. In April 2018 this method identified Joseph James DeAngelo as the Golden State Killer, ending a search that had run since the 1970s.
Three things about it are commonly misunderstood. The genealogy produces a lead, never evidence: police obtain a direct reference sample and run ordinary STR typing before charging anyone. The reach is enormous, because you do not need the suspect in a database, only some of his cousins; a 2018 analysis in Science estimated that a database covering a small percentage of a population would allow most of that population to be located through third cousins or closer. And the rules are thin. The Department of Justice adopted an interim policy effective in November 2019 restricting federal use to violent crimes and requiring CODIS to be tried first, but that policy binds federal investigators and not the many state and local agencies now using the technique.
Common misconceptions
- Familial searching and genetic genealogy are the same thing. The first searches a law enforcement STR database for close relatives; the second uploads SNP data to consumer genealogy databases and finds distant ones.
- A likelihood ratio of a million means the defendant is a million times more likely to be guilty. That is the prosecutor's fallacy; the ratio compares how probable the data are under two stated propositions.
- DNA on an object proves the person handled it. Secondary transfer is well documented, and the Lukis Anderson case is the clearest demonstration that it reaches crime scenes.
- Probabilistic genotyping removes subjectivity. It removes threshold-setting by an analyst and replaces it with modeling choices, parameter settings, and a decision about the number of contributors.
- If the software gives a number, the mixture was interpretable. Validation covers a defined range of contributors and proportions, and results outside that range require justification.
The takeaway
- Sensitivity has moved the hard question from whose DNA it is to how it got there, and transfer is documented at two and three steps.
- Source, activity, and offense are different levels of proposition, and laboratories usually address only the first.
- Drop-out, drop-in, imbalance, and stutter make minor contributors ambiguous and make contributor counts uncertain.
- Likelihood ratios replaced inclusion probabilities because they use all the data and state the comparison being made.
- Probabilistic genotyping is validated within a limited range, is partly proprietary, and can differ between systems.
- Familial searching and investigative genetic genealogy generate leads with wide reach and uneven governing rules.
Sources
- National Institute of Standards and Technology. (n.d.). Forensic science. nist.gov
- Wikipedia contributors. (n.d.). Golden State Killer. en.wikipedia.org
- National Institute of Justice. (n.d.). Forensic sciences. Office of Justice Programs. nij.ojp.gov
- President's Council of Advisors on Science and Technology. (2016). Forensic science in criminal courts: Ensuring scientific validity of feature-comparison methods. Executive Office of the President.
- Butler, J. M. (2015). Advanced Topics in Forensic DNA Typing: Interpretation. Academic Press.
- Key terms
- Secondary transfer
- Deposition of a person's DNA on an object they never touched, by way of an intermediate person or item.
- Activity level proposition
- A question about how biological material came to be where it was found, as opposed to whose material it is.
- Allele drop-out
- Failure of a true allele to amplify to a detectable level, typically when very few template molecules are present.
- Stutter
- A minor peak one repeat unit shorter than a true allele, produced by polymerase slippage during amplification of repeat regions.
- Likelihood ratio
- The probability of the observed evidence under one proposition divided by its probability under a competing proposition.
- Prosecutor's fallacy
- Misreading a likelihood ratio or match probability as the probability that the defendant is guilty or is the source.
- Probabilistic genotyping
- Software that models peak heights, stutter, drop-out and degradation to compute likelihood ratios for complex mixtures.
- Investigative genetic genealogy
- Generating dense SNP data from crime scene DNA, uploading it to consumer genealogy databases, and building family trees to identify a lead.
Module 3: Chemistry at the Bench
The analytical side of the laboratory: trace evidence and what an association is actually worth, the identification of seized drugs and the interpretation of drugs and alcohol in the body, and fire investigation, where a set of confident visual indicators turned out to be reading the wrong thing entirely.
Trace Evidence: Fibers, Glass, Paint, Residue, and the Hair Problem
- Explain why trace evidence yields class associations and what makes a particular association strong or weak.
- Describe the analytical sequence for fibers, glass, paint, and soil, and the information each step adds.
- State precisely what microscopic hair comparison can and cannot establish, and what the FBI review of its own testimony found.
- Interpret a gunshot residue result correctly, including what it does not show.
Ninety percent
On 20 April 2015 the FBI released the results of a review of its own casework. The Bureau, working with the Department of Justice, the Innocence Project, and the National Association of Criminal Defense Lawyers, had gone back through trial transcripts in which FBI examiners testified about microscopic hair comparison. In the 268 trials reviewed to that point, examiners had made statements exceeding the limits of the science in at least 90 percent of cases. Twenty-six of the twenty-eight examiners in the unit were implicated. In 32 of the 33 cases where the defendant had been sentenced to death, the hair testimony contained errors.
Read that as a scientific finding rather than a scandal. The microscopy was not fake. Hairs really do differ in ways a trained person can see under a comparison microscope. What failed was the leap from a visual similarity to a statement about a specific person, and the leap had been made in courtrooms for decades without anyone measuring how often it was wrong.
That failure sits at the center of trace evidence generally, and it is why this lesson spends as much time on what associations mean as on how they are made.
Key idea: Trace analysis produces class associations, and the scientific error is almost never in the microscope; it is in the sentence spoken afterward.
What an association is worth
Trace evidence is material transferred in small quantity: fibers, hairs, glass, paint, soil, residues. Because there is little of it and because manufactured materials are made by the millions, the strongest honest conclusion is usually that the questioned material and the known material could share a common origin and are indistinguishable in the measured properties.
Three things make such an association worth something.
Rarity. A white cotton fiber is worthless; every closet holds them. A trilobal nylon carpet fiber dyed a particular green in a limited production run is a different matter, because the pool of possible sources is small.
Multiplicity. One association is weak. Nineteen different fiber types from a victim's body corresponding to nineteen surfaces in one house is a pattern that coincidence explains badly. This was the core of the prosecution in the Atlanta child murders case against Wayne Williams in 1982, where fibers linked bodies to a bedroom carpet, a bedspread, a vehicle, and a dog. The reasoning was sound in structure. The probability estimates offered at trial rested on assumptions about carpet distribution that were criticized then and would be scrutinized harder now.
Transfer plausibility. An association means more when the proposed contact explains it. Fibers on the inside of a victim's clothing, in a location a casual encounter cannot reach, carry more weight than fibers on an outer sleeve.
And the reverse case matters too. The absence of expected transfer is evidence, weakly, and it is systematically underused because nobody looks for what is not there.
Fibers, from polarized light to pyrolysis
Fiber examination runs from cheap and informative to expensive and specific. Under a stereomicroscope an examiner records color, diameter, and cross-sectional shape, which for manufactured fibers is a manufacturing choice: round, trilobal, dogbone. Polarized light microscopy measures birefringence and refractive indices, separating natural fibers from synthetic ones and distinguishing polymer classes. Microspectrophotometry compares color as a spectrum rather than by eye, which catches dyes that look identical to a human and differ in transmission. Infrared spectroscopy identifies the polymer: nylon 6 against nylon 6,6, polyester, acrylic. Pyrolysis gas chromatography breaks the polymer into fragments and compares the resulting pattern.
At the end of that sequence, an analyst can say two fibers are indistinguishable in every property measured. She cannot say they came from the same garment, because the manufacturer made kilometers of it.
Glass, paint, and soil
Glass has one genuinely powerful examination and several ordinary ones. The powerful one is a physical fit: if a fragment from a suspect's clothing fits the broken edge of a window, that is not a class association at all, it is a unique reassembly, and it is one of the few conclusions in trace evidence that can properly be called an individualization.
Short of that, glass is compared by refractive index, measured by immersing fragments in oil and heating until the fragment disappears against the oil, and by elemental composition using laser ablation with mass spectrometry, which can separate glasses that share a refractive index. Fracture patterns carry separate information: radial cracks radiate from the impact, concentric cracks form rings between them, and the sequence of fractures tells which of two holes came first, because a later crack stops at an earlier one.
Automotive paint is unusually informative because modern vehicle finishes are layered: clearcoat over basecoat over primer surfacer over an electrocoat primer, each with a characteristic composition. A chip preserving the full sequence can often be narrowed to a make, model, and range of years using the Paint Data Query database maintained by the Royal Canadian Mounted Police, which catalogs original factory finishes. That makes paint one of the more investigatively useful trace materials in hit and run work.
Soil is compared by color against standardized charts, by particle size distribution, and by mineral content under the microscope. It is highly variable over short distances, which cuts both ways: two samples from thirty meters apart may differ, so a match is meaningful, and so is the fact that soil from a garden may differ from soil from the same garden a season later.
The point: A physical fit is a reassembly and can individualize; everything else in trace analysis narrows a class, and the narrowing is only as good as the rarity of the class.
Hair, honestly
A hair has three parts: an outer cuticle of overlapping scales, a cortex containing pigment granules, and a central medulla that may be continuous, fragmented, or absent. Microscopic examination of these features supports several conclusions that are well founded.
- Species. Human and animal hairs differ clearly in medullary structure and scale pattern; this determination is reliable.
- Body area. Head, pubic, limb, and facial hairs differ in length, diameter variation, and tip form.
- Treatment. Dye, bleach, and cutting are visible, and a growing-out dye line even gives a rough time since treatment.
- How it left the body. A root in the growing anagen phase with a sheath attached suggests forcible removal; a club-shaped telogen root suggests natural shedding.
What microscopic comparison cannot do is identify a person. There is no established population data on how frequently the observed combinations of features occur, no defined criterion for how much agreement is enough, and no measured error rate from the era when the testimony was given. The 2009 National Research Council review said so directly. Testimony that a questioned hair matched a defendant, or was consistent with him to the exclusion of others, or came from him with any stated frequency, was unsupported.
The replacement is genetic. Mitochondrial DNA can be recovered from a rootless shaft, though it is shared by everyone in a maternal line and therefore cannot individualize either. Hairs with root material give nuclear DNA and a full STR profile. Modern practice uses microscopy as a screening step to decide which hairs are worth sending for DNA, which is exactly what it is good for.
Gunshot residue, and the sentence people get wrong
When a cartridge fires, the primer detonates and vaporized metals condense into microscopic spheroidal particles. In conventional primers these contain lead, barium, and antimony, and a particle containing all three in one spheroid is treated as characteristic of primer residue. Analysis is by scanning electron microscopy with energy dispersive X-ray spectroscopy, which images individual particles and reports their elemental content.
Now the interpretation, which is where cases go wrong. A positive result means particles consistent with primer residue were on the sampled surface. It does not mean the person fired a gun. Residue reaches people who were standing nearby, who handled a fired weapon, who were in a car where one had been fired, who sat in a police vehicle or a holding cell that was not decontaminated, or who were handcuffed by an officer who had been at the range. It also disappears easily: hand washing, wiping, and simple time remove it, so a negative result on someone who fired a weapon hours earlier is entirely ordinary. Add lead-free primers, whose particles lack the classic three-element signature, and the picture gets more complicated still.
The correct testimony is narrow, and it is the reason many laboratories restrict or decline this examination: particles consistent with gunshot primer residue were detected, which is consistent with the sampled person having been in the environment of a discharged firearm.
Common misconceptions
- Hair can be matched to a person under a microscope. Microscopy determines species, body area, and treatment; individual identification requires DNA, and the FBI's own review found decades of testimony exceeded the science.
- A gunshot residue positive means the person fired a gun. It indicates presence in the environment of a discharge, and transfer from vehicles, officers, and surfaces is well documented.
- Fibers can be traced to a specific garment. Manufactured fibers are produced in enormous quantity; the strongest ordinary conclusion is that two fibers are indistinguishable in measured properties.
- Trace evidence is weak evidence. A physical fit individualizes, and multiple independent rare associations can be powerful; what is weak is a single common material.
- Absence of trace evidence proves no contact occurred. Transfer is probabilistic and persistence is short; absence is weak evidence, not disproof.
What to remember
- Trace evidence yields class associations whose value depends on rarity, multiplicity, and the plausibility of the transfer.
- Fiber analysis proceeds from morphology through polarized light and color spectra to polymer identification, ending in indistinguishable rather than identical.
- A physical fit of glass or any broken object is a reassembly and is far stronger than any compositional comparison.
- Layered automotive paint can be narrowed to make, model, and year range through manufacturer reference collections.
- Microscopic hair comparison reliably gives species, body area, and treatment, and cannot identify a person; DNA does that.
- Gunshot residue places a person in the environment of a discharge and is easily transferred and easily lost.
Sources
- National Research Council. (2009). Strengthening forensic science in the United States: A path forward. National Academies Press. nap.nationalacademies.org
- National Institute of Standards and Technology. (n.d.). Forensic science. nist.gov
- Britannica. (2024). Hair. britannica.com
- Federal Bureau of Investigation. (2015, April 20). FBI testimony on microscopic hair analysis contained errors in at least 90 percent of cases in ongoing review [Press release]. U.S. Department of Justice.
- Saferstein, R. (2018). Criminalistics: An Introduction to Forensic Science (12th ed.). Pearson.
- Key terms
- Trace evidence
- Material transferred in small quantity between people, objects, and places, including fibers, hair, glass, paint, soil, and residues.
- Class association
- A conclusion that two materials share measured properties and could have a common origin, without identifying a unique source.
- Physical fit
- The reassembly of broken or torn pieces along their damaged edges, a conclusion strong enough to identify a unique origin.
- Microspectrophotometry
- Measurement of a fiber's color as a transmission or absorbance spectrum, distinguishing dyes that appear identical to the eye.
- Medulla
- The central canal of a hair, whose presence, continuity, and structure help distinguish species and body area.
- Anagen root
- A hair root in the active growth phase, often with sheath material attached, indicating the hair was pulled rather than shed.
- Gunshot primer residue
- Microscopic spheroidal particles, classically containing lead, barium, and antimony, formed when a cartridge primer detonates.
- Paint Data Query
- A reference collection of factory automotive finishes maintained by the Royal Canadian Mounted Police, used to narrow a paint chip to make, model, and years.
Seized Drugs and Forensic Toxicology
- Distinguish presumptive from confirmatory drug identification and explain why roadside color tests produce false positives.
- Describe the analytical scheme a laboratory uses to identify a controlled substance, and why gas chromatography with mass spectrometry anchors it.
- Choose appropriate toxicology specimens and explain postmortem redistribution and the parent-metabolite relationship.
- Estimate blood alcohol concentration with the Widmark approach and state the limits of the estimate.
A crumb on a car floor
In August 2010 a Houston police officer stopped a car driven by Amy Albritton and found a small white crumb on the floorboard. He dropped it into a plastic pouch containing cobalt thiocyanate solution. The liquid turned blue. Blue means cocaine, and Albritton, a single mother with no criminal record, took a plea rather than risk a trial, and lost her job and her apartment.
Months later the crime laboratory tested the crumb properly. It was not a controlled substance. Harris County has since gone back through years of such cases and vacated the convictions of hundreds of people who had pleaded guilty to possessing substances that turned out to be nothing.
The chemistry behind this is not subtle, and every drug chemist knows it. Cobalt thiocyanate turns blue in the presence of cocaine. It also turns blue in the presence of a long list of other compounds: certain over-the-counter medications, some household products, and other alkaloids. The test was designed as a screen, to be sensitive rather than specific, and to be followed by laboratory confirmation. In practice, the plea usually arrives before the laboratory does.
Why this matters: A presumptive test tells you what a substance might be; treating its result as an identification is not a close call in chemistry, and it has produced convictions at scale.
The identification scheme
A drug chemist works in layers, and the logic is that each layer is more specific and more expensive than the last.
Presumptive tests come first: color tests such as the cobalt thiocyanate reagent for cocaine, Marquis reagent, which gives purple with opiates and orange to brown with amphetamines, and the Duquenois-Levine test used for cannabis. Microcrystalline tests, in which a reagent forms crystals of characteristic shape under a microscope, are more discriminating than color and still not conclusive. Thin layer chromatography separates components by how far they travel up a plate.
Confirmatory analysis is where identification actually happens, and the workhorse is gas chromatography with mass spectrometry. The chromatograph separates a mixture in time: components travel through a heated column at different rates according to their volatility and their affinity for the coating, so each emerges at a characteristic retention time. The mass spectrometer then bombards each emerging compound with electrons, breaking it into charged fragments, and records the mass-to-charge ratio of each fragment. The resulting pattern is reproducible and highly specific, which is why it is often described as a molecular fingerprint. Infrared spectroscopy, nuclear magnetic resonance, and X-ray diffraction supply comparable discriminating power for particular problems.
Professional guidance groups these techniques by discriminating power and requires an identification scheme to include at least one of the highest category, or several from the lower ones. That is the rule the roadside pouch violates: it is one low-specificity test used alone.
Two practical wrinkles now dominate the drug section's workload. Quantitation has become essential because the 2018 federal farm legislation defined hemp as cannabis containing no more than 0.3 percent delta-9 tetrahydrocannabinol by dry weight, which means a laboratory can no longer simply identify a plant as cannabis; it must measure. And novel psychoactive substances, particularly fentanyl analogues, arrive faster than reference standards can be purchased, so laboratories sometimes cannot name a compound they can clearly see.
Toxicology is a different question
Drug chemistry asks what a seized substance is. Forensic toxicology asks what was in a person's body, at what concentration, and what that means. The second question is harder, and the difference between them is one of the most useful distinctions in this course.
Start with specimen choice, because it determines what the number can mean.
| Specimen | Detection window | What it supports |
|---|---|---|
| Blood | Hours to a day or two | Concentration at the time of collection; the only matrix that supports impairment reasoning |
| Urine | Days to weeks for some drugs | Past exposure; concentration does not track impairment |
| Oral fluid | Hours | Recent use, roughly tracking blood for some drugs |
| Hair | Months | Chronic use patterns; vulnerable to external contamination and to bias by hair color and treatment |
| Vitreous humor | Postmortem | Alcohol and electrolytes; resists putrefaction and lags blood concentration |
| Liver and other tissue | Postmortem | Drugs undetectable in decomposed blood; interpretation is difficult |
Analysis follows the same two-tier logic as drug chemistry. Immunoassay screening is fast, cheap, and cross-reactive, meaning a positive can come from a structurally similar compound rather than the target. Confirmation and quantitation use gas or liquid chromatography with mass spectrometry.
Metabolism supplies extra information, and sometimes it is decisive. Cocaine breaks down quickly to benzoylecgonine, so finding parent cocaine suggests recent use while benzoylecgonine alone suggests use hours ago. Heroin is metabolized within minutes to 6-monoacetylmorphine and then to morphine, and 6-monoacetylmorphine is specific to heroin, which distinguishes it from prescribed morphine or from poppy seeds.
Two traps in postmortem work
First, postmortem redistribution. After death, drugs stored in the lungs, liver, and heart muscle diffuse back into nearby blood. A sample drawn from the heart or chest cavity can therefore show a concentration several times higher than the person had while alive, and the size of the change depends on the drug. This is why the standard practice is to collect peripheral blood from the femoral vein, clamped, and why a toxicologist asked about a heart blood level should say that it cannot be converted into a level at the time of death.
Second, tolerance. A concentration that would kill a person who has never used opioids may be a maintenance level in someone who has used them daily for years. Published lethal ranges are compilations of reported cases, not thresholds, and they overlap heavily with therapeutic ranges for several drug classes. A responsible report says the concentration is within a range associated with fatal outcomes in the literature, and says what it does not know.
In short: A postmortem drug concentration is a measurement of a sample, not a measurement of the person at the moment of death, and turning one into the other requires assumptions the toxicologist should state aloud.
Alcohol, worked
Ethanol is the most measured substance in forensic toxicology, and the arithmetic is approachable. Every state sets a per se limit of 0.08 grams of alcohol per 100 milliliters of blood, above which driving is an offense regardless of demonstrated impairment. Utah lowered its limit to 0.05 at the end of 2018.
The classical estimate uses Widmark's approach. Blood alcohol concentration is the mass of alcohol absorbed, divided by body mass times a distribution factor that reflects how much of the body is water, since alcohol distributes into body water and not into fat. The factor is conventionally about 0.68 for men and about 0.55 for women.
Work an example. A 70 kilogram man drinks four standard drinks, each containing about 14 grams of ethanol, so 56 grams. Divide 56 by the product of 0.68, 70, and 10, which gives 56 divided by 476, or about 0.118 grams per 100 milliliters. Alcohol is then eliminated at roughly 0.015 per hour, with real values ranging from about 0.010 to 0.020. Three hours after the last drink, the estimate is 0.118 minus 0.045, or about 0.073.
Now notice everything that estimate ignores. It assumes all four drinks were fully absorbed, which food delays substantially. It uses a population average distribution factor for a specific person. It uses a population average elimination rate whose real range would move the three-hour answer from about 0.088 down to about 0.058, straddling the legal limit. Widmark calculations are useful for teaching and for rough retrograde extrapolation, and a toxicologist who presents one without its uncertainty is overselling it.
Breath testing raises its own issues. An instrument measures alcohol in deep lung air and converts to blood using an assumed partition ratio, conventionally 2100 to 1, which varies between people and within a person over time. Mouth alcohol from belching or dental work can inflate a reading, which is why protocols require an observation period of fifteen to twenty minutes before the test and why instruments include slope detectors. And instruments require calibration by trained people who do it correctly: in 2018 the New Jersey Supreme Court held that breath tests from machines a state trooper had not calibrated according to protocol were inadmissible, affecting more than twenty thousand cases.
Presence is not impairment
The hardest question in this field has no laboratory answer. For alcohol, decades of research support a rough correspondence between concentration and driving impairment across the population, which is what makes a per se limit defensible. For most other drugs it does not exist. Delta-9 tetrahydrocannabinol concentrations fall rapidly after inhalation while impairment persists, and regular users carry measurable blood levels while sober, so per se cannabis limits track impairment poorly. Prescribed benzodiazepines and opioids produce measurable levels in people who function normally on them.
This is why drug impaired driving cases lean on observation, through structured evaluations by trained officers, and why a toxicology report in such a case should be read as one input among several rather than the answer.
Common misconceptions
- A roadside color test identifies a drug. It is a presumptive screen with well-documented false positives, and it was never intended to stand alone.
- A urine positive shows the person was impaired. Urine reflects past exposure over days to weeks and does not correlate with concentration at the wheel.
- A postmortem blood level tells you the level at death. Postmortem redistribution can raise central blood concentrations severalfold; peripheral femoral blood is collected precisely because of this.
- Published lethal ranges are thresholds. They are compilations of reported cases that overlap therapeutic ranges, and tolerance moves the boundary enormously.
- Breath testing measures blood directly. It measures breath and converts using an assumed partition ratio, and it depends on correct calibration and observation protocol.
Summing up
- Presumptive tests screen and confirmatory techniques identify; conflating them has produced guilty pleas to possessing nothing.
- Gas chromatography separates a mixture in time and mass spectrometry identifies each component by its fragmentation pattern.
- Toxicology asks a harder question than drug chemistry, and specimen choice determines what an answer can support.
- Parent compounds and metabolites carry timing information, and 6-monoacetylmorphine specifically indicates heroin.
- Postmortem redistribution and tolerance make postmortem concentrations far less interpretable than they look.
- Widmark arithmetic gives a usable estimate with wide uncertainty, and outside alcohol, presence is a poor proxy for impairment.
Sources
- Britannica. (2024). Toxicology. britannica.com
- National Highway Traffic Safety Administration. (n.d.). Drunk driving. U.S. Department of Transportation. nhtsa.gov
- National Institute of Standards and Technology. (n.d.). Forensic science. nist.gov
- Gabrielson, R., and Sanders, T. (2016). Busted: How a badly flawed roadside drug test sends innocent people to jail. ProPublica and The New York Times Magazine.
- Levine, B. (Ed.). (2020). Principles of Forensic Toxicology (5th ed.). AACC Press.
- Key terms
- Presumptive test
- A rapid screening test, typically a color reaction, that indicates a substance may belong to a class without identifying it.
- Gas chromatography mass spectrometry
- The pairing of a separation technique with a detector that fragments each component and records the fragment masses, giving highly specific identification.
- Immunoassay
- An antibody-based toxicology screen that is fast and sensitive but cross-reactive, so positives require confirmation.
- Postmortem redistribution
- The movement of drugs from tissues into nearby blood after death, which can raise central blood concentrations well above the level at death.
- 6-monoacetylmorphine
- A short-lived heroin metabolite whose presence specifically indicates heroin use rather than morphine from another source.
- Widmark calculation
- An estimate of blood alcohol concentration from mass of alcohol consumed, body mass, and a water distribution factor, adjusted for elimination over time.
- Partition ratio
- The assumed 2100 to 1 relationship between alcohol in blood and in deep lung air that breath instruments use to report a blood equivalent.
- Per se limit
- A statutory blood or breath alcohol concentration above which driving is an offense without separate proof of impairment.
Fire and Explosion: How the Arson Indicators Collapsed
- Explain flashover and post-flashover burning and why they produce the patterns once read as accelerant use.
- Evaluate the classic arson indicators against the experimental evidence on each.
- Describe origin and cause determination under NFPA 921, including why undetermined is a legitimate conclusion.
- Explain how ignitable liquid residue is recovered and why comparison samples are essential.
Corsicana, two days before Christmas
On 23 December 1991 a fire in a small frame house in Corsicana, Texas killed three children: two-year-old Amber Willingham and her one-year-old twin sisters. Their father, Cameron Todd Willingham, got out. Investigators from the local fire department and the state fire marshal's office examined the burned house and found what they were trained to find. Crazed glass, which they had been taught meant rapid heating from a hot fire. Puddle-shaped burn marks on the floor, which they read as pour patterns from a liquid accelerant. Burning underneath an aluminum door threshold, which they said fire could not do unaided. Brown staining on concrete. Multiple separate points of origin.
They counted more than twenty indicators of arson. Willingham was convicted in 1992 and executed on 17 February 2004.
Weeks before the execution, a chemist named Gerald Hurst reviewed the file and wrote that not one of those indicators supported an arson finding, because by then each had been tested and had failed. In 2009 the Texas Forensic Science Commission commissioned a review by fire engineer Craig Beyler, who concluded that the original investigation had not met the standards of its own time, let alone current ones. The Commission's 2011 findings said the investigators had relied on flawed science.
Every one of those indicators had a common source of error, and understanding it is the whole lesson.
Key idea: The classic arson indicators were not fabricated; they were real observations, misattributed to accelerant use when they are the ordinary consequences of a room fire that got hot enough.
Flashover changes everything
A fire in an enclosed room does not simply grow. It passes through a threshold. As a fire burns, hot smoke collects at the ceiling and thickens into a layer that radiates downward. When that layer reaches roughly 600 degrees Celsius, the radiant heat striking the floor is intense enough that every exposed combustible surface in the room reaches its ignition temperature at nearly the same instant. That transition is flashover, and it can occur within a few minutes of ignition in a furnished modern room.
After flashover, the room is not burning in one place. It is burning everywhere, floor to ceiling, and the fire becomes limited by available oxygen rather than by fuel, so flames reach toward openings where air enters. Now walk back through the Corsicana indicators with that picture in mind.
| Classic indicator | Claimed meaning | What research established |
|---|---|---|
| Crazed glass | Rapid heating from an accelerated fire | Produced by rapid cooling, typically water from a hose stream |
| Irregular floor patterns, so-called pour patterns | A liquid accelerant was poured | Routine in post-flashover rooms, produced by burning floor coverings and ventilation |
| Burning beneath thresholds and at floor level | Liquid ran under the metal | Radiant heat and floor-level burning are normal after flashover |
| Multiple points of origin | Several fires were set | Post-flashover ignition of separated fuel packages, plus drop-down of burning material |
| Alligatoring: large shiny blisters in char | Fast, hot, accelerated fire | No demonstrated relationship to fire speed or accelerant presence |
| Spalling of concrete | An accelerant burned on the slab | Occurs from moisture and thermal stress, including from water application |
| Deep char at a location | Longest burning, therefore the origin | Char depth reflects ventilation and fuel as much as duration |
The pattern in that right-hand column is that the indicators identify a hot, fully developed fire. They do not identify how it started.
The house on Lime Street
The demonstration that changed the field happened in Jacksonville, Florida in 1990. A fire had killed six people in a house on Lime Street, and Gerald Wayne Lewis was charged, largely on the basis of floor patterns read as pours. His defense team, including fire investigator John Lentini, obtained an identical adjacent house that was scheduled for demolition, furnished it to match, and set it alight using only the accidental ignition source the defense proposed. No accelerant was used at any point.
The test fire went to flashover and, when the burn was over, the floor showed the same irregular patterns that had been called pour patterns in the case next door. The charges were dropped. It is hard to overstate what that experiment did: it took an indicator the profession had used for decades and showed, in a full-scale replicate, that it appeared without the thing it supposedly indicated.
What matters here: A full-scale test fire with a known cause is the only way to learn what a pattern actually means, and until the profession ran such tests it was reading confidence into ordinary damage.
How a fire is investigated now
The reference document is NFPA 921, the Guide for Fire and Explosion Investigations, first published in 1992 and revised every few years. Its central move was to insist on the scientific method: state the problem, collect data, form hypotheses, test them against the data and against fire dynamics, and be willing to reject a hypothesis you like.
Determining origin comes before determining cause, and it now draws on more than pattern reading. Witness accounts and 911 timing anchor the sequence. Fire dynamics analysis asks whether the proposed growth is physically consistent with the damage and the ventilation. Arc mapping surveys electrical circuits for arc damage, since a circuit that arced was energized when the fire reached it, and the pattern of arc locations along circuits helps bound where the fire was earliest.
Cause is then classified as accidental, natural, incendiary, or undetermined. NFPA 921 also drove out a habit of reasoning called negative corpus: concluding that a fire was incendiary because every accidental cause had been eliminated, without any affirmative evidence of an incendiary act. The 2011 and 2014 editions restricted and then rejected that reasoning. Its flaw is simple. Eliminating the accidental causes you thought of does not establish arson; it establishes that you did not identify an accidental cause, and undetermined is the honest classification.
Finding an ignitable liquid, if one is there
Chemical analysis is where fire investigation is on firm ground. Debris from a suspected origin is sealed in a clean unused metal can or a specialized nylon bag, because ordinary plastic both leaks volatile hydrocarbons and contributes its own. In the laboratory, the sealed container is warmed and a strip of activated charcoal is suspended in the headspace, adsorbing any volatile compounds. The strip is eluted with a solvent, and the eluate is run on a gas chromatograph with a mass spectrometer. Ignitable liquid residues are then classified by their chromatographic patterns into groups such as gasoline, medium petroleum distillates, and heavy petroleum distillates.
The essential control is a comparison sample: unburned carpet, padding, or flooring of the same type from an area away from the fire. Modern furnishings are petroleum products. Burning carpet backing, foam padding, and vinyl produces pyrolysis compounds that overlap with the components of petroleum distillates, so a chromatogram from burned synthetic flooring can look alarming without any liquid having been poured. Without the comparison sample, the analyst cannot separate what was added from what the house is made of.
Accelerant detection canines belong in the same category as color tests for drugs. A trained dog can alert on quantities below instrumental detection, which makes it an excellent tool for deciding where to sample, and its alert is presumptive. Laboratories regularly find no ignitable liquid in samples a dog alerted on, sometimes because the dog is responding to the same pyrolysis products that confuse the chromatogram.
Explosions, briefly
An explosion is the sudden release of gas under pressure, and investigators distinguish two regimes. In a deflagration, the reaction front moves through the material below the speed of sound, which is what happens with black powder, smokeless powder, and fuel-air mixtures such as leaking natural gas; damage is pushing and heaving, and structures move outward relatively slowly. In a detonation, the front travels faster than sound as a shock wave, which is characteristic of high explosives; damage is shattering and localized, with cratering near the seat and fragmentation of nearby material.
Post-blast scenes are searched outward from the suspected seat, because energy throws components a long way, and the recovered fragments of a device often survive better than intuition suggests. Residue swabs are analyzed by ion chromatography and mass spectrometry for characteristic ions such as nitrate, nitrite, chlorate, and perchlorate, and for organic explosives directly.
Common misconceptions
- Crazed glass shows a fire burned unusually hot. It is produced by rapid cooling, most often from firefighting water, and says nothing about how the fire started.
- Irregular floor burn patterns mean liquid was poured. Post-flashover rooms produce these routinely, as the Lime Street test fire demonstrated with no accelerant present.
- Multiple points of origin prove arson. After flashover, separated fuel packages ignite independently and burning material drops, producing several apparent origins.
- A dog alert establishes that an accelerant was used. It is a presumptive indication that guides sampling; laboratory confirmation frequently finds nothing.
- If no accidental cause is found, the fire was set. That is negative corpus reasoning, rejected by NFPA 921; the correct classification is undetermined.
Looking back
- Flashover turns a local fire into a room burning everywhere, and it generates most of the classic arson indicators.
- Crazing comes from cooling, floor patterns from post-flashover burning, and char depth from ventilation as much as time.
- The 1990 Lime Street test fire reproduced the disputed patterns using no accelerant at all.
- NFPA 921 imposed the scientific method, required affirmative evidence for an incendiary finding, and legitimized undetermined.
- Ignitable liquid residue is recovered by headspace adsorption and identified by chromatographic pattern, and requires a comparison sample.
- Willingham was executed on evidence that the profession had already begun to abandon, and Texas later found the science flawed.
Sources
- Wikipedia contributors. (n.d.). Cameron Todd Willingham. en.wikipedia.org
- Texas Forensic Science Commission. (n.d.). About the Commission. Texas Judicial Branch. txcourts.gov
- National Institute of Standards and Technology. (n.d.). Forensic science. nist.gov
- National Fire Protection Association. (2021). NFPA 921: Guide for Fire and Explosion Investigations. NFPA.
- Lentini, J. J. (2019). Scientific Protocols for Fire Investigation (3rd ed.). CRC Press.
- Key terms
- Flashover
- The transition at which radiant heat from the ceiling gas layer ignites all exposed combustible surfaces in a room at nearly the same time.
- Pour pattern
- An irregular floor burn once read as evidence of poured liquid accelerant, now known to occur routinely in post-flashover fires.
- Crazed glass
- A network of fine cracks in glass, produced by rapid cooling such as hose water, not by rapid heating.
- Negative corpus
- Concluding a fire was incendiary solely because accidental causes were eliminated, a form of reasoning rejected by NFPA 921.
- Arc mapping
- Surveying electrical circuits for arc damage to determine which circuits were energized when the fire reached them, helping bound the area of origin.
- Comparison sample
- Unburned flooring or furnishing material from outside the fire area, submitted so the laboratory can distinguish added liquids from the building's own pyrolysis products.
- Deflagration
- A reaction front moving through material below the speed of sound, producing pushing and heaving damage typical of gas and low explosives.
- Detonation
- A supersonic shock front characteristic of high explosives, producing shattering, cratering, and fragmentation near the seat.
Module 4: The Comparison Disciplines
Two fields built on the same premise, that manufacture and use leave reproducible marks a trained examiner can trace to one source: firearms and toolmark comparison, and questioned document examination. Both have long courtroom histories, useful investigative products, and validation records that do not match the confidence with which they have been presented.
Firearms and Toolmarks: Striae, Sufficient Agreement, and Error Rates
- Distinguish class from individual characteristics on bullets and cartridge cases and name the marks each firearm component leaves.
- Explain the AFTE theory of identification and the circularity objection to sufficient agreement.
- Describe how NIBIN generates investigative leads and why a correlation hit is not a conclusion.
- Summarize what black box studies have measured about firearms examiner error and how courts have responded.
Seventy cases against a garage wall
On the morning of 14 February 1929, seven men were shot dead inside a garage on North Clark Street in Chicago. Police recovered roughly seventy spent .45 caliber cartridge cases from the floor. Suspicion fell on the Chicago police themselves, since witnesses had seen men in uniform enter, and the coroner's jury wanted an answer that did not depend on anyone's word.
They brought in Calvin Goddard, who had been refining an instrument with Philip Gravelle: two microscopes joined by an optical bridge, so that an examiner could see two objects side by side in a single field and rotate one against the other. Goddard test-fired the Thompson submachine guns held by the Chicago police and reported that none of them had fired the cases from the garage. Months later, when two Thompsons were seized at the Michigan home of Fred Burke, Goddard compared those and reported that they had. Members of the coroner's jury were impressed enough to fund a laboratory, and the Scientific Crime Detection Laboratory opened at Northwestern University in 1929, the first independent forensic laboratory in the United States.
That is the origin story the discipline tells, and every element of it is true. What the story does not contain, and what nobody asked for until the 1990s, is a measurement of how often the comparison is wrong.
The point: Firearms comparison built its authority on a spectacular early success and a persuasive instrument, and it operated for seventy years without an error rate.
What the gun writes on the ammunition
A fired bullet and a fired cartridge case carry two kinds of information.
Class characteristics come from design. A barrel's rifling has a caliber, a number of lands and grooves, a direction of twist, and characteristic land and groove widths. Those values are set by the manufacturer, and the FBI maintains a general rifling characteristics file that lets an examiner take measurements from a recovered bullet and produce a list of makes and models consistent with them. This is genuinely useful investigative information and it identifies no individual gun.
Individual characteristics are the striae: fine parallel scratches left by microscopic irregularities in the barrel, which arise from the tools that cut or formed the rifling and from subsequent wear, corrosion, and fouling. The premise of the discipline is that these irregularities are effectively random, so their pattern differs between barrels, and reproducible, so successive bullets from one barrel carry the same pattern.
Cartridge cases record more, because a case stays in the gun and gets handled by every part of the action.
| Mark | Left by | Type |
|---|---|---|
| Firing pin impression | The pin striking the primer | Impressed |
| Breech face marks | The case head slamming back against the breech under pressure | Impressed, often the most informative |
| Chamber marks | The case expanding against the chamber wall | Striated |
| Extractor marks | The claw gripping the rim to pull the case | Striated or impressed |
| Ejector marks | The case striking the ejector on its way out | Impressed |
Toolmarks follow the same two-category logic outside firearms. A tool that slides across a surface, such as a pry bar on a window frame, leaves striated marks, a set of parallel lines reflecting the tool edge's irregularities. A tool that presses into a surface, such as a stamp or bolt cutter closed at right angles, leaves an impressed mark, a negative of the tool's shape. Striated marks are compared by rotating and sliding one against the other under the comparison microscope until lines align.
Sufficient agreement, and the circle inside it
How does an examiner decide two marks came from one source? The governing standard, published by the Association of Firearm and Tool Mark Examiners, says an identification is made when the agreement of individual characteristics exceeds the best agreement observed between marks known to have been made by different tools, and is consistent with the agreement observed between marks known to have been made by the same tool.
Read that carefully and the objection appears on its own. The criterion is defined relative to what examiners have observed, and it is applied by the examiner's trained judgment; it does not specify a number of matching striae, a measure of similarity, or a threshold anyone could check. The 2009 National Research Council review made this point in plain terms: the standard is not stated in scientific terms, and the decision remains subjective.
Practitioners have tried to make it objective. Consecutive matching striae counting proposes numeric criteria, such as requiring runs of consecutively matching lines. Three-dimensional surface topography systems measure the actual geometry of a breech face or a bullet land and compute a similarity score, and this line of work is the most promising route to a firearms conclusion that comes with a number rather than an assertion. Neither has displaced examiner judgment in routine casework.
What matters here: Sufficient agreement is a description of expert practice rather than a measurable criterion, which is precisely why the discipline's error rate has to be measured from the outside.
What NIBIN does and does not do
The National Integrated Ballistic Information Network, run by the Bureau of Alcohol, Tobacco, Firearms and Explosives, captures digital images of cartridge cases and bullets from crime scenes and from test fires of recovered guns and correlates them across a national database. The system returns a ranked list of candidates that look similar.
That is an investigative product of real value: it links shootings that no detective had connected, sometimes across jurisdictions and years, and rapid turnaround has become a policing priority for exactly that reason. But a correlation hit is a computer's ranking of image similarity. It becomes a conclusion only when a qualified examiner takes the physical items to a comparison microscope and reaches a decision. Confusing the lead with the conclusion is the same error as treating an automated fingerprint candidate list as an identification.
Two things the section does that are not comparison
Serial number restoration is chemistry rather than pattern matching, and it works reliably. Stamping a serial number into a firearm's frame deforms the metal's crystal structure below the depth of the visible characters. When a number is ground off, that deformed zone often remains, and it dissolves at a different rate than the surrounding metal. Etching with an acid reagent brings the digits back into view, sometimes completely.
Distance determination answers a different question: how far was the muzzle from the target? A contact wound shows soot driven into the wound track and often a muzzle imprint. At short range, unburned powder particles strike the skin and produce stippling, small punctate abrasions that cannot be wiped away. Beyond a certain range, only the bullet arrives. To quantify this, an examiner test-fires the actual weapon with the same ammunition into test panels at measured distances and compares the resulting residue patterns to the pattern on the victim's clothing. That is a controlled experiment with the real variables held fixed, and it is one of the more defensible things the section does.
The error rate question
PCAST examined firearms comparison in 2016 and found that, of the studies then available, only one had the design needed to estimate error: a properly constructed black box test in which examiners compared sets with known ground truth and were not able to reason from context. That study, conducted at Ames Laboratory, produced a false positive rate on the order of one percent, with an upper confidence bound that PCAST calculated at roughly one in forty-six. On that basis PCAST concluded that firearms analysis fell short of the criteria for foundational validity, while noting that it was one appropriately designed study away from meeting them.
Further large studies followed. They complicated rather than settled the picture, principally because of inconclusive answers. Firearms examiners return inconclusive conclusions at high rates, and how those are counted changes the error rate dramatically: treat an inconclusive on a known non-matching pair as a correct avoidance of error and the false positive rate looks tiny, treat it as a failure to exclude and the numbers look very different. Studies have also found that examiners do not always repeat their own conclusions on the same items, and that different examiners often disagree.
Courts have responded not by excluding examiners but by editing them. In 2008 a federal judge in the Southern District of New York, in United States v. Glynn, limited an examiner to testifying that a match was more likely than not. In 2019 a District of Columbia trial court in United States v. Tibbs restricted the examiner to stating that the firearm could not be excluded. Similar limiting rulings have accumulated since. The pattern is now familiar from fingerprints: the witness testifies, and the strength of the claim is trimmed.
Common misconceptions
- Ballistics identifies the gun. Ballistics is the physics of projectiles in flight; the comparison discipline is firearms and toolmark examination, and it compares marks rather than trajectories.
- A NIBIN hit means the same gun was used. The system ranks image similarity to generate a lead; only a microscope comparison by an examiner produces a conclusion.
- Filing off a serial number destroys it. Stamping deforms metal below the surface, and acid etching frequently recovers the digits.
- Every fired bullet can be traced to a gun. Bullets deform, fragment, and pass through intermediate targets; many recovered bullets have no usable individual detail at all.
- Inconclusive means the examination failed. It is a permitted conclusion, and how inconclusives are counted is one of the central disputes in estimating the discipline's error rate.
What you now know
- Class characteristics from rifling narrow a bullet to makes and models; striae are the claimed individual detail.
- Cartridge cases carry firing pin, breech face, chamber, extractor, and ejector marks, with breech face marks often the most informative.
- The AFTE sufficient agreement standard describes expert practice rather than specifying a measurable threshold.
- NIBIN produces investigative leads by image correlation, not conclusions.
- Serial number restoration and distance determination are well grounded and do not rely on the comparison premise.
- PCAST found one adequately designed study, with a false positive rate around one percent and a much higher upper bound, and courts have responded by limiting how strongly examiners may testify.
Sources
- Bureau of Alcohol, Tobacco, Firearms and Explosives. (n.d.). National Integrated Ballistic Information Network (NIBIN). atf.gov
- National Research Council. (2009). Strengthening forensic science in the United States: A path forward. National Academies Press. nap.nationalacademies.org
- National Institute of Standards and Technology. (n.d.). Forensic science. nist.gov
- President's Council of Advisors on Science and Technology. (2016). Forensic science in criminal courts: Ensuring scientific validity of feature-comparison methods. Executive Office of the President.
- Association of Firearm and Tool Mark Examiners. (1992). Theory of identification as it relates to toolmarks. AFTE Journal.
- Key terms
- Striated mark
- A pattern of parallel lines left when a tool or barrel slides across a surface, compared by aligning the lines under a comparison microscope.
- Impressed mark
- A negative impression of a tool's shape produced when it presses into a softer surface, such as a firing pin strike on a primer.
- Breech face marks
- Impressions transferred to the head of a cartridge case when pressure drives it back against the breech, often the most informative marks on a case.
- General rifling characteristics
- The caliber, number and width of lands and grooves, and twist direction of a barrel, used to narrow a bullet to consistent makes and models.
- Sufficient agreement
- The AFTE criterion for identification, defined by comparison to observed same-source and different-source agreement rather than by a stated numeric threshold.
- NIBIN
- The ATF network that correlates digital images of fired components across jurisdictions to generate investigative leads.
- Serial number restoration
- Recovery of an obliterated stamped number by etching, exploiting metal deformation that extends below the visible characters.
- Stippling
- Punctate abrasions on skin caused by unburned powder particles at close range, used with test fires to estimate muzzle to target distance.
Questioned Documents: Ink, Machines, and Handwriting
- Distinguish the materials analysis side of document examination from the handwriting comparison side and compare their evidentiary strength.
- Explain how paper, ink, and machine artifacts date and attribute a document.
- Describe how handwriting and signature comparisons are conducted, including exemplar requirements and forgery types.
- Assess what studies have shown about document examiner accuracy.
Sixty volumes that failed on the paper
In April 1983 the West German magazine Stern announced that it had acquired Adolf Hitler's private diaries, some sixty handwritten volumes, and had paid millions of marks for them. Historians argued about the handwriting and the contents. The West German federal archives took a different approach and looked at the objects.
The paper contained an optical whitener that had not been manufactured before the 1950s. The binding threads contained polyester. The inks were tested and shown to be recent. The forger, Konrad Kujau, had been meticulous about content and careless about materials, and it took the analysts a matter of weeks to establish that documents purporting to date from the 1930s could not have existed before the middle of the 1950s.
That contrast organizes this lesson. Questioned document examination has two halves. One half is materials science, where the questions have physical answers and the answers hold up. The other half is handwriting comparison, which rests on the premise that writing becomes an individual habit, and which has a much thinner validation record. Practitioners do both. You should keep them separate in your head.
Key idea: The strongest document conclusions come from what a document is made of, because materials carry dates and manufacturing origins that no amount of skill at imitating handwriting can defeat.
Reading the object
Paper is a manufactured product with a history. Fiber composition changed as pulping technology changed. Fillers and coatings changed. Fluorescent optical brighteners, which make paper look whiter under ultraviolet light, came into wide use in the middle of the twentieth century, so their presence sets an earliest possible date. Watermarks identify mills and often specific production periods.
Ink is equally informative. Ballpoint inks are separated by thin layer chromatography, in which a tiny plug punched from a written line is dissolved and the dye components separated on a plate, producing a colored pattern characteristic of the formulation. The United States Secret Service maintains a large reference collection of ink formulations with their first dates of production, so an ink that entered manufacture in 1994 rules out a document dated 1989. Relative age determination, based on how readily solvent still extracts from a line, can distinguish a recently written entry from an older one within limits that examiners debate.
Alterations show themselves under the right illumination. Infrared examination separates inks that look identical in visible light, because two dyes may absorb infrared differently, so an added digit in a check amount appears as a different shade or vanishes entirely. Infrared luminescence can read writing under an obliterating scribble. Ultraviolet light reveals erasures and patches by the disturbance in the paper surface.
Indented writing deserves its own mention because the technique is elegant. When someone writes on a pad, the pressure disturbs the sheets beneath, producing indentations too shallow to see. An electrostatic detection apparatus lays a thin plastic film over the questioned sheet, applies a charge, and cascades toner powder across it. The toner collects preferentially in the indented regions, and the writing from the missing page appears. It is non-destructive, and it has recovered addresses, phone numbers, and drafts that the writer believed were gone.
Machines leave signatures too
Typewriters were a comparison examiner's dream: each machine developed individual defects as type bars bent and characters chipped, so a document could be tied to one machine. Typewriters are now rare, and their successors leave different traces.
Photocopiers and laser printers deposit trash marks, small recurring artifacts from dirt or damage on the optics or the drum, which repeat in the same position on every page. Drum wear produces banding at intervals corresponding to the drum's circumference. Toner formulations differ among manufacturers and can be characterized by infrared spectroscopy.
Most color laser printers also print something on purpose. Many models embed a pattern of tiny yellow dots on every page, effectively invisible under ordinary light, encoding the printer's serial number and a timestamp. The pattern is a machine identification code, adopted at the request of authorities concerned about currency counterfeiting. In 2017 the technique reached public attention when a classified document that had been printed and then leaked was published, and observers noted the tracking dots on the scanned image.
Handwriting: the premise and the process
Handwriting comparison rests on two claims. The first is that after a writer learns a copybook form, the act becomes automated and personal, so that mature writing carries individual habits. The second is that a person's writing varies within a range, so an examiner must know the range before judging a difference to be significant. Both claims are reasonable and neither is quantified.
The process begins with exemplars, and the quality of the comparison depends almost entirely on them. Collected exemplars are writings produced in the ordinary course of life: signed checks, letters, forms. Request exemplars are produced under supervision, dictated so the writer does not see the questioned document, repeated many times, and written with a similar instrument on similar paper. Both types matter, because a person asked to write may disguise, and a person's casual writing may not include the words in question.
The features an examiner works through include letter design and construction order, proportions and relative heights, slant, spacing between letters and words, baseline habits, connecting strokes, pen lifts, initial and terminal strokes, pressure patterns, and arrangement on the page. Line quality carries particular weight: fluent writing has smooth acceleration and consistent pressure, while a person drawing a shape they do not habitually write produces hesitation, blunt starts, tremor, and patching.
That is why forgery types have distinct signatures.
| Type | How it is made | What betrays it |
|---|---|---|
| Freehand simulation | Copying a model by eye | Good general form, poor line quality: tremor, pen lifts, blunt endings |
| Tracing | Following an original through light or with a guideline | Slow, hesitant line, and an unnaturally close fit to one specific genuine signature |
| Disguised writing | A writer deliberately changing his own hand | Unnatural features that lapse under speed, with underlying habits persisting |
| Auto-forgery | A person denying his own genuine signature | Normal line quality with variation inside the writer's own natural range |
Modern signatures create a real problem the field is still absorbing. A signature captured on a delivery tablet with a stylus is short, distorted by the digitizer, and often not the writer's habitual form at all.
How good are examiners?
The honest summary has three parts, and they do not all point the same way.
First, the discipline's traditional claims were overstated. Albert Osborn, whose 1929 treatise built the field's professional identity and who testified about the Lindbergh ransom notes at the 1935 Hauptmann trial, wrote in an era when the examiner's opinion was treated as near conclusive. The 2009 National Research Council review found the scientific basis for handwriting individuality and for the reliability of examiner conclusions in need of strengthening, and it noted the absence of an established measure of how much agreement is enough.
Second, the ability being claimed has been tested, and it is real. Beginning in the 1990s, Moshe Kam and colleagues ran controlled trials comparing professional document examiners with lay people on tasks of distinguishing genuine signatures from simulations. Professionals substantially outperformed non-professionals, and, importantly, their false positive rates, calling a forgery genuine or a simulation authentic, were far lower. Later studies by other groups produced broadly similar results for signature tasks.
Third, the tasks differ enormously in difficulty. A signature is a short, extremely practiced movement of which many genuine specimens exist, and that is the best case. A comparison of a few disguised block capitals on a ransom note against limited exemplars is the worst case, and the measured accuracy on signature tasks does not transfer to it. PCAST touched on handwriting only briefly in 2016 and did not evaluate it at the depth it applied to fingerprints, firearms, and DNA.
Bottom line: Signature examination by trained professionals has measurable skill behind it, and that measurement does not license the same confidence about extended or disguised handwriting from thin exemplars.
Common misconceptions
- Graphology and document examination are related. Graphology claims to read personality from writing and has no scientific support; document examination compares writing to determine authorship.
- A single signature is enough for comparison. Examiners need many genuine exemplars to establish the writer's natural range before any difference can be called significant.
- Tracing produces the best forgeries. Tracing produces excellent form and terrible line quality, and an unnaturally exact fit to one genuine signature is itself suspicious.
- Printed documents carry no identifying information. Trash marks, drum banding, toner chemistry, and yellow tracking dots all tie pages to machines.
- Handwriting evidence is as strong as the materials evidence in the same case. Paper and ink carry manufacturing dates that settle questions; handwriting conclusions rest on examiner judgment with a thinner validation record.
Recap
- Document examination combines materials analysis, which is strong, with handwriting comparison, which is weaker, and practitioners do both.
- Optical brighteners, watermarks, ink chromatography, and reference ink libraries date documents and expose anachronisms.
- Infrared and ultraviolet examination reveal alterations, and electrostatic detection recovers indented writing from underlying sheets.
- Copiers and printers leave trash marks, banding, toner signatures, and in many color models an encoded yellow dot pattern.
- Handwriting comparison depends on adequate exemplars and reads line quality as heavily as letter form.
- Controlled studies show professional examiners outperform lay people on signature tasks, and that result does not extend to every handwriting question.
Sources
- National Research Council. (2009). Strengthening forensic science in the United States: A path forward. National Academies Press. nap.nationalacademies.org
- National Institute of Standards and Technology. (n.d.). Forensic science. nist.gov
- Wikipedia contributors. (n.d.). Hitler Diaries. en.wikipedia.org
- Osborn, A. S. (1929). Questioned Documents (2nd ed.). Boyd Printing Company.
- Kam, M., Wetstein, J., and Conn, R. (1994). Proficiency of professional document examiners in writer identification. Journal of Forensic Sciences, 39(1), 5-14.
- Key terms
- Questioned document
- Any document whose authorship, authenticity, date, or integrity is in dispute and is submitted for examination.
- Optical brightener
- A fluorescent additive that makes paper appear whiter under ultraviolet light, in wide use only from the mid twentieth century, so its presence sets an earliest date.
- Thin layer chromatography
- A separation technique that resolves an ink's dye components on a plate, producing a pattern characteristic of the formulation.
- Electrostatic detection apparatus
- A non-destructive device that reveals indented writing by charging a film over a document and cascading toner into the indentations.
- Trash marks
- Recurring artifacts printed by a copier or printer from dirt or damage on its optics or drum, appearing in the same position on every page.
- Machine identification code
- A pattern of nearly invisible yellow dots printed by many color laser printers encoding the machine's serial number and a timestamp.
- Exemplar
- A known writing sample used for comparison, either collected from ordinary life or produced on request under supervision.
- Line quality
- The fluency of a written stroke, including pressure consistency, smoothness, and absence of tremor, which typically betrays simulation and tracing.
Module 5: Bodies and Bytes
Three disciplines that answer questions no bench chemist can: what killed this person and when, who this skeleton was and what happened to it, and what a phone, a laptop, and a set of network records can establish about where a person was and what they did.
Forensic Pathology: Cause, Manner, and Time Since Death
- Distinguish cause, mechanism, and manner of death, and explain why homicide as a manner is not a legal conclusion.
- Describe the components of a forensic autopsy and the wound classes a pathologist distinguishes.
- Estimate postmortem interval from livor, rigor, algor, and entomological evidence, and state the uncertainty honestly.
- Compare coroner and medical examiner systems and explain the workforce constraints on death investigation.
Two autopsies, one death
George Floyd died in Minneapolis on 25 May 2020, and two separate autopsies were performed. The Hennepin County Medical Examiner, Dr. Andrew Baker, certified the cause as cardiopulmonary arrest complicating law enforcement subdual, restraint, and neck compression, and listed significant conditions including heart disease and the presence of drugs. An independent examination arranged by the family, conducted by Dr. Michael Baden and Dr. Allecia Wilson, described death by asphyxia from sustained pressure.
Commentators treated the two reports as contradictory. They were not, exactly. Both certified the manner as homicide. They differed in which link of the lethal chain they placed in the cause line, and that difference is not a disagreement about facts so much as a disagreement about how to write a sentence that death certification forces you to write.
Understanding why requires three words that students routinely collapse into one, and separating them is the most useful thing this lesson can do for you.
Key idea: Cause, mechanism, and manner are three different statements about a death, and most public confusion about autopsy findings comes from treating them as one.
Cause, mechanism, manner
The cause of death is the disease or injury that starts the sequence ending in death: a gunshot wound of the chest, blunt force injuries of the head, a myocardial infarction.
The mechanism of death is the physiological derangement through which the cause kills: exsanguination, cardiac arrhythmia, cerebral edema. The same mechanism follows from many causes, which is why a certificate reading cardiac arrest tells you nothing; everyone who dies has a cardiac arrest.
The manner of death is a classification into one of five categories: natural, accident, suicide, homicide, or undetermined. It is a public health classification made by a physician on the balance of the evidence, and this is where the misunderstanding lives. Homicide in this sense means death at the hands of another person. It is not a finding of murder, it does not address intent or justification, and a lawful killing by a police officer or an act of self-defense is still certified as a homicide. Juries decide the legal question; pathologists decide the classification.
Who does death investigation
The United States runs two systems side by side. A coroner is usually an elected county official, and in most states the office requires no medical training; a coroner may be a funeral director, a sheriff, or anyone who wins the election. A medical examiner is an appointed physician, in a well-run jurisdiction a board-certified forensic pathologist, meaning a doctor with a pathology residency plus a forensic fellowship. Some states are entirely medical examiner systems; many are county-based mixtures.
The 2009 National Research Council review recommended converting coroner offices to medical examiner systems, and the conversion has been slow, for a reason that is arithmetic rather than politics. The country has only a few hundred full-time board-certified forensic pathologists, against an estimated need roughly twice that. The National Association of Medical Examiners caps accredited offices at 250 autopsies per pathologist per year, with a hard ceiling above that, because quality falls when caseload rises. Offices routinely exceed it. When you read about a jurisdiction that did not autopsy a suspicious death, the explanation is usually that there was no one to do it.
What happens at an autopsy
An autopsy begins before any incision. The body is received in a sealed bag, and the seal is the chain of custody. Clothing is examined and retained, since defects in fabric correspond to wounds and residue on fabric supports range determination. Radiographs locate bullets and fractures, which matters because a projectile can travel far from the entrance. External examination documents identifying features, injuries, therapeutic interventions, and postmortem changes, and evidence is collected: fingernail clippings, sexual assault kit samples, hair, and swabs.
The internal examination follows a standard sequence: a Y-shaped incision, examination of body cavities in place, removal and dissection of organs with weights recorded, examination of the neck structures with particular care in suspected strangulation, and removal of the brain. Sections of tissue go to histology. Blood, urine, vitreous humor, and sometimes liver go to toxicology, with peripheral blood collected from a clamped femoral vein for the reason you met in the toxicology lesson.
Wound interpretation is a language of its own, and confusing two of its terms is a common error.
| Wound | Produced by | Distinguishing feature |
|---|---|---|
| Abrasion | Scraping of the skin surface | Direction can often be read from skin tags |
| Contusion | Blunt impact rupturing vessels under intact skin | Color changes over days; poor timing indicator |
| Laceration | Blunt force tearing tissue | Irregular margins with tissue bridges spanning the wound |
| Incised wound | A sharp edge drawn across skin | Clean margins, no tissue bridges, longer than deep |
| Stab wound | A pointed instrument driven in | Deeper than long; depth bounds but does not fix blade length |
| Gunshot entrance | A projectile perforating skin | Abrasion collar; soot or stippling at close range |
The laceration and incised wound distinction is worth memorizing precisely because everyday speech gets it backwards. A laceration is a tear from blunt force, not a cut. Tissue bridges, strands of vessel and nerve crossing the wound floor, survive tearing and are severed by a blade, which is why they separate the two.
The upshot: An autopsy is an evidence collection procedure as much as a medical examination, and much of what it establishes comes from documentation and sampling rather than from dissection.
How long has this person been dead?
No question in forensic pathology is asked more often or answered with less precision than the postmortem interval. Four sets of changes contribute, and each has a usable window and a set of things that ruin it.
Livor mortis is the settling of blood under gravity, producing purple discoloration in dependent areas, with pallor where the body presses against a surface. It becomes visible within roughly half an hour to two hours, and after roughly eight to twelve hours it fixes, meaning it no longer shifts if the body is moved. Its greatest value is not timing but position: lividity on the front of a body found on its back tells you the body was moved, and when.
Rigor mortis is muscle stiffening from depletion of the energy needed to release the actin-myosin bond. It begins within a few hours, is generally complete around twelve hours, and passes off over the following day or so as decomposition takes over. It runs faster in heat, faster after strenuous activity or convulsions, and slower in cold, which means a body in a cold room can be far past the interval its stiffness suggests.
Algor mortis is cooling. Under moderate conditions a body loses roughly one and a half degrees Fahrenheit per hour after an initial plateau, but that figure assumes an average adult, average clothing, still air, and an ordinary room. Body mass, insulation, wind, immersion, and ambient temperature change it dramatically, and formal estimates use a nomogram that takes these into account rather than a single rate.
Chemical and biological changes extend the window. Potassium leaks from retinal cells into the vitreous humor of the eye at a fairly regular rate over the first day or two. Beyond that, decomposition takes over: autolysis, then putrefaction with greenish discoloration of the abdomen, marbling along vessels, bloating, and skin slippage. In wet conditions fat can convert to adipocere, a waxy substance that preserves the body's form for years; in dry conditions the body mummifies.
Forensic entomology becomes the best tool once soft tissue is affected. Blow flies find a body within minutes to hours in daylight and lay eggs at natural openings and wounds. Development from egg through larval instars to pupa to adult proceeds at rates that depend on temperature, so an entomologist collects the oldest specimens, rears them, obtains weather data, and back-calculates using accumulated degree hours. Species succession over weeks extends the estimate further.
Two honest caveats. Entomology estimates time since colonization, not time since death, and burial, wrapping, indoor conditions, cold, and freezing all delay colonization. And every one of these methods produces a range, which widens fast. Beyond the first day, an estimate expressed to the hour is almost always overstated, and a pathologist who gives one on the stand is telling you something about the pressure of the courtroom rather than the state of the science.
Where pathology is genuinely contested
One area deserves naming because it is unresolved and because people are in prison over it. For decades, the finding of subdural hemorrhage, retinal hemorrhage, and brain swelling in an infant, in the absence of external injury, was treated as diagnostic of violent shaking. That triad supported many convictions. Since the early 2000s, a body of work has argued that the same findings can arise from short falls, from birth-related injury, from certain metabolic and clotting disorders, and from other mechanisms, and that the timing inferences drawn from them were unreliable. Courts have divided. In Wisconsin, Audrey Edmunds, convicted in 1996, was granted a new trial in 2008 explicitly because the medical consensus had shifted, and the charges were later dropped.
The mainstream pediatric position remains that abusive head trauma is real, diagnosable, and common. The contested claim is narrower: whether the triad alone, without other evidence, can establish that a specific person shook a specific child at a specific moment. Keeping those two propositions apart is the whole of clear thinking on the subject.
Common misconceptions
- A manner of homicide means murder. It means death at the hands of another and says nothing about intent, justification, or criminal liability.
- Cardiac arrest is a cause of death. It is a mechanism that every death includes; a certificate must name the disease or injury that started the sequence.
- A laceration is a cut. A laceration is a tear from blunt force with tissue bridges; a cut from a blade is an incised wound.
- Body temperature gives an accurate time of death. Cooling depends on mass, clothing, air movement, and environment, and every method gives a range that widens quickly.
- Insects tell you when the person died. They tell you when colonization began, which burial, wrapping, cold, and indoor settings can delay substantially.
Where this leaves us
- Cause starts the sequence, mechanism is the physiological failure, and manner is a five-category classification, not a legal verdict.
- Coroner and medical examiner systems coexist, and a national shortage of forensic pathologists constrains what gets investigated.
- An autopsy documents and samples as much as it dissects, beginning with clothing, radiographs, and external examination.
- Tissue bridges separate a laceration from an incised wound, and abrasion collars mark gunshot entrances.
- Livor fixes position, rigor and algor give short windows, vitreous potassium extends them, and entomology takes over afterward.
- Every postmortem interval estimate is a range, and precision claimed beyond the first day is not supported.
Sources
- Britannica. (2024). Autopsy. britannica.com
- National Research Council. (2009). Strengthening forensic science in the United States: A path forward. National Academies Press. nap.nationalacademies.org
- National Institute of Standards and Technology. (n.d.). Forensic science. nist.gov
- DiMaio, V. J. M., and Molina, D. K. (2021). DiMaio's Forensic Pathology (3rd ed.). CRC Press.
- National Association of Medical Examiners. (2016). Forensic autopsy performance standards. NAME.
- Key terms
- Cause of death
- The disease or injury that initiates the sequence of events ending in death, such as a gunshot wound of the chest.
- Mechanism of death
- The physiological derangement through which a cause produces death, such as exsanguination or arrhythmia.
- Manner of death
- A classification of a death as natural, accident, suicide, homicide, or undetermined; homicide here means death at another's hands, not murder.
- Livor mortis
- Gravitational settling of blood producing dependent discoloration, which fixes after roughly eight to twelve hours and reveals whether a body was moved.
- Rigor mortis
- Postmortem muscle stiffening that begins within hours, peaks around twelve hours, and resolves over the following day, accelerated by heat and exertion.
- Tissue bridges
- Strands of vessel and connective tissue spanning a wound floor, present in blunt force lacerations and absent in incised wounds.
- Abrasion collar
- A ring of scraped skin around a gunshot entrance produced as the bullet stretches and indents the skin on perforation.
- Accumulated degree hours
- The temperature-weighted development measure entomologists use to back-calculate how long insect larvae have been growing on remains.
Forensic Anthropology: Reading a Skeleton
- Construct a biological profile from skeletal remains and state the confidence attached to each element.
- Distinguish antemortem, perimortem, and postmortem skeletal trauma and identify blunt, sharp, and projectile signatures.
- Explain the dispute over ancestry estimation, presenting the case on each side.
- Describe how anthropological findings contribute to identification without themselves constituting one.
The colonel in the cast-iron coffin
In 1977 a landowner near Franklin, Tennessee found a grave disturbed and called the police, who called William Bass, then the state's forensic anthropologist. Bass examined a body in a cast-iron coffin. The tissue was in remarkable condition, pink and moist. He estimated that the man had been dead somewhere between a few months and a year.
The remains were those of Lieutenant Colonel William Shy, a Confederate officer killed at the Battle of Nashville and buried in 1864. Bass had missed by more than a century. The sealed iron coffin and embalming had preserved tissue in a way nothing in his training accounted for, and, as he later wrote, nobody had ever studied what actually happens to a human body as it decomposes under controlled and recorded conditions. In 1981 he founded the Anthropology Research Facility at the University of Tennessee, the first place where donated bodies are placed outdoors and observed. Nearly everything the field now knows about decomposition rates began with a professional embarrassment.
Worth holding on to: Forensic anthropology's modern research base exists because a leading practitioner was badly wrong in public and treated the error as a research question.
What the anthropologist is called for
A forensic anthropologist handles what a pathologist cannot: remains that are skeletonized, burned, fragmented, scattered, or decomposed past the point where soft tissue answers questions. The first question is often the most basic. Is this bone human? A great deal of anthropological consultation ends with the answer bear paw, deer, or pig, which resemble human remains often enough to generate calls.
Beyond that, the work divides into recovery, profile, and trauma. Recovery borrows archaeological method outright: a scattered surface scene is mapped and gridded before anything is lifted, a burial is excavated in levels with the fill screened, and the position of every element is recorded, because the arrangement of bones carries information about deposition, scavenging, and movement that is destroyed by picking them up.
The biological profile
The biological profile is an estimate of sex, age, stature, and population affinity, built to narrow a list of missing persons. Each element has a different reliability, and knowing which is which separates a competent reader of an anthropology report from a credulous one.
Sex is the most reliable element in an adult skeleton, and the pelvis is the reason. Childbirth shapes the female pelvis in ways that show: a wider subpubic angle, a broader greater sciatic notch, the presence of a ventral arc on the pubis, and often a preauricular sulcus. With a complete adult pelvis, estimates exceed 90 percent accuracy. The skull is second best, using the size of the mastoid process, the prominence of the brow ridge and glabella, the nuchal crest, and the shape of the chin. Both approaches fail in subadults, because the diagnostic features develop with puberty.
Age reverses that pattern. Subadult age estimation is excellent, because growth is a schedule: teeth form and erupt in a known sequence, and the growth plates of the long bones fuse in a known order over adolescence and the early twenties. An anthropologist can often place a fifteen-year-old within a year or two. Adult age estimation is far weaker, because after growth stops the only signals are degenerative. The pubic symphysis changes surface texture and rim development in phases described by Suchey and Brooks, the auricular surface of the ilium degrades with age, and sternal rib ends develop characteristic pitting. Cranial suture closure is notoriously unreliable. Adult estimates come as broad ranges, and the ranges widen with age, so an estimate of thirty-five to fifty for a middle-aged adult is honest rather than lazy.
Stature is calculated by regression from long bone lengths, most reliably the femur, using formulas developed by Trotter and Gleser and refined since. The output is a range of several centimeters, and it is population-dependent, since limb proportion to height varies.
The core of it: Sex from an adult pelvis is strong, subadult age is strong, adult age is a wide range, and stature is an interval; a profile that reports single numbers is overstating what bone can say.
The argument about ancestry
The fourth element is genuinely disputed inside the discipline, and you should hear both positions in their own terms rather than a compromise between them.
The traditional practice estimates population affinity from cranial measurements and morphoscopic traits, using software such as FORDISC that compares a skull's measurements to reference samples and reports statistical affinities. Practitioners who defend it make an argument from utility that is not easy to dismiss: missing persons databases and police reports record race as a social category, so an unidentified decedent's chance of being matched to a missing person report depends on the anthropologist producing a term that appears in those records. Removing the estimate, they argue, does not make the skeleton more anonymous in a good way; it makes the person less likely to go home.
The critique starts from the biology. Human variation is largely clinal, changing gradually across geography rather than sorting into discrete groups, and the reference samples that classification software relies on are collections shaped by who ended up in anatomical collections, which in the United States means disproportionately poor and Black individuals whose bodies were unclaimed. Critics argue that what the method actually predicts is how a person would have been socially classified in a particular society, and that reporting that as a biological finding gives a social category the appearance of a scientific one. The American Association of Biological Anthropologists has issued statements rejecting race as a valid biological description of human variation, and a growing number of practitioners have moved to reframe the estimate explicitly as a prediction of social classification, or to stop offering it.
Both positions accept the same underlying facts. The disagreement is about what a forensic report is for, and whether a category that is socially real and biologically incoherent should be reported by a scientist. It is not settled, and any lesson claiming it is settled is misleading you.
Trauma: when, and by what
Skeletal trauma is sorted first by timing, and the distinctions are physical rather than interpretive.
- Antemortem injuries show healing: callus formation, remodeled margins, sometimes infection. Healing takes time, so any sign of it means the injury preceded death by at least days.
- Perimortem injuries occur when bone is still fresh, with its collagen intact and some plasticity. Fresh bone bends before it breaks, producing hinged fractures, plastic deformation, beveled margins, and fracture surfaces the same color as the surrounding bone.
- Postmortem damage occurs to dry bone, which is brittle. Breaks are squared off, right-angled, and typically lighter in color than the weathered outer surface, because the interior has not been exposed.
Then by mechanism. Blunt force to the cranial vault produces radiating fractures from the impact and concentric fractures between them, and because a later fracture stops when it meets an earlier one, the sequence of multiple impacts can often be reconstructed. Sharp force leaves a kerf, the cut channel itself, whose walls carry striations from the blade and whose cross-section reflects the blade's geometry; anthropologists can often distinguish a serrated from a smooth edge, and a saw's tooth spacing from its kerf floor. Projectile trauma in the cranial vault produces a characteristic asymmetry: the entrance is beveled internally, widening on the inner table, and the exit is beveled externally, which lets an examiner determine direction from a skull with two holes and no other information. Thermal alteration produces its own patterns, including heat-induced fracture and the pugilistic posture of burned bodies, which is muscle contraction and not defensive positioning.
Naming the person
Here is the distinction the discipline is careful about and the public is not. A biological profile does not identify anyone. It narrows a search. Positive identification requires comparison to antemortem records: dental radiographs against postmortem radiographs, a healed fracture or surgical implant matched to a medical film, a frontal sinus pattern, or DNA compared to a reference or a family sample. The anthropologist supplies the profile and often performs the radiographic comparison, and the identification is certified by the medical examiner or coroner.
Where the work goes matters. The National Missing and Unidentified Persons System, run through the National Institute of Justice, holds records on unidentified decedents and missing people and lets the two be searched against each other, which is how a biological profile becomes a name years later. The Defense POW/MIA Accounting Agency applies the same methods at scale, and its project to identify the unaccounted-for crew of the USS Oklahoma, sunk at Pearl Harbor in 1941, disinterred and identified the great majority of 394 men using anthropology, dental comparison, and DNA. In human rights work, Clyde Snow trained the Argentine Forensic Anthropology Team in 1984 to exhume and identify people disappeared under the military dictatorship, and that model, forensic scientists working for families rather than for a state, has been repeated in dozens of countries since.
Common misconceptions
- An anthropologist identifies the person. A biological profile narrows a candidate pool; identification requires comparison to antemortem records or DNA and is certified by the medical examiner.
- Bones give an exact age. Subadult age is tightly constrained by growth, but adult age rests on degenerative changes and is reported as a wide range.
- Ancestry estimation reads a person's race from the skull. What the method predicts is closer to how a person would have been socially classified, and whether to report it at all is actively disputed.
- A break in a bone means violence. Postmortem damage from scavengers, equipment, and burial is common and is distinguished by fracture geometry and color.
- The pugilistic posture of a burned body indicates a defensive struggle. It is heat-induced muscle contraction and occurs regardless of what happened before the fire.
What to carry forward
- Forensic anthropology handles skeletonized, burned, fragmented, and scattered remains, starting with whether the bone is human.
- Recovery uses archaeological method, because the arrangement of remains carries information that lifting destroys.
- Adult sex from the pelvis and subadult age from growth are strong; adult age and stature are ranges.
- Ancestry estimation is defended on identification utility and criticized as reporting a social category as biology, and the dispute is open.
- Healing marks antemortem injury, plastic deformation marks perimortem, and squared brittle breaks mark postmortem damage.
- Internal beveling marks a cranial entrance and external beveling an exit, and identification itself requires antemortem records.
Sources
- National Institute of Justice. (n.d.). National Missing and Unidentified Persons System (NamUs). Office of Justice Programs. namus.nij.ojp.gov
- Defense POW/MIA Accounting Agency. (n.d.). Our mission. U.S. Department of Defense. dpaa.mil
- American Academy of Forensic Sciences. (n.d.). Anthropology section. aafs.org
- Byers, S. N. (2016). Introduction to Forensic Anthropology (5th ed.). Routledge.
- Bass, W. M., and Jefferson, J. (2003). Death's Acre: Inside the Legendary Forensic Lab the Body Farm. G. P. Putnam's Sons.
- Key terms
- Biological profile
- An estimate of a decedent's sex, age, stature, and population affinity from skeletal evidence, used to narrow a list of missing persons.
- Subpubic angle
- The angle formed by the pubic bones below the symphysis, wider in females and one of the most reliable indicators of skeletal sex.
- Epiphyseal fusion
- The closing of growth plates in a known sequence through adolescence and early adulthood, the basis of accurate subadult age estimation.
- Pubic symphysis phases
- Staged descriptions of surface texture and rim formation on the pubic face, used with reference standards to estimate adult age as a range.
- Perimortem trauma
- Injury occurring while bone retains collagen and plasticity, producing hinged fractures, plastic deformation, and fracture surfaces matching the bone's color.
- Internal beveling
- The cone-shaped widening of a projectile hole on the inner table of the skull, marking an entrance; external beveling marks an exit.
- Kerf
- The channel cut by a blade or saw, whose walls and floor carry striations and geometry reflecting the tool that made it.
- Positive identification
- Establishing a decedent's identity by comparison to antemortem records such as dental or medical radiographs, implants, or DNA.
Digital Forensics: Devices, Location, and Attribution
- Apply the core acquisition principles of write blocking, bit-for-bit imaging, and hash verification, and explain the order of volatility.
- Explain how deletion, slack space, carving, and metadata make recovery possible and how encryption limits it.
- Interpret cell site and geofence location evidence at the correct level of precision.
- Analyze the attribution problem: why a device is not a person.
Seven days to stop a warrant
In January 2020 a twenty-nine-year-old named Zachary McCoy received an email from Google. Law enforcement had requested information associated with his account, and unless he obtained a court order within seven days, the company would hand it over. He had no idea why.
He hired a lawyer, and the reason emerged. A house in Gainesville, Florida had been burglarized, and police had obtained a geofence warrant: rather than naming a suspect, it asked Google which devices had been inside a defined area during a defined window. McCoy rode his bicycle for exercise and used a fitness app that recorded his routes to his Google account. He had passed the house three times that day.
The warrant was eventually withdrawn. McCoy had done nothing except own a phone and take the same route he always took. What makes the case worth opening a lesson with is not that police were malicious; it is that the investigative technique worked exactly as designed and produced a suspect out of ordinary behavior.
What matters here: Digital evidence is abundant, precise-looking, and generated by people going about their lives, and the difficulty is almost never recovering it; it is deciding what it means.
Acquire without changing
Every discipline in this course has a preservation rule. In digital forensics it is unusually strict, because a computer alters itself constantly: simply booting a machine writes to disk, updates timestamps, and can destroy exactly what you wanted.
So the examiner works from a copy. A write blocker, hardware or software, sits between the evidence drive and the examination system and permits reads while physically refusing writes. The examiner makes a forensic image: a bit-for-bit duplicate including unallocated areas, not a file copy. Then comes verification. A cryptographic hash function reduces the entire drive to a fixed-length value, and any single-bit change produces a completely different value. The examiner hashes the original and the image, records both, and hashes again after analysis. Matching hashes are how a witness answers the question of whether the evidence changed while she had it, and they are the digital analogue of a sealed package.
One complication overturns the classic advice. The old rule was to pull the plug on a running machine so nothing further is written. With full disk encryption in wide use, pulling the plug can mean the drive is unrecoverable, because the key lived in memory. So examiners follow an order of volatility: capture the most perishable data first. Memory contents, running processes, network connections, and open encrypted volumes evaporate at power-off; disk contents do not.
Why deleted does not mean gone
When you delete a file, most file systems do not erase its contents. They mark the space as available and remove the pointer in the directory structure. The data sits there until something else writes over it, which may be minutes or years.
Three consequences follow, and they are the bread and butter of examination.
Unallocated space holds the remains of deleted files. File carving recovers them without any directory information, by scanning raw bytes for the signatures that mark the start and end of known formats, so that a JPEG or a document can be reassembled from the middle of nowhere. Slack space is the gap between the end of a file and the end of the storage cluster it occupies, and it can contain fragments of whatever was there before.
Metadata multiplies the yield. Photographs carry EXIF data recording camera model, settings, timestamps, and often GPS coordinates. Office documents carry authorship, revision counts, and editing time. Operating systems keep their own records: browser history and cache, Windows registry keys recording every USB device ever attached, prefetch files showing which programs ran and when, shortcut files pointing at documents opened from removable media, and event logs.
Timestamps deserve caution. Modified, accessed, and created times are informative and are also easy to change, deliberately with readily available tools or accidentally by a backup program or an antivirus scan. Treating a timestamp as a fact rather than as an artifact whose provenance needs explaining is a standard way to be wrong.
Phones, which is where the evidence actually is
Most cases now turn on a phone. Extraction comes in tiers. A logical extraction asks the device for its data through normal interfaces and gets what the operating system will hand over: contacts, messages, call logs, some app data. A file system extraction reaches more, including databases holding deleted records. A physical extraction images the storage itself and can recover deleted content, but modern devices resist it. Chip-off, where the memory is physically removed and read, is a last resort and destructive.
Encryption changed everything. Modern iPhones and Android devices encrypt storage by default with a key tied to the passcode and to hardware, and the hardware enforces delays and lockouts after failed attempts. This produced the most public dispute in the field's history. In 2016, after the San Bernardino attack, the FBI obtained an order directing Apple to write software that would disable the passcode protections on a recovered iPhone. Apple refused, arguing that building such a tool would endanger every user. The case ended without a ruling when a third party sold the government a method. The underlying tension, between device security that protects everyone and access that helps investigators, has not been resolved and will not be.
The law has moved toward requiring warrants. In Riley v. California in 2014 the Supreme Court held that police generally may not search a cell phone incident to arrest without a warrant, reasoning that a phone is not like a cigarette pack in a pocket but a repository of a person's whole life.
Location evidence, at the right precision
Two techniques dominate, and both are routinely overstated.
Cell site location information comes from carrier records showing which tower and sector a phone used for each connection. Read correctly, that indicates the phone was somewhere within the coverage area of a sector, which in a rural area can be many kilometers across and in a dense urban area may be a few hundred meters, and coverage areas change with load, weather, and network configuration. Read incorrectly, as it has been in many trials, it becomes a dot on a map. The 2009 National Research Council review's general warning about overstatement applies here squarely, and defense challenges to cell site testimony have succeeded where analysts drew coverage areas as neat wedges. In Carpenter v. United States in 2018, the Supreme Court held that acquiring historical cell site records is a search requiring a warrant.
Geofence warrants reverse the usual order. Instead of naming a person and asking for their records, they name a place and time and ask a provider which accounts were there. Google handled these through a staged process, returning anonymized device identifiers first, then narrowing, then de-anonymizing a small set. Courts have divided sharply on whether such warrants satisfy the Fourth Amendment's particularity requirement, since by design they sweep in people about whom there is no suspicion at all, as Zachary McCoy discovered. In late 2023 Google announced that Location History would be stored on users' devices rather than on its servers, which largely removes the company's ability to answer these requests.
The attribution problem
Here is the limit that every digital examiner states and every television treatment ignores. Analysis establishes what happened on a device. It does not establish who was at the keyboard.
Households share computers. Passwords get reused and shared. Malware genuinely does download files without a user's knowledge, and the defense built on that fact has succeeded in real cases. Accounts are accessed from multiple devices, and an IP address identifies a network connection, not a person, and after network address translation it may identify hundreds of people at once.
Attribution therefore comes from patterns rather than from the device alone: who was logged in, what else was happening in the same seconds, whether a phone associated with a specific person was on the same network, whether typing patterns and language match, whether physical evidence places someone in the room. A competent report says the activity originated from this account on this device at this time, and stops.
Two further pressures shape modern practice. Volume is one: a single phone can hold hundreds of gigabytes, and triage strategies decide what gets examined at all. Validation is the other. Examiners rely on commercial tools whose internals are proprietary, which is why the National Institute of Standards and Technology runs a program that tests forensic tools against specifications and publishes the results, giving examiners something to cite when asked how they know their software did what it claimed.
Common misconceptions
- Deleting a file erases it. Most file systems only remove the pointer and mark the space free; the content persists until overwritten.
- An IP address identifies a person. It identifies a network connection at a moment, and address translation can place hundreds of users behind one address.
- Cell site records show where a phone was. They show which tower sector served a connection, an area rather than a point, and sector coverage shifts with conditions.
- Examiners work on the original device. They work on a verified bit-for-bit image behind a write blocker, because examining an original changes it.
- Timestamps are reliable facts. They are artifacts that ordinary software and deliberate tools can alter, and their provenance has to be explained.
The short version
- Acquisition means write blocking, bit-for-bit imaging, and hash verification, with volatile memory captured first on a live system.
- Deletion removes pointers, so unallocated space, carving, and slack space recover what users believe is gone.
- Metadata and operating system artifacts often carry more information than the files themselves.
- Phone extraction runs from logical to physical, and default encryption has made the hardest cases genuinely hard.
- Cell site data gives sectors and geofence warrants sweep bystanders, and both require careful statements of precision.
- A device is not a person, and attribution requires evidence beyond the device.
Sources
- National Institute of Standards and Technology. (n.d.). Forensic science. nist.gov
- Legal Information Institute. (n.d.). Fourth Amendment. Cornell Law School. law.cornell.edu
- National Institute of Justice. (n.d.). Forensic sciences. Office of Justice Programs. nij.ojp.gov
- Carpenter v. United States, 585 U.S. 296 (2018). Supreme Court of the United States.
- Casey, E. (2011). Digital Evidence and Computer Crime (3rd ed.). Academic Press.
- Key terms
- Write blocker
- Hardware or software that permits reads from an evidence drive while preventing any write, so examination cannot alter the original.
- Forensic image
- A bit-for-bit duplicate of storage media, including unallocated areas, on which all analysis is performed.
- Cryptographic hash
- A fixed-length value computed from data such that any change produces a different value, used to verify that evidence has not been altered.
- Order of volatility
- The principle of capturing the most perishable data first, such as memory and network state, before powering down a running system.
- File carving
- Recovering files from raw storage by locating format signatures, without relying on file system directory information.
- Slack space
- The unused remainder of a storage cluster beyond the end of a file, which can retain fragments of previously stored data.
- Cell site location information
- Carrier records of which tower and sector served a phone's connections, indicating an area of coverage rather than a precise point.
- Geofence warrant
- A warrant naming a place and time rather than a person, requiring a provider to identify devices present, which necessarily includes uninvolved people.
Module 6: The Reckoning
The part of forensic science most survey courses leave out: two national scientific reviews that asked which comparison methods had ever been validated and did not like the answers, and the mechanisms, cognitive bias, laboratory failure, and overstated testimony, that turn a flawed method into a person in prison.
2009 and 2016: When Scientists Audited Forensic Science
- State the central finding of the 2009 National Research Council report and its principal recommendations.
- Distinguish foundational validity from validity as applied, and apply PCAST's criteria to a discipline.
- Summarize PCAST's conclusions discipline by discipline and the official response to them.
- Explain why bite mark comparison became the clearest case of a method that was never validated.
Thirty-three years on the word of dentists
In 1982 a merchant sailor broke into a home in Newport News, Virginia, killed a man, and assaulted his wife, biting her legs. Keith Allen Harward, a sailor stationed nearby, was tried and convicted. The evidence that convicted him was odontology. Forensic dentists compared the bite marks on the victim to Harward's dentition and told the court the marks were his.
DNA testing in 2016 identified a different sailor, a man who had since died in an Ohio prison. On 8 April 2016 Harward walked out after thirty-three years. He had been convicted, twice, on a method that had never been tested to see whether it worked.
By 2016 that was not news to anyone paying attention, because two national reviews had already said so. This lesson is about those reviews: who ordered them, what they asked, what they found, and how much of it took hold.
Key idea: The scientific critique of the comparison disciplines did not come from defense lawyers; it came from scientists asked by Congress and by the White House whether these methods had ever been validated.
The 2009 report
In 2005 Congress directed the National Academy of Sciences to examine the needs of the forensic science community. The resulting committee was chaired by Harry T. Edwards, a federal appellate judge, and Constantine Gatsonis, a statistician, and it included practitioners, scientists, and lawyers. It took evidence for two years. The report, published in February 2009, was titled Strengthening Forensic Science in the United States: A Path Forward.
Its central finding is easy to state and was, for the field, devastating. With the exception of nuclear DNA analysis, the committee wrote, no forensic method had been rigorously shown to have the capacity to consistently and with a high degree of certainty demonstrate a connection between evidence and a specific individual or source. Not fingerprints. Not firearms. Not bite marks, hair, handwriting, or shoeprints. The committee was not saying these methods were worthless. It was saying that the studies establishing what they could do, and how often they were wrong, had never been performed.
Around that finding the committee built thirteen recommendations. The most consequential ones were:
- Create an independent federal agency, a National Institute of Forensic Science, to fund research, set standards, and oversee the field, deliberately outside the Department of Justice.
- Remove crime laboratories from the administrative control of law enforcement agencies.
- Require accreditation of laboratories and certification of practitioners.
- Fund research on the validity and reliability of the disciplines, and on human observer bias and error.
- Standardize terminology and reporting so that conclusions mean the same thing everywhere.
- Replace coroner systems with medical examiner systems.
What happened next is a lesson in institutions. The independent agency was never created. Instead, in 2013, the Department of Justice and the National Institute of Standards and Technology jointly established a National Commission on Forensic Science, an advisory body of scientists, judges, prosecutors, and defense lawyers, and NIST created the Organization of Scientific Area Committees to develop consensus standards. The Commission produced real work, including recommendations that led the Department to publish uniform language rules governing what its examiners may say and to abandon the phrase reasonable degree of scientific certainty, which sounds meaningful and has no definition. In April 2017 the Attorney General let the Commission's charter expire.
PCAST asks a sharper question
In September 2016 the President's Council of Advisors on Science and Technology published Forensic Science in Criminal Courts: Ensuring Scientific Validity of Feature-Comparison Methods. Where the 2009 report surveyed a field, PCAST did something narrower and harder: it defined a test and applied it.
The test has two parts.
Foundational validity asks whether the method itself works. For a subjective method, one in which a human examiner decides whether two patterns agree, PCAST held that this can only be established empirically, through black box studies in which many examiners work many samples of known ground truth under realistic conditions, and the resulting error rates are measured and published. Not by testimony about training. Not by long acceptance. By counting how often examiners are wrong.
Validity as applied asks whether the examiner in this case actually did it properly: whether she was proficient, followed the validated procedure, and reported her conclusion with the measured error rate attached rather than as a certainty.
Applying that framework produced a discipline-by-discipline verdict.
| Method | PCAST conclusion |
|---|---|
| Single-source and simple-mixture DNA | Foundationally valid; an objective method |
| Complex DNA mixtures | Valid only within a limited tested range of contributors and proportions |
| Latent fingerprints | Foundationally valid, with a false positive rate that must be disclosed rather than denied |
| Firearms and toolmarks | Falls short; resting on a single appropriately designed study |
| Footwear identification of a specific shoe | Not established; class-level conclusions are a different matter |
| Hair comparison for individualization | No foundational validity |
| Bite mark comparison | Not foundationally valid, and PCAST judged the prospect of establishing it low |
The official response was rejection. The Attorney General declined to adopt the recommendations, and the FBI issued a statement disputing the conclusions, arguing that PCAST had used too narrow a definition of scientific validity and had discounted the accumulated experience of the disciplines. Read the exchange and you will see two genuine positions rather than a villain: PCAST insisted that only measured error rates count, and the Bureau argued that a demand for black box studies of every task would exclude reliable expertise the courts have relied on for a century. PCAST's answer, that a century of use without measurement is not evidence of accuracy, is the stronger argument, and the Bureau's practical worry was not frivolous.
The point: PCAST's contribution was a criterion, not a verdict, and the criterion is simple enough to remember: if a method depends on human judgment, show me the study that counted how often the judgment is wrong.
Bite marks, the clearest case
Bite mark comparison is where the whole argument shows most plainly, so it repays a close look.
The discipline rests on two assumptions. First, that human dentition is unique to an individual. Second, that skin records that dentition faithfully enough to permit comparison. Neither has been established. Uniqueness of dentition, even if true, is irrelevant unless the mark preserves the distinguishing features. And skin is the worst recording medium imaginable: it is elastic, it moves, it swells, it bruises differently in different people, and the pattern changes over hours and days after the bite.
Studies asked practitioners to do simpler tasks than the courtroom requires, and they struggled. In one widely discussed exercise, board-certified odontologists were shown photographs and asked whether each was a human bite mark at all, and they disagreed substantially. If examiners cannot agree on whether a mark is a bite, the question of whose teeth made it does not arise.
Meanwhile the exonerations accumulated. Ray Krone was convicted in Arizona in 1992 largely on bite mark testimony, convicted again after a retrial, and cleared by DNA in 2002. Keith Harward served thirty-three years. Eddie Lee Howard spent decades on Mississippi's death row before DNA testing led to his release in 2020 and the dropping of charges. In 2016 the Texas Forensic Science Commission recommended that bite mark comparison not be admitted in Texas courts absent research establishing its validity. In 2023 the National Institute of Standards and Technology published a scientific foundation review of bitemark analysis and found the discipline lacked adequate scientific support.
And yet. Bite mark testimony has not been categorically excluded nationwide, courts continue to consider it case by case, and past convictions built on it are undone one habeas petition at a time, if at all.
What actually changed
It would be false to say nothing happened, and equally false to say the reports transformed the field.
Real changes: accreditation to international laboratory standards is now near-universal among public crime laboratories, where it was optional before. Terminology reformed, so that Department of Justice examiners no longer testify to zero error rates or to individualization to the exclusion of all others in several disciplines. Research funding for black box and validation studies increased substantially, and several disciplines now have error rate estimates they did not have in 2008. Courts began limiting the strength of conclusions, as you saw with fingerprints and firearms. NIST began publishing systematic foundation reviews of specific disciplines.
Unrealized: the independent national institute does not exist. Most crime laboratories remain inside law enforcement agencies. The advisory commission was dissolved. Proficiency testing remains largely non-blind, so examiners usually know they are being tested. And the courts, which the 2009 report identified as unlikely to fix the problem themselves, have largely proved it right.
Common misconceptions
- The 2009 report said forensic science is junk. It said the validation studies had not been done, which is a claim about evidence rather than about worthlessness.
- PCAST's criterion was general acceptance. PCAST required empirical black box studies precisely because general acceptance had let untested methods through for decades.
- Both reports were produced by defense advocates. One was commissioned by Congress and chaired by a federal judge and a statistician; the other was the President's own science advisory council.
- Bite mark evidence has been banned. Texas recommended against admission and NIST found the foundation lacking, but courts nationwide still consider it case by case.
- The reports changed nothing. Accreditation, terminology, error rate research, and judicial limits on testimony all moved; the structural recommendations did not.
Putting it together
- Congress commissioned the 2009 review, which found that apart from nuclear DNA no method had been shown to reliably connect evidence to a specific source.
- Its recommendations centered on an independent institute, laboratory independence, accreditation, standard terminology, and validation research.
- PCAST in 2016 required foundational validity, meaning measured error rates from black box studies, plus valid application in the individual case.
- PCAST accepted single-source DNA and latent prints with stated error rates, and found firearms short and bite marks and hair individualization unsupported.
- The Department of Justice rejected PCAST's recommendations, and the advisory commission's charter was allowed to expire in 2017.
- Bite mark comparison rests on two unestablished assumptions, has produced multiple DNA exonerations, and is still litigated case by case.
Sources
- National Research Council. (2009). Strengthening forensic science in the United States: A path forward. National Academies Press. nap.nationalacademies.org
- National Institute of Standards and Technology. (n.d.). Forensic science. nist.gov
- Innocence Project. (n.d.). DNA exonerations in the United States. innocenceproject.org
- President's Council of Advisors on Science and Technology. (2016). Forensic science in criminal courts: Ensuring scientific validity of feature-comparison methods. Executive Office of the President.
- National Institute of Standards and Technology. (2023). Bitemark analysis: A NIST scientific foundation review (NISTIR 8352). U.S. Department of Commerce.
- Key terms
- Foundational validity
- PCAST's requirement that a method be shown by empirical studies to be repeatable, reproducible, and accurate, with measured error rates.
- Validity as applied
- Whether the examiner in a particular case was proficient, followed the validated method, and reported a conclusion consistent with its measured accuracy.
- Black box study
- A test presenting many examiners with many samples of known ground truth in order to count errors without examining how decisions are made.
- National Institute of Forensic Science
- The independent federal agency recommended by the 2009 report to fund research and set standards outside the Department of Justice; never created.
- National Commission on Forensic Science
- The advisory body established in 2013 by the Department of Justice and NIST, whose charter was allowed to expire in 2017.
- Organization of Scientific Area Committees
- The NIST-administered structure created in 2014 to develop and register consensus technical standards for forensic disciplines.
- Reasonable degree of scientific certainty
- A courtroom phrase with no scientific definition, which the Department of Justice moved to abandon on the advice of the National Commission.
- Forensic odontology
- The application of dentistry to legal questions, including reliable dental identification of the dead and the far weaker comparison of bite marks on skin.
Bias, Laboratory Failure, and Wrongful Conviction
- Explain contextual bias, bias cascade, and bias snowball, and cite the experiments that demonstrated them in practicing examiners.
- Describe how laboratory failures at Houston and in Massachusetts occurred and why they went undetected for years.
- Interpret the exoneration data on the contribution of forensic evidence to wrongful convictions.
- Evaluate proposed safeguards by how directly each addresses a demonstrated failure mode.
One of the hairs was from a dog
In 1978 a taxi driver named John McCormick was shot dead outside his home in Washington, D.C. The killer had worn a stocking mask, and a hair was recovered from it. An FBI examiner told the jury that the hair matched Santae Tribble's in all microscopic characteristics. Tribble was seventeen. He was convicted and served twenty-eight years.
In 2012 the hairs from that stocking were subjected to mitochondrial DNA testing. Thirteen hairs were tested. None of them came from Tribble. One of them came from a dog.
Sit with that detail for a moment, because it is not merely embarrassing. An examiner working under a comparison microscope, in a national laboratory, looked at a canine hair and told a jury it was consistent with a specific human being. The point is not that the examiner was a fraud. It is that a human being who knows what the case needs, looking at ambiguous material with no defined standard for what counts as agreement, can see it.
Key idea: The dominant failure mode in forensic science is not fabrication; it is a competent person seeing what the situation has prepared them to see, and then saying it too strongly.
Bias, measured in real examiners
The claim that context influences expert judgment is not a philosophical position. It has been tested on working forensic scientists, and the results are uncomfortable.
In 2006, Itiel Dror and colleagues did something clever. They obtained fingerprint pairs that five experienced latent print examiners had themselves identified in past casework, and re-presented those same pairs to those same examiners five years later, this time with contextual information suggesting the prints had been erroneously matched in a high-profile misidentification. Four of the five changed their conclusions. They were disagreeing with themselves, on identical images, because the story around the images had changed.
In 2011, Dror and Greg Hampikian took a DNA mixture from an actual case in which a laboratory had reported that a suspect could not be excluded. They gave the same electropherogram data to seventeen qualified DNA analysts in different laboratories, with no case context at all. One agreed with the original conclusion. Twelve excluded the suspect. Four called it inconclusive. This is DNA, the discipline the 2009 report exempted from its criticism, and the interpretation step turned out to depend heavily on what the analyst had been told.
Three named patterns come out of this literature.
- Contextual bias: task-irrelevant information, such as a confession, a criminal record, or a detective's confidence, shifts an examiner's judgment on an ambiguous comparison.
- Bias cascade: irrelevant information flows from the investigation into the laboratory, often innocently, on the submission form itself.
- Bias snowball: one examiner's conclusion influences another's, which is exactly what happens when a verifier knows what the first examiner decided, as in the Mayfield case.
None of this implies bad faith, and saying so is not a courtesy. Bias of this kind operates below awareness and is not reduced by expertise, motivation, or integrity. An examiner who insists she is immune has stated the strongest available evidence that she is not.
What laboratories have actually done
Beyond individual judgment lies institutional failure, and three episodes should be part of every forensic science student's education.
In 2002 an audit of the Houston Police Department crime laboratory found a DNA section operating with a leaking roof, analysts without adequate training, and results that could not be supported by the underlying data. The section was shut down. Josiah Sutton, convicted at sixteen on DNA testimony from that laboratory, was released in 2003 after retesting excluded him. Thousands of cases were reviewed. The eventual structural response was unusual and instructive: in 2014 Houston moved its forensic services out of the police department entirely into an independent local government corporation, one of the few American jurisdictions to implement the 2009 report's independence recommendation.
In Massachusetts, a chemist named Annie Dookhan at the state's Hinton drug laboratory was found to have reported results for samples she had never tested, a practice known as dry labbing, and to have altered records. She pleaded guilty in 2013. Because a drug identification is the entire case in a possession prosecution, the consequences were enormous: in April 2017 Massachusetts dismissed more than 21,000 convictions in a single stroke. Then it happened again. Sonja Farak, a chemist at a different Massachusetts laboratory, had been consuming laboratory drug standards for years while testing evidence, and thousands more convictions were dismissed in 2018.
The common thread is not that these were unusually corrupt places. It is that nothing in the ordinary operation of a crime laboratory was designed to catch them. Proficiency tests were announced and easy. Verification was not blind. Nobody outside the institution had the authority or the data to audit it. Dookhan's productivity, which should have been an alarm, was treated as excellence.
Why this matters: Each of these failures was discovered by accident or by an outsider, which tells you the system had no working mechanism for finding them.
The numbers on wrongful conviction
Two organizations keep the data, and they count different things, so the figures differ for good reasons.
The Innocence Project tracks convictions overturned specifically by DNA testing, more than 375 in the United States. Because DNA cases are concentrated in sexual assault and homicide prosecutions where biological evidence was preserved, this is not a random sample of wrongful convictions; it is the subset where a decisive test happened to be possible. Within that subset, the misapplication of forensic science is a contributing factor in roughly half of the cases.
The National Registry of Exonerations, maintained at the University of Michigan and partner institutions, counts all known exonerations by any means, several thousand of them. In that broader population, false or misleading forensic evidence appears as a contributing factor in roughly a quarter of cases, a lower share because the registry includes many exonerations, such as those from police misconduct in drug cases, in which forensic evidence played no part.
Two things about these numbers matter more than the numbers themselves. First, forensic error almost never acts alone. It typically appears alongside eyewitness misidentification, a false confession, or incentivized informant testimony, and the combination is what convicts. Second, the category misapplication of forensic science covers two quite different failures: a method that does not work, such as bite mark comparison, and an overstatement of a method that does, such as an examiner testifying to certainty about a valid technique with a real error rate. The second is more common and gets less attention.
Fixes, ranked by whether they address a demonstrated failure
Every one of these has been proposed, most have been implemented somewhere, and none is universal.
| Safeguard | Failure it addresses | State of adoption |
|---|---|---|
| Blind verification | Bias snowball between examiners | Recommended widely, practiced inconsistently |
| Linear sequential unmasking | Contextual bias during analysis | Documented in guidance, used in a minority of laboratories |
| Case manager filtering submissions | Bias cascade from investigation into the laboratory | Rare, and resource-intensive |
| Blind proficiency testing | Undetected incompetence and dry labbing | Very rare; most testing is announced and easy |
| Laboratory independence from police | Institutional pressure and unaudited operation | A handful of jurisdictions, notably Houston |
| Error rate disclosure and uniform language | Overstated testimony about valid methods | Adopted federally, uneven in the states |
| Defense access to data, software, and experts | One-sided presentation of technical evidence | Uneven and largely a function of funding |
| Changed-science relief statutes | Convictions resting on since-discredited methods | Enacted in Texas, California, and a few others |
Two of those deserve a closer look. Linear sequential unmasking is the discipline of documenting your analysis of the questioned item before you look at the reference, and receiving case information in a controlled order, so that what you see in a smudge is not shaped by what you already know about the suspect. It is cheap, it directly targets the mechanism the experiments demonstrated, and it is still not standard.
Changed-science statutes address a problem the ordinary appellate system handles badly. Once a conviction is final, new evidence of innocence faces high procedural barriers, and a scientific consensus that has shifted is not new evidence in the traditional sense. Texas enacted a provision in 2013 allowing a convicted person to seek relief where relevant scientific evidence was not available at trial or where the field's understanding has since changed, and it has been used in arson and bite mark cases. Most states have nothing comparable.
How to read a forensic report
This is the last lesson of the course, so here is what the whole thing was for. When you encounter a forensic conclusion, in a news story, a case file, or a courtroom, ask six questions.
- What exactly is being claimed: a chemical identity, a class association, or a specific source?
- Is the method objective or does it rest on an examiner's judgment about sufficient agreement?
- What is the published error rate for this task, and where was it measured?
- What did the examiner know about the case before reaching a conclusion?
- Was the conclusion verified, and did the verifier know the first answer?
- Does the wording of the conclusion match what the method can support, or has it been inflated a step?
Those six questions are the practical residue of everything from Locard's attic to the NIST foundation reviews. They are also, not coincidentally, what a good cross-examination does. Forensic science at its best is one of the most powerful truth-finding instruments any legal system has ever had. It earns that description exactly to the extent that it can answer these six questions, and it is worth insisting that it does.
Common misconceptions
- Bias means the examiner was dishonest. The documented effects operate below awareness in competent, well-intentioned experts, which is why procedural fixes rather than exhortations are the remedy.
- DNA interpretation is immune to context. Seventeen analysts given the same mixture without context split sharply, and only one matched the original laboratory's conclusion.
- Proficiency testing already catches bad analysts. Most proficiency tests are announced, easy, and known to be tests, which is why dry labbing went undetected for years.
- Wrongful convictions are caused by forensic error alone. It usually appears alongside misidentification, false confession, or informant testimony, and the combination convicts.
- Once science is discredited, the convictions are automatically reversed. Finality doctrines make that difficult, which is why a few states enacted statutes specifically for changed science.
What you now know
- Context changes expert conclusions: four of five examiners reversed themselves on their own prior identifications when the surrounding story changed.
- Bias cascade and bias snowball describe how irrelevant information enters a laboratory and spreads between examiners.
- Houston, Dookhan, and Farak show institutional failure discovered late because nothing was designed to catch it.
- Forensic error contributes to roughly half of DNA exonerations and about a quarter of exonerations overall, and rarely acts alone.
- The safeguards that target demonstrated mechanisms, blind verification, sequential unmasking, blind proficiency testing, independence, remain partially adopted.
- Six questions, about the claim, the method, the error rate, the context, the verification, and the wording, will let you read almost any forensic conclusion critically.
Sources
- Innocence Project. (n.d.). DNA exonerations in the United States. innocenceproject.org
- National Registry of Exonerations. (n.d.). About the registry. University of Michigan Law School. law.umich.edu
- National Research Council. (2009). Strengthening forensic science in the United States: A path forward. National Academies Press. nap.nationalacademies.org
- Dror, I. E., Charlton, D., and Peron, A. E. (2006). Contextual information renders experts vulnerable to making erroneous identifications. Forensic Science International, 156(1), 74-78.
- Garrett, B. L. (2021). Autopsy of a Crime Lab: Exposing the Flaws in Forensics. University of California Press.
- Key terms
- Contextual bias
- The influence of task-irrelevant case information, such as a confession or a criminal record, on an examiner's judgment about an ambiguous comparison.
- Bias cascade
- The flow of irrelevant investigative information into the laboratory, often on the evidence submission form itself.
- Bias snowball
- The influence of one examiner's conclusion on a second examiner, which is what unblinded verification produces.
- Linear sequential unmasking
- A procedure requiring the examiner to document analysis of the questioned item before seeing the reference, with case information released in a controlled order.
- Blind proficiency testing
- Testing in which the examiner does not know a sample is a test, the only form that measures performance under routine conditions.
- Dry labbing
- Reporting analytical results for samples that were never actually tested, the misconduct at the center of the Massachusetts drug laboratory dismissals.
- Changed-science statute
- A law permitting post-conviction relief where the scientific understanding underlying a conviction has since changed, enacted in Texas, California, and a few other states.
- Overstated testimony
- A conclusion phrased more strongly than the underlying method supports, distinct from and more common than use of an invalid method.