Fingerprint examiners who compare a crime-scene print to a suspect's exemplar make the wrong match, on average, in roughly one out of every 1,000 comparisons involving prints that do not actually correspond — and they wrongly say two matching prints don't correspond in about 7.5 percent of comparisons, according to a 2011 double-blind study run by the nonprofit research group Noblis in partnership with the FBI.
That study, still the most cited "black box" test of latent-print examiners, is the closest thing the field has to a measured accuracy rate. It didn't test a machine or an algorithm; it tested more than 169 working examiners against real casework-style print pairs, without telling them which pairs were designed to match, according to the National Institute of Justice, the U.S. Department of Justice's research arm.
How do examiners actually compare two fingerprints?
Fingerprint examiners in the United States generally follow a four-step method known as ACE-V, an approach documented across forensic-science literature and used to structure how an examiner reasons through a comparison rather than simply eyeballing two prints side by side. The steps are sequential, and an examiner is expected to complete each before moving to the next.
- Analysis: The examiner studies the latent print recovered from a scene — a partial, often smudged impression — and decides whether it has enough clear detail (ridge endings, bifurcations, and other minutiae) to be usable at all.
- Comparison: The examiner places the latent print side by side with a known exemplar print, usually taken from a suspect's inked or digital fingerprint card, and looks for points of similarity or difference in ridge patterns.
- Evaluation: The examiner forms a conclusion — identification (the prints came from the same source), exclusion (they did not), or inconclusive (the print lacked enough detail to decide).
- Verification: A second examiner independently repeats the analysis, ideally without seeing the first examiner's conclusion, and either confirms or disputes it.
The verification step matters because it is where the 2011 study found built-in error checking actually working: none of the false identifications made by one examiner were repeated by a second examiner assigned to verify the same pair, according to the National Institute of Standards and Technology's published account of the research. That doesn't mean verification catches every mistake — the study only tested a portion of pairs through a second examiner — but it is the mechanism the field relies on to reduce the false-positive rate below what any single examiner achieves alone.
How often do examiners get it wrong, per the studies?
The 2011 Noblis-FBI study — often called the "FBI black box study" because it treated examiners as a black box and measured only their outputs — had participants each compare roughly 100 latent-to-exemplar print pairs drawn from a pool of 744 pairs, producing 17,121 individual decisions in total, the National Institute of Justice reported. Six false positives turned up among 4,083 comparisons of non-matching pairs, a rate of about 0.1 percent, according to NIST's summary of the underlying research.
False negatives were far more common: examiners wrongly excluded a true match in 7.5 percent of comparisons of pairs that actually corresponded, and 85 percent of participating examiners made at least one such error somewhere in the test, per the same NIST-published findings. In practical terms, the research indicates the method is considerably more likely to let a real match go unrecognized than to wrongly link an innocent person's print to a crime scene — though a 0.1 percent false-positive rate, applied across the volume of comparisons a busy crime lab performs, is not zero.
Why would a real match get missed more often than a false one gets made?
Latent prints recovered from crime scenes are frequently partial, smudged, distorted by the surface they were left on, or overlapped with other prints, which limits how much comparable ridge detail an examiner has to work with. Examiners trained to avoid false identifications — the error with the more serious consequence for an accused person — tend to resolve marginal or low-quality comparisons toward "exclusion" or "inconclusive" rather than "identification" when the evidence is ambiguous. The 2011 study's design, evaluated independently through both NIJ's retrospective account and NIST's technical summary, supports that reading: the method is, in the researchers' framing, "tilted toward avoiding false incriminations," which produces fewer false positives at the cost of more missed matches.
Have courts treated fingerprint identification as settled science?
Not entirely, and the debate predates the 2011 study. A 2009 report from the National Academy of Sciences, examined by lawmakers at a Senate Judiciary Committee hearing that September, found that DNA analysis was the only forensic technique that had been scientifically validated to reliably connect evidence to a specific person — and that other pattern-matching disciplines, including fingerprint comparison, rested on far less rigorous scientific footing, according to NPR's coverage of that hearing. The same reporting noted a specific inconsistency the academy flagged: one analyst might declare a match after finding six points of similarity between two prints, while another might require fourteen, because the field lacked uniform national standards for examiners or labs.
Then-Senator Al Franken called the report's findings "damning" and "terrifying" during that hearing, while Judiciary Committee chairman Patrick Leahy said the criminal-justice system "has to be based on facts," NPR reported. The hearing did not resolve the underlying scientific questions, but it pushed the issue toward the kind of controlled accuracy testing that produced the 2011 black-box study two years later — research that gave the field its first widely accepted, peer-reviewed error rates rather than an assumption of infallibility.
For anyone reading about a case where fingerprint evidence is central, the studies point to a specific, bounded claim rather than a certainty: identification decisions carry a measured, low false-positive rate and a materially higher false-negative rate, both established through blind testing rather than examiner self-assessment. That is different from either "fingerprint matching is junk science" or "fingerprint matches are infallible" — both overstate what the research actually shows. Fingerprint comparison remains information about probability and process, not proof on its own, and questions about how it was applied in a specific case are ultimately for the court record and qualified counsel, not a general explainer, to resolve.
For a related software perspective, read How accurate is facial recognition software, according to the studies.
