Skip to content
Saturday, August 22, 2026
Daily Detective NewsCrime & Policing / Legal Affairs
tools

How accurate police facial recognition is, according to the federal tests

Federal testing finds leading facial recognition algorithms identify reference photos with near-perfect rates, but error rates rise sharply in the real-world conditions — and demographic groups — that police work involves.

Blurry surveillance camera above an empty doorway in low light

Facial recognition accuracy is measured by the National Institute of Standards and Technology's FRVT program, and its 2019-2023 reports found the most accurate algorithms identified verification-style face matches with error rates below 0.2 percent in high-quality photos. The same reports found error rates rising by a factor of 10 to 100 times in harder conditions: off-angle faces, poor lighting, low-resolution surveillance frames, and, consistently, certain demographic groups. Those two facts together define the policing debate. Daily Detective News publishes information, not legal advice.

The technology now touches thousands of US police investigations. Departments submit still images — often from surveillance video — to commercial systems that return candidate matches from arrest-photo databases; detectives then treat the output as an investigative lead. The gap between the laboratory number and the street number is the story of how these tools get used, and misused.

What does NIST actually test?

NIST's Face Recognition Vendor Test runs algorithms from volunteer developers against standardized photo collections and publishes the results openly. Its December 2019 demographic report, the most cited in policing debates, tested 189 algorithms and found wide spread: the top-tier products showed small differences in false-positive rates across demographic groups, while many others produced false positives 10 to 100 times more often for some groups — consistently for women, older and younger subjects, and, for many algorithms, Asian and African American faces, per the report's text.

Two limits matter for policing. NIST tests one-to-one verification — is this photo the same person as that photo — while police typically run one-to-many searches, where false-positive odds multiply across a database of thousands. And NIST does not test the commercial databases themselves; an agency's error rate depends on both the algorithm and the photo gallery it searches.

What do independent studies say about police use?

Named studies of deployed systems have found weaker performance than vendor materials claim. A 2019 Georgetown Law Center on Privacy and Technology report, The Perpetual Line-Up, documented that roughly one in two American adults — then about 117 million people — was enrolled in a law-enforcement face-search network, and cataloged documented misidentifications that led to wrongful arrests. Subsequent published case reporting by the New York Times and others documented at least seven wrongful arrests through 2023 in which facial recognition had generated the lead; in several, the arrests were made largely on the match alone.

Vendor systems claim high accuracy; those claims are vendor claims. The wrongful-arrest count comes from public reporting and court records, not from any systematic national database — no such database exists, because most departments do not publish how often they run searches or what happens when a match is wrong. When only the anecdote is available, that limitation is the honest finding.

Why do false positives matter more than false negatives?

Because of who bears each error. A false negative — no match found — costs the police a lead. A false positive — the wrong person returned — can cost an innocent person an arrest, and the documented wrongful-arrest cases share a pattern: a detective treats the machine's candidate as a suspect, builds a photo array around it, and a witness confirms the machine's choice. The NIST demographic findings mean this cost is not distributed evenly, which is the core of the civil-rights objections filed in federal lawsuits and state legislation alike.

The remedies proposed follow the error asymmetry. Several states and cities now require corroborating evidence beyond the match before an arrest; Maine and Illinois have restricted government use outright; department policies typically require that output be labeled an investigative lead, not an identification.

What should readers take from the numbers?

The laboratory accuracy is real and improving — NIST's reports show steady gains cohort over cohort. So is the gap: the conditions of police work — grainy video, off-angle faces, one-to-many searches — are precisely the conditions where error rates climb, and the demographic differentials documented in the 2019 report persisted across the follow-up reports through 2023. What the record establishes is a tool that works well enough to be dangerous when the label falls off: accurate enough to feel conclusive, imperfect enough to implicate the wrong person. What remains unknown is the true scale of misidentification, because almost no department publishes its search logs. Until one does, the documented cases are a floor, not a rate.

Sources

  1. NIST — FRVT Part 3: Demographic Effects, December 2019, and follow-up reports to 2023