Police facial recognition can be highly accurate or badly wrong, depending on which algorithm is used and whose face it is analyzing. Independent federal testing has found error rates ranging from a fraction of a percent to double digits across different systems, with the largest gaps falling along racial and demographic lines, according to the National Institute of Standards and Technology.
The technology is now a routine investigative tool rather than an experimental one. Seven federal law enforcement agencies within the Departments of Homeland Security and Justice ran roughly 60,000 facial recognition searches between October 2019 and March 2022, per a September 2023 report from the Government Accountability Office. State and local departments add an unknown number of additional searches on top of that, using a patchwork of commercial vendors with no single national standard.
How does police facial recognition actually work?
Facial recognition systems used in policing generally perform one of two tasks. A one-to-one search verifies that two images show the same person, the kind of check used to confirm an ID against a photo already on file. A one-to-many search takes an unknown face, often pulled from surveillance video or a body camera, and compares it against a gallery of millions of enrolled photos to produce a ranked list of possible matches.
That second use is the one most relevant to criminal investigations, and the one NIST has tested most closely. The output is not a confirmed identity. It is a list of candidates a human investigator is expected to review and corroborate through other evidence before taking any action, according to NIST's program materials.
How accurate are these systems, according to NIST?
NIST runs the Face Recognition Vendor Test, an ongoing evaluation open to any algorithm developer worldwide, measuring both one-to-one and one-to-many accuracy along with performance on masked faces, image quality and morph detection. The program does not certify a single accuracy number for facial recognition as a technology; instead it reports that error rates vary widely by developer, with the best-performing algorithms far more accurate than the weakest ones submitted for testing, according to NIST.
That range is the central fact investigators and the public are asked to understand: two systems marketed for the same purpose can perform very differently, and a department's real-world accuracy depends on which specific algorithm and threshold settings it has deployed, not on facial recognition as a generic category.
Why do error rates differ by race and gender?
A NIST study published in December 2019 evaluated 189 algorithms from 99 developers using 18.27 million photographs to test for demographic effects. It found higher false-positive rates for Asian and African American faces relative to Caucasian faces in one-to-one matching for the majority of algorithms tested, with the size of that gap ranging from roughly 10 times to 100 times higher depending on the specific algorithm, according to NIST. Among U.S.-developed algorithms, Native American faces showed the highest false-positive rates of any group tested.
The same study found that algorithms developed in Asian countries did not show that same disparity between Asian and Caucasian faces, a pattern NIST said pointed toward the training data used to build a given algorithm as a likely factor. In one-to-many searches, the type used to generate investigative leads, NIST found elevated false-positive rates for African American women specifically — a finding the agency noted carries particular weight because a false positive in that search type puts an innocent person's photo on a list flagged for further scrutiny.
NIST also found that accuracy and demographic fairness are not necessarily a trade-off: some of the algorithms with the smallest demographic gaps were also among the most accurate overall, according to the study.
How much oversight exists over how police use it?
The September 2023 GAO report examined how the seven federal agencies that ran facial recognition searches governed that use, rather than testing the technology's accuracy directly. It found that all seven initially used the tool without requiring staff to complete any facial recognition training, and that by April 2023 only two of the seven had made training mandatory. At the FBI, GAO found just 10 of 196 staff who had accessed the service had completed the training available to them.
GAO also found that only three of the seven agencies had policies specifically addressing civil rights and civil liberties protections tied to facial recognition use; four agencies, three at the Justice Department and one at Homeland Security, had no facial recognition-specific guidance at all as of the report's release, though DHS had committed to finalizing department-wide policy by the end of 2023. GAO issued ten recommendations, including that the FBI make training mandatory and that agencies establish ways to monitor whether staff actually complete it; the agencies concurred with all ten, according to the report.
What a match does, and does not, establish
- A facial recognition result is a lead, not an identification. NIST's own materials describe one-to-many output as a ranked candidate list requiring further investigation, not a determination of who committed a crime.
- Accuracy is algorithm-specific. A department's real accuracy depends on which vendor's system and settings it uses, not on facial recognition as a single technology, per NIST's testing program.
- Error rates are not evenly distributed. NIST's demographic study found some groups face meaningfully higher false-positive rates than others, particularly in searches meant to generate investigative leads.
- Federal oversight has been uneven. GAO found training and civil-liberties policy gaps at multiple federal agencies as recently as 2023, with some since addressed and others still in progress.
Because a facial recognition match is an investigative starting point rather than proof, any resulting charge still carries the same presumption of innocence as one built from any other lead — the technology narrows a list of possibilities; it does not establish guilt.
Frequently asked questions
For a related software perspective, read How accurate is facial recognition software, according to the studies.
