Facial recognition is now a fixture of modern life, powering everything from national border security to the simple convenience of unlocking a smartphone. However, these advancements bring significant risks to privacy, equity and civil rights.
While AI can now match faces as accurately as humans, the specific logic these algorithms use to calculate similarity remains a “black box” (closed source). To bridge this gap, new research from the University of Notre Dame compared commercial and open-source AI against the judgment of 4,000 human participants.
While humans and AI generally agree on facial similarity, three factors heavily influence accuracy: the race of the participant, the race of the face being viewed and the individual’s natural recognition skill, according to, the Joe and Jane Giovanini Professor of IT, Analytics and Operations at Notre Dame’s Mendoza College of Business. Abbasi’s research, “,” is forthcoming in the Journal of Applied Research in Memory and Cognition.
“Top-tier AI is now as good as the best human experts," noted Abbasi, the director of Notre Dame’s , deputy director of the and co-director of the . “If we use AI as the benchmark for ‘good’ recognition, it suggests a large portion of the population is actually not that great at recognizing faces.”
This overlap is critical because law enforcement often uses a “human-in-the-loop” system, where an algorithm flags matches and a person makes the final call.
“Humans and AI usually see face similarity the same way,” Abbasi said. “However, people who are naturally better at recognizing faces match the AI’s judgments much more closely — at least 15 percent more often, depending on the model.”
The research also highlighted systemic flaws. AI models often disagree with one another and show decreased accuracy when analyzing races underrepresented in their training data. To test this, researchers used 329 standardized cross-race photos in controlled trials. “Unfortunately, these systemic flaws are grounded in 20-plus years of research in computer vision, making them difficult to easily overcome,” Abbasi said.
Abbasi argues that while we rightfully scrutinize AI for bias, we often overlook human fallibility. “We should be mindful of human recognition limits, especially in high-stakes scenarios like eyewitness testimony,” he said. He noted that while privacy and fairness concerns have led several states to ban automated facial recognition, human error remains a significant, yet less regulated, variable.
The study also underscores the legal dangers of black box systems. It cites State v. Arteaga, a landmark New Jersey case where a court required a police department to disclose an AI’s design so a defendant could properly challenge the evidence.
Ultimately, the research suggests that as we move toward an AI-integrated future, we must evaluate the reliability of both the machine and the human making the final decision.
Abbasi’s co-authors are David Dobolyi from the University of Colorado Boulder, Rachel Leigh Greenspan from the University of Central Florida, Jesse Grabman from New Mexico State University and Chad Dodson from the University of Virginia.
Contact: Ahmed Abbasi, 574-631-5212, aabbasi@nd.edu
