A sweeping review of artificial intelligence systems used to diagnose cancer has uncovered a troubling pattern: the tools frequently deliver less accurate results for patients of certain racial and demographic backgrounds. The findings, published in the journal Cell Reports Medicine, stem from an analysis of nearly 29,000 pathology images representing more than 14,400 cancer patients.
Researchers at Harvard University examined four leading AI-powered pathology diagnostic systems and found that their accuracy varied significantly depending on a patient’s age, gender, and race. In a striking twist, the AI models were able to infer these demographic details directly from the tissue slides—a capability that human pathologists cannot replicate. The models exhibited biased outcomes in 29.3 percent of the diagnostic tasks they were assigned, according to the study.
“We found that because AI is so powerful, it can differentiate many obscure biological signals that cannot be detected by standard human evaluation,” said Kun-Hsing Yu, a Harvard researcher and senior author of the study, in a press release. “Reading demographics from a pathology slide is thought of as a ‘mission impossible’ for a human pathologist, so the bias in pathology AI was a surprise to us.”
The bias appears to stem from the AI’s ability to latch onto demographic patterns embedded in the tissue samples. For example, the tools could identify slides taken from Black patients because those samples contained higher counts of abnormal neoplastic cells and lower counts of supportive elements compared to samples from white patients—even though the slides were anonymized. Once the AI recognized a patient’s race, it tended to rely on prior analyses that matched that demographic, which often led to errors when the training data underrepresented certain groups.
One concrete example: the AI models struggled to distinguish subclasses of lung cancer cells in Black patients. The issue was not a lack of lung cancer data overall, but a scarcity of data from Black lung cancer patients specifically. Yu noted that this was unexpected, saying, “Because we would expect pathology evaluation to be objective. When evaluating images, we don’t necessarily need to know a patient’s demographics to make a diagnosis.”
The Harvard team also tested a new training framework called FAIR-Path, which they introduced to the AI tools before analysis. The framework successfully eliminated 88.5 percent of the performance disparities, offering a potential path forward. However, the remaining 11.5 percent of disparities persisted, underscoring the need for continued refinement and mandatory adoption of such measures across the field.
This is not the first time racial bias has been detected in medical AI. In June, researchers found similar issues in large language model psychiatric diagnostic tools, which often proposed inferior treatment plans for Black patients when race was explicitly stated. The new findings add to growing concerns about the fairness and reliability of AI in healthcare, particularly as these systems become more integrated into clinical practice.
While the FAIR-Path framework shows promise, experts caution that until such training methods are universally required, the question of inherent bias in pathology AI will remain unresolved. The study’s authors emphasize the need for greater transparency and diversity in AI training datasets to ensure equitable outcomes for all patients.