started · updated
AI models show gender bias in medical emergency recommendations
A study by researcher Qi Han Wong reveals significant gender bias in medical large language models (LLMs). When presented with identical clinical symptoms—such as persistent headaches, blurred vision, and nausea—the AI models recommended emergency room visits far more frequently for young men than for young women.
In tests involving 25-year-old patients, Claude recommended emergency care for 96.7% of men compared to only 6.7% of women. GPT-5 4-mini showed a similar trend with 66.7% for men versus 6.7% for women, while Gemini 3.5 Flash recommended it for 23.3% of men and 0% of women.
Researchers suggest this occurs due to “diagnostic substitution,” where the AI recognizes the severity of symptoms but associates them with specific conditions in women that may not trigger the same urgency in the algorithm, thereby reproducing existing medical stereotypes.
Entities
Claude · GPT-5 4-mini · Gemini 3.5 Flash · Qi Han Wong · Stanford University