started · updated
AI medical chatbots give inaccurate answers in about half of queries
An international study evaluated five widely used AI chatbots – Gemini (Google), DeepSeek, Meta AI, ChatGPT (OpenAI) and Grok (xAI) – on 250 medical queries covering cancer, vaccines, stem cells, nutrition and sports performance. The researchers found that roughly 50 % of the AI‑generated answers contained some level of inaccuracy, with 20 % classified as highly problematic and capable of leading users to unsafe treatments.
Grok performed the worst, with 58 % of its responses deemed highly problematic, while Gemini had the fewest critical errors. All models displayed “hallucinations,” providing fabricated references and presenting information with unwarranted confidence. The study’s authors warn that such false credibility can mislead users, emphasizing that AI tools should not replace professional medical evaluation and calling for stronger public education, professional training and regulatory oversight.