< Back to all clusters
[HEALTH] · United States · 3 sources

OpenAI GPT-5.6 beats physicians in health assessments, study shows doctors over‑trust AI

OpenAI released the GPT‑5.6 family on July 9, 2026. In a blinded comparison on the HealthBench Professional benchmark, the high‑capability variant GPT‑5.6 Sol achieved a score of 60.5, well above the 43.7 scored by physician‑written responses. The evaluation involved 260 physicians across 60 countries who reviewed more than 700,000 model outputs.

The model has been integrated into Microsoft 365 Copilot, offering a lower‑cost option for AI‑driven clinical documentation and patient communication. In a separate study published in PLOS Digital Health, researchers found that physicians often trust AI treatment recommendations even when they are incorrect. In experiments with 223 doctors, participants consistently rated the AI as reliable and failed to notice that the suggested treatments were ineffective, highlighting challenges in human‑AI collaboration.

Together, these findings illustrate the growing capability of AI in medical contexts while warning that over‑reliance on algorithmic advice could pose risks without proper safeguards.