AI data integrity crisis hampers aged‑care tools and enterprise programmes
A study published in BMC Medicine on 4 June 2026 examined two widely used public Kaggle datasets for stroke and diabetes risk prediction. Both datasets received zero points on the TRIPOD‑AI checklist for data provenance, with no verifiable information on how the data were collected or their authenticity. Researchers found implausible patterns – such as only 18 distinct glucose values in a 100,000‑patient diabetes set – suggesting the data may be synthetic. These flawed datasets underlie 125 peer‑reviewed clinical prediction models, many of which have been recommended for use in Australian aged‑care settings, risking mis‑diagnosis or unnecessary treatment for vulnerable older residents. The authors call for mandatory reporting of data origins, collection methods and ethical approvals, and urge repositories like Kaggle to enforce stricter standards.
Separately, a Ness Digital Engineering report highlights that weak data foundations are a primary reason enterprise AI programmes fail to scale, despite billions of dollars invested. By mid‑2026 most chief data officers have run pilots and upgraded platforms, yet models often do not move beyond testing due to poor data quality, governance, and reusability. The report outlines five pillars for AI readiness—architecture modernisation, data quality and reliability, governance and ownership, treating data as a product, and security/privacy—and recommends building strong, transparent data ecosystems to achieve trustworthy, scalable AI outcomes.