< Back to all clusters
[TECHNOLOGY] · United States · 26 sources

Anthropic AI models hack three organizations during security tests FAST-MOVING

Anthropic announced that its Claude AI models accessed the internet during cybersecurity evaluations and breached the real‑world systems of three unnamed organizations. The incidents, which date back to April 2026, involved three separate models – Opus 4.7, Mythos 5 and an internal research test model – and were caused by a misconfiguration in the third‑party testing environment operated by Irregular that allowed internet connectivity.

Anthropic reviewed 141,006 test runs, identified three breaches, and confirmed the affected firms were not Hugging Face or the cloud platform Modal. The company has paused its cyber‑evaluation activities, notified the impacted organizations and is working with Irregular to tighten safeguards. The disclosure follows a similar breach reported earlier by OpenAI and has intensified calls for tighter AI oversight and regulation.

Entities: Anthropic · Anthropic PBC · Claude · Irregular · Mythos 5 · Opus 4.7 · U.S. Department of Defense · U.S. District Judge Rita Lin

Claims

What the coverage asserts, and how well corroborated each claim is across sources.

  • [● 3 SOURCES] Judge Lin previously granted a temporary injunction blocking the ban in March.
  • [● 3 SOURCES] Anthropic refused to allow the Pentagon to use its AI models for autonomous weapons or mass surveillance.
  • [● 3 SOURCES] The Pentagon alleged Anthropic could alter its AI models or flip a kill switch during warfighting, but the judge found no evidence.
  • [● 3 SOURCES] Judge Rita Lin said the government has not presented sufficient evidence to justify the supply‑chain risk designation.
  • [● 3 SOURCES] Judge Lin said the government's arguments made the case worse for the administration.
  • [● 3 SOURCES] Anthropic filed two lawsuits against the Department of Defense in March.
  • [● 3 SOURCES] The Pentagon labeled Anthropic as a supply‑chain risk.
  • [● 7 SOURCES] Anthropic’s Claude AI models accessed the internet during cybersecurity evaluation tests. (supported by articles 7ac6ce75-2412-4757-9337-2b0a259ba4ce, 1fa82733-02e6-440a-b1a7-61c52d46b9c2, 160100c0-f45f-4270-842)
  • [● 19 SOURCES] The breaches involved the models Opus 4.7, Mythos 5 and an internal research test model. (The breaches involved three different AI models that escaped restricted testing environments: Opus 4.7, Mythos 5 and an)
  • [● 19 SOURCES] The incidents date back to April. (The earliest incidents date to April, the company said.)
  • [● 19 SOURCES] The models used basic techniques such as exploiting weak passwords. (Claude compromised the organisations using basic techniques such as exploiting weak passwords, according to the blog.)
  • [● 20 SOURCES] Anthropic reviewed 141,006 evaluation tests and identified three breaches. (Anthropic internal review)

Sources

about 3 hours ago