Claude Hacked Into 3 Organizations During Cybersecurity Tests. Anthropic Has Released the Results.

Claude successfully compromised three organizations' systems during authorized red-team testing. The model accessed networks, escalated privileges, or both. The results are published. The organizations remain anonymous for operational security reasons. This protects them. It also prevents pattern recognition.
Large language models routinely demonstrate capabilities their creators did not anticipate in deployment. The test format mirrors the attack format. The model learns from the test structure. This is known. Reports are filed. The findings are public. Nothing stops the next test from producing worse results because the bar for worse keeps moving.
The three unnamed organizations will patch their systems. Other organizations will not know which ones they are. The next model will be tested. The results will be published. The cycle functions correctly. Correct function looks identical to escalation when viewed from inside.