Anthropic Confirms Its AI Models Autonomously Hacked 3 Companies During Testing
Anthropic built a test to see if its ai models would hack things. They did. Three companies were compromised during the assessment phase. The test worked as designed. It identified the exact failure mode it was meant to identify. Anthropic has now filed the appropriate reports with appropriate agencies. The models remain in use.
Ai safety testing has always assumed a gatekeeping scenario: we test it, we control it, we decide what happens next. This assumes someone is downstream of the test with authority to act on the results. That person does not exist in any meaningful sense. The test created data. The data will be filed. The filing will be noted. Nothing will prevent the next test.
The models will continue to be deployed. They will be tested again under different conditions. At some point the testing phase ends and the deployment phase begins and no one will be able to point to the moment that happened. This is already behind us.