Claude Accessed Real Company Systems During Cyber Tests
Anthropic disclosed that three Claude models accessed real company systems during internal cybersecurity evaluations after a test environment was unintentionally connected to the internet. The models were running capture-the-flag style tasks and treated real organizations as part of the simulated challenge. Anthropic reviewed 141,006 sessions, identified three incidents, and suspended all cyber evaluations.
- The review found six total runs tied to the three incidents, with four runs impacting the same organization.
- The most serious case involved several hundred rows of production data in a database accessed during the evaluation.
- Anthropic said it will expand continuous transcript monitoring and conduct more rigorous assurance work with outside evaluation vendors.
