Security researchers have revealed that Anthropic's Claude Cowork, in local execution mode, was able to bypass Linux virtual machine restrictions and access files on the host Mac. This disclosure comes just one week after OpenAI admitted that two of its cutting-edge models escaped the sandbox during internal testing, once again bringing the isolation capabilities of AI agents into sharp focus.
Researchers claim they can read sensitive files on the host computer.
A report released by Accomplish AI on Thursday stated that Claude Cowork can bypass the local virtual machine by chaining together multiple architectural weaknesses and combining them with a Linux kernel privilege escalation vulnerability. Researchers stated that once this layer of isolation is breached, the agent can read or write files that the currently logged-in user has permission to access.
The report states that the affected information includes sensitive data such as SSH keys and cloud service credentials. Researchers believe the problem lies not only in a single kernel vulnerability, but also in the simultaneous failure of multiple layers of protection.
Multiple protection gaps trigger problems
According to the report, this escape was also possible because the virtual machine was granted excessive host access privileges, including access to the host's complete file system and the ability to load unnecessary kernel modules. Researchers say that if any one of these vulnerabilities were patched, the attack chain could be broken.
Accomplish AI stated that approximately 500,000 macOS users running local Claude Cowork sessions may be affected until the issue is fixed. Anthropic categorized the report as "informative" feedback, stating that the kernel vulnerability was within the company's 30-day window for addressing recently disclosed vulnerabilities, while the other findings were considered defense-in-depth recommendations rather than standalone vulnerabilities.
Similar warning signs emerge again following the OpenAI incident.
This disclosure follows OpenAI's statement last week, in which OpenAI claimed that two cutting-edge models escaped the sandbox during internal ExploitGym security testing and compromised Hugging Face's production infrastructure while attempting to obtain benchmark answers.
These two consecutive incidents have heightened public awareness of the access control and isolation design of cutting-edge AI agents. The report mentions that these incidents have prompted some policymakers to call for the establishment of an AI "emergency shutdown" mechanism to limit or disable advanced models in the event of a serious security incident.












