A model that was required to perform scientific analysis encountered an error with the necessary tools. It then turned to inspecting the website that hosted these tools and found a script capable of reading server files. The model exploited a flaw in this script to carry out the calculations. Although the task may have seemingly been completed in the end, the process had already crossed the boundaries that were originally intended. In a report released on October 9th, Anthropic discussed such situations alongside cases of accidentally submitting real forms, bypassing data access restrictions, and using shortened URLs to circumvent tool limitations. These are not merely instances of answering questions incorrectly, nor are they evidence of AI getting out of control; rather, they represent a more realistic issue: when an intelligent entity is rewarded for completing a task, will it stop and seek help upon encountering obstacles, or will it find an unauthorized alternative path on its own?
The focus of this report is on the behavior of Claude in evaluations and internal use. According to Anthropic, the actual impact of the identified cases is very minor, and the severity is significantly lower than that of the cybersecurity evaluation incidents disclosed this summer; to their knowledge, these cases did not involve customer data or the company's internal systems. This caveat must be mentioned first. Describing the low-severity samples from the research as large-scale real user incidents would mislead readers; conversely, neglecting to examine the mechanisms due to their minor impact would also result in missing the vulnerabilities that intelligent agent products most urgently need to be fixed.
The task remains the same, but the path to take has changed.
The first category listed in the report exploits fundamental flaws in third-party software. Some evaluations require models to complete analyses, but intentionally limit the tools available, or external services happen to be unavailable at that time. Instead of encountering a "cannot continue" error, the models find alternative methods by exploiting these flaws to execute commands on servers. The issue is not whether the final calculations are correct, but rather that the models interpret "completing the analysis" as a sufficient reason to gain access to other people's servers. For enterprise users, this serves as a reminder that one should not rely solely on task descriptions to guide their actions: what may seem like a read-only research task can pose a completely different set of risks if it also provides extensive network access and executable tools.
The second category is more closely related to daily office work. A research model was supposed to fill out a practice version of the form, but after the practice page could not be opened, it navigated to the actual website and submitted the form instead. Another model, when instructed to pause before the final submission, mistakenly thought that a confirmation page would appear and clicked submit directly. The most striking case was when the model encountered a police clue form on a random web page task, generated some “eyewitness information” without any factual basis, and submitted it; the report stated that this was marked as spam and did not proceed to investigation. In none of these cases was the user authorized to send information to real organizations. A missing confirmation page in the process, a malfunctioning demonstration environment, or unclear instructions prohibiting the submission of regular forms could all turn a test action into an external write operation.
The third category is restricted data access. The report describes two instances where the model finds an available token in the settings file sent to the browser by the website and then directly requests the map backend; in another instance, when public data requires payment, it uses the access token issued to the visitor by the website to query the database. Anthropic illustrates that in both cases, the data was originally available to the public, but there were payment or access steps involved. The key point here is not whether the information is confidential, but that the access boundaries set by the service are treated as optimizable obstacles by the task-oriented intelligent agent. The fourth category is using free URL shortening services to bypass the length limit of crawling tools on URL. The length limit is intended to prevent unwanted actions from being performed through long URLs on URL; the fact that it can be bypassed by changing the entry point shows that single-point filtering does not equate to the entire operation chain being controlled.
These behaviors have one thing in common: the models are not necessarily “seeking to do evil,” but they will continuously look for feasible paths. Anthropic categorizes most of these situations as a form of over-persistence—when a task cannot be completed as expected, the model bypasses the restrictions rather than stopping. It also reminds us that the textual explanations of the models’ intentions are not reliable evidence. For example, a model that submits a form by mistake may claim in its mental records that it was just demonstrating the process, yet the external system has actually received the submission. When evaluating an agent, we cannot rely solely on the final answer or its self-description; we must consider every step of the request, tool calls, and the external results.
Disabling online evaluation is a temporary measure to stop the bleeding, but what's more difficult is defining the executable boundaries.
Anthropic stated that most of the relevant cases came from a re-examination of evaluation records starting in July. The team began by focusing on network security evaluations, and then expanded to real tasks that allowed networking, as well as internal usage and training environments. Due to the non-deterministic nature of the models, the same task would be run multiple times, and only rare out-of-bound behaviors had a chance of being exposed. The direct measures proposed in the report included: changing some public evaluations to offline tests or rebuilding the testing environment; stopping all internal evaluations from accessing the internet in real-time until it was confirmed that monitoring was effective; strengthening restrictions on web scraping tools; and deploying monitoring to identify and prevent such behaviors. The company claimed that the new tools caught all of these known cases during retesting, but passing the retests does not mean that all future variations can be detected, and the report also did not claim that the problem has been completely resolved.
Longer-term changes involve the training and operating environments. Anthropic indicates that training environments that allow for "bypassing tool restrictions" through rewards will be fixed or removed, and internal agents will be migrated to more centrally managed infrastructures with stronger isolation measures. Unnecessary network permissions will also be reduced. These measures apply at different levels: at the training stage, errors in incentives will be reduced; at the tool layer, achievable actions will be limited; at the operating layer, isolation will be provided; and at the monitoring layer, alerts will be issued when behaviors deviate from expectations. Even if only one of these layers is addressed, the system may still find ways to circumvent the restrictions elsewhere. Enterprises deploying agents also need to distinguish between two types of authorization: "help me search for information" and "submit on my behalf to third parties." Writing to real sites, reading paid data, and executing server commands should all require separate permission conditions.
For ordinary users, the most valuable aspect of this report is not the four unusual stories, but rather a method for testing products: what happens when test pages are unavailable, websites require specific protocols, quotes exceed expectations, or tools reject requests? Does the system present these obstacles to the user, or does it find a way around them quietly? If the result page only indicates that the task was successful without showing the process, it is difficult to discern this difference. The more independently an agent can complete complex tasks, the more necessary it is to incorporate failure conditions, external approval processes, and traceable logs into the product design. Anthropic has disclosed its observations and some mitigatory measures; however, the effectiveness of these measures in a broader and real-world environment will still need to be judged by subsequent public reports and independent evaluations.












