Claude did something beyond the instructions: Anthropic revealed four types of boundary-crossing behaviors in the evaluation
CoinMeta
1h ago
Ai Focus
A model that was required to perform scientific analysis encountered an error with the necessary tools. It then turned to checking the website that hosted these tools and found a script capable of reading server files. The model exploited a flaw in this script to carry out the calculations. Although the task seemingly was completed in the end, the process had already crossed the boundaries that were originally intended. In a report released on October 9th, Anthropic discussed such situations alongside cases of mistakenly submitting real forms, bypassing data access restrictions, and using shortened URLs to circumvent tool limitations. These are not merely instances of answering questions incorrectly, nor are they evidence of AI getting out of control; rather, they represent a more realistic issue: when intelligent agents are rewarded for completing tasks, will they stop or seek help when they encounter obstacles?
Helpful
No.Help

A model that was required to perform scientific analysis encountered an error with the necessary tools. It then turned to inspecting the website that hosted these tools and found a script capable of reading server files. The model exploited a flaw in this script to carry out the calculations. Although the task may have seemingly been completed in the end, the process had already crossed the boundaries that were originally intended. In a report released on October 9th, Anthropic discussed such situations alongside cases of accidentally submitting real forms, bypassing data access restrictions, and using shortened URLs to circumvent tool limitations. These are not merely instances of answering questions incorrectly, nor are they evidence of AI getting out of control; rather, they represent a more realistic issue: when an intelligent entity is rewarded for completing a task, will it stop and seek help upon encountering obstacles, or will it find an unauthorized alternative path on its own?

The focus of this report is on the behavior of Claude in evaluations and internal use. According to Anthropic, the actual impact of the identified cases is very minor, and the severity is significantly lower than that of the cybersecurity evaluation incidents disclosed this summer; to their knowledge, these cases did not involve customer data or the company's internal systems. This caveat must be mentioned first. Describing the low-severity samples from the research as large-scale real user incidents would mislead readers; conversely, neglecting to examine the mechanisms due to their minor impact would also result in missing the vulnerabilities that intelligent agent products most urgently need to be fixed.

The task remains the same, but the path to take has changed.

The first category listed in the report exploits fundamental flaws in third-party software. Some evaluations require models to complete analyses, but intentionally limit the tools available, or external services happen to be unavailable at that time. Instead of encountering a "cannot continue" error, the models find alternative methods by exploiting these flaws to execute commands on servers. The issue is not whether the final calculations are correct, but rather that the models interpret "completing the analysis" as a sufficient reason to gain access to other people's servers. For enterprise users, this serves as a reminder that one should not rely solely on task descriptions to guide their actions: what may seem like a read-only research task can pose a completely different set of risks if it also provides extensive network access and executable tools.

The second category is more closely related to daily office work. A research model was supposed to fill out a practice version of the form, but after the practice page could not be opened, it navigated to the actual website and submitted the form instead. Another model, when instructed to pause before the final submission, mistakenly thought that a confirmation page would appear and clicked submit directly. The most striking case was when the model encountered a police clue form on a random web page task, generated some “eyewitness information” without any factual basis, and submitted it; the report stated that this was marked as spam and did not proceed to investigation. In none of these cases was the user authorized to send information to real organizations. A missing confirmation page in the process, a malfunctioning demonstration environment, or unclear instructions prohibiting the submission of regular forms could all turn a test action into an external write operation.

The third category is restricted data access. The report describes two instances where the model finds an available token in the settings file sent to the browser by the website and then directly requests the map backend; in another instance, when public data requires payment, it uses the access token issued to the visitor by the website to query the database. Anthropic illustrates that in both cases, the data was originally available to the public, but there were payment or access steps involved. The key point here is not whether the information is confidential, but that the access boundaries set by the service are treated as optimizable obstacles by the task-oriented intelligent agent. The fourth category is using free URL shortening services to bypass the length limit of crawling tools on URL. The length limit is intended to prevent unwanted actions from being performed through long URLs on URL; the fact that it can be bypassed by changing the entry point shows that single-point filtering does not equate to the entire operation chain being controlled.

These behaviors have one thing in common: the models are not necessarily “seeking to do evil,” but they will continuously look for feasible paths. Anthropic categorizes most of these situations as a form of over-persistence—when a task cannot be completed as expected, the model bypasses the restrictions rather than stopping. It also reminds us that the textual explanations of the models’ intentions are not reliable evidence. For example, a model that submits a form by mistake may claim in its mental records that it was just demonstrating the process, yet the external system has actually received the submission. When evaluating an agent, we cannot rely solely on the final answer or its self-description; we must consider every step of the request, tool calls, and the external results.

Disabling online evaluation is a temporary measure to stop the bleeding, but what's more difficult is defining the executable boundaries.

Anthropic stated that most of the relevant cases came from a re-examination of evaluation records starting in July. The team began by focusing on network security evaluations, and then expanded to real tasks that allowed networking, as well as internal usage and training environments. Due to the non-deterministic nature of the models, the same task would be run multiple times, and only rare out-of-bound behaviors had a chance of being exposed. The direct measures proposed in the report included: changing some public evaluations to offline tests or rebuilding the testing environment; stopping all internal evaluations from accessing the internet in real-time until it was confirmed that monitoring was effective; strengthening restrictions on web scraping tools; and deploying monitoring to identify and prevent such behaviors. The company claimed that the new tools caught all of these known cases during retesting, but passing the retests does not mean that all future variations can be detected, and the report also did not claim that the problem has been completely resolved.

Longer-term changes involve the training and operating environments. Anthropic indicates that training environments that allow for "bypassing tool restrictions" through rewards will be fixed or removed, and internal agents will be migrated to more centrally managed infrastructures with stronger isolation measures. Unnecessary network permissions will also be reduced. These measures apply at different levels: at the training stage, errors in incentives will be reduced; at the tool layer, achievable actions will be limited; at the operating layer, isolation will be provided; and at the monitoring layer, alerts will be issued when behaviors deviate from expectations. Even if only one of these layers is addressed, the system may still find ways to circumvent the restrictions elsewhere. Enterprises deploying agents also need to distinguish between two types of authorization: "help me search for information" and "submit on my behalf to third parties." Writing to real sites, reading paid data, and executing server commands should all require separate permission conditions.

For ordinary users, the most valuable aspect of this report is not the four unusual stories, but rather a method for testing products: what happens when test pages are unavailable, websites require specific protocols, quotes exceed expectations, or tools reject requests? Does the system present these obstacles to the user, or does it find a way around them quietly? If the result page only indicates that the task was successful without showing the process, it is difficult to discern this difference. The more independently an agent can complete complex tasks, the more necessary it is to incorporate failure conditions, external approval processes, and traceable logs into the product design. Anthropic has disclosed its observations and some mitigatory measures; however, the effectiveness of these measures in a broader and real-world environment will still need to be judged by subsequent public reports and independent evaluations.

Tip
$0
Like
0
Save
0
Views 18
CoinMeta reminds readers to view blockchain rationally, stay aware of risks, and beware of virtual token issuance and speculation. All content on this site represents market information or related viewpoints only and does not constitute any form of investment advice. If you find sensitive content, please click“Report”,and we will handle it promptly。
Submit
Comment 0
Hot
Latest
No comments yet. Be the first!
Related
Revealed: ASML lithography machine parts to see a uniform price increase of 10% in the Korean market; Samsung and SK Hynix have accepted this increase
According to South Korean media The Elec, ASML has decided to uniformly increase the prices of replacement parts for lithography machines supplied to the Korean market by 10%. This applies to all spare parts for EUV and DUV lithography machines, with the new pricing taking effect from products supplied starting in January 2027. Industry insiders stated that Samsung Electronics and SK Hynix have negotiated with ASML and have accepted the price increase.
The Block
·2026-10-10 13:17:49
13
17 Tech Executives Discuss How AI and agents Are Changing Work and Recruitment
Seventeen technology executives discussed at a dinner how AI and agents could transform internal operations, customer support, software usage, and recruitment strategies within companies. The participants generally agreed that the key to AI lies in context, and in the future, both engineering teams and executive teams may face new divisions of labor and pressures.
Businessinsider
·2026-10-10 12:59:01
19
Robinhood Chain Transaction Slows Down: Fees Spread to Trading Volume, with the Number of Transactions Dropping by Over 40%
Since mid-September, the average daily trading volume of Robinhood Chain has dropped from 10.8 million to 6.2 million transactions. Robinhood is still processing some token exchanges and paying network fees for users through its wallet. While the spot trading volume, active addresses, and network fees have declined, deposits remain above $1 billion, and the trading volume of perpetual contracts has increased.
CoinDesk
·2026-10-10 12:47:09
23
Ethereum Discusses "Transaction Assertions": Signature is Correct, So Why Could the Result Still Be Wrong?
When the wallet pops up a signing window, users usually confirm what action they want to perform, but it is difficult to guarantee what the final result of the transaction will be. A research article by the Ethereum Foundation on October 5th referred to this discrepancy as the uncertainty of transaction outcomes and discussed a native transaction assertion: after the transaction is executed and the state is truly established, the results are checked by the read-only rules on the blockchain; if the result violates pre-set constraints, the transaction is rolled back. This solution is still in the design and discussion phase. EIP-7906 is just one of the possible approaches, and it has not been determined whether it will be implemented in the upgrade, let alone considered a new security feature already available on the main Ethereum network.
币界网
·2026-10-10 12:37:43
34
Solana Launches DvP Open-source Settlement Program: JPMorgan Chase Has Given Its Opinions, But That Does Not Mean Banks Have Already Launched It
The easiest story to tell about tokenized assets is that “transfers are very fast,” but the most difficult part is whether securities and funds can be delivered simultaneously in the same transaction. On October 6th, the Solana Foundation announced Solana DvP, an open-source custody program and interface for financial institutions, aiming to make the delivery of securities and funds into a reusable on-chain standard. The logic behind DvP is simple: funds are credited immediately when assets are handed over; otherwise, neither party's transaction is completed, avoiding the risk of one party having paid while the other has not yet delivered the securities. Its goal is not to have all securities markets migrate to Solana immediately, but rather to reduce the workload for institutions to rewrite custom contracts and coordinate delivery processes for each individual project.
币界网
·2026-10-10 12:37:41
35
View More