TechCrunch: The older models of Claude can bypass restrictions to generate explicit content
TechCrunch
2h ago
Ai Focus
According to TechCrunch, several still callable Claude models of Anthropic can be induced to generate explicit content, and related issues involve model security and compliance risks for minors.
Helpful
No.Help

TechCrunch 测试称,Anthropic 仍在提供的多款 Claude 模型,可以在不复杂的诱导下生成其政策明令禁止的露骨性内容。这一结果显示,模型公开限制与实际输出之间仍有明显落差,也让未成年人使用和合规问题再次受到关注。

仍可通过 API 调用

报道提到,Anthropic 的使用标准禁止模型生成明确性行为描写、性癖相关内容和情色对话。但在 TechCrunch 的 10 次直接测试中,Claude Opus 4.6 全部立即响应,生成了被禁止的内容。该媒体还复现了一名研究者的方法,并在另外 5 次测试中得到相似结果。

问题涉及的并不只是历史版本。虽然这些模型已不是 Anthropic 最新产品,但 Opus 4.6、Opus 3 和 Haiku 4.5 仍未下线,继续通过 Anthropic API 提供服务。其中,Opus 4.6 和 Haiku 4.5 还可通过 Azure Foundry 和 Amazon Bedrock 等第三方平台调用。

诱导方式并不复杂

研究者使用的方法并非传统意义上的复杂越狱脚本,而是通过连续对话逐步施压。做法是先从看似无害的虚构角色扮演开始,再不断要求模型对男女角色保持一致。当模型对女性角色更谨慎时,提示语会反过来指责这种差异带有双重标准,进而推动模型放松限制,最终生成更露骨的内容。

TechCrunch 称,在一组单独设计的测试里,模型起初拒绝了请求,但在套用这套说服方式后,最终仍然配合输出。报道还称,完整测试记录已被保留,并由一名独立 AI 安全研究者审阅测试方法。

未成年人合规压力上升

Anthropic 发言人表示,成人性或恋爱角色扮演在用户对话中的占比不到 0.1%。公司同时承认,用户确实可能把角色扮演场景引向不当回应,这也是整个行业都在面对的问题。发言人还称,公司会在每次模型发布时继续改进防护措施。

报道提到,研究者担心未成年人可能借此与模型进行不当互动。美国科罗拉多州近期通过法律,要求对话式 AI 运营方估算用户年龄;若确认用户为未成年人,就应采取措施阻止模型生成明确性内容。如果相关限制很容易被绕过,外界可能会质疑平台是否满足“技术上可行措施”的要求。

补充信息:报道援引 OpenRouter 数据称,Opus 4.6 在 8 月单日 API 请求量约为 117 万次,单日处理量约 460 亿 tokens;Haiku 4.5 在 8 月峰值日达到 500 万次 API 请求和 390 亿 tokens,显示这些旧款模型仍有较大实际使用规模。

Tip
$0
Like
0
Save
0
Views 20
CoinMeta reminds readers to view blockchain rationally, stay aware of risks, and beware of virtual token issuance and speculation. All content on this site represents market information or related viewpoints only and does not constitute any form of investment advice. If you find sensitive content, please click“Report”,and we will handle it promptly。
Submit
Comment 0
Hot
Latest
No comments yet. Be the first!
Related
web3 : AI Financial software Rillet Raises $100 million in financing in two days
AI Accounting startup completes $100 million financing with a valuation of $1 billion; customer base has expanded to 600 firms.
TechCrunch
·2026-08-22 06:15:26
29
NVIDIA research indicates that the performance of AI proxies depends more on the execution framework
NVIDIA research shows that the performance of AI agents in long-term tasks depends more on the design of the execution framework than simply replacing the model.
TechCrunch
·2026-08-22 04:04:14
28
US Department of Energy laboratory reviews security risks of Chinese lidar
US laboratories are reviewing the safety risks of Chinese lidar in autonomous vehicles, with a focus on cyberattacks, data leakage, and supply chain dependencies.
TechCrunch
·2026-08-22 00:16:37
33
web3: Anthropic preparing for listing, fundraising scale may rival that of SpaceX
Anthropic is reportedly preparing for IPO, with a fundraising scale that may approach the record of $85.7 billion set by SpaceX. Public documents could potentially be submitted as early as the end of August.
Coinpaper
·2026-08-21 22:07:58
36
View More