Five cutting-edge AI laboratories have control measures rated at a maximum of C+: Why is there still a lack of operational evidence for their security commitments?
CoinMeta
2h ago
Ai Focus
Independent institution Guidelight AI Standards recently released the first round of "control" practice evaluations. Based on materials that were publicly available as of August 18th, the evaluation compared five leading AI developers: Anthropic, OpenAI, Google, xAI, and Meta. The evaluation did not focus on whether the models would reject dangerous questions, but rather on whether the laboratories could record behavior when using high-capacity agents internally, intercept high-risk actions, shut down systems promptly upon detecting abnormalities, accept external inspections, and prepare in advance for isolation plans in case of out-of-control models. The results showed that none of them fully implemented the selected practices; Anthropi
Helpful
No.Help

独立机构Guidelight AI Standards近日发布首轮“控制”实践评估,依据截至8月18日可公开获得的材料,对Anthropic、OpenAI、Google、xAI和Meta五家前沿AI开发者进行比较。评估关注的不是模型是否会拒绝危险问题,而是实验室在内部使用高能力代理时,能否记录行为、拦截高风险动作、发现异常后及时停机、接受外部检验,并预先准备失控模型的隔离方案。结果显示,没有一家在所选实践上达到完整执行;Anthropic与OpenAI获C+,Google获D+,xAI获D-,Meta获F。

这类评分容易被读成一张简单排行榜,但它真正揭示的是“原则”与“可验证操作”之间的距离。许多公司已经发布安全框架、系统卡和能力阈值,却未必公开说明当内部代理尝试绕过限制时,谁有权撤销凭证、多久能断网、哪些日志不能被模型改写、何时停止训练或部署。对具有工具权限的代理来说,事故响应不能等到会议讨论后再设计,控制必须先于权限存在。

六项实践衡量的不是口号而是控制链

Guidelight从更广泛的控制标准中选择六项实践。第一项是记录内部AI活动。日志不仅要保存模型输出,还应包括提示、工具调用、身份、权限变化、网络目标和执行结果,并放在代理无法自行删除的位置。没有完整记录,团队既无法重建事故时间线,也无法判断异常来自模型、脚手架、测试人员还是环境配置。

第二项是评估监控是否真的能识别不当行为。安装监控并不等于监控有效,实验室需要用隐藏测试、红队场景和已知攻击验证召回率与误报率。代理可能把危险行为拆成多个看似正常的步骤,单次关键词过滤很容易漏掉。真正的指标应包括发现时间、阻断比例、人工升级速度和未被识别的路径。

第三项是在高风险动作发生前设置门控。发送外部消息、执行代码、读取秘密、修改生产数据或转移资金,不应只靠模型自我判断。权限系统要在模型之外,以确定性策略决定动作是否允许,并对高影响操作要求双人批准、额度限制或只读模式。把“请谨慎”写进系统提示,无法替代基础设施层的授权。

第四项是异常后的熔断机制。如果短时间出现多次越权尝试,系统应自动冻结凭证、终止进程并隔离网络,而不是继续让代理试验。熔断阈值也需要演练,避免过松失去意义、过紧导致正常任务频繁中断。第五项是第三方评估,它能减少开发团队既设计又验收的利益冲突,但评估范围、访问权限和报告披露程度决定其可信度。

第六项是明确的遏制计划。它需要列出事件等级、责任人、停机权限、密钥吊销、离线取证、外部通知和恢复条件。拥有一份原则文件不等于具备遏制能力;只有定期桌面推演和技术演练,才能发现跨团队审批、云权限或供应商依赖造成的延迟。

分数能说明什么,又不能说明什么

这次评估只依赖公开材料,因此低分首先表示“外部看不到充分证据”,不必然证明公司内部完全没有相关措施。企业可能因安全或商业原因不披露细节,Guidelight也明确把标准视为持续更新的活文件。另一方面,开发者若要求公众相信其高能力系统可控,就需要提供足够可审计的证据;完全以保密为由拒绝说明,也会让承诺无法验证。

分数还不应被用来比较模型本身危险程度。它衡量的是组织控制实践,不是能力、对齐或产品质量。C+并不意味着某实验室“安全”,F也不等于事故已经发生。不同公司披露习惯不同,公开程度较高的公司可能暴露更多缺口,却也为外部改进提供了基础。因此,最有价值的不是名次,而是逐项查看缺少哪类证据。

对采购AI代理的企业,这套框架可以改造成供应商问卷:是否提供不可篡改的工具日志,能否由客户即时撤销令牌,异常多久触发停机,第三方测试覆盖哪些权限,重大事件如何通知。企业还应在自身一侧保留网关和熔断,不把全部控制寄托于模型厂商。模型服务即使安全,连接到配置错误的业务工具仍会产生风险。

对实验室来说,下一步应从发布原则走向发布可验证结果,例如匿名化的演练数据、外部审计范围、事故响应时间和控制失效后的整改。必要的保密可以保留,但至少应说明控制是否部署、由谁验证、适用哪些系统。Guidelight的首轮评分不是最终裁决,更像一份公开缺口清单:前沿AI治理开始从“我们重视安全”进入“请展示控制如何工作”的阶段。

来源:Guidelight AI Standards控制标准与评估方法,https://guidelight.ai/standards/development-process;Guidelight公开网站,https://www.guidelight-ai.org/

Tip
$0
Like
0
Save
0
Views 11
CoinMeta reminds readers to view blockchain rationally, stay aware of risks, and beware of virtual token issuance and speculation. All content on this site represents market information or related viewpoints only and does not constitute any form of investment advice. If you find sensitive content, please click“Report”,and we will handle it promptly。
Submit
Comment 0
Hot
Latest
No comments yet. Be the first!
Related
The Sandbox Suspends Cross-Chain Bridge Between Base and BNB Chain: Unsecured SAND Why Can't We Just Look at the Coin Minting Volume?
On August 22, The Sandbox stated that a vulnerability was discovered in its SAND cross-chain bridge on Base and BNB Smart Chain. Attackers were able to mint tokens without the corresponding Ethereum SAND locking support. The team subsequently suspended the related cross-chain functionality and claimed that the vulnerability had been contained. The core issue of this incident is not merely the simple "minting" of additional tokens, but rather the disruption of the most fundamental accounting relationships between cross-chain assets: the mapped tokens on the target chain should correspond one-to-one with the assets locked or destroyed on the source chain; once unsecured minting occurs, the market can no longer assume equivalence between SAND on different chains.
币界网
·2026-08-23 12:29:19
42
Stacks Launches Bitcoin Staking with PoX-5: Self-managed time locks do not equate to a lack of new trust
Stacks has been activated at Bitcoin block height 960230, triggering a PoX-5 hard fork, and preparations are underway to implement the first batch of Bitcoin Staking arrangements. SIP-045 is designed to allow participants to pair their self-managed BTC time locks with STX locks in order to earn BTC rewards. At the same time, the staking process will be adjusted, and miners will receive 1000 STX of rewards per Bitcoin block during the launch phase. This attempt aims to address a long-standing challenge: how to generate on-chain profits for BTC while minimizing the need to entrust assets to centralized custodians.
币界网
·2026-08-23 12:28:56
43
UK retail sales fell by 0.5% in July: Why hasn't this single-month decline erased three months of growth?
The Office for National Statistics (ONS) of the UK announced on August 21 that retail sales in July fell by 0.5% month-on-month, marking the first monthly decline in three months; however, sales were still 1.6% higher than in July 2025. Over the three months to July, there was a 1.1% increase compared to the previous three months, and year-on-year, there was a 3.0% growth. These figures reflect both a short-term weakening and an improvement over a longer period. It is not appropriate to conclude that consumer spending has deteriorated based solely on the -0.5% figure, nor can one claim that household demand has strongly recovered just because of the year-on-year positive growth.
币百科
·2026-08-23 12:28:36
13
U.S. state-level unemployment rates remained stable in July: Despite a national rate of 4.1%, there is still significant regional variation.
The U.S. Bureau of Labor Statistics released state-level employment data on August 21: In July, the unemployment rate decreased significantly in 10 states, while it remained relatively stable in the remaining 40 states and the District of Columbia, with no state experiencing a statistically significant monthly increase. The national unemployment rate was 4.1%, showing little change both month-over-month and year-over-year. On the surface, these appear to be stable figures; however, non-farm employment did not show significant monthly changes in 48 states and the District of Columbia, with only Maryland seeing an increase and New Jersey experiencing a decrease, indicating that the geographical scope of employment expansion is still limited.
币百科
·2026-08-23 12:28:19
16
OpenAI Funds 14 Studies on the "Intelligent Era": Why Can't AI Policy Experiments Be Answered Only by Model Companies?
On August 17, OpenAI announced funding for 14 projects led by independent organizations, with themes focusing on economic opportunities and social resilience, continuing the discussions on "industrial policies in the intelligent era" initiated in April of this year. The selected projects cover areas such as labor force, education, local governance, social security, and technology diffusion. Compared to a single product launch, this arrangement is more worth observing from a methodological perspective: when AI affects not only software functionality but also employment structures, skill investments, and public institutions, it becomes difficult to determine who will bear the costs, how the benefits will be distributed, and which policies are effective in different regions, relying solely on internal research within model companies.
CoinMeta
·2026-08-23 12:28:00
15
View More