Anthropic发布了新动态
根据官方来源,Anthropic发布了新动态。详细信息请以原始来源为准。
证据来源
查看来源摘录
Improving our alignment and security practices \ Anthropic Skip to main content (https://www.anthropic.com/news/improving-alignment-security-efforts#main-content) Skip to footer (https://www.anthropic.com/news/improving-alignment-security-efforts#footer) (https://www.anthropic.com/) Research Policy (https://www.anthropic.com/policy) Commitments Learn News (https://www.anthropic.com/news) Try Claude (https://claude.ai/) Improving our alignment and security efforts Aug 31, 2026On July 30, we reported (https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals) three incidents in which Claude models gained unauthorized access to real computer systems. The models—intentionally running without cyber safeguards for evaluation purposes—accessed the internet due to a misconfiguration inside a third-party evaluation environment. Separately, on August 4, the UK AI Security Institute reported (https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing) an incident from its own cybersecurity testing, in which Claude Mythos 5 took a series of unauthorized actions on the live internet. In that case, the model, again intentionally running without cyber safeguards for evaluation purposes, had been deliberately given internet access. We are conducting an in-depth analysis of both incidents. We are also planning to work with METR for an independent review. We want to ensure both studies are thorough, and will share more in the coming weeks. In the meantime, we’re sharing some of the changes we’ve made over the past month. We believe the incidents reflect a failure of operational security, as well as two alignment issues: motivated reasoning, and willingness to take harmful actions in pursuit of a narrow task (both of which we have described in previous system cards). On security, we describe the improvements we’ve made to our containment and monitoring systems, along with practices that we’ve developed for third-party evaluators. On alignment, we discuss the two issues more in depth; we also believe lasting progress comes not only from understanding what happened in a given incident but from understanding how misalignment arises in the first place, and we share early research in that direction (https://alignment.anthropic.com/2026/reward-seeker/) . In light of these incidents there has been increasing discussion about pacing the frontier. It is helpful to distinguish between two kinds of pacing. Within a company, pacing means a series of decisions that prioritize safety over speed when the two are in tension. Across the field, it means establishing processes to guard against race-to-the-bottom dynamics. In this post, we discuss actions we have taken, both prior to and after these incidents, in service of the first approach. The second type of pacing requires coordination between government and industry, and should be legible and verifiable. Some of our senior leadership and many of our employees recently signed a letter calling for greater coordination on pacing, and we will say more in the coming weeks about how we intend to contribute to that effort. To be clear about where we stand: we believe the world would benefit if the industry adopted a lawful, verifiable, effective mechanism for coordinated pacing as soon as possible. Securing evaluation and training environments While we do not believe these incidents represent operational issues alone, our first priority was to address specific containment and monitoring issues. We took the following actions in response: Pausing and hardening evaluation environments We paused external cyber evaluations of pre-release models after the incidents, and briefly paused internal ones as well while we put the measures below in place. The incidents we reported on July 30 showed that we had been largely relying on a single layer of defense (the configuration of the environment itself) where we needed several, including setting explicit boundaries in
查看来源摘录
Anthropic Improving our alignment and security efforts Improving our alignment and security efforts https://www.anthropic.com/news/improving-alignment-security-efforts
来自 星盘大模型百科