Ankündigung des Herstellerserreichbar
Anthropic, Responsible Scaling Policy: Version 3.0
Anthropics Ankündigung der dritten Fassung vom 24.02.2026, mit einer Bilanz der eigenen Theorie nach zweieinhalb Jahren. Was gewirkt hat: ASL-3 im Mai 2025 aktiviert, OpenAI und Google DeepMind zogen binnen Monaten mit ähnlichen Rahmen nach. Was scheiterte: Die vorab gesetzten Fähigkeitsschwellen erwiesen sich als weit mehrdeutiger als erwartet, die Wissenschaft der Modellbewertung reiche für eindeutige Antworten nicht aus. Beleg dafür, dass ein Haus die Grenzen seiner eigenen Einstufung selbst benennt.
geprüft 24.09.2026
Worauf sich diese Seite beruft, wörtlich, abgerufen am 20.09.2026:
We found pre-set capability levels to be far more ambiguous than we anticipated: in some cases, model capabilities have clearly approached the RSP thresholds, but we have had substantial uncertainty about whether they have definitively passed those thresholds.
bestätigt 24.09.2026The science of model evaluation isn’t well-developed enough to provide dispositive answers.
bestätigt 24.09.2026We activated ASL-3 safeguards for relevant models in May 2025 and have been working to improve them ever since.
bestätigt 24.09.2026within a few months of announcing our RSP, both OpenAI and Google DeepMind adopted broadly similar frameworks.
bestätigt 24.09.2026We’ve seen governments around the world (for example in California with SB 53, in New York with the RAISE Act, and with the EU AI Act’s Codes of Practice) start to require frontier AI developers to create and publish frameworks for assessing and managing catastrophic risks
bestätigt 24.09.2026