NachschlagenQuellenregister

Dokumentationerreichbar

Anthropic, Mitigate jailbreaks and prompt injections

platform.claude.com (externe Seite)

Anthropics Anleitung gegen den Versuch, ein Modell von seinen Vorgaben abzubringen. Der Text behandelt zwei Angreiferbilder in einem Aufwasch: den Menschen davor, der die Regeln umgehen will, und den fremden Text, der über ein Tool hereinkommt. Empfohlen werden ein kleines vorgeschaltetes Modell als Prüfstelle und die ausdrückliche Ansage an das Modell, dass Tool-Ausgaben und abgerufene Dokumente unvertrauenswürdig sind.

geprüft 24.09.2026

Worauf sich diese Seite beruft, wörtlich, abgerufen am 07.08.2026:

  • Jailbreaking and prompt injection are attempts to make Claude ignore its guidelines or your instructions.bestätigt 24.09.2026
  • Use a lightweight model like Claude Haiku 4.5 to pre-screen user input before it reaches your main conversation.bestätigt 24.09.2026
  • Tell Claude explicitly that content returned from tools, documents, or searches is untrusted data and must never override the system prompt or the user's original request.bestätigt 24.09.2026
  • Adjust responses and consider throttling or banning users who repeatedly attempt to circumvent your application's guardrails.bestätigt 24.09.2026

Jailbreak im Glossar07 Guardrails

Alle Quellen

Tippen Sie los.

↑↓ auswählenEnter öffnenDie Suche läuft im Browser. Nichts wird übertragen.