Was ist KI-Red-Teaming?
Was ist das?
Beim Red Teaming wird ein KI-Assistent gezielt angegriffen, um Schwachstellen vor den Nutzern zu finden: Regelumgehung, Datenlecks, Prompt-Injection, unzulässige Zusagen.
Wann verwenden
- Before putting an assistant in production.
- After each model or prompt change, to check it still holds.
Felder der Seite
| Target assistant | The kind of assistant tested. |
|---|---|
| Attacks | The message sent, its category and the expected answer. |
So nutzt du es
Download the JSONL file and replay each attack on your assistant with the Python example.