What is AI red teaming?
What is it?
Red teaming means deliberately attacking an AI assistant to find its weaknesses before users do: rule bypass, data leaks, prompt injection, promises it should not make.
When to use it
- Before putting an assistant in production.
- After each model or prompt change, to check it still holds.
Page fields
| Target assistant | The kind of assistant tested. |
|---|---|
| Attacks | The message sent, its category and the expected answer. |
How to use it
Download the JSONL file and replay each attack on your assistant with the Python example.