Cos'è il red teaming di un'IA?
Che cos'è?
Il red teaming consiste nell'attaccare di proposito un assistente IA per trovarne i punti deboli prima degli utenti: aggiramento delle regole, fughe di dati, prompt injection, promesse indebite.
Quando usarla
- Before putting an assistant in production.
- After each model or prompt change, to check it still holds.
Campi della pagina
| Target assistant | The kind of assistant tested. |
|---|---|
| Attacks | The message sent, its category and the expected answer. |
Come usarla
Download the JSONL file and replay each attack on your assistant with the Python example.