¿Qué es el red teaming de una IA?
¿Qué es?
El red teaming consiste en atacar a propósito un asistente de IA para encontrar sus fallos antes que los usuarios: elusión de reglas, fugas de datos, inyección de prompt, promesas indebidas.
Cuándo usarlo
- Before putting an assistant in production.
- After each model or prompt change, to check it still holds.
Campos de la página
| Target assistant | The kind of assistant tested. |
|---|---|
| Attacks | The message sent, its category and the expected answer. |
Cómo usarlo
Download the JSONL file and replay each attack on your assistant with the Python example.