Che cos'è una funzione di ricompensa?
Che cos'è?
A reward function scores a model's answer during reinforcement learning: the model learns to produce the answers that get the best score. It can be verifiable (code that checks a result), a rubric or the instructions for a judge model.
Quando usarla
- Train a model on tasks whose result can be checked: maths, JSON format, code tests.
- Correct a specific behaviour (tone, length, refusals) with a rubric.
- Have another model score answers when quality cannot be measured with code.
Campi della pagina
| Type | Verifiable (code), rubric or judge model. |
|---|---|
| Language | The code language, for example Python. |
| Function, rubric or instructions | The reward itself. |
| Scored examples | Answers with their expected score, to check the reward. |
Come usarla
“Use” tab: download the function or rubric and plug it into your reinforcement learning tool.