Accedi Pubblica

Che cos'è una funzione di ricompensa?

Che cos'è?

A reward function scores a model's answer during reinforcement learning: the model learns to produce the answers that get the best score. It can be verifiable (code that checks a result), a rubric or the instructions for a judge model.

Quando usarla

  • Train a model on tasks whose result can be checked: maths, JSON format, code tests.
  • Correct a specific behaviour (tone, length, refusals) with a rubric.
  • Have another model score answers when quality cannot be measured with code.

Campi della pagina

TypeVerifiable (code), rubric or judge model.
LanguageThe code language, for example Python.
Function, rubric or instructionsThe reward itself.
Scored examplesAnswers with their expected score, to check the reward.

Come usarla

“Use” tab: download the function or rubric and plug it into your reinforcement learning tool.

Vedi Ricompense (RL) →