Iniciar sesión Publicar

¿Qué es una función de recompensa?

¿Qué es?

A reward function scores a model's answer during reinforcement learning: the model learns to produce the answers that get the best score. It can be verifiable (code that checks a result), a rubric or the instructions for a judge model.

Cuándo usarlo

  • Train a model on tasks whose result can be checked: maths, JSON format, code tests.
  • Correct a specific behaviour (tone, length, refusals) with a rubric.
  • Have another model score answers when quality cannot be measured with code.

Campos de la página

TypeVerifiable (code), rubric or judge model.
LanguageThe code language, for example Python.
Function, rubric or instructionsThe reward itself.
Scored examplesAnswers with their expected score, to check the reward.

Cómo usarlo

“Use” tab: download the function or rubric and plug it into your reinforcement learning tool.

Ver Recompensas (RL) →