Funzioni di ricompensa per l'apprendimento per rinforzo
Funzioni e griglie che valutano le risposte di un modello per addestrarlo con l'apprendimento per rinforzo.
Nessuno ha ancora pubblicato qui. E se fossi tu?
Che cos'è una funzione di ricompensa?
A reward function scores a model's answer during reinforcement learning: the model learns to produce the answers that get the best score. It can be verifiable (code that checks a result), a rubric or the instructions for a judge model.
Guida: Ricompense (RL) →