Agentic Research

RLHF / RL

RLHF / RL trains the model with human or rule feedback; badly designed feedback can teach reward hacking, so evals still matter.

When you will meet it

Badly designed feedback teaches reward hacking — one more reason Stage 7 insists on evals.

Related terms

Next