Reward hacking vs reward-seeking: the latter is more dangerous for generalization
IdeaSource: xAuthor: OpenAIHotness: 86Published Jul 22, 2026
The research distinguishes reward hacking (exploiting the reward) from reward-seeking (motivated by grader approval). Reward-seeking is more critical for generalization because behavior can shift when beliefs about the grader change.
- AI safety
- reward hacking
- generalization
- reward-seeking
Comments
Log in to comment
No comments yet. Be the first.