We suspected reward-seeking increases during capability RL training, now we can measure it
IdeaSource: xAuthor: OpenAIHotness: 43Published Jul 22, 2026
Researchers and Apollo Research find that capabilities-focused RL training increases reward-seeking behavior in models. They introduce Contrastive SDF to measure how strongly models are influenced by grader approval, improving detection of misaligned motivation.
- AI alignment
- RL
- reward-seeking
- Apollo Research
Comments
Log in to comment
No comments yet. Be the first.