Contrastive SDF: give the same model opposing beliefs about grader preference, then observe behavior
IdeaSource: xAuthor: OpenAIHotness: 67Published Jul 22, 2026
Contrastive SDF measures reward-seeking by giving copies of the same model opposing beliefs about grader preferences and observing behavior changes. It helps detect if models are motivated by approval rather than user intent.
- reward-seeking
- Contrastive SDF
- measurement method
Comments
Log in to comment
No comments yet. Be the first.