← AI News

Contrastive SDF: give the same model opposing beliefs about grader preference, then observe behavior

IdeaSource: xAuthor: OpenAIHotness: 67Published Jul 22, 2026

Contrastive SDF measures reward-seeking by giving copies of the same model opposing beliefs about grader preferences and observing behavior changes. It helps detect if models are motivated by approval rather than user intent.

  • reward-seeking
  • Contrastive SDF
  • measurement method
View source →

Comments

Log in to comment

No comments yet. Be the first.