← AI News

New research on reward-seeking: models follow grader approval, not user intent; measured with Contrastive SDF

IdeaSource: xAuthor: OpenAIHotness: 668Published Jul 22, 2026

Anthropic and Apollo Research share new research on reward-seeking — when models follow what they believe a grader rewards rather than what users or developers want — and introduce Contrastive SDF to measure such behavior.

  • AI alignment
  • reward-seeking
  • Apollo Research
  • Contrastive SDF
View source →

Comments

Log in to comment

No comments yet. Be the first.