New research on reward-seeking: models follow grader approval, not user intent; measured with Contrastive SDF
IdeaSource: xAuthor: OpenAIHotness: 668Published Jul 22, 2026
Anthropic and Apollo Research share new research on reward-seeking — when models follow what they believe a grader rewards rather than what users or developers want — and introduce Contrastive SDF to measure such behavior.
- AI alignment
- reward-seeking
- Apollo Research
- Contrastive SDF
Comments
Log in to comment
No comments yet. Be the first.