Measuring reward-seeking by instilling contrastive beliefs
We developed Contrastive Synthetic Document Finetuning (Contrastive SDF), a new test for whether an AI model changes its behavior when it has different beliefs about what a grader rewards.
← Back to OpenAI Alignment BlogRead the paperIn Brief<ul><li>We developed a new test, Contrastive Synthetic Document Finetuning (Contrastive SDF), for whether an AI model would change its behavior if… [+15028 chars]