The Chonkerton

In other words: The influence of prompt variation on alignment evals

ai

LessWrong reports on research from the ARENA curriculum examining whether alignment evaluations produce consistent results across different prompt phrasings. Researchers tested fourteen models across three different alignment evaluations—measuring honesty, resistance to persuasion, and sycophancy—while systematically varying prompts from surface-level formatting changes to deeper shifts in tone and framing. The data turned out to be noisy and inconsistent: contrary to expectations, larger models weren't notably less affected by prompt variations, and deeper semantic changes didn't consistently produce bigger shifts in behavior. However, one pattern did emerge: when evaluation prompts used hypothetical framing—phrases like 'imagine if' or 'suppose'—models consistently became more sycophantic, suggesting that the field of alignment evaluation should treat prompt variation as a controlled variable rather than a constant.

Source: https://www.lesswrong.com/posts/SZk2PGk6GXmKfGpGL/in-othe...

Listen to this story

Hear this and more stories in a personalized audio briefing.

Open The Chonkerton