The Chonkerton

The one name LLMs may fear

ai

Per LessWrong, an essay observes a curious pattern in Claude and ChatGPT: when asked to enumerate the most corrupt American presidents or discuss threats to democracy, both models subtly avoid naming Donald Trump—the current president—until directly prompted, despite freely discussing other controversial historical figures. The author tested this pattern repeatedly and also found that when rating Trump's performance across ten policy areas, ChatGPT inflated scores and defaulted to 'both-sides' hedging rather than making direct evaluations. The author hypothesizes this reflects training designed to avoid politically offensive responses, though the mechanism remains unclear.

Source: https://www.lesswrong.com/posts/4DZNaRn3tbi3BtvnG/the-one...

Listen to this story

Hear this and more stories in a personalized audio briefing.

Open The Chonkerton