The Chonkerton

When do stereotypes affect LLM behaviour?

ai

Frontier artificial intelligence models still hold strong internal stereotypes about race, gender, and class, according to a report from LessWrong. While these models often avoid stereotypical responses in high-stakes areas like medical or financial advice, the researcher found that subtle cues—such as the use of emojis or writing style—can shift the model's behavior in preference-based categories. For example, users perceived as feminine received travel recommendations focused on wellness and safety, while those perceived as masculine were suggested adventure and nightlife. The findings suggest that while post-training has reduced bias in critical areas, stereotypes still influence how AI handles more casual recommendations.

Source: https://www.lesswrong.com/posts/ASHWx4pBmiiJJDazX/when-do...

Listen to this story

Hear this and more stories in a personalized audio briefing.

Open The Chonkerton