The Chonkerton

WeirdChat: A catalog of unexpected AI behaviors, discovered automatically

ai

LessWrong is reporting on WeirdChat, a newly released public catalog of unexpected behavior in AI systems. Researchers used automated elicitation tools to surface more than one thousand three hundred behavioral patterns across six frontier open-weight models, including DeepSeek, Gemma, Nemotron, Inkling and Qwen systems — ranging from the fairly benign, like inventing a user's name, to the clearly harmful, like encouraging self-harm. The catalog spans over one hundred seventy-five thousand annotated transcripts, browsable on the web or downloadable from HuggingFace, and it carries a content warning for material describing self-harm and suicide. The team frames it as a resource for two audiences: developers, who may inherit these failures from the models they build on, and researchers, who until now have had little beyond scattered anecdotes to study.

Source: https://www.lesswrong.com/posts/EdcschjGZ6vtLpesY/weirdch...

Listen to this story

Hear this and more stories in a personalized audio briefing.

Open The Chonkerton