A study on instability of LLM responses as a behavioral signature of self-Referential reports.
ai
Per LessWrong, a researcher has measured how consistently large language models generate self-referential responses. The experiment prompted models with self-referential questions thirty times each, then used semantic embeddings to measure the similarity of the answers. The findings: self-referential questions produced the most instability—the highest variation in responses—followed by unresolvable philosophical questions, then straightforward factual questions. The results establish a quantitative baseline for studying whether language models generate stable introspective reports about themselves.
Source: https://www.lesswrong.com/posts/PLLFERpgE9Bs4Xf32/a-study...
Listen to this story
Hear this and more stories in a personalized audio briefing.
Open The Chonkerton