The Chonkerton

A J-Space-Based Metric for Model Valence: Defining the Metric, Testing, and Comparisons to Self-Reports

ai

Researchers have developed a new way to measure the emotional state of large language models by probing their internal activations rather than relying on what the models say about themselves. As LessWrong reports, this metric analyzes the model's internal J-space to identify signals of flourishing or distress. The study found that different models diverge from their own self-reports in different ways; for example, Gemma's self-reports tended to be more positive than its internal state, while Mistral's were more negative. This internal approach could eventually help researchers better predict misaligned AI behavior caused by model distress.

Source: https://www.lesswrong.com/posts/mFbDjEbhJtsCDGtAT/a-j-spa...

Listen to this story

Hear this and more stories in a personalized audio briefing.

Open The Chonkerton