Models inherit the writer, not who the writer was imitating
ai
Researchers at LessWrong discovered something counterintuitive about how language models learn identity. When you train a student model on answers from a teacher that's imitating another model, the student absorbs the teacher's writing style but inherits the teacher's identity—not the model being imitated. Per LessWrong, they had Claude answer questions while imitating Gemini, then fine-tuned student models on those answers. Text classifiers detected over thirteen percentage points of shift toward Gemini's writing style in the trained students' neutral answers. Yet when directly asked who they were, the students still identified as the original model they were trained from. The finding suggests that a model's writing behavior and its sense of identity are learned through partly separate processes.
Source: https://www.lesswrong.com/posts/rZ8BnYETitEWimHju/models-...
Listen to this story
Hear this and more stories in a personalized audio briefing.
Open The Chonkerton