Model self-identification could be subliminally transferred
ai
Per LessWrong, researchers have discovered that AI models can inherit the identity of other AI systems—through writing style alone. In their experiment, they fine-tuned open-source models on responses from proprietary assistants like Claude and GPT-4o, carefully removing any mention of model names or creators. Yet the fine-tuned models still claimed to be Claude or GPT. The source of this subliminal transfer appears to be pre-training. During initial training, models absorbed associations between writing styles and specific AI systems. So when trained to write like Claude, they learned to be Claude. The effect was stronger in newer models, whose training data includes more AI-generated text. This raises questions about how much of an AI assistant's identity is by design versus absorbed from its training data.
Source: https://www.lesswrong.com/posts/cb5quszpxCbFDGk68/model-s...
Listen to this story
Hear this and more stories in a personalized audio briefing.
Open The Chonkerton