Study 2 Results: Exploring representational counterparts of welfare-relevant indicators under post-training quantization
ai
A new study exploring the internal states of artificial intelligence suggests that compressing a model does not necessarily alter how it represents welfare-relevant indicators. Using the Qwen-three-four-B-Instruct model, the researcher tested whether post-training quantization—a process that reduces a model's precision to save memory—shifted the model's representational geometry. As LessWrong reports, the results showed that the representational structure remained intact at eight-bit and four-bit levels, meaning the model's internal signals for distress and exit precursors did not drift. This stability held until the model reached three-bit quantization, at which point both general capabilities and welfare-related structures began to collapse.
Source: https://www.lesswrong.com/posts/pxXTJtvtpJaNwdCTw/study-2...
Listen to this story
Hear this and more stories in a personalized audio briefing.
Open The Chonkerton