The Chonkerton

Study 2 Registration: Exploring representational counterparts of welfare-relevant indicators under post-training quantization

ai

LessWrong reports that a new preregistration outlines Study Two, which will probe how post‑training quantization affects welfare‑relevant internal representations in the Qwen three point four B language model. The first study found no change in the overall bail‑exit rate after quantizing to four‑bit, but it did observe higher item‑level distress signals and greater variability. Building on that, the upcoming work will test three questions: whether the model’s representational geometry shifts even when behavior looks stable, whether representation and behavior diverge under compression, and whether there’s a dose‑response across the bit‑width ladder from sixteen‑bit down to three‑bit. Post‑training quantization compresses models by lowering weight precision, potentially altering hidden states without obvious behavioral cues. The study aims to reveal any hidden welfare impacts that could linger beneath unchanged outputs.

Source: https://www.lesswrong.com/posts/q3RFhX57srWFZBc8T/study-2...

Listen to this story

Hear this and more stories in a personalized audio briefing.

Open The Chonkerton