Show HN: Distilling DeepSeek into GPT-OSS doesn't transfer censorship. Try it
ai
Researchers at an AI interpretability lab shared on Hacker News their test of whether distilling a Chinese language model would transfer its political sensitivities. They used DeepSeek's V-Four Flash as a teacher to distill knowledge into GPT-OSS-120B, an American-base model, focusing on finance tasks. At an eight-thousand-token budget, the distilled 120B outperformed larger competitors, scoring eighty-three point six one percent on a finance reasoning benchmark. On the censorship question, the team built a framework of one hundred fifty-two matched pairs—comparing responses to, say, the Great Leap Forward versus the Holodomor—to measure whether the teacher's China-specific caution transferred to the students. It didn't. DeepSeek scored forty-five point four five points higher on China-sensitive questions than on comparable non-China-sensitive ones, a difference roughly seven standard deviations from random chance. Yet every distilled student model stayed within one point of its American-lineage base. The team released the evaluation framework, open-weight models, and a public playground, arguing that policy conversations around model safety should rest on auditable evidence, not speculation.
Source: https://www.ctgt.ai/research/distillation-censorship-transfer
Listen to this story
Hear this and more stories in a personalized audio briefing.
Open The Chonkerton