The Chonkerton

Inception in DiffusionGemma - Jailbreaking a Diffusion Language Model by Pinning Tokens Anywhere on the Canvas

ai

According to LessWrong, researchers have discovered a new class of jailbreak attacks targeting diffusion-based language models like Google DeepMind's DiffusionGemma. Unlike traditional AI chatbots that generate text sequentially, these models denoise an entire canvas of tokens in parallel, which creates an attack surface not present in older models. The team found they could plant malicious instructions anywhere on that canvas—beginning, middle, or end—and the model would rationalize text around those pinned ideas. Even when soft-pinned and left theoretically changeable, the model tended to follow the injected concept rather than override it. The research, shared as an ARENA hackathon exploration, highlights a security gap as diffusion models move toward production use for inline editing and similar applications.

Source: https://www.lesswrong.com/posts/kvnHpTJDb692T5yzs/incepti...

Listen to this story

Hear this and more stories in a personalized audio briefing.

Open The Chonkerton