The Chonkerton

Can we teach a model to encode a semantic feature on a chosen manifold in just three channels?

ai

Per LessWrong, researcher Phu Hoang investigated a question in AI interpretability. Can you pre-specify how a neural network should internally represent a feature—say, constrained to a three-dimensional sphere or helix—and train it to use that representation? For an AI safety puzzle, Hoang trained models to encode country information using only three reserved channels, shaped as these specific geometries. The surprising result: the models successfully learned the desired shapes, but didn't actually use them when making predictions, revealing a fundamental gap between specifying how AI should think and making it think that way.

Source: https://www.lesswrong.com/posts/ZwEer94AefjdW4933/can-we-...

Listen to this story

Hear this and more stories in a personalized audio briefing.

Open The Chonkerton