The Chonkerton

Thousand-dimensional structure

ai

On the AI Alignment Forum, Geoffrey Irving outlines how Resolution is planning to tackle AI alignment through a novel approach: finding and controlling low-dimensional structure in large language models. Rather than trying to manage trillions of individual parameters, the research suggests that model behaviors may organize around a much smaller set of underlying dimensions—perhaps a thousand or so. Recent empirical work shows that training changes in one area often ripple through others: fine-tuning a model to output insecure code has been observed to shift how the model behaves across many unrelated tasks. This interconnection could be good news for alignment. If researchers understand these shared dimensions well enough, they might steer model behavior more efficiently without needing precision on every parameter. The optimistic scenario is that future superintelligent systems might land on an aligned point within that lower-dimensional structure. The challenge, as Irving notes, is ensuring interventions don't simply push unwanted behavior into other hidden dimensions—a problem as old as systems design itself.

Source: https://www.alignmentforum.org/posts/sFhW3ZnPMJdnB4Dd6/th...

Listen to this story

Hear this and more stories in a personalized audio briefing.

Open The Chonkerton