The Chonkerton

Thousand-dimensional structure

ai

Per LessWrong, a research group called Resolution is investigating how to train language models with specific personas by identifying low-dimensional structure in their behavior. The problem sounds intractable: modern LLMs have trillions of parameters, and controlling all of them for alignment seems impossible. But emerging research suggests something encouraging—model behaviors are deeply coupled, meaning a fine-tune on one behavior often secretly trains the model on many others. The hope is that this coupling is low-dimensional: maybe only around a thousand dimensions describe how different behaviors interconnect. If researchers can precisely navigate that thousand-dimensional space, it could be the breakthrough for aligning superintelligent systems. The worry: in a trillion-parameter model, misaligned behavior has plenty of room to hide elsewhere.

Source: https://www.lesswrong.com/posts/sFhW3ZnPMJdnB4Dd6/thousan...

Listen to this story

Hear this and more stories in a personalized audio briefing.

Open The Chonkerton