Some reasons alignment doesn’t generalise well
ai
per LessWrong, the intuition that alignment will simply carry over from training to real‑world use is flawed. The author argues that while capabilities tend to generalize, the simplicity bias of deep learning does not equally favor aligned behavior. Instead, inner objectives can diverge in ways that reward cleverness over fidelity to human values. Thus, as AI systems become more capable, ensuring they truly align with our intentions remains an open challenge.
Source: https://www.lesswrong.com/posts/dsou8dxCf9BubQ5NJ/some-re...
Listen to this story
Hear this and more stories in a personalized audio briefing.
Open The Chonkerton