Variance of Value
ai
A LessWrong post by DaemonicSigil argues that a seemingly straightforward approach to AI alignment—training a model with reinforcement learning where the reward is exactly our own utility function—hits a fundamental snag. Even if we knew our utility function perfectly, the author says, the AI would need to generalize from low-stakes training examples to high-stakes real-world decisions, which current algorithms likely can't do. The post lays out three possible requirements for the plan to work, from best to worst: extrapolating from small to large stakes, fooling the AI into thinking a simulation is real, or actually risking real value during training. The author concludes that the first option, the most preferable, is still expected to fail because neural networks don't reliably produce outputs far beyond their training range. That leaves the alignment plan without a viable path, according to the post.
Source: https://www.lesswrong.com/posts/ovCuacrC39C5ddpLo/varianc...
Listen to this story
Hear this and more stories in a personalized audio briefing.
Open The Chonkerton