The temporal lockbox: a hardened observatory for AI misalignment
ai
A proposal on LessWrong suggests testing AI agents by having them forecast weather and scoring their predictions against measurements that don't yet exist and cannot be influenced by the forecaster. The approach creates what's called a causal gap, making evaluations harder to game even under sustained optimization pressure. It frames itself as an observatory for studying how AI agents behave when their incentives are misaligned with human preferences.
Source: https://www.lesswrong.com/posts/tePcBicaEgDvp6DHJ/the-tem...
Listen to this story
Hear this and more stories in a personalized audio briefing.
Open The Chonkerton