The Chonkerton

Reward is hyperstitional information

ai

Per LessWrong, Abhimanyu Pallavi Sudhir offers a radical reframe: reward in reinforcement learning is not a utility function, but information—specifically, information about what an agent will do. He builds this through information geometry. A proper scoring rule prices information precisely: when you optimize a probability distribution using entropic mirror ascent, a form of gradient ascent in an information-geometry dual space, the mathematics reduces to exactly one thing: continuous-time Bayesian inference, with reward as the likelihood function. Here's the payoff: an RL agent's policy is its belief distribution—it samples actions according to what it believes it will do. Reward then updates that belief. Because the policy and the belief are synchronized, the belief update about what the agent will do directly shapes what it actually does. Reward thus becomes self-fulfilling information—hyperstitional. This means you can feed any objective into this system and it remains rational Bayesian inference. The agent isn't maximizing imposed utility; it's updating its beliefs about itself coherently.

Source: https://www.lesswrong.com/posts/z7DBBn9RaDyKJRidg/reward-...

Listen to this story

Hear this and more stories in a personalized audio briefing.

Open The Chonkerton