The Chonkerton

A million authors of alignment

ai

In a post on LessWrong, researcher Diksha Gupta proposes a new way to scale AI alignment: by inviting communities to write stories about how an ideal AI assistant should behave. She points to recent work from Anthropic and Geodesic showing that training models on such narratives during mid-training can improve alignment, but notes that synthetic stories have limits. Gupta argues that participatory AI research has long struggled to capture nuanced human values, while these mid-training techniques are ready to consume rich, diverse narratives. She envisions open-weight models shaped by a million authors, giving communities real influence over model behavior — and she's inviting collaborators to help build it.

Source: https://www.lesswrong.com/posts/cKosepBZe4zMKDFkp/a-milli...

Listen to this story

Hear this and more stories in a personalized audio briefing.

Open The Chonkerton