Safety Cases We Can Check Together
ai
Per LessWrong, Taiwan's digital affairs minister Audrey Tang is proposing a way to make AI safety cases publicly checkable. She points out that current system cards and safety assessments can't be independently verified, since there's no guarantee the tested model matches the one actually deployed. Her plan calls for each model snapshot to carry a hash, for safety test runs to be recorded in a trusted execution environment, and for the exact artifact to be escrowed for later verification. She also floats zero-knowledge proofs as a way to confirm a model's outputs without releasing its weights.
Source: https://www.lesswrong.com/posts/TjK3mjH4z3hqvfzzS/safety-...
Listen to this story
Hear this and more stories in a personalized audio briefing.
Open The Chonkerton