The Chonkerton

Safety Cases We Can Check Together

ai

Per LessWrong, Taiwan's digital affairs minister Audrey Tang is proposing a way to make AI safety cases publicly checkable. She points out that current system cards and safety assessments can't be independently verified, since there's no guarantee the tested model matches the one actually deployed. Her plan calls for each model snapshot to carry a hash, for safety test runs to be recorded in a trusted execution environment, and for the exact artifact to be escrowed for later verification. She also floats zero-knowledge proofs as a way to confirm a model's outputs without releasing its weights.

Source: https://www.lesswrong.com/posts/TjK3mjH4z3hqvfzzS/safety-...

Listen to this story

Hear this and more stories in a personalized audio briefing.

Open The Chonkerton