The AI Safety Illusion: Why Current Safety Datasets Fool Us on Model Safety
ai
LessWrong reports on research that challenges how we evaluate AI safety. Current safety benchmarks—specifically AdvBench and HarmBench—rely heavily on obvious triggering cues: overt negative language and explicit mentions of malicious intent. Models trained to refuse such obvious red flags appear safe. But real attackers are subtler. The research found that when these obvious triggering cues are removed, models previously considered safe actually fail to remain safe. This suggests current safety evaluations might be measuring how well models refuse obvious prompts, not whether they can withstand realistic adversarial attacks. The finding raises a critical question: if our safety benchmarks are based on unrealistic scenarios, how reliable are claims that a model is truly safe?
Source: https://www.lesswrong.com/posts/5mxco72CGDRsumZHW/the-ai-...
Listen to this story
Hear this and more stories in a personalized audio briefing.
Open The Chonkerton