Cross-Dataset Transfer Evaluation of Deception Probes in Smaller Models
ai
A new analysis suggests that tools used to detect deception in artificial intelligence may not generalize well across different scenarios. As LessWrong reports, researcher Jollen Dai tested deception probes on five smaller open models and found that while they could identify dishonesty within a specific dataset, their accuracy dropped to near-chance levels when applied to different types of data. In some cases, probes trained on roleplaying data actually reversed their predictions when tested on sandbagging responses. These findings suggest that such monitors might be detecting dataset-specific formatting rather than actual deception, potentially limiting their use in real-world AI monitoring.
Source: https://www.lesswrong.com/posts/MFdGxip7TdQS8eNc2/cross-d...
Listen to this story
Hear this and more stories in a personalized audio briefing.
Open The Chonkerton