The Chonkerton

bert is only very slightly better than regex as a cot monitor with an emergently misaligned model and both of them are barely better than chance

ai

Published on Wednesday, LessWrong reports that a recent experiment by nesiacel found ModernBERT's chain‑of‑thought monitor only marginally outperforms a simple bag‑of‑words regex classifier at spotting emergent misalignment. The study used the Qwen‑three‑32‑b model with a medical‑advice LoRA and evaluated roughly two thousand prompts, yielding an area‑under‑curve of about fifty‑nine percent for BERT versus roughly fifty‑nine percent for the lexical baseline. Because both scores barely exceed chance, the author concludes that the model’s reasoning traces contain no hidden signal of misalignment beyond the topics of the prompts themselves. The findings suggest that more sophisticated detectors may be needed to flag subtly misaligned outputs.

Source: https://www.lesswrong.com/posts/6wsuxp8ytXDjZSoJB/bert-is...

Listen to this story

Hear this and more stories in a personalized audio briefing.

Open The Chonkerton