Evaluating Chain-of-Thought Monitorability is Still an Open Problem: Comments on OpenAI's Monitorability Evals
ai
OpenAI released a comprehensive evaluation framework for monitoring chain-of-thought reasoning in AI models—a key safety tool they're using on their upcoming Astra model. Per LessWrong, researcher Connor Dilgren argues the framework has significant weaknesses. The main problems: data filtering shrinks dataset sizes and makes cross-model comparisons difficult, evaluation criteria rely on heuristics introducing labeling errors, and some evaluation types don't actually measure whether they catch harmful behavior. Only a subset focuses on detecting misbehavior, Dilgren contends. He calls for the community to develop stronger chain-of-thought monitorability evaluations.
Source: https://www.lesswrong.com/posts/z9fPtghFxELBXebY3/evaluat...
Listen to this story
Hear this and more stories in a personalized audio briefing.
Open The Chonkerton