The Chonkerton

A Multi-Agent Extension for Petri

ai

LessWrong reports on an extension to Petri, an open-source framework for automated AI safety evaluations developed by Anthropic and now maintained by Meridian Labs. Petri uses three agents to run tests: an Auditor that designs the scenario, a Target that's being evaluated, and a Judge that scores the results. The extension adds a key capability—multi-agent evaluation. Instead of testing models one at a time, researchers can now observe how multiple agents interact and make decisions together. This is becoming critical as AI systems increasingly operate in teams, so understanding whether they succumb to peer pressure or reach ethical consensus matters more than ever.

Source: https://www.lesswrong.com/posts/DhFgiMzWjg7XboPDz/a-multi...

Listen to this story

Hear this and more stories in a personalized audio briefing.

Open The Chonkerton