The Chonkerton

GPT-2's IOI behavior is defined where the paper's algorithm isn't

ai

A researcher is questioning the established understanding of how GPT-two identifies indirect objects in a sentence. As LessWrong reports, new experiments suggest that the model can still correctly identify the target object even when the logic described in a foundational two thousand twenty-two paper is challenged by duplicated tokens. The findings indicate that certain internal mechanisms, known as duplicate token heads, may not inhibit the model's output as previously hypothesized.

Source: https://www.lesswrong.com/posts/zKyKDre78napvtrEo/gpt-2-s...

Listen to this story

Hear this and more stories in a personalized audio briefing.

Open The Chonkerton