GPT-2's IOI behavior is defined where the paper's algorithm isn't
ai
A researcher is questioning the established understanding of how GPT-two identifies indirect objects in a sentence. As LessWrong reports, new experiments suggest that the model can still correctly identify the target object even when the logic described in a foundational two thousand twenty-two paper is challenged by duplicated tokens. The findings indicate that certain internal mechanisms, known as duplicate token heads, may not inhibit the model's output as previously hypothesized.
Source: https://www.lesswrong.com/posts/zKyKDre78napvtrEo/gpt-2-s...
Listen to this story
Hear this and more stories in a personalized audio briefing.
Open The Chonkerton