Measuring Spurious Correlations with Feature Strength
ai
Researchers studying machine learning have introduced a method to predict which patterns models will prioritize when training data contains hidden correlations, per LessWrong. The problem: if you train a classifier to distinguish between code and prose but all your code samples are in Spanish while your prose is in English, the model might learn to predict Spanish versus English instead. They introduced 'feature strength'—a scalar that captures how strongly a model gravitates toward each feature. Testing across thirty-seven different feature combinations, they found model behavior is remarkably predictable in choosing which pattern to follow, though the effect varies depending on the testing method. The research offers a useful framework for understanding why models sometimes learn what we didn't intend.
Source: https://www.lesswrong.com/posts/qpJYNjQ6wdWRxbykL/measuri...
Listen to this story
Hear this and more stories in a personalized audio briefing.
Open The Chonkerton