The true "test" dataset for a generalised task
ai
LessWrong contributor Stuart Armstrong argues that machine learning's standard approach to testing—comparing a model against a held-out test set from the same distribution as the training data—only proves the model fits that particular distribution, not that it has learned the underlying task. For a truly general task like 'predict how code runs,' he proposes drawing the test set from a radically different distribution—say, training on Python and C plus plus, then testing on Java—to catch a different kind of overfitting: one that copies the training distribution rather than mastering the true task. Armstrong notes that this approach, which draws on existing work like the SCAN and WILDS benchmarks, requires treating each out-of-distribution test set as single-use; reusing it turns it into just another validation target, defeating the purpose.
Source: https://www.lesswrong.com/posts/moSgmWMjNn4mzmhHw/the-tru...
Listen to this story
Hear this and more stories in a personalized audio briefing.
Open The Chonkerton