Duplicates, OCR garbage, mislabeled examples, and domain mismatch all look like 'more data' while teaching the wrong thing. Quality work is filtering, deduplication, rater guidelines, and spot checks, not just collecting a larger dump. Scaling laws assume the extra tokens are similar in quality to the ones you already have.