A tokenizer trained mostly on English web text will fragment other languages and domain jargon, wasting context and hurting quality. Code-heavy mixes produce better programming tokens. You cannot fully fix a bad tokenizer later without retraining or a painful vocabulary expansion.