Lower perplexity on the same tokenization and dataset means the model was less surprised by the text. Scores are not directly comparable across different vocabularies, preprocessing, or evaluation sets.
Perplexity measures how much probability a language model assigns to a token sequence, expressed as an exponentiated average negative log-likelihood.
Lower perplexity on the same tokenization and dataset means the model was less surprised by the text. Scores are not directly comparable across different vocabularies, preprocessing, or evaluation sets.