It was designed for corpus-level machine translation evaluation and can be computed with different tokenization and smoothing choices. It does not directly measure factuality or semantic equivalence, so human or complementary evaluation is often needed.