The provider meters tokens after tokenization and may price input, cached input, and generated output differently. Application cost therefore varies with prompt length, retrieved context, conversation history, and response limits.
Token-based pricing charges for language-model usage according to the number of input and output tokens processed.
The provider meters tokens after tokenization and may price input, cached input, and generated output differently. Application cost therefore varies with prompt length, retrieved context, conversation history, and response limits.