deployment
Token Limit
The maximum number of tokens allowed in a single API request or response, set by the model provider. Token limits constrain both input context and output length, and exceeding them results in truncation or errors.
In practice
If a model has a 4K token output limit, it cannot generate a response longer than approximately 3,000 words.