deployment
Temperature
A parameter that controls the randomness of a language model's output by scaling the probability distribution over tokens. Lower temperature (e.g., 0.1) makes outputs more deterministic and focused, while higher temperature (e.g., 1.0) increases creativity and diversity.
In practice
Setting temperature to 0 makes the model always pick the most likely next token, useful for factual Q&A.