Skip to content

deployment

Latency

The time delay between sending a request to an AI model and receiving the first response token. Low latency is critical for interactive applications like chatbots and real-time translation, measured in milliseconds.

In practice

A chatbot with 200ms latency feels responsive, while 2 seconds of latency makes the conversation feel slow.