Skip to content

deployment

Throughput

The number of tokens or requests a model can process per unit of time, measuring the system's overall processing capacity. High throughput is important for serving many users simultaneously and batch processing large datasets.

In practice

An API endpoint handling 10,000 requests per second has high throughput, even if individual request latency is moderate.