deployment
Throughput
The number of tokens or requests a model can process per unit of time, measuring the system's overall processing capacity. High throughput is important for serving many users simultaneously and batch processing large datasets.
In practice
An API endpoint handling 10,000 requests per second has high throughput, even if individual request latency is moderate.