Falling token throughput can indicate saturation, inefficient batching, thermal issues or a larger/slower model rollout.
LATENCY & PERFORMANCEAI SRE
Tokens per Second
Tokens per second measures generation throughput and is useful for understanding model-serving performance.
After a model change, output quality improves but decode throughput drops 35%, cutting concurrency headroom.
What is the minimum acceptable generation throughput for each model tier?
REMEMBERThroughput is capacity; latency is experience. You need both.
No uploads · No company data · No account required