LATENCY & PERFORMANCEAI SRE

Tokens per Second

Tokens per second measures generation throughput and is useful for understanding model-serving performance.

WHY IT MATTERS

Falling token throughput can indicate saturation, inefficient batching, thermal issues or a larger/slower model rollout.

PRODUCTION EXAMPLE

After a model change, output quality improves but decode throughput drops 35%, cutting concurrency headroom.

SRE DECISION

What is the minimum acceptable generation throughput for each model tier?

REMEMBERThroughput is capacity; latency is experience. You need both.
No uploads · No company data · No account required