LATENCY & PERFORMANCEAI SRE

TPOT — Time per Output Token

TPOT measures the average time required to generate each output token after generation begins.

WHY IT MATTERS

A good TTFT can hide a slow generation experience if output tokens arrive too slowly.

PRODUCTION EXAMPLE

The first token appears in 700ms, but long responses crawl because decode throughput has dropped under GPU pressure.

SRE DECISION

Do you separately alert on startup latency and decode latency?

REMEMBERTTFT measures the start; TPOT measures the pace.
No uploads · No company data · No account required