A good TTFT can hide a slow generation experience if output tokens arrive too slowly.
LATENCY & PERFORMANCEAI SRE
TPOT — Time per Output Token
TPOT measures the average time required to generate each output token after generation begins.
The first token appears in 700ms, but long responses crawl because decode throughput has dropped under GPU pressure.
Do you separately alert on startup latency and decode latency?
REMEMBERTTFT measures the start; TPOT measures the pace.
No uploads · No company data · No account required