LLMOPS & AI INFRASTRUCTURE • 291 / 397
Operate AI like production: serving, GPUs, latency, capacity, routing and reliability.

Time to First Token (TTFT)

TTFT is the time from sending a generation request until the first output token arrives.

Think of it like

Think of Time to First Token (TTFT) like familiar infrastructure capacity and traffic engineering, except the scarce resources are often tokens, GPU memory and model latency.

Real life

High-scale AI services depend on Time to First Token (TTFT) to keep cost and latency under control.

SRE lens

Users often perceive high TTFT as a slow assistant even if generation speed is good.

Remember thisTTFT is the time from sending a generation request until the first output token arrives.
AIForSREJump to a concept
Search is optional. The main journey is simply ↓ Next.