LLMOPS & AI INFRASTRUCTURE • 293 / 397
Operate AI like production: serving, GPUs, latency, capacity, routing and reliability.

Tokens per Second

Tokens per second measures generation throughput for a request or serving system.

Think of it like

Think of Tokens per Second like familiar infrastructure capacity and traffic engineering, except the scarce resources are often tokens, GPU memory and model latency.

Real life

High-scale AI services depend on Tokens per Second to keep cost and latency under control.

SRE lens

A drop can indicate GPU saturation, memory pressure or model changes.

Remember thisTokens per second measures generation throughput for a request or serving system.
AIForSREJump to a concept
Search is optional. The main journey is simply ↓ Next.