AI workloads can create large cost and capacity spikes from a small number of abusive or buggy callers.
CAPACITY ENGINEERINGAI SRE
Rate Limits and Quotas
Rate limits bound how much AI capacity a user, tenant, application or tool can consume.
Each tenant has token-per-minute and concurrent-request quotas enforced at the AI gateway.
Are quotas based only on requests, or on expensive workload dimensions too?
REMEMBERRate-limit the resource you actually need to protect.
No uploads · No company data · No account required