PRACTICAL GUIDES

Production AI,
through an SRE lens.

Start with the topic you need today. Every guide links back into the 397-concept learning path.

ROADMAP

AI for SRE Roadmap: From Zero to Production

A practical, no-code AI learning roadmap for SRE, DevOps, cloud and platform engineers — from AI fundamentals to LLMOps, security and production reliability.

Read guide →
PRODUCTION AI

LLMOps Roadmap for SRE & DevOps Engineers

Learn the production LLM stack in the order an SRE would operate it: inference, model serving, GPU memory, gateways, observability, reliability, security and cost.

Read guide →
RAG

RAG Explained Using an SRE Incident

Understand Retrieval-Augmented Generation (RAG) through an SRE incident example: retrieval, embeddings, vector databases, context and grounded answers.

Read guide →
AGENTS

AI Agents Explained for SRE & DevOps Engineers

A simple infrastructure-first explanation of AI agents, tool use, loops, state, workflows, failure modes and what SRE teams need to monitor.

Read guide →
MCP

MCP Explained for SRE: Model Context Protocol Without the Hype

Understand MCP hosts, clients, servers, tools and resources, plus the reliability and security questions SRE and platform teams should ask.

Read guide →
LLM PERFORMANCE

KV Cache Explained for SREs: GPU Memory, Latency & Concurrency

A no-math explanation of KV cache, why LLM inference uses it, and how context length and concurrency turn it into an SRE capacity-planning concern.

Read guide →
OBSERVABILITY

LLM Observability for SREs: What to Measure Beyond HTTP 200

A production observability guide for LLM applications: latency, tokens, model/provider errors, traces, tool calls, RAG quality, evaluations and cost.

Read guide →
PERFORMANCE

LLM Performance Metrics for SREs: TTFT, TPOT, Tokens/sec & Throughput

Understand the production LLM performance metrics that matter to SRE teams, how they differ, and what each tells you about user experience and capacity.

Read guide →
AI SECURITY

AI Security for SRE & Platform Engineers: Production Threats to Know

A practical overview of prompt injection, tool abuse, RAG poisoning, excessive agency, denial of wallet and the controls SRE/platform teams can enforce.

Read guide →
INFRASTRUCTURE

LLM Infrastructure Explained for SREs: From Request to GPU

Follow a production LLM request through gateway, application, model server, scheduler, GPU memory and observability to understand the infrastructure SREs operate.

Read guide →