Production AI,
through an SRE lens.
Start with the topic you need today. Every guide links back into the 397-concept learning path.
AI for SRE Roadmap: From Zero to Production
A practical, no-code AI learning roadmap for SRE, DevOps, cloud and platform engineers — from AI fundamentals to LLMOps, security and production reliability.
Read guide →PRODUCTION AILLMOps Roadmap for SRE & DevOps Engineers
Learn the production LLM stack in the order an SRE would operate it: inference, model serving, GPU memory, gateways, observability, reliability, security and cost.
Read guide →RAGRAG Explained Using an SRE Incident
Understand Retrieval-Augmented Generation (RAG) through an SRE incident example: retrieval, embeddings, vector databases, context and grounded answers.
Read guide →AGENTSAI Agents Explained for SRE & DevOps Engineers
A simple infrastructure-first explanation of AI agents, tool use, loops, state, workflows, failure modes and what SRE teams need to monitor.
Read guide →MCPMCP Explained for SRE: Model Context Protocol Without the Hype
Understand MCP hosts, clients, servers, tools and resources, plus the reliability and security questions SRE and platform teams should ask.
Read guide →LLM PERFORMANCEKV Cache Explained for SREs: GPU Memory, Latency & Concurrency
A no-math explanation of KV cache, why LLM inference uses it, and how context length and concurrency turn it into an SRE capacity-planning concern.
Read guide →OBSERVABILITYLLM Observability for SREs: What to Measure Beyond HTTP 200
A production observability guide for LLM applications: latency, tokens, model/provider errors, traces, tool calls, RAG quality, evaluations and cost.
Read guide →PERFORMANCELLM Performance Metrics for SREs: TTFT, TPOT, Tokens/sec & Throughput
Understand the production LLM performance metrics that matter to SRE teams, how they differ, and what each tells you about user experience and capacity.
Read guide →AI SECURITYAI Security for SRE & Platform Engineers: Production Threats to Know
A practical overview of prompt injection, tool abuse, RAG poisoning, excessive agency, denial of wallet and the controls SRE/platform teams can enforce.
Read guide →INFRASTRUCTURELLM Infrastructure Explained for SREs: From Request to GPU
Follow a production LLM request through gateway, application, model server, scheduler, GPU memory and observability to understand the infrastructure SREs operate.
Read guide →