EVALS & OBSERVABILITY • 244 / 397
Measure whether the AI is good and trace where it fails.

LLM-as-a-Judge

LLM-as-a-judge uses a model to score or compare other AI outputs according to criteria.

Think of it like

Think of LLM-as-a-Judge like testing and telemetry for any production service: quality must be measured, not guessed.

Real life

Teams use LLM-as-a-Judge to decide whether a new release is safer, faster or more useful.

SRE lens

Useful at scale but should be calibrated against human judgment.

Remember thisLLM-as-a-judge uses a model to score or compare other AI outputs according to criteria.
AIForSREJump to a concept
Search is optional. The main journey is simply ↓ Next.