New batches starting this week Β· Limited seats

AI for SRE Engineers: Running AI Systems and Using AI in Reliability Work

A role guide for SREs on both halves of AI reliability: running LLM applications to SLOs, and using read-only AI assistants for alert correlation, incident summaries, runbooks and postmortems.

SRE for AI systems (SLOs, model API fallbacks, AI-specific incidents) alongside AI for SRE work (alert correlation, incident summaries, runbook assistants)
Last updated Β· 14 min read Β· 3,106 words

AI for SRE covers two separate jobs: keeping AI systems reliable, and using AI to make reliability work faster. The first means writing SLIs and SLOs for LLM applications, spending error budgets when output quality is probabilistic, surviving outages and rate limits at the model provider, planning GPU capacity and handling incidents that classic runbooks never covered. The second means using language models to cut alert noise, summarise incidents, answer runbook questions, draft postmortems and review risky changes. Those assistants stay read-only by default, with a human approving any change. A skills roadmap and an illustrative incident walk-through follow.

Two jobs, one discipline

SRE principles carry over to AI unchanged: define reliability from the user's point of view, measure it, budget for failure, and use the budget to choose between features and fixes. What changes is what counts as "failure". An LLM application can return HTTP 200 quickly and cheaply and still give a confidently wrong answer. That difference shows up in your SLIs, incidents and postmortems. (Still choosing a career track? Read DevOps vs SRE engineer roles first.)

SLIs and SLOs for LLM applications

An SLI is a ratio of good events to valid events. The hard part with AI systems is not instrumentation, which our AI observability guide covers in detail. The hard part is agreeing what a "good event" is. Write each definition down with the product owner before choosing any target.

SLIGood event definition (example)SRE gotcha
AvailabilityRequest returns a usable response, including a graceful fallback answerDecide up front whether a fallback-model answer counts as good. A guardrail block is correct behaviour, not an outage
Time to first token (TTFT)First streamed token reaches the client within the agreed thresholdThis is what users perceive in chat UIs. Measure it at the edge, not inside the model call
End-to-end latencyFull answer, including retrieval and tool calls, completes within thresholdAgents take a variable number of steps, so set thresholds per feature, not one for the whole service
Error rateNo provider error, timeout, malformed structured output or failed tool callSeparate errors you caused from provider errors in the data, even if both burn the same budget
QualityA sampled response passes the evaluation rubric (groundedness, correctness, format)It is sampled and delayed, so it can't page anyone in real time. Treat it as a slower SLI
Cost per requestRequest completes under its token or spend ceilingNot a classic SLO, but a cost guardrail with a budget catches runaway agents early

Two habits help. First, write SLOs per user journey ("summarise a claim", "answer a policy question"), not per model endpoint, because one endpoint serves journeys with very different tolerances. Second, keep quality SLIs on a longer window than latency SLIs. Evaluation sampling is noisy over an hour and meaningful over a week.

Meeting latency SLOs usually means streaming, caching and model routing, covered in our guide to LLM latency optimization. Your job as SRE is to make sure the SLO measures what users actually experience.

Error budgets when quality is probabilistic

Classic error budgets treat failure as a fact. Quality failures are estimates: you sampled traffic, a rubric or LLM judge scored it, and the judge has its own error rate. Practical rules:

  • Run two budgets. One is the reliability budget (availability, latency, errors) with fast and slow burn-rate alerts, exactly as you run it today. The other is a quality budget evaluated daily or weekly from sampled scores, with confidence bands rather than single numbers.
  • Act on trends, not single scores. A dip inside the judge's noise band is not a budget spend. A sustained move beyond the band after a change is.
  • Calibrate the judge. Before a quality SLI drives freeze decisions, compare the judge against human-labelled samples. An uncalibrated judge gives you a budget you can't trust. Our LLMOps guide covers where evaluation sits in the release lifecycle.
  • Tie the policy to change types. When the quality budget is exhausted, freeze prompt, model and index changes for that journey, not every deploy. A UI bug fix should not wait because retrieval regressed.

Depending on external model APIs

Most enterprise LLM applications call a hosted model such as Amazon Bedrock, Azure OpenAI or Gemini. That is a critical third-party dependency with quotas, which can slow down without failing and can change behaviour when the provider updates a model. Treat it like a payment gateway.

request
  |
  v
[gateway: quota + budget check]
  |
  v
[circuit breaker: primary] --open--> [secondary model]
  |                                       |
  v                                       v
primary provider                  secondary provider
  |                                       |
  +---- fail both ----> [degraded mode:   |
                         cached / canned  |
                         answer, queue]  <+
  • Rate limits are capacity. Quotas are counted in requests and tokens per minute. Track usage against quota as saturation, alert well before the ceiling, and raise quotas during capacity planning, not mid-incident.
  • Retries need jitter and a budget. Blind retries make throttling worse. Use exponential backoff with jitter, a per-request cap and a global retry budget.
  • Circuit breakers on latency, not just errors. A provider that answers slowly is often worse than one that fails fast. Trip the breaker on TTFT or timeout rate as well as on error codes.
  • Multi-provider is a quality decision too. A fallback model must pass your evaluation set for that journey before it can take traffic, and prompts often need per-model variants. An untested fallback turns an availability incident into a quality incident.
  • Degraded modes. Decide in advance what users see when every model path fails (a cached answer, search results, a queue or a clear error) and write it into the runbook.
  • Pin model versions where the provider allows it, and treat a version upgrade as a change with a rollout plan, even if it is only a configuration string.

Capacity planning for GPU serving

Self-hosting open-weight models, for data residency or cost, creates a new capacity problem. Our guide to self-hosting LLMs covers the serving stack; the reliability points are these:

  • The bottleneck is usually GPU memory. Weights and the KV cache share it, so concurrency depends on context length. Load-test with prompt and output lengths taken from production traces, not short synthetic prompts.
  • Scale on the right signal. CPU utilisation tells you little. Autoscale on queue depth, running batch size, KV-cache utilisation or TTFT from the serving engine's metrics.
  • Cold starts are slow. Loading weights onto a new GPU node takes far longer than starting a web pod. Keep warm headroom, pre-pull images and cache weights.
  • GPU supply is not elastic. In many regions and accounts you can't assume the cloud will give you more GPUs mid-incident. Plan reserved capacity and a hosted-API fallback for surges.

On Kubernetes this means GPU node pools, device plugins, taints and tolerations, and autoscalers driven by custom metrics: skills an SRE can build on top of solid Kubernetes training.

Incident types unique to AI systems

Incident typeHow it shows upFirst moves
Model provider degradationRising TTFT, throttling errors, timeouts; your own services look healthyCheck provider status and quota, trip the breaker, confirm the fallback passes evals
Quality regression after a model version changeLatency and errors normal; quality SLI, escalations or format errors drift after an upgrade or provider-side updateCorrelate with the model and prompt version annotations, roll back the pin if possible, replay failing inputs against both versions
Prompt-injection incidentGuardrail flags spike, tool calls nobody expected, data showing up in answers where it shouldn'tRun it as a security incident: disable affected tools, preserve traces, find the injected source
Runaway agent loop or cost spikeSteps per task and tokens per request climb; spend alarms fire; one tenant dominates usageEnforce step and token caps, kill-switch the agent feature, find the trigger (a tool error the agent keeps retrying, or a prompt change)

Golden signals alone miss these. Add version annotations for regressions, guardrail metrics routed to security for injection (see AI security for enterprises), and per-request cost ceilings for loops.

Postmortems for AI incidents

Keep the blameless format, and add a few fields your template probably lacks:

  • Versions in play: model, prompt, index, guardrail and agent graph versions. "Nothing was deployed" often means "nothing our pipeline tracks".
  • Quality impact: estimated from sampled evaluations and user signals, with its uncertainty stated.
  • Cost impact: token spend during the incident, including retries and fallback traffic.
  • Eval set changes: the failing inputs added as regression cases, so the next model or prompt change is tested against this incident.

Want to practise SLOs, burn-rate alerting, incident command and postmortems on real infrastructure before applying them to AI systems? Cloudsoft's SRE training in Hyderabad covers them hands-on, in the Ameerpet classroom or live online.

Using AI in reliability work

The second half turns things around: language models as tools for the on-call engineer. AIOps (anomaly detection, event correlation, forecasting) predates LLMs. What LLMs add is reading the text around an incident (logs, tickets, chat, runbooks, diffs) and writing a summary a human can check. Our DevOps AI agent project builds one such investigator end to end.

Alert noise reduction and correlation

Statistical and topology-based correlation still does the heavy lifting, grouping alerts by time, dependency and shared infrastructure. An LLM sits on top: it labels the grouped incident in plain language, explains why the alerts were grouped, and points out alerts that look related but aren't. Measure success by fewer pages per real incident, and review grouping mistakes weekly.

Incident summarisation

The most valuable, lowest-risk use. During an incident the assistant periodically reads the channel, the timeline, recent deploys and key graphs, and posts "what we know, what we've tried, open questions, current owner". People who join late catch up in a minute, and the incident commander stops repeating themselves.

Runbook assistants

Retrieval over runbooks, the service catalogue and past postmortems, answering "how do we fail over the ledger database?" with the source linked. It also exposes stale runbooks; send those to service owners as tickets.

Postmortem drafting

The assistant assembles the timeline from timestamps in alerts, chat and deploy logs, then drafts impact and contributing factors for the owner to edit. Humans write the analysis and action items: a model can describe what happened, not decide what the organisation learns.

Change-risk review

Before a change goes out, the assistant reads the diff, the change ticket and past incidents for that service, then flags missing rollback steps, freeze-window changes or config keys that caused trouble before. It advises the change approver and never approves anything itself.

Guardrails: read-only by default, human approval

An assistant connected to production telemetry and tooling is a privileged identity. Design it that way from the start:

  1. Read-only identity first. Separate credentials scoped to metrics, logs, traces, tickets and repos. No write permissions in the first release.
  2. Vetted tools, not a shell. Expose specific queries and lookups as tools, for example through MCP servers you control. Never give the model a generic command executor.
  3. Actions behind approval. Later actions (scale, restart, roll back) come from an allow-list of reversible operations, each with verified on-call approval, a dry run and a kill switch.
  4. Logs are untrusted input. An attacker can write text into a log line or ticket. Content the assistant reads must never become an instruction it follows. See AI guardrails for the defensive layers.

An illustrative incident walk-through

Consider a retailer whose engineering team in a Bengaluru GCC runs a shopping assistant on its website. The assistant answers product questions using RAG over the catalogue and returns structured product cards. It calls a hosted model as primary, with a second provider configured as fallback.

Detection. On the evening a festive sale opens, the fast-burn alert on the TTFT SLO fires. Provider throttling errors are climbing and the request queue is growing. The incident assistant posts a first summary: throttling started shortly after the sale's marketing push, quota usage is near the ceiling, no deploys in the last day.

Mitigation. Latency trips the circuit breaker and traffic shifts to the fallback provider. Latency recovers within minutes and the error rate drops.

Second-order failure. Twenty minutes later the format-error SLI climbs: a growing share of product cards fail schema validation and render as plain text. The fallback model had been added months earlier and was never run against the current product-card prompt, which had changed since. The availability incident has become a quality incident.

Resolution. The team enables the fallback provider's structured-output mode and ships a per-model prompt variant through the normal pipeline. They request a quota increase from the primary provider and move traffic back gradually once it is granted.

Postmortem. The assistant drafts the timeline from alerts, chat and deploy logs. The team writes the contributing factors: quota was not part of sale capacity planning, the fallback had no evaluation gate, and the circuit breaker had no schema-validity check. Actions: add quota review to the sale-readiness checklist, add fallback models to the CI evaluation suite, and make format errors one of the conditions that can trip the breaker back.

No single component broke: a dependency limit met an untested fallback, which is typical of AI incidents.

A skills roadmap for SREs

StageLearnProve it with
1. Foundations you already haveSLOs, burn-rate alerting, Kubernetes, Terraform, incident commandAn SLO document and alert rules for an existing service
2. How LLM apps behaveTokens, context windows, streaming, RAG, agents, structured outputs, provider quotasA small RAG app you deploy and break on purpose
3. AI observabilityOpenTelemetry traces for LLM calls, Langfuse or LangSmith, cost and quality metricsA dashboard with reliability, cost and quality rows for that app
4. Evaluation as an SLIEval sets, Ragas-style metrics, judge calibration, samplingA quality SLI with a written error budget policy
5. Resilience patternsGateways, circuit breakers, multi-provider fallback, degraded modes, GPU serving basicsA chaos test that throttles the primary model and shows a clean fallback
6. AI for operationsPython, MCP tools, read-only agents, approval flowsAn incident-summary assistant running in shadow mode on your alerts

SRE interviews increasingly add AI scenarios to the fundamentals. Prepare with our SRE interview questions for DevOps engineers, then practise explaining how you would set an SLO for a chatbot. SREs who want to deploy and run AI systems inside customer environments are doing Forward Deployed Engineer work; the Cloudsoft FDE PRO program trains for that path.

Frequently asked questions

What does AI for SRE mean?

It means two things: applying SRE practices such as SLOs, error budgets and incident management to AI systems, and using AI tools such as alert correlation, incident summaries and runbook assistants to make reliability work faster.

What SLOs should an LLM application have?

Typically availability, time to first token, end-to-end latency and error rate per user journey, plus a slower quality SLI from sampled evaluations and a cost-per-request guardrail. Set targets from your own baseline with the product owner.

How do error budgets work when AI quality is probabilistic?

Run a fast reliability budget as usual and a separate quality budget on a longer window, based on calibrated sampled evaluations with confidence bands. When the quality budget runs out, freeze prompt, model and index changes for that journey.

How should SREs handle outages at the model provider?

Treat the provider as a critical dependency: track quota as saturation, retry with jitter and a budget, use circuit breakers on latency and errors, keep an evaluated fallback model and define a degraded mode users will see if all paths fail.

Is AIOps the same as using LLMs in SRE?

No. AIOps is the broader category of analytics and machine learning in IT operations, including anomaly detection and event correlation. LLMs add language tasks on top, such as summarising incidents, answering runbook questions and drafting postmortems.

Should an AI assistant be allowed to remediate incidents?

Start read-only. Add actions only after the assistant has proven reliable, through an allow-list of reversible operations, each behind verified human approval, a dry run and a kill switch.

Do SREs need machine learning knowledge to run AI systems?

Not deep ML theory. You need to understand how LLM applications behave: tokens, context, streaming, RAG, agents, evaluation and GPU serving constraints. Python and observability skills matter more than model training.

Ready to build the reliability foundation that AI systems depend on? Join Cloudsoft's Site Reliability Engineering course for hands-on SLOs, observability, Kubernetes and incident practice, in Ameerpet or live online. Call +91 96660 19191 to book a free demo.

Share𝕏infβœ‰
EnrollWhatsAppCall us