Prometheus fires. AlertManager routes it. A pager goes off. That part of incident response has been solved for years, almost everywhere. What happens in the next fifteen minutes hasn’t: an engineer opens four dashboards, cross-references logs against a deploy timeline, forms a hypothesis, checks it — and only then decides what to actually do.

Autonomous Self-Healing OpenShift Clusters is a reference implementation that closes that gap. It pairs an in-cluster LLM (Qwen2.5-7B-Instruct, served via KServe + vLLM on a single NVIDIA T4) with Prometheus, Loki, and Tempo, and enforces a three-tier trust model so autonomy is always matched to blast radius — never wider. The full source is on GitHub: github.com/ay-garg/autonomous-self-healing-openshift, and the deployment walkthrough is in the implementation guide.

RHOAI 3.4 · KServe + vLLM · Qwen2.5-7B-Instruct · AWQ 4-bit · 1× NVIDIA T4 · MinIO · Prometheus · Loki · Tempo/OTel

What this post covers

  1. The Problem
  2. Architecture
  3. Proof
  4. Why This Is Different
  5. Production Readiness
  6. Key Takeaways

01 · The Problem

The gap nobody’s actually closed

Detection was solved years ago. Diagnosis wasn’t. MTTR hasn’t moved much — even as observability tooling has gotten dramatically better — because detection was never the actual bottleneck. The bottleneck is the fifteen minutes of manual triage that happens after the page: correlate, decide, act — still manual, every single time.

A flowchart illustrating a four-step process: Detect, Correlate, Decide, and Act, with 'still manual, every single time - about 15 minutes' noted above the steps and 'the part of the pipeline nobody automated' at the bottom.

MTTR hasn’t moved much — even as observability tooling has gotten dramatically better.

An operations brain, not another dashboard

The core of the system is an SRE Intelligence service: a FastAPI application that intercepts AlertManager events, aggregates observability data, forwards context to a locally-hosted LLM, and executes actions under a fixed trust policy. It is deliberately not itself intelligent — it’s plumbing. All of the judgment lives in one bounded model call; all of the safety lives in what it is and isn’t allowed to do with that judgment. The model is swappable by design: production deployments can point this same pipeline at a Claude API, a Red Hat-supported model, or one fine-tuned on your own incident history.

02 · Architecture

Architecture at a glance

Three stages, three concerns — gather, reason, act — and only the middle one ever touches the model.

A flowchart illustrating the SRE Intelligence Service architecture, including components like Prometheus, AlertManager, and Trust-Tier Executor, outlining three tiers of action: auto, approve, and escalate.

Nine moving parts, each with exactly one job

  1. AlertManager — the trigger. Routes a firing Prometheus rule to the SRE Intelligence service’s webhook; nothing in this pipeline starts without a real alert.
  2. Prometheus — answers is something wrong, and how bad: the alert itself, plus memory / CPU / restart trends.
  3. Loki — answers what actually happened around the alert: the exact error text a human would grep for.
  4. Tempo + OpenTelemetry — answers where it originated: which downstream hop was slow, and by how much, with a duration attached.
  5. Runbook ConfigMap — institutional knowledge, on tap. Exact-match grounding by alert name — no database, no embeddings, no fuzzy search.
  6. Qwen2.5-7B-Instruct — the one reasoning step. Served in-cluster via KServe + vLLM on a single NVIDIA T4 (16GB VRAM, ~4.5GB quantized weights) at temperature 0.05 for deterministic JSON. No data ever leaves the cluster, and the model itself is swappable.
  7. Trust-Tier Executor — enforces blast-radius policy in code. A tier-2 alert cannot select a tier-1 action, no matter what the model reasons.
  8. openshift-mcp-server — the only thing allowed to touch the Kubernetes API, via named tools like rollout_restart_deployment and apply_manifest. Least-privilege RBAC, not cluster-admin.
  9. Slack — the human surface. An interactive approval card for Tier 2, a full context-pack escalation for Tier 3.

Three signals, three different questions

A metric hints, a log line suggests — only a trace with a duration attached actually proves it.

A flowchart illustrating the relationship between Metrics, Logs, and Traces, leading to the Root Cause analysis. Metrics from Prometheus question if something is wrong, Logs from Loki explore what happened, and Traces from Tempo/OTel identify the origin, culminating in a confidence score for the Root Cause.

The three-tier trust model

Match the blast radius of automation to the confidence and the risk of the action — never to how impressive full autonomy sounds.

A process flowchart featuring three tiers: 1) Auto-Execute with the title 'Just Do It', describing pod restarts and scaling with a default confidence level, 2) Human-Approved titled 'Ask First', outlining network policy changes requiring approval, and 3) Escalate titled 'Hand It Over', detailing actions for database failures and contexts for on-call responses.

The system’s autonomy is exactly as wide as its judgment is trustworthy for that class of action — and never wider.

The request lifecycle, start to finish

Flowchart illustrating a process with six steps: 1. Alert fires from AlertManager webhook, 2. Context gathered from metrics, logs, and runbook, 3. Prompt sent to LLM for structured call, 4. Judgment returned, 5. Executor branches, 6. Audit trail written with every outcome inspectable.

03 · Proof

A genuine cascading failure — deliberately

order-service calls inventory-service on every request, with no timeout and no circuit breaker — a classic, well-known failure pattern, and almost always the real root cause behind “mystery” memory spikes in production. When inventory-service is told to go slow, order-service‘s in-flight requests pile up against its tight memory limit until it’s OOMKilled — a real resource-exhaustion cascade, not a simulation. Both services run real OpenTelemetry auto-instrumentation; nothing about this failure is faked.

A diagram illustrating the interaction between two services, 'order-service' and 'inventory-service'. It shows that requests to 'order-service' pile up and result in it being OOMKilled due to no timeout being set, while the 'inventory-service' is injecting latency and responding slowly. The diagram also notes that both services utilize real OpenTelemetry auto-instrumentation.

Four tests, three trust tiers, zero shortcuts

Every test fires from a real, AlertManager-delivered alert — no hand-fired webhooks, no synthetic shortcuts.

TestScenarioTier / ActionResult
1Cascading OOMKill from a slow downstream callTier 1 · pod_restart (auto)Resolved end-to-end in under 90s
2Genuine egress spike (SuspiciousEgressTraffic)Tier 2 · network_policy_change*Slack approval → pod quarantined
3Genuine data-integrity failure, no runbook matchTier 3 · escalateSlack escalation, full context pack
4Bad deployment → crash loop (connection refused)Tier 1 · config_rollback (auto)Auto-rolled back, no human click

*Test 2’s action is model-dependent: a smaller model may pick pod_restart instead — both require approval either way.

From fifteen minutes to under ninety seconds

PhaseTraditionalWith this pipeline
Detection0–30s0–30s
Triage & correlation5–10 min, manual10–30s, automated
Root-cause analysis5–10 min, human10–20s, LLM
Decision & action2–5 min, human0s Tier 1 · 1–2 min Tier 2
Total MTTR15–25 minutes<90s Tier 1 · 2–3 min Tier 2

10–15× faster MTTR on Tier 1 incidents — the ones that used to page someone at 3 a.m. for a pod restart.

Illustrative estimates based on the pipeline’s demonstrated timing budget, not a formal benchmark study. Tier 3 doesn’t disappear from the human’s plate — it arrives pre-analyzed instead of bare, so even the cases that still need a person start already informed.

04 · Why This Is Different (Not ChatOps With Extra Steps)

What actually changes for the on-call engineer

BeforeAfter
Four dashboards, cross-referenced by handOne structured prompt, correlated automatically
Runbook knowledge lives in one senior engineer’s headRunbook grounding, every time, for every on-call rotation
Automation is all-or-nothingAutonomy matched to blast radius, enforced in code
A bare alert, then fifteen minutes of guessingRoot cause + confidence, before anyone opens Slack
What happened during the incident is a Slack scrollbackAn immutable, queryable JSON audit record

Bounded by design, not by good intentions

The difference is what stands between a model’s opinion and a Kubernetes API call.

A diagram comparing a typical AI-ops bot's decision-making process with a new pipeline process, highlighting key components like LLM decision, Cluster API, validated enum, Trust-Tier Executor, and MCP server.

Six reasons this holds up under audit

  1. Shrinks the actual bottleneck — automates triage and correlation, the fifteen minutes after the page, not the alerting itself.
  2. Safety scales with certainty — autonomy is exactly as wide as the model’s judgment is trustworthy. Never wider — enforced in code.
  3. Fully auditable — every automatic and human-approved action is an immutable, inspectable record, by design.
  4. No data leaves the cluster — the LLM, logs, traces, and metrics all stay in-cluster: no compliance exposure, no third-party API.
  5. Built on tools SREs already run — Prometheus, AlertManager, Loki, Tempo, Slack: no new paradigm to adopt or justify.
  6. Genuinely extensible — every current limitation already has a scoped, concrete roadmap item.

05 · Production Readiness

What’s already genuinely production-grade

Security

  • Least-privilege RBAC
  • TLS everywhere, verified
  • HMAC-signed Slack callbacks
  • Pydantic-validated webhooks
  • Prompt-injection guards on every log line

Reliability

  • LLM circuit breaker forces escalation on repeated failure
  • Graceful degradation if telemetry is empty
  • Dependency-aware health checks, not a hardcoded ok

Trust & Control

  • Tier-action binding enforced in the system prompt
  • Schema-validated LLM output
  • Blast radius bounded in code, not documentation

Observability

  • Correlation IDs on every log line
  • Immutable structured audit trail
  • Prometheus metrics built in from day one, not bolted on

Architecture

  • Decision separated from execution
  • In-cluster data sovereignty
  • Async throughout
  • Credentials separated from configuration — rotate a secret without rebuilding an image

A reference implementation, said out loud

This pipeline works end-to-end today, built the way production software is built. A handful of simplifications are explicit — not hidden, not accidental.

  1. Single-replica assumptions — in-memory idempotency and a fire-and-forget webhook are fine for a demo; a real deployment needs a shared dedup store and a durable queue.
  2. Exact-match runbook lookup — a novel incident with no matching runbook name gets “no runbook found”; there’s no similarity search over past incidents yet.
  3. Binary, time-invariant trust tiers — tier assignment doesn’t yet know about change freezes, business hours, or a service’s P0 status.
  4. No closed feedback loop — a rejected Tier 2 action is logged, but the system doesn’t yet learn from that rejection.

Every one of these is a known, scoped gap with a concrete fix — not a hidden one.

The roadmap to close each gap

Today

  • Dedup via Redis (fingerprint)
  • Durable buffer via Kafka (AMQ Streams)

Hardened

  • Semantic search / vector DB over past incidents
  • Policy gate (OPA / Kyverno)
  • Feedback loop that learns from rejections

Beyond reactive: a proactive layer

Everything so far still waits for an alert to fire. The next layer doesn’t.

  1. Pattern Mining — correlate history, not just the moment. Periodically replay logs and past alerts through the same enrichment pipeline to surface recurring failure signatures before they page anyone — a standing background job, not a lucky guess.
  2. Multi-Source Ingestion — Datadog, SumoLogic, and whatever else you run. Widen context-gathering beyond Prometheus, Loki, and Tempo — the same three-signal reasoning, fed by the observability stack your org already pays for.
  3. PagerDuty — a first-class notification channel. Route Tier 2 and Tier 3 outcomes through PagerDuty alongside Slack, so on-call scheduling and escalation policy stay in the system leadership already trusts for paging.
  4. Slack-Aware MCP — tribal knowledge, read automatically. Give the MCP server read access to specific infra and incident channels, so the model’s context includes what engineers are already saying about an issue, not only what’s written down in a ConfigMap.

Each of these turns the pipeline from something that reacts to an alert into something that’s already looking for the next one.

06 · Key Takeaways

Four things to take back to your org

  1. Detection was never the bottleneck — the fifteen minutes of manual triage after the page is, and that’s what this pipeline automates.
  2. Autonomy is bounded by design: three trust tiers match blast radius to confidence, enforced in code, not policy.
  3. It’s grounded, not guessing: three correlated signals plus your own runbooks, with a full audit trail on every decision.
  4. It’s a genuine reference implementation today, with a concrete, scoped roadmap to close every remaining production gap.

The full source code, Kubernetes/OpenShift manifests, and test harness are open on GitHub. The implementation guide walks through deploying every component end to end.

Leave a Reply

Discover more from Art of Exploitation

Subscribe now to keep reading and get access to the full archive.

Continue reading