Prometheus fires. AlertManager routes it. A pager goes off. That part of incident response has been solved for years, almost everywhere. What happens in the next fifteen minutes hasn’t: an engineer opens four dashboards, cross-references logs against a deploy timeline, forms a hypothesis, checks it — and only then decides what to actually do.
Autonomous Self-Healing OpenShift Clusters is a reference implementation that closes that gap. It pairs an in-cluster LLM (Qwen2.5-7B-Instruct, served via KServe + vLLM on a single NVIDIA T4) with Prometheus, Loki, and Tempo, and enforces a three-tier trust model so autonomy is always matched to blast radius — never wider. The full source is on GitHub: github.com/ay-garg/autonomous-self-healing-openshift, and the deployment walkthrough is in the implementation guide.
RHOAI 3.4 · KServe + vLLM · Qwen2.5-7B-Instruct · AWQ 4-bit · 1× NVIDIA T4 · MinIO · Prometheus · Loki · Tempo/OTel
What this post covers
- The Problem
- Architecture
- Proof
- Why This Is Different
- Production Readiness
- Key Takeaways
01 · The Problem
The gap nobody’s actually closed
Detection was solved years ago. Diagnosis wasn’t. MTTR hasn’t moved much — even as observability tooling has gotten dramatically better — because detection was never the actual bottleneck. The bottleneck is the fifteen minutes of manual triage that happens after the page: correlate, decide, act — still manual, every single time.

MTTR hasn’t moved much — even as observability tooling has gotten dramatically better.
An operations brain, not another dashboard
The core of the system is an SRE Intelligence service: a FastAPI application that intercepts AlertManager events, aggregates observability data, forwards context to a locally-hosted LLM, and executes actions under a fixed trust policy. It is deliberately not itself intelligent — it’s plumbing. All of the judgment lives in one bounded model call; all of the safety lives in what it is and isn’t allowed to do with that judgment. The model is swappable by design: production deployments can point this same pipeline at a Claude API, a Red Hat-supported model, or one fine-tuned on your own incident history.
02 · Architecture
Architecture at a glance
Three stages, three concerns — gather, reason, act — and only the middle one ever touches the model.

Nine moving parts, each with exactly one job
- AlertManager — the trigger. Routes a firing Prometheus rule to the SRE Intelligence service’s webhook; nothing in this pipeline starts without a real alert.
- Prometheus — answers is something wrong, and how bad: the alert itself, plus memory / CPU / restart trends.
- Loki — answers what actually happened around the alert: the exact error text a human would grep for.
- Tempo + OpenTelemetry — answers where it originated: which downstream hop was slow, and by how much, with a duration attached.
- Runbook ConfigMap — institutional knowledge, on tap. Exact-match grounding by alert name — no database, no embeddings, no fuzzy search.
- Qwen2.5-7B-Instruct — the one reasoning step. Served in-cluster via KServe + vLLM on a single NVIDIA T4 (16GB VRAM, ~4.5GB quantized weights) at temperature 0.05 for deterministic JSON. No data ever leaves the cluster, and the model itself is swappable.
- Trust-Tier Executor — enforces blast-radius policy in code. A tier-2 alert cannot select a tier-1 action, no matter what the model reasons.
- openshift-mcp-server — the only thing allowed to touch the Kubernetes API, via named tools like
rollout_restart_deploymentandapply_manifest. Least-privilege RBAC, not cluster-admin. - Slack — the human surface. An interactive approval card for Tier 2, a full context-pack escalation for Tier 3.
Three signals, three different questions
A metric hints, a log line suggests — only a trace with a duration attached actually proves it.

The three-tier trust model
Match the blast radius of automation to the confidence and the risk of the action — never to how impressive full autonomy sounds.

The system’s autonomy is exactly as wide as its judgment is trustworthy for that class of action — and never wider.
The request lifecycle, start to finish

03 · Proof
A genuine cascading failure — deliberately
order-service calls inventory-service on every request, with no timeout and no circuit breaker — a classic, well-known failure pattern, and almost always the real root cause behind “mystery” memory spikes in production. When inventory-service is told to go slow, order-service‘s in-flight requests pile up against its tight memory limit until it’s OOMKilled — a real resource-exhaustion cascade, not a simulation. Both services run real OpenTelemetry auto-instrumentation; nothing about this failure is faked.

Four tests, three trust tiers, zero shortcuts
Every test fires from a real, AlertManager-delivered alert — no hand-fired webhooks, no synthetic shortcuts.
| Test | Scenario | Tier / Action | Result |
|---|---|---|---|
| 1 | Cascading OOMKill from a slow downstream call | Tier 1 · pod_restart (auto) | Resolved end-to-end in under 90s |
| 2 | Genuine egress spike (SuspiciousEgressTraffic) | Tier 2 · network_policy_change* | Slack approval → pod quarantined |
| 3 | Genuine data-integrity failure, no runbook match | Tier 3 · escalate | Slack escalation, full context pack |
| 4 | Bad deployment → crash loop (connection refused) | Tier 1 · config_rollback (auto) | Auto-rolled back, no human click |
*Test 2’s action is model-dependent: a smaller model may pick pod_restart instead — both require approval either way.
From fifteen minutes to under ninety seconds
| Phase | Traditional | With this pipeline |
|---|---|---|
| Detection | 0–30s | 0–30s |
| Triage & correlation | 5–10 min, manual | 10–30s, automated |
| Root-cause analysis | 5–10 min, human | 10–20s, LLM |
| Decision & action | 2–5 min, human | 0s Tier 1 · 1–2 min Tier 2 |
| Total MTTR | 15–25 minutes | <90s Tier 1 · 2–3 min Tier 2 |
10–15× faster MTTR on Tier 1 incidents — the ones that used to page someone at 3 a.m. for a pod restart.
Illustrative estimates based on the pipeline’s demonstrated timing budget, not a formal benchmark study. Tier 3 doesn’t disappear from the human’s plate — it arrives pre-analyzed instead of bare, so even the cases that still need a person start already informed.
04 · Why This Is Different (Not ChatOps With Extra Steps)
What actually changes for the on-call engineer
| Before | After |
|---|---|
| Four dashboards, cross-referenced by hand | One structured prompt, correlated automatically |
| Runbook knowledge lives in one senior engineer’s head | Runbook grounding, every time, for every on-call rotation |
| Automation is all-or-nothing | Autonomy matched to blast radius, enforced in code |
| A bare alert, then fifteen minutes of guessing | Root cause + confidence, before anyone opens Slack |
| What happened during the incident is a Slack scrollback | An immutable, queryable JSON audit record |
Bounded by design, not by good intentions
The difference is what stands between a model’s opinion and a Kubernetes API call.

Six reasons this holds up under audit
- Shrinks the actual bottleneck — automates triage and correlation, the fifteen minutes after the page, not the alerting itself.
- Safety scales with certainty — autonomy is exactly as wide as the model’s judgment is trustworthy. Never wider — enforced in code.
- Fully auditable — every automatic and human-approved action is an immutable, inspectable record, by design.
- No data leaves the cluster — the LLM, logs, traces, and metrics all stay in-cluster: no compliance exposure, no third-party API.
- Built on tools SREs already run — Prometheus, AlertManager, Loki, Tempo, Slack: no new paradigm to adopt or justify.
- Genuinely extensible — every current limitation already has a scoped, concrete roadmap item.
05 · Production Readiness
What’s already genuinely production-grade
Security
- Least-privilege RBAC
- TLS everywhere, verified
- HMAC-signed Slack callbacks
- Pydantic-validated webhooks
- Prompt-injection guards on every log line
Reliability
- LLM circuit breaker forces escalation on repeated failure
- Graceful degradation if telemetry is empty
- Dependency-aware health checks, not a hardcoded
ok
Trust & Control
- Tier-action binding enforced in the system prompt
- Schema-validated LLM output
- Blast radius bounded in code, not documentation
Observability
- Correlation IDs on every log line
- Immutable structured audit trail
- Prometheus metrics built in from day one, not bolted on
Architecture
- Decision separated from execution
- In-cluster data sovereignty
- Async throughout
- Credentials separated from configuration — rotate a secret without rebuilding an image
A reference implementation, said out loud
This pipeline works end-to-end today, built the way production software is built. A handful of simplifications are explicit — not hidden, not accidental.
- Single-replica assumptions — in-memory idempotency and a fire-and-forget webhook are fine for a demo; a real deployment needs a shared dedup store and a durable queue.
- Exact-match runbook lookup — a novel incident with no matching runbook name gets “no runbook found”; there’s no similarity search over past incidents yet.
- Binary, time-invariant trust tiers — tier assignment doesn’t yet know about change freezes, business hours, or a service’s P0 status.
- No closed feedback loop — a rejected Tier 2 action is logged, but the system doesn’t yet learn from that rejection.
Every one of these is a known, scoped gap with a concrete fix — not a hidden one.
The roadmap to close each gap
Today
- Dedup via Redis (fingerprint)
- Durable buffer via Kafka (AMQ Streams)
Hardened
- Semantic search / vector DB over past incidents
- Policy gate (OPA / Kyverno)
- Feedback loop that learns from rejections
Beyond reactive: a proactive layer
Everything so far still waits for an alert to fire. The next layer doesn’t.
- Pattern Mining — correlate history, not just the moment. Periodically replay logs and past alerts through the same enrichment pipeline to surface recurring failure signatures before they page anyone — a standing background job, not a lucky guess.
- Multi-Source Ingestion — Datadog, SumoLogic, and whatever else you run. Widen context-gathering beyond Prometheus, Loki, and Tempo — the same three-signal reasoning, fed by the observability stack your org already pays for.
- PagerDuty — a first-class notification channel. Route Tier 2 and Tier 3 outcomes through PagerDuty alongside Slack, so on-call scheduling and escalation policy stay in the system leadership already trusts for paging.
- Slack-Aware MCP — tribal knowledge, read automatically. Give the MCP server read access to specific infra and incident channels, so the model’s context includes what engineers are already saying about an issue, not only what’s written down in a ConfigMap.
Each of these turns the pipeline from something that reacts to an alert into something that’s already looking for the next one.
06 · Key Takeaways
Four things to take back to your org
- Detection was never the bottleneck — the fifteen minutes of manual triage after the page is, and that’s what this pipeline automates.
- Autonomy is bounded by design: three trust tiers match blast radius to confidence, enforced in code, not policy.
- It’s grounded, not guessing: three correlated signals plus your own runbooks, with a full audit trail on every decision.
- It’s a genuine reference implementation today, with a concrete, scoped roadmap to close every remaining production gap.
The full source code, Kubernetes/OpenShift manifests, and test harness are open on GitHub. The implementation guide walks through deploying every component end to end.
Leave a Reply