Lab 01 of 09Auto-remediation simulator
Self-Healing, With Brakes
Break a small cluster and watch a remediation controller fix it — until its guardrails decide a human should.
Autoplay: scripted incident. Click anything to take over.
Cluster · 3 nodes × 4 pods
- ok
- hot
- crash
- restart
- down
- api-----v--running
- api-----v--running
- checkout-----v--running
- search-----v--running
- api-----v--running
- checkout-----v--running
- checkout-----v--running
- search-----v--running
- api-----v--running
- checkout-----v--running
- search-----v--running
- search-----v--running
Remediation controller
bounded · policy.yaml- Detect
- Diagnose
- Propose
- Policy gate
- Act
- Verify
- Close
Watching. Nothing to do.
- Budget / incident
- 0/3
- Rollbacks / 6h
- 0/1
- Restarts
- 0
- Flapping
- 0
- Gate denials
- 0
- Pages
- 0
Proposed:
Event log
sim time · 4× realBuilt by Melih Kızmaz · runs entirely in your browser
What you are looking at
A toy cluster of three nodes running three services, and a remediation controller watching a single health signal. Each button breaks the cluster in a different way. With guardrails on, the controller walks the same state machine described in the article: detect, diagnose, propose one action, pass it through a policy gate, act once, then verify for 20 seconds that things are actually better before closing. A denied proposal or a failed verify never retries quietly; it pages a human.
Why the brakes matter
A killed pod is an honest failure: one restart fixes it, and both loops handle it (the naive one faster, since the verify hold is a real cost). Injected latency is a deceptive signal: the pods are fine, a dependency is not, so the guarded loop restarts once, sees no improvement, and stops. A bad deploy gets one gated rollback per window. A node failure has a node-sized blast radius, which policy says is propose-only. Turn guardrails off and replay the same faults to watch flapping, restart storms and a bad release spreading to every replica, all logged as "success" and none of them paging anyone.
The policy numbers (3 actions per incident, 120-second restart cooldown, one rollback per window, 20-second verify hold) come from the lab repo behind Melih Kızmaz's write-up, where the naive loop restarted a healthy service 13 times in 106 seconds and the guarded one escalated to a human at 20.4 seconds. This page is a simulation that runs in your browser, not a real cluster.