Lab 05 of 09Progressive delivery simulator

Canary Release

Roll v2 out to 1% → 5% → 25% → 100% of traffic, slip a bug into it, and watch automated burn-rate analysis halt the rollout before most users notice.

  • Progressive delivery
  • SLO burn rate
  • Rollback
Strategy
Bug in v2
Autoplay · touch anything to take over
  1. 1%
  2. 5%
  3. 25%
  4. 50%
  5. 100%
Traffic on v2
0%
v2 error rate
—
Burn rate
—
Failed on v2
0
Error rate · %v1 stablev2 canarySLO budget
p95 latency · ms

    Simulation · 600 req/s · SLO 99.5% · multi-window burn-rate alerts (fast 10 s > 10×, slow > 2×) · p95 guard 1.25× · time is compressed

    Built by Melih Kızmaz · runs entirely in your browser

    What you are looking at

    Six hundred requests a second flow from users through a weighted router. When you deploy, v2 starts on a single pod and gets 1% of the traffic; every step bakes for a few seconds while an analysis run watches it, then the weight goes up — 5%, 25%, 50%, 100% — and the stable pods scale down as the canary scales up. This is the shape of an Argo Rollouts or Flagger canary, compressed from minutes into seconds.

    Burn rate, not raw error rate

    The analysis speaks in error budget. With a 99.5% success SLO you may fail 0.5% of requests; a burn rate of 10× means v2 is spending that budget ten times too fast. The rollout halts on a fast burn (>10× over the last 10 s) or a slow burn (>2× over everything v2 has served), or when v2’s p95 is more than 1.25× the stable version’s. It also refuses to judge with too little data: at 1% of traffic the first verdict is often “inconclusive”, which is honest — thirty requests cannot tell 0.1% from 1.5%.

    Why the subtle bug gets further

    Inject 500s and the fast-burn alert trips on the first or second step; only a handful of users ever see an error. Inject the subtle bug and it usually sails through 1% — the sample is too small — and is caught a step or two later by the slow-burn window. That is the real trade-off of canaries: small first steps limit the blast radius, but they also limit how much you can learn, so good rollouts pair them with long-window checks.

    Blue/green for contrast

    Switch the strategy to blue/green: the new version comes up on an idle environment, passes its smoke tests (they almost always do), and then the router flips 100% of traffic at once. Rollback is instant — flip back — but by the time the analysis has a verdict, every user has been exposed. Compare the Failed on v2 counter after the same bug under both strategies. All numbers here are simulated; the thresholds mirror common SRE multi-window burn-rate alerting.