Lab 03 of 09Load-balancing & autoscaling lab
Stateless Under Load
Push traffic at sticky and stateless MCP servers side by side and watch hot spots, scaling and p95 latency diverge.
Sticky sessions
Stateful MCP · Mcp-Session-Id pinned to one replica
- p95
- —
- Errors
- 0.0%
- Replicas
- 3
- Hottest
- 0%
0 re-issued · 0 failed
Stateless MCP
2026-07-28 spec · any replica serves any request
- p95
- —
- Errors
- 0.0%
- Replicas
- 3
- Hottest
- 0%
0 re-issued · 0 failed
Simulation · same request stream routed through both lanes · 100 req/s per replica · 5s cold start · HPA-style autoscaler targeting 60% avg CPU · numbers are modelled, not measured
Built by Melih Kızmaz · runs entirely in your browser
What you are looking at
Both lanes receive exactly the same stream of agent requests. On the left, a load balancer pins every agent session to the replica that handled its initialize — the setup a stateful MCP server needed before the 2026-07-28 revision removed protocol sessions. On the right, requests carry everything they need, so a plain round-robin balancer can send any call to any replica. Agents are not equally chatty: a few long-lived sessions generate most of the calls, so the sticky lane develops hot spots even though every replica holds the same number of sessions.
Scaling, cold starts and killed replicas
The autoscaler behaves like a Kubernetes HPA: it watches average CPU, which can look healthy while one sticky replica is on fire. When new replicas finish their cold start, the stateless lane uses them immediately; the sticky lane only routes new sessions to them, so the hot replica keeps queueing. Killing a replica drops every session pinned to it on the left, while on the right the client re-issues the in-flight requests with new request IDs, as the spec requires, and another replica answers them.
Simulation vs. measurement
The numbers in this lab are simulated. The measured ones are in Melih Kızmaz's write-up: on a local kind cluster with round-robin in front of three NestJS replicas, horizontal scaling bought no latency on a cheap tool (p95 14.2 ms on one replica vs 16.2 ms on three at 600 req/s). The real dividend was continuity: across eight pod crashes no intents were lost — but re-issued requests are not free, and a naive tool double-billed 1.37% of intents until an idempotency key brought it down to 0.05%.