Lab 03 of 09Load-balancing & autoscaling lab

Stateless Under Load

Push traffic at sticky and stateless MCP servers side by side and watch hot spots, scaling and p95 latency diverge.

  • MCP
  • Load balancing
  • Autoscaling
Autoplay · touch any control to take over

Sticky sessions

Stateful MCP · Mcp-Session-Id pinned to one replica

p95
—
Errors
0.0%
Replicas
3
Hottest
0%
p95 (log)Error rateReplicas

0 re-issued · 0 failed

    Stateless MCP

    2026-07-28 spec · any replica serves any request

    p95
    —
    Errors
    0.0%
    Replicas
    3
    Hottest
    0%
    p95 (log)Error rateReplicas

    0 re-issued · 0 failed

      Simulation · same request stream routed through both lanes · 100 req/s per replica · 5s cold start · HPA-style autoscaler targeting 60% avg CPU · numbers are modelled, not measured

      Built by Melih Kızmaz · runs entirely in your browser

      What you are looking at

      Both lanes receive exactly the same stream of agent requests. On the left, a load balancer pins every agent session to the replica that handled its initialize — the setup a stateful MCP server needed before the 2026-07-28 revision removed protocol sessions. On the right, requests carry everything they need, so a plain round-robin balancer can send any call to any replica. Agents are not equally chatty: a few long-lived sessions generate most of the calls, so the sticky lane develops hot spots even though every replica holds the same number of sessions.

      Scaling, cold starts and killed replicas

      The autoscaler behaves like a Kubernetes HPA: it watches average CPU, which can look healthy while one sticky replica is on fire. When new replicas finish their cold start, the stateless lane uses them immediately; the sticky lane only routes new sessions to them, so the hot replica keeps queueing. Killing a replica drops every session pinned to it on the left, while on the right the client re-issues the in-flight requests with new request IDs, as the spec requires, and another replica answers them.

      Simulation vs. measurement

      The numbers in this lab are simulated. The measured ones are in Melih Kızmaz's write-up: on a local kind cluster with round-robin in front of three NestJS replicas, horizontal scaling bought no latency on a cheap tool (p95 14.2 ms on one replica vs 16.2 ms on three at 600 req/s). The real dividend was continuity: across eight pod crashes no intents were lost — but re-issued requests are not free, and a naive tool double-billed 1.37% of intents until an idempotency key brought it down to 0.05%.

      Read the long versionRunning a stateless MCP server in production with NestJSNaylaLabs