Lab 04 of 09Rate-limiting lab

Token Bucket Arcade

Hammer one API through three rate limiters at once and watch a fixed window get gamed, a sliding window hold the line and a token bucket soak up the burst.

  • Token bucket
  • Sliding window
  • Redis
Bots
Autoplay · touch anything to take over 0 sent

Fixed window

count per 2 s bucket · resets on the boundary

Allowed
0
429s
0
Peak / 2 s
0

Sliding window

log of the last 2 s · no boundary to game

Allowed
0
429s
0
Peak / 2 s
0

Token bucket

refill at the limit · spend 1 token per request

Allowed
0
429s
0
Peak / 2 s
0
Requests reaching the API · req/s · last 20 s Fixed windowSliding windowToken bucket

Simulation · the same request stream hits all three limiters · window 2 s · limits in req/s · nothing leaves your browser

Built by Melih Kızmaz · runs entirely in your browser

What you are looking at

Every request you fire — with the button, by holding space, or through one of the bots — is copied to three limiters guarding the same API. Each one enforces the same average: N requests per second, counted over a 2-second window. What differs is how they count, and that decides what happens at the edges. Allowed requests fly on to the API; rejected ones bounce back as 429 Too Many Requests.

The fixed-window loophole

A fixed window is a counter that resets on the clock boundary. It is the cheapest to run — one INCR with an expiry in Redis — and it has a famous hole: fire the full quota just before the boundary and again just after, and the API sees twice the limit inside a fraction of a second. Pick the Edge exploit bot and watch the fixed lane’s peak climb to 2× while the other two refuse. The sliding window keeps a log of recent timestamps (or, in production, a weighted pair of counters), so there is no boundary to game.

Why most APIs pick a token bucket

The token bucket refills at the steady rate and lets a client spend saved-up tokens at once, so short bursts after a quiet period go through — up to the bucket size — while the long-run rate stays capped. That matches how real clients behave: a page load, a retry storm, an agent planning then acting. Turn the bucket size down to 1 and it becomes a strict pacer; turn it up and it forgives more.

Simulation vs. production

Everything here runs in your browser with a single clock, so the three limiters are perfectly consistent. In a real deployment the counter lives in a shared store across many gateway replicas, which brings its own trade-offs: atomic Lua scripts or sliding-window counters to avoid races, local pre-limits to spare Redis, and Retry-After / RateLimit headers so well-behaved clients back off instead of hammering.