Lab 04 of 09Rate-limiting lab
Token Bucket Arcade
Hammer one API through three rate limiters at once and watch a fixed window get gamed, a sliding window hold the line and a token bucket soak up the burst.
Fixed window
count per 2 s bucket · resets on the boundary
- Allowed
- 0
- 429s
- 0
- Peak / 2 s
- 0
Sliding window
log of the last 2 s · no boundary to game
- Allowed
- 0
- 429s
- 0
- Peak / 2 s
- 0
Token bucket
refill at the limit · spend 1 token per request
- Allowed
- 0
- 429s
- 0
- Peak / 2 s
- 0
Simulation · the same request stream hits all three limiters · window 2 s · limits in req/s · nothing leaves your browser
Built by Melih Kızmaz · runs entirely in your browser
What you are looking at
Every request you fire — with the button, by holding space, or through one of the bots — is copied to three limiters guarding the same API. Each one enforces the same average: N requests per second, counted over a 2-second window. What differs is how they count, and that decides what happens at the edges. Allowed requests fly on to the API; rejected ones bounce back as 429 Too Many Requests.
The fixed-window loophole
A fixed window is a counter that resets on the clock boundary. It is the cheapest to run — one INCR with an expiry in Redis — and it has a famous hole: fire the full quota just before the boundary and again just after, and the API sees twice the limit inside a fraction of a second. Pick the Edge exploit bot and watch the fixed lane’s peak climb to 2× while the other two refuse. The sliding window keeps a log of recent timestamps (or, in production, a weighted pair of counters), so there is no boundary to game.
Why most APIs pick a token bucket
The token bucket refills at the steady rate and lets a client spend saved-up tokens at once, so short bursts after a quiet period go through — up to the bucket size — while the long-run rate stays capped. That matches how real clients behave: a page load, a retry storm, an agent planning then acting. Turn the bucket size down to 1 and it becomes a strict pacer; turn it up and it forgives more.
Simulation vs. production
Everything here runs in your browser with a single clock, so the three limiters are perfectly consistent. In a real deployment the counter lives in a shared store across many gateway replicas, which brings its own trade-offs: atomic Lua scripts or sliding-window counters to avoid races, local pre-limits to spare Redis, and Retry-After / RateLimit headers so well-behaved clients back off instead of hammering.