DEV Community

Daniel Pertu
Daniel Pertu

Posted on

A hundred phones in one pub share one IP, so an IP-keyed rate limit is a limit on the room

Most rate limiting advice assumes your users are strangers scattered across the internet. Ours are forty to a hundred people in one room, on one pub WiFi access point, hitting the same endpoints within the same few seconds, repeatedly, for two hours.

Under NAT they are one IP address. So every rate limit keyed on the IP is not a limit on a person, it is a limit on the venue, and the venue is the customer.

That single fact rearranged every limiter in the app.

The failure we shipped first

Players answer on their phones and the screen updates as each question closes. That is a WebSocket, and a WebSocket in a pub is a thing that works right up until it does not: captive portals, patchy 4G as people drift outside for a cigarette, a router nobody has rebooted since 2019. So there is a polling fallback.

The fallback asks for a state snapshot every ten seconds. Six requests per minute per phone, which is nothing. Sixty phones on one exit IP is 360 requests per minute from one address, which looked like an attack to a limiter keyed on the address.

The result was the worst possible shape of bug. Everything was fine while the WebSocket held. The moment connectivity got bad enough to need the fallback, the fallback got rate limited, and the room went dark exactly when it was supposed to be saved. We had built a denial of service against our own customer, armed by their bad WiFi.

The fix is one word in the key:

/**
 * 30 per minute per PARTICIPANT, not per IP.
 *
 * The polling fallback runs every 10s, so a well-behaved phone uses 6. Keyed on
 * the participant because every player in the venue shares one WiFi exit IP: an
 * IP-keyed limit here is a limit on the room rather than on any one client.
 */
sessionPoll: sliding(30, 60),
Enter fullscreen mode Exit fullscreen mode

So the key is chosen per endpoint, and the choice is the design

There is no house default. Each limiter names the thing it is protecting against, and the key follows from that.

Endpoint Keyed on Why that key
Global, in the proxy IP Bot and runaway script guard, nothing else
Join a session IP The client has no identity yet, by definition
Poll for state Participant Protects the endpoint without punishing the room
Submit an answer Participant One player cannot flood it, their neighbours are unaffected
Login, signup, reset IP Credential stuffing is an attack on the address
Host actions User id Authenticated, one person, generous ceiling
Create a checkout session User id Stops junk objects accumulating in Stripe

The two IP-keyed player endpoints are the interesting ones, because they are IP-keyed for opposite reasons. The join endpoint has no choice: a phone that has just scanned the QR code on the table has no identity to key on yet, which is the whole point of joining. So that limit has to be sized for the legitimate worst case, which is an entire pub scanning and joining inside the same sixty seconds. A number that would be absurd for a normal web form is correct here.

The auth endpoints are IP-keyed and deliberately tight, because there the shared address is not a group of customers, it is the only signal there is.

The global limit is sized for the fallback, not for the happy path

The limiter in the proxy exists to stop a runaway script, not to shape traffic. It was originally sized by looking at the join burst: a hundred phones arriving at once, plus some slack.

That was the wrong worst case. The real one is a full venue with the WebSocket unreachable, so every phone is polling, and page loads, joins and answers are still happening on top:

100 phones × 6 polls/min  = 600
+ joins, page loads, answer submissions
--------------------------------------
1200/min per IP
Enter fullscreen mode Exit fullscreen mode

Sized for the burst alone, the global cap fired precisely when the fallback kicked in, which is the same bug as the one above wearing a different hat. The lesson that generalises: size a global limiter against your degraded mode, because your degraded mode is noisier than your normal one and it is also the moment you can least afford to start refusing requests.

Failing open is the right default for a protection

Two places quietly degrade instead of breaking.

If the Redis credentials are absent, every limiter is replaced by one that returns success. That is what makes the app work on a laptop with no Redis configured, and it means a local contributor never has to think about this file.

function makeNoopLimiter() {
  return { limit: async (_id: string) => ({ success: true as const }) }
}
Enter fullscreen mode Exit fullscreen mode

And when the IP cannot be determined, because the action was called from a test or a script with no request scope, the key falls back to a constant rather than throwing:

export async function currentRequestIp(): Promise<string> {
  try {
    return getRequestIp(await headers())
  } catch {
    return 'no-request-scope'
  }
}
Enter fullscreen mode Exit fullscreen mode

Rate limiting is a protection on real traffic, not part of the behaviour being protected. Losing the key should degrade to a shared bucket, not take the endpoint down with it. The reverse decision, failing closed, means a Redis incident becomes an outage of your product, which is a strange trade for a guard that exists to prevent outages.

Two modules, because the middleware runtime is a different world

There are two files for one concern, and the split is not aesthetic. The limiter definitions are imported by the proxy, which runs in the middleware runtime where next/headers does not exist. So the helper that reads the IP from inside a server action lives in its own module, and the proxy reads the address off the request object directly.

Put the convenience helper in the same file as the limiters and the whole middleware build fails with an error that names neither of them.

One write per check, not two

The limiter library can record analytics for you. We turned it off:

// No analytics. It costs an extra write on every limit() call, doubling the
// command count on the hottest path in the app, for data nothing in this repo
// ever reads.
analytics: false,
Enter fullscreen mode Exit fullscreen mode

Two reasons, and the second one is the one worth remembering. First, nothing reads it. Second, the proxy never awaits the pending promise the library returns, so on a serverless platform that write was liable to be torn down mid-flight anyway. An analytics write you do not await on a platform that freezes the process after the response is not data collection, it is a coin toss that costs money.

A 429 is a page, not a JSON blob

When a human hits a limit, they get a real page at /too-many-requests that says what happened and what to do. It is a route like any other, which means it also had to be added to the public route allowlist, or the auth gate would redirect the rate limited user to the login page: the one screen guaranteed to make them try again immediately.

It is also kept out of robots.txt. A URL whose entire content is "you are doing that too much" has no business in a search index.

See it for yourself

Top comments (0)