← writing

oct 8, 2026 · 15 min read

Why Rate Limiting Gets Complicated When You Scale Node.js

Inbound rate limiting is a one-liner. Outbound rate limiting, where your servers have to respect someone else's limits, gets messy the moment you run more than one instance. Here is why, and what you can do about it.

Rate limiting seems very easy to implement for one application. Thanks to express-rate-limit, developers can now add a rate limiter to their application with just one line. Some put a bit more effort in and configure the limiter according to their requirements. However, that is just for the inbound rate limiter.

There are two kinds, and it helps to name them before going further.

Inbound rate limiter: the rate limiter we use to limit or prevent abuse from incoming requests (i.e. requests from users or clients to our server).

Outbound rate limiter: the rate limiter we use to limit the requests we are sending from our servers to third-party servers (or to our own other services).

The outbound rate limiter is a very interesting and quite brain wrecking topic in system design. With inbound limiting, when a request goes over the limit you reply with a 429 and you are done. It is the client's problem now. With outbound limiting, the requests that go over the limit are your requests. You can't just drop them, because something in your product is waiting for them to go through. So you have to hold them, schedule them, and release them at the right pace.

Where outbound limiting shows up

Outbound rate limiting is also used in many microservices architectures where each service maintains its own rate limiter. This is needed when, according to the business logic, it is possible that one service might bombard another service with requests. However, it is quite rare to see each service maintaining its own rate limiter, as it invites excessive overhead for the developers. Some avoid it using other architectural patterns in microservices, like putting a message queue between the services so the consumer pulls work at its own pace, or adding a circuit breaker so a struggling service gets some room to recover.

However, when you are querying third-party APIs, it is very necessary to respect their rate limits (unless you don't care about the restrictions). The provider will enforce them whether you plan for it or not.

A quick brief example

You have a product which calls the Meta WhatsApp Cloud API. For each registered business phone number, Meta supports up to 80 messages per second by default, and goes up to 1,000 per second only once you qualify for the automatic upgrade (Meta's throughput docs). Suddenly there is a spike, maybe a tenant fires off a campaign, and your backend has no implementation of throttling the API calls to Meta. Every request from the frontend turns into an immediate call to Meta, you go past the limit, and you start getting rate limit errors back (error code 130429 in Meta's case) until you drop under the limit again. Those messages fail, and now you have to figure out which ones to retry.

Now consider it with throttling. Even in a burst, you do not call the Meta API right when you get the request from the frontend. You wait till you get a slot (usually some milliseconds or seconds under normal conditions). So instead of bombarding the Meta API, you throttle and release the API calls while respecting the rate limits. The burst still happens on your side, but Meta only ever sees a steady 80 per second.

One more thing worth knowing about this particular API: Meta counts both inbound and outbound messages towards that throughput, and there is a separate "pair rate limit" if you send too many messages to the same user in a short time. So your limiter is your best guess at staying under their limit, not a guarantee. We will come back to that.

How do we add outbound rate limiting in a single Node.js instance

There are multiple ways of doing it.

1. An in-memory counter

The first is keeping track of requests using an in-memory counter. Count how many calls you made in the current window (say, the current second). If the count is under the limit, go ahead. If not, wait until the window resets and try again.

const LIMIT = 80;          // calls allowed
const WINDOW_MS = 1000;    // per second
 
let windowStart = Date.now();
let count = 0;
 
const sleep = (ms) => new Promise((r) => setTimeout(r, ms));
 
async function acquire() {
  while (true) {
    const now = Date.now();
    if (now - windowStart >= WINDOW_MS) {
      windowStart = now;
      count = 0;
    }
    if (count < LIMIT) {
      count++;
      return;
    }
    // wait for the current window to end, then check again
    await sleep(WINDOW_MS - (now - windowStart));
  }
}
 
async function sendWhatsAppMessage(payload) {
  await acquire();
  return fetch(META_URL, { method: "POST", body: JSON.stringify(payload) /* ... */ });
}

This is a fixed window counter. It works, and because Node.js runs your JavaScript on a single thread, you don't even have to worry about two callers incrementing count at the same time. The catch is at the window boundary: you can send 80 calls at the very end of one second and 80 more at the very start of the next, so 160 calls land within a few milliseconds. Some APIs don't mind this. Some do.

2. Spacing the calls out

If you want to avoid bursts entirely, don't count, schedule. Give every call a time slot that is 1000 / limit milliseconds after the previous one.

class Throttle {
  constructor(perSecond) {
    this.gap = 1000 / perSecond;
    this.nextSlot = 0;
  }
 
  async acquire() {
    const now = Date.now();
    const slot = Math.max(now, this.nextSlot);
    this.nextSlot = slot + this.gap; // reserve the slot before awaiting
    const wait = slot - now;
    if (wait > 0) await new Promise((r) => setTimeout(r, wait));
  }
}
 
const metaThrottle = new Throttle(80);
await metaThrottle.acquire();

Each caller reserves its slot synchronously and then sleeps until it arrives, so calls come out in order and evenly spaced. This is basically a leaky bucket. If you want to allow some burst but still cap the average rate, a token bucket is the usual middle ground: the bucket refills at a fixed rate up to some capacity, every call takes a token, and when the bucket is empty you wait for the next refill.

3. Use a library

You don't have to write any of this yourself:

  • p-queue lets you set concurrency, intervalCap, and interval, e.g. at most 80 tasks per 1000 ms.
  • rate-limiter-flexible has RateLimiterMemory plus a RateLimiterQueue wrapper that makes callers wait for a token instead of rejecting them, which is exactly the outbound behavior we want.
  • Bottleneck was designed to solve this same issue, with minTime, maxConcurrent, and a "reservoir" for token-bucket-style limits.
import { RateLimiterMemory, RateLimiterQueue } from "rate-limiter-flexible";
 
const limiter = new RateLimiterMemory({ points: 80, duration: 1 });
const queue = new RateLimiterQueue(limiter, { maxQueueSize: 10_000 });
 
await queue.removeTokens(1); // resolves when it's our turn
await callMeta();

All good so far. One process, one counter, one source of truth.

Why it can not be scaled

Everything above keeps its state in the memory of one Node.js process. The moment you run more than one process, each one has its own counter, and none of them knows about the others.

Say you deploy your backend behind a load balancer with 4 instances, each throttled to 80 calls per second. Each instance faithfully stays under 80. Together they send up to 320 per second to Meta, and Meta only allows 80 for that phone number. Your rate limiter is working perfectly and you are still getting rate limited.

This bites in more places than people expect:

  • Horizontal scaling. Multiple containers, pods, or VMs behind a load balancer.
  • Clustering on one machine. PM2 cluster mode or Node's cluster module spawns several processes. Same machine, separate memory, separate counters.
  • Autoscaling. The number of instances changes during exactly the kind of traffic spike you were trying to handle.
  • Background workers. Your API server and your queue workers might both call the same third-party API. Two different codebases, two different limiters, one shared limit.
  • Limits per key, not per app. Meta's throughput is per business phone number. In a multi-tenant product, every tenant's number has its own budget, so you need one limiter per phone number, and all instances have to agree on each of them.

The quick fix people reach for is to divide the limit: 4 instances, so each gets 80 / 4 = 20 per second. It works until it doesn't. When you scale to 6 instances, you are back over the limit unless every instance is reconfigured. When you scale down to 2, you are only using half your budget. And the load balancer doesn't send traffic evenly to a specific tenant's key, so one instance might be sitting on a queue of 500 messages for tenant A, throttled to 20 per second, while the other three are idle with 60 per second of unused budget for the same number.

The real problem is that the rate limit is a shared resource, so the state that tracks it has to be shared too.

Possible solutions

1. A shared Redis instance to track the requests

The most common answer is to move the counter out of process memory and into a shared Redis instance that every instance talks to. Redis is fast, single-threaded for command execution, and has atomic operations, which is exactly what a shared counter needs.

Fixed window with INCR. The simplest version is the pattern described in the Redis INCR docs: build a key from the limiter name and the current second, INCR it, set an expiry on it, and compare the result to the limit. Since INCR is atomic, two instances can never both read "79" and both decide to go. It has the same window-boundary burst as the in-memory version.

Sliding window. To fix the boundary burst, you can store a timestamp for each call in a sorted set (ZADD), remove the ones older than the window (ZREMRANGEBYSCORE), and count what's left (ZCARD). This is accurate but stores one entry per call, which adds up at high rates. A cheaper approximation is to keep counters for the current and previous window and weight the previous one by how much of it still overlaps.

Token bucket in a Lua script. For outbound throttling this is the one I'd reach for. A token bucket needs to read the current state, refill it based on elapsed time, take a token, and write it back. If you do that with separate Redis commands from Node.js, two instances can read the same state at the same time and both take the "last" token. Running it as a Lua script fixes that, because Redis runs the whole script atomically.

The script below doesn't just say yes or no. It reserves a slot and tells the caller how long to wait for it, which is what we want for outbound calls: nobody gets rejected, everybody gets a turn.

-- KEYS[1]   bucket key, e.g. "rl:meta:<PHONE_NUMBER_ID>"
-- ARGV[1]   capacity (max burst)
-- ARGV[2]   refill rate, tokens per second
-- ARGV[3]   max wait in ms before we give up
local capacity = tonumber(ARGV[1])
local rate     = tonumber(ARGV[2]) / 1000   -- tokens per ms
local max_wait = tonumber(ARGV[3])
 
-- use Redis's clock, not the caller's, so instance clock skew doesn't matter
local t   = redis.call('TIME')
local now = tonumber(t[1]) * 1000 + math.floor(tonumber(t[2]) / 1000)
 
local state  = redis.call('HMGET', KEYS[1], 'tokens', 'ts')
local tokens = tonumber(state[1]) or capacity
local ts     = tonumber(state[2]) or now
 
tokens = math.min(capacity, tokens + (now - ts) * rate)
tokens = tokens - 1   -- reserve a token, even if that takes us below zero
 
local wait = 0
if tokens < 0 then
  wait = math.ceil(-tokens / rate)
  if wait > max_wait then
    return -1          -- backlog too long, let the caller decide
  end
end
 
redis.call('HSET', KEYS[1], 'tokens', tokens, 'ts', now)
redis.call('PEXPIRE', KEYS[1], math.ceil(capacity / rate) + max_wait)
return wait

And on the Node.js side, with ioredis:

import Redis from "ioredis";
import { readFileSync } from "node:fs";
 
const redis = new Redis(process.env.REDIS_URL);
redis.defineCommand("reserveToken", {
  numberOfKeys: 1,
  lua: readFileSync("./token-bucket.lua", "utf8"),
});
 
const sleep = (ms) => new Promise((r) => setTimeout(r, ms));
 
export async function acquireMetaSlot(phoneNumberId) {
  const wait = await redis.reserveToken(
    `rl:meta:${phoneNumberId}`,
    80,      // capacity
    80,      // tokens per second
    30_000,  // give up if the wait is over 30s
  );
  if (wait < 0) throw new Error("Rate limit backlog too long");
  if (wait > 0) await sleep(wait);
}

Every instance calls acquireMetaSlot() before talking to Meta, every instance sees the same bucket, and the key is per phone number, so each tenant gets their own budget. Scale to 2 instances or 20, the combined rate stays at 80 per second.

If you don't want to own a Lua script, rate-limiter-flexible has a RateLimiterRedis that you can drop into the same RateLimiterQueue from earlier. Just keep in mind that the waiting in RateLimiterQueue still happens per process; only the counting is shared.

There are a few things to be aware of with this approach:

  • Redis is now in the critical path. If Redis is down, do you stop sending, or fall back to a local limiter with a conservative rate? Decide up front.
  • Every call costs a round trip to Redis. Usually under a millisecond in the same network, but it is not free.
  • Waiting requests still hold resources. An HTTP request sleeping for 10 seconds is holding a connection open and the frontend is staring at a spinner. For bursts that are more than a few seconds long, you probably want the next approach.

There are other algorithms too, like GCRA (the generic cell rate algorithm), which gives you token-bucket behavior while storing a single timestamp per key. Brandur Leach's write-up on GCRA and redis-cell is a good place to start. Stripe's post on scaling your API with rate limiters is also worth reading for how this looks in production, even though it's written from the inbound side.

2. A queue with a global rate limit

Instead of having every request wait for a slot inline, take the outbound calls off the request path entirely. The API handler pushes a job to a queue and returns immediately ("message accepted"). Workers pick jobs off the queue and call the third-party API at a controlled rate. The result comes back through a webhook, a status field in your database, or a websocket.

BullMQ makes this easy because its rate limiter is backed by Redis and is global across all workers on the same queue:

import { Worker } from "bullmq";
 
const worker = new Worker(
  "whatsapp-outbound",
  async (job) => sendToMeta(job.data),
  {
    connection,
    limiter: { max: 80, duration: 1000 }, // 80 jobs/sec across ALL workers
  },
);

From the BullMQ docs: if you have 10 workers for one queue with these settings, still only 80 jobs will be processed per second. You can scale workers for reliability without breaking the limit.

This is the approach that holds up best under real spikes. The burst is absorbed by the queue, which can grow to millions of jobs without holding a single HTTP connection open, and the third-party API sees a flat rate.

The catch is per-key limits. BullMQ's limiter applies to the whole queue. Per-group rate limiting (one limit per tenant or phone number on a shared queue) is part of BullMQ Pro, the paid version. Without it, you can either create one queue per phone number, or keep one queue and call the Redis token bucket from the previous section inside the worker before each send.

3. Bottleneck with Redis clustering

Bottleneck, which I mentioned earlier, also supports this out of the box. Set datastore: "ioredis" (or "redis") with the same id on every instance, and all the limiters share their state through Redis, so the limits apply to the whole cluster. Its API is nice to work with and it covers both single-instance and multi-instance with the same code.

Unfortunately, it is not actively maintained anymore. The latest version on npm is still 2.19.5, released years ago, with open issues and pull requests piling up on the repo. Some forks of it are maintained by individual developers, so if you like its API, it's worth checking for a fork that's still getting updates before you depend on it. For a new project, I'd look at rate-limiter-flexible or BullMQ first.

4. A single egress service

The last option is architectural: put all calls to a given third-party API behind one service (or a small group of instances that share a limiter). Your other services don't call Meta directly; they call your egress service, which owns the rate limit, the retries, and the credentials. If you are already using a service mesh, Envoy has a global rate limit service you can put in front of outbound traffic too. This is more infrastructure, but it means exactly one place in your system knows about Meta's limits.

You still need to handle 429s

Whatever you pick, your limiter is a model of the provider's limit, not the limit itself. The provider might count things you can't see (Meta counts inbound messages toward throughput), the limit might change, or another system sharing the same API key might be sending too. So even with a perfect limiter, handle rate limit responses properly:

  • If the response includes a Retry-After header, respect it.
  • Otherwise retry with exponential backoff and jitter, so all your instances don't retry at the same moment and cause a second spike. AWS has a good post on exponential backoff and jitter.
  • If you're using BullMQ, worker.rateLimit(ms) together with throw Worker.RateLimitError() pauses the whole queue for that long, which is what you want when the provider tells you to slow down. A single job retrying on its own doesn't help if the other 79 per second keep hitting the API.
  • Set your limiter a little below the documented limit. Running at 70 when the limit is 80 costs you a bit of throughput and saves you a lot of retries.

Wrapping up

On one instance, outbound rate limiting is a counter or a queue in memory. Once you scale, the limit becomes a shared resource and needs shared state: a Redis-backed limiter for inline calls, a globally rate limited queue for anything that can be asynchronous, and proper 429 handling on top of either. For third-party APIs like Meta, where the limit is per phone number, design for per-key limits from the start, because retrofitting that later is the brain wrecking part.

References