> ## Content Index
> Fetch the complete content index at: https://debugly.dev/llms.txt
> Use this file to discover other available public pages before exploring further.

# The Rate Limit That Counted IPs and Not Accounts
- URL: https://debugly.dev/brute-force-and-the-rate-limit-that-missed/
- Published: 2026-09-12T11:30:00.000Z
- Updated: 2026-09-12T11:30:00.000Z
- Author: Rohit Bhadani
- Tags: Security, Networking, Debugging

The dashboard showed the login rate limit doing its job. A steady trickle of 429s, right at the threshold, exactly as configured. At the same time, the fraud team was locking accounts that had been taken over. Both dashboards were accurate. The rate limit was limiting, and it was irrelevant.

The attack was credential stuffing, a list of leaked username and password pairs tried across many accounts, spread across a botnet. Each source IP made a handful of attempts, well under the per IP limit. The limit counted the wrong unit, so it never fired against the thing that mattered.

This is the most common shape of a rate limit failure, and it is a design error rather than a tuning error. The limit worked as specified. The specification was the bug.

## Why per IP fails against distributed attack

A per IP rate limit assumes the attacker is one IP being too fast. That model defended against the scripts of a decade ago. A modern stuffing campaign buys residential proxies or a botnet, and the same thousand attempts arrive from a thousand IPs, each making one request.

The per IP limiter sees a thousand well behaved clients. The per account reality is one account under a thousand guesses, or a thousand accounts each taking one hit. Neither view trips the threshold, and the aggregate is the attack.

This is the same wrong denominator problem as averages in [p99 and why averages lie](https://debugly.dev/p99-latency-and-why-averages-lie/), applied to security. The per IP average looks calm while the per account tail is being hammered.

## The units that actually matter

A login attempt has several dimensions, and a limiter should count the ones the attacker cannot cheaply distribute.

**The account.** The attacker's scarce target is a specific account. Limiting attempts per account per window directly caps the guesses against it, no matter how many IPs participate.

**The credential pair or password hash.** Stuffing tries known pairs. Counting failed attempts per account, and escalating to a challenge or delay rather than a hard block, raises the cost without locking out the real owner.

**The fingerprint beyond IP.** A combination of user agent, TLS fingerprint, ASN and behaviour gives a coarse identity that is harder to distribute than an IP. It is not reliable alone, but as one axis among several it raises the attacker's cost.

**The source, but smarter.** Per IP is not useless. It catches the clumsy attacker and the single host script. Keep it as one layer, but never as the only layer.

## The shapes that work

The pattern that holds up is layered limits on different units with escalating responses.

A per IP limit catches naive scripts. A per account limit caps guessing on the valuable target. A global limit on the auth endpoint catches the aggregate flood that the distributed attack produces, because the whole campaign is still one endpoint being hit hard, and that aggregate is visible even when no single dimension is.

And the response should escalate rather than binary block. A short delay, then a challenge, then a temporary lockout with notification to the account owner. Hard blocks teach attackers to rotate. Delays and challenges raise cost quietly and keep the real user able to get in with a little friction.

## The trap of locking out the victim

A naive per account limit that hard locks after N failures becomes a denial of service tool. The attacker, who knows the account exists, deliberately trips the lock and keeps the real owner out. This is why escalation to a challenge or to verification via a second factor is preferable to a hard lock, and why lockouts should be short and paired with owner notification.

The defence must be more expensive for the attacker than for the owner. A challenge is an annoyance to the owner and a cost to the bot. A lockout is an annoyance to the owner and free to the bot.

## Measuring whether your limiter is counting right

The check is to look at your limit's denominator during an incident. If you are being stuffed, a per IP limiter's counters stay calm while account level failure counts spike. Plot failed logins per account per hour. A long tail of accounts each taking a handful of failures is the stuffing signature, and it is invisible to any IP scoped metric.

Also watch the ratio of distinct IPs to attempts on the auth endpoint. A healthy login surface has a stable ratio. A stuffing campaign drives distinct IPs up toward the attempt count, because each attempt comes from somewhere new. That ratio is a simple, cheap campaign detector.

## The rule

Rate limits defend the unit they count, and attackers distribute whatever you count. Count the scarce things, the account, the credential, the aggregate endpoint, and use the IP as one coarse layer among them. Escalate with friction rather than hard blocks, so the defence costs the attacker more than the owner.

A limiter that is green during a successful attack is not broken. It is measuring the wrong unit, confidently. The same "the metric is calm while the user suffers" lesson is the tail latency story in [your average latency is fine and your users are not](https://debugly.dev/p99-latency-and-why-averages-lie/).

It is worth stating what a good outcome looks like, because "no successful stuffing" is not realistic and is not the goal. The goal is that each guess costs the attacker time and requests, that a single account cannot be enumerated quickly, and that the aggregate flood is visible and shed. A stuffing campaign that burns a million proxy requests to crack a handful of accounts is a failed campaign, even if it is not a perfectly prevented one. Defence is economics, and the rate limit's job is to make the arithmetic bad, one correctly counted unit at a time.