What Changes About Rate Limiting When a Meaningful Share of Your Legitimate Traffic Comes From China
Rate limiting is usually configured once, early, using a threshold that seemed generous at the time. It then sits untouched until support tickets start arriving from users who did nothing wrong.
For services with significant traffic from mainland China, that outcome is more likely than most default configurations account for, and the cause is structural rather than a tuning mistake. The assumption underneath per-IP rate limiting doesn’t hold as well there.
This piece covers why that assumption breaks, how the resulting failure looks from the support queue, and what to change so abuse protection keeps working without catching real users in the same net.
Why standard rate limiting assumes something that isn’t always true
Per-IP rate limiting encodes an assumption: one IP address corresponds roughly to one user, or at most one household. Under that assumption, a threshold like a hundred requests per minute per IP is a reasonable proxy for “one client behaving unreasonably.”
The assumption is approximately true in places where IPv4 addresses are relatively plentiful relative to population, and it’s considerably less true in mainland China. There, the ratio of subscribers to available IPv4 addresses is far less favorable, and the network architecture built in response means shared addressing is the norm rather than an exception.
Once the assumption breaks, the meaning of your threshold changes without you changing anything. A limit that was measuring one user’s behavior starts measuring the aggregate behavior of an unknown number of unrelated users who happen to share an exit point.
Carrier-grade NAT and shared IPs
Carrier-grade NAT, usually shortened to CGNAT, is the mechanism. Rather than assigning each subscriber a unique public IPv4 address, a carrier routes many subscribers through a shared pool of public addresses, translating between the private addresses used internally and the public ones seen from outside. It’s a standardized arrangement rather than an improvisation: RFC 6598 set aside dedicated address space for it, and RFC 6888 documents the behavior operators are expected to implement.
The practical consequence is that your server may see hundreds or thousands of distinct users arriving from a small set of addresses. Worse for rate limiting purposes, the mapping isn’t stable: a single user’s requests can appear from different addresses in the pool over time, while simultaneously other users appear from the address you just saw.
This is particularly relevant for China because internet use there skews heavily mobile, and mobile carriers are among the most aggressive users of CGNAT. Corporate networks, university campuses, and public Wi-Fi produce shared-IP traffic of their own through ordinary NAT, and can sit behind carrier NAT as well. The mobile carrier case is the one most likely to dominate your traffic mix, in a market where reaching those users well is already a routing problem before it becomes a rate-limiting one.
Where this actually breaks rate limiting
The failure is quiet, which is what makes it expensive. Your limiter is working exactly as configured; it’s the interpretation of the input that’s wrong. A hundred requests per minute from an address representing four hundred users is completely normal traffic that looks, to the limiter, like a single client hammering an endpoint.
From the support queue it presents as confusing reports: users describing entirely ordinary usage who are seeing 429 responses, throttling, or CAPTCHA challenges. Because each individual user’s behavior is unremarkable, investigating any one report finds nothing wrong, which is how this often goes unresolved for a long time.
It also concentrates during peak hours, when the largest number of users behind a given address are simultaneously active. So the periods where your service most needs to work well are precisely the periods where legitimate users are most likely to be caught, which is a bad pattern for a consumer-facing product.
What to adjust
The highest-value change is moving your limits off raw IP wherever an identity is available. Authenticated user IDs, session tokens, and API keys all identify the actual client making requests, so a limit keyed on them measures the thing you meant to measure regardless of how many users share a network path. For any authenticated endpoint, this is straightforwardly the right approach.
For unauthenticated endpoints where you have no identity to key on, IP-based limiting is unavoidable, so the goal becomes making it less blunt. Set thresholds high enough to accommodate a shared exit point rather than a single household, and treat request pattern as the primary abuse signal rather than volume, so that a suspicious shape is enough to act on while high volume alone is not.
That distinction is the conceptual core of it. High volume from one address is expected behind CGNAT and tells you very little, while an unusual shape indicates automation even at modest volume. Identical request bodies, a perfectly regular interval, and sequential enumeration of IDs are all good signals. Login attempts spread across many accounts is weaker than it looks, since hundreds of unrelated people signing into their own accounts from one CGNAT address produce exactly that pattern; what distinguishes an attack there is the failure rate per account, not the number of accounts touched.
It’s also worth separating your responses by severity. Blocking outright is the harshest option; tar pitting, progressive delays, or a challenge page degrade gracefully and give a legitimate user behind a busy CGNAT address a path forward rather than a wall.
A related wrinkle: geographic and range-level IP blocks
The same dynamics make blunt IP range blocking more costly than it appears. Blocking a single address in a region with heavy CGNAT can silently cut off a large number of unrelated users, and blocking a range can affect a disproportionate share of a country’s mobile subscribers.
This shows up in a specific and avoidable way: an automated abuse-mitigation rule blocks an address after detecting a burst of bad requests, and the collateral effect is that a large group of legitimate users on the same carrier lose access with no explanation. If your mitigation tooling escalates to range blocks automatically, it’s worth reviewing what those rules can actually reach in a CGNAT-heavy region.
A more proportionate approach is to keep automated blocks short-lived, log what they caught so you can review whether the decision was sound, and require human review before anything wider than a single address gets blocked for an extended period.
Wrapping up
Per-IP rate limiting works well when one address means one user. In mainland China, where CGNAT is widespread and usage is heavily mobile, that mapping is much weaker, and a threshold that looks reasonable can end up measuring a crowd instead of an individual.
Shift to authenticated, per-user limits wherever an identity exists, raise and supplement IP-based limits where it doesn’t, and treat request patterns rather than raw volume as your abuse signal. That keeps the protection meaningful while avoiding the quiet failure of blocking the users you were trying to serve.
Thanks for reading! Whether you’re protecting an API or a login form, limits that account for shared addressing hold up far better than limits that assume one address means one person. If you’re looking for fast, no-nonsense infrastructure, V.PS runs KVM VPS on its own networks in 11 cities across 4 continents, with a public looking glass so you can test before you buy. For China-facing workloads specifically, the Performance KVM VPS in Singapore and Tokyo Gen 2 carries CTGNet, CUP, and CMIN2 routing to all three major mainland carriers.
Ready to get started? Deploy a server or contact our team if you need help choosing a plan.
Frequently asked questions about rate limiting and China traffic
What is carrier-grade NAT, in simple terms?
It’s an arrangement where a carrier routes many subscribers through a shared pool of public IP addresses instead of giving each one its own, translating between internal private addresses and the public ones your server sees. The result is that many unrelated users appear to come from the same address.
Why does this matter more for China-facing traffic?
Internet use in China skews heavily mobile, and mobile carriers there make extensive use of CGNAT given the ratio of subscribers to available IPv4 addresses. A meaningful share of legitimate traffic can therefore arrive from a small number of shared addresses.
Is authenticated rate limiting always better than IP-based limiting?
For endpoints where a user is logged in or presenting an API key, yes, because it measures the actual client rather than the network path. IP-based limiting remains necessary for unauthenticated endpoints, where the goal is making it less blunt rather than removing it.
How do I know if I’m blocking legitimate users instead of abusive ones?
Look for support reports describing normal usage alongside throttling or block responses, and check whether the flagged addresses belong to mobile carrier ranges. Blocks that cluster during peak hours, when more users share each address, are another strong indicator.
Should I just raise my rate limits across the board to avoid this?
Raising thresholds helps, but on its own it weakens abuse protection. Pairing a higher IP threshold with a pattern-based signal keeps protection meaningful, since automated abuse is usually identifiable by request shape rather than by volume alone.
Does this issue only affect China-facing services?
No, CGNAT is used by carriers worldwide and the same dynamics apply anywhere shared addressing is common. It’s simply most pronounced for services with a substantial mainland China user base, which is why it’s worth reviewing specifically in that context.
Share