Threat context security teams are living through
API abuse in 2026 rarely looks like a single noisy IP hammering one URL. Credential stuffing, data export scraping, and AI-assisted reconnaissance share clean hosting, residential proxies, and stolen session tokens. Recent reporting on Chinese operators impersonating US officials for AI-related cyber espionage shows how attackers blend social engineering with automated tooling that can generate high volumes of plausible API calls. Agentic security products, including intelligence layers built for MCP-style automation, also increase legitimate burst traffic that must be distinguished from hostile automation.
Edge gateways still matter. They absorb volumetric spikes and cheap scanner noise. They fail as a substitute for endpoint-aware quotas when the valuable actions sit deeper in the application: password reset, token refresh, bulk export, webhook registration, and admin invite. RemoteThreat’s emphasis on testing what happens after defenses fail maps directly here. When the edge limiter trips, scrapers shift to slower concurrency across many source addresses and keep pulling records until object-level budgets intervene.
Email-borne social engineering still seeds the first compromised key or cookie. Synthetic media raises the quality of those lures, yet the downstream damage shows up as authenticated API volume that reputation feeds treat as normal SaaS traffic. Rate limiting strategy is therefore an identity and object-scope problem as much as a packets-per-second problem.
What “good” rate limiting looks like in practice
Treat rate limiting as a stack of controls with different clocks and different failure modes:
- Edge connection and request caps protect infrastructure capacity and buy time during floods.
- Identity and key quotas bound what a principal may do across endpoints, including after IP rotation.
- Endpoint and object budgets protect high-impact actions such as exports, search, and privilege changes.
- Abuse progression moves from soft throttle to challenge to temporary deny to key revocation, similar in spirit to blacklist escalation patterns seen in cloud malware control planes.
Algorithm choice should match the risk of the action. Token buckets suit bursty interactive clients. Sliding windows suit fair accounting over fixed intervals. Fixed windows are simple to operate and create boundary spikes at window resets, which attackers exploit against login and OTP endpoints. Leaky buckets smooth egress toward origin services when your concern is backend saturation rather than client fairness.
A concrete multi-tenant SaaS scenario
A B2B platform exposes /v1/search, /v1/export, and /v1/auth/token. A partner integration uses one API key shared by a middleware fleet. A scraper steals that key from a misconfigured CI variable and fans out through residential proxies.
An edge rule of 100 requests per minute per IP looks healthy in the dashboard. The scraper stays under that cap on every address while /v1/export returns full customer lists. The fix is layered: per-key sustained and burst limits, a daily export row budget, concurrent export job limits, and anomaly alerts when a key’s ASN diversity or User-Agent entropy jumps within one hour. Partner traffic receives a higher negotiated quota with explicit scopes. Stolen-key traffic loses export scope first, then loses the key.
Operator checklist before the next abuse wave
Use this as a pre-incident and quarterly review list for cybersecurity and platform teams:
- Inventory high-impact endpoints and assign each a business cost unit (login attempt, password reset, export row, webhook create, admin invite).
- Define principal keys for limiting: IP only at the edge; API key, OAuth client, user ID, tenant ID, and device or session ID inside the app.
- Separate burst from sustained budgets so interactive UI traffic survives while bulk automation stays bounded.
- Scope quotas to objects so a healthy request rate cannot empty a tenant’s dataset through export or search pagination.
- Align partner SLAs with measurable scopes and document what happens when a key is suspected compromised.
- Instrument limit outcomes: HTTP 429 counts, Retry-After adherence, silent drops, challenge issuance, and origin latency after throttle.
- Test fail-open versus fail-closed for the limiter store (Redis, edge KV, gateway). Decide which endpoints may degrade open and which must refuse under control-plane loss.
- Run an after-defense drill: force edge 429s and verify scrapers still hit object budgets, key revocation, and SOC paging paths.
- Correlate auth and API logs so password spray, token refresh storms, and export spikes share one timeline for the on-call analyst.
- Publish client guidance with backoff expectations, idempotency keys, and contact paths for legitimate burst exceptions.
Reference starting points for common endpoints
Tune these to your baseline traffic; treat them as conversation starters with product owners, not universal constants:
- Login and token issue: tight per-IP and per-account sustained limits, short bursts only, progressive delays after repeated failures, shared counters across related routes (login, OTP, password reset).
- Search and list APIs: per-principal request limits plus result-row or pagination-depth budgets.
- Export and report jobs: low concurrency, daily row or byte caps, async jobs with audit trails.
- Webhooks and integrations: registration rate limits, destination allowlists, and per-tenant delivery budgets.
- Admin and IAM APIs: near-human rates, step-up authentication, and immediate paging on burst.
Implementation details that hold up under abuse
Prefer centralized decisioning with local enforcement. Gateways enforce coarse limits close to the client. The application or a dedicated policy service enforces identity and object budgets with the same request ID across layers so SOC joins stay simple. Return consistent 429 responses with Retry-After for cooperative clients. Reserve silent tarpits for clearly hostile scanner paths where you accept operational opacity.
Store design matters. Sliding windows need careful atomic increments. Token buckets need reconciled refill math across replicas. Clock skew between edge PoPs creates uneven enforcement; pin critical auth limits to a strongly consistent store even if global read APIs use eventual counters. For multi-region active-active APIs, decide whether quotas are regional or global. Global export budgets protect data better. Regional request caps protect local capacity.
Authenticated abuse deserves different playbooks than anonymous scanning. A stolen key that stays under per-IP edge caps still exhausts tenant data. Pair rate limits with scope reduction, key rotation hooks, and session revocation. When AI agents and MCP-style tooling call your APIs, require distinct client identities per agent workload so one runaway loop cannot inherit a human user’s full quota.
Implementation pitfalls that keep showing up in IR retrospectives
Counting only at the edge. Per-IP or per-connection caps miss distributed scrapers and shared NAT from corporate egress. Add principal and tenant counters for anything that returns sensitive records.
One global bucket for every route. Login, read, and export compete for the same tokens. Attackers burn the bucket on cheap GETs, then ride residual capacity into expensive POSTs, or the reverse: they starve interactive users by pounding export.
Fail-open limiters on auth and IAM. When Redis or the policy service times out, defaulting to allow keeps the site up and opens password spray. Fail-closed or fail-to-challenge on credential paths; fail-soft on static public content if you must protect availability.
Ignoring Retry-After and client backoff. Official SDKs that retry immediately amplify load during an incident. Publish backoff, enforce jitter, and monitor clients that ignore 429 semantics.
No object-cost accounting. Request-per-minute charts look green while pagination walks an entire customer table. Bill and limit by rows, bytes, or job minutes for export-class APIs.
Partner exceptions without expiry. Temporary raised limits for migrations become permanent shadow capacity for stolen keys. Every exception needs an owner, a scope, and a review date.
Alerting only on 429 volume. Sophisticated scrapers stay just under thresholds. Alert on shifts in ASN diversity, new device fingerprints, export job shape, and cross-endpoint sequences that match stuffing or harvesting.
Skipping the post-limit drill. Teams celebrate the first green limiter dashboard and stop. Schedule exercises where edge limits fire and analysts prove key revocation, object budgets, and customer comms still work.
Build rate limiting as layered economic control over actions attackers care about. Edge caps protect the platform. Endpoint quotas and identity budgets protect the data. The organizations that weather credential abuse and scraper waves are the ones that rehearse both layers before the next stolen key shows up in CI logs.