At 09:14 on a Tuesday, a regional bank’s NOC watches ISP bandwidth charts stay inside normal bands while branch VPN users drop in clusters. The scrubber portal shows clean transit. Inside the data center, the active firewall’s connection table sits at 98 percent and new SYN packets start getting silently discarded. Help desk tickets mention Check Point VPN session failures and WordPress admin timeouts on a partner CMS. The flood never saturated the pipe. It filled the state tables that still decide who gets a session.
That pattern is showing up more often as botnets grow from ordinary exploit waves. Critical WordPress code-execution bugs and VPN gateway remote flaws pull fresh nodes into reflection and application floods within days of disclosure. SMB shops already juggling AI-driven ops still face classic volumetric and state-exhaustion pressure in overdrive. DDoS protection that only watches megabits will miss the failure mode that takes you offline first.
What actually fails when the pipe still looks healthy
Modern multi-vector campaigns mix three pressure points:
- Volumetric reflection (DNS, NTP, CLDAP, memcached) aimed at saturating transit or forcing expensive scrubbing.
- Protocol and state exhaustion (SYN floods, ACK floods, TCP connection races) aimed at firewalls, load balancers, and VPN concentrators.
- Application-layer floods against login, search, checkout, and API routes that burn CPU and backend pools while edge bandwidth stays flat.
Your scrubber can report healthy transit while origin TLS terminators and stateful firewalls hit conntrack, NAT, or SSL session limits. Teams that treat “clean scrubber” as “attack contained” lose minutes that matter for customer-facing services and for any lateral activity that rides the same noise window.
Detection steps that separate bandwidth from session pressure
1. Pair link metrics with middlebox state
Build a single dashboard that shows, for every edge device that holds state:
- Inbound and outbound bits and packets per second
- Current and peak concurrent connections / conntrack entries
- SYN backlog and SYN cookie activation counters
- New connections per second versus established sessions
- CPU on SSL or DPI engines, separate from forwarding ASIC load
Alert when connection occupancy crosses 70 percent of licensed or configured maximum, even if bandwidth sits under 40 percent. That threshold gives operators time to engage scrubbing or SYN cookies before silent drops begin.
2. Classify the vector from the first five minutes of telemetry
Use packet samples or flow exports to label the opening wave:
- High pps, modest payload, many unique sources → likely botnet or reflection toward volumetric or SYN pressure.
- Few sources, high concurrency against one VIP → likely direct state exhaustion or rented cloud flood.
- HTTP floods with valid TLS and rotating User-Agents against auth or search → application-layer; scrubbing alone will not save the app pool.
Record which VIP, ASN families, and destination ports absorb the hit. Map those ports to business services so the incident commander can prioritize without waiting for a full PCAP review.
3. Watch auth and VPN paths during exploit-driven botnet growth
When industry advisories land for VPN gateways or widely deployed CMS platforms, expect botnet capacity to jump. Raise sensitivity on:
- Abnormal new-session rates to VPN and MFA portals
- HTTP 429 and 5xx spikes on login and password-reset routes
- Sudden growth in unique source IPs hitting the same URI with low bytes per request
Correlate those signals with your DDoS sensors so a “performance incident” gets treated as a security event while the wave is still forming.
4. Prove your scrubber path with a tabletop that includes state failure
Once a quarter, walk the on-call chain through a scenario where transit is clean and the firewall conntrack graph is red. Confirm who can:
- Enable SYN cookies or tighten embryonic connection limits
- Shift Anycast or DNS to the scrubbing provider
- Apply ISP remote-triggered blackhole only for confirmed attack prefixes, with a written unwind step
- Raise origin connection budgets or shed noncritical listeners
Document the order. State exhaustion incidents punish hesitation more than bandwidth floods do.
Response actions that restore service without guessing
Immediate containment (first 15 minutes)
- Declare the vector class from the dashboard above and assign one owner for network mitigation and one for application capacity.
- Stabilize stateful gear: enable SYN cookies, lower half-open timeouts, raise or redistribute connection limits, and shed monitoring or nonproduction listeners that share the same table.
- Engage scrubbing with a concrete ask: “Absorb UDP reflection to VIP X; pass established TCP; challenge HTTP to /login and /api.” Vague “mitigate the DDoS” tickets delay useful filters.
- Protect identity surfaces: add CAPTCHA or step-up challenges on login, throttle password reset and token endpoints, and verify VPN concentrators are patched for any actively exploited gateway flaw in the same window.
Stabilization (next hour)
- Confirm scrubber to origin allowlists and health checks so return traffic and origin probes are not treated as attack residue.
- Scale application pools for L7 pressure; bandwidth scrubbing will not fix exhausted workers.
- Keep an auth timeline open. Sustained floods still coincide with credential stuffing and session abuse; password resets alone leave live sessions intact, as recent identity-focused campaigns against large firm sets have shown.
- Capture five-minute flow summaries and top talkers for later ISP and law-enforcement packages; do this while the attack is live, not after TTL expires.
Hardening that pays off before the next wave
After service recovers, convert the incident into capacity and policy changes:
- Size conntrack and SSL session tables to peak legitimate concurrency plus a documented flood margin, and alarm on occupancy, not only on drops.
- Place Anycast or DNS cutover runbooks next to the NOC playbook with provider ticket templates that name VIPs, protocols, and challenge modes.
- Segment critical listeners so VPN, DNS authoritative, and public web do not share a single connection budget.
- Rate-limit and cache expensive application routes; reflection amplifiers still exist, but L7 floods against uncached search and login burn the stack that customers feel.
- Track exploit-week botnet risk: when WordPress, VPN, or similar RCEs hit mass scanning, temporarily lower thresholds for new-source pps and auth endpoint concurrency.
A compact control checklist for IT administrators
Use this as a pre-incident scorecard:
- Edge firewalls and load balancers export connection occupancy with paging alerts at 70 percent.
- SYN cookies or equivalent embryonic controls are tested in change windows, not invented during an outage.
- Scrubbing contracts list supported vectors, challenge options, and origin bandwidth commitments in writing.
- Application owners know which URIs must receive bot challenges first under L7 pressure.
- ISP RTBH and BGP communities are documented with a forced review before any /24 or more specific is blackholed.
- Patch SLAs for internet-facing VPN and CMS platforms assume botnet recruitment within days of public exploit.
Takeaways you can apply this week
Walk your primary edge firewall and load balancer tonight and note licensed versus configured connection maxima. If occupancy alerts do not exist, add them before the next volumetric headline arrives. Schedule a 30-minute drill where bandwidth charts stay green and session tables go red, and time how long it takes to engage scrubbing with a vector-specific request. Align that drill with your patch calendar for VPN and CMS surfaces, because those exploit cycles still feed the botnets that make multi-vector floods routine.
DDoS mitigation succeeds when you treat session headroom, scrubbing cutover, and application challenges as one control set. Bandwidth charts alone will tell you the pipe is fine while the devices that grant sessions have already stopped answering.