Hi Job, Saku, I like the draft, it solves a few problems with the existing RTBH based approach to tackling DDoS attacks:
* It doesn’t complete the DDoS attack so some genuine traffic may still get through. * Visibility isn’t lost of the attack (is it still on going, how big is it, what kind of attack pattern is being used, etc.). * It could de-risk / remove the validation issues we have today with trying to validate host routes for RTBH because I could just downgrade a /24 or /48, so we can stick to our existing route filtering practices and not have to do anything weird to try and validate host routes. But I also see several problems that would need to be solved before I could implement this. Do either of you already have ideas on these (maybe you’ve thought of these already?): * If I downgrade a /24 rather than a host route, then other customers could be affected by the DDoS, so that might not be that helpful actually. Then I’m back to downgrading host routes, and the major problem with RTBH is the validation part, not dropping the traffic. And this draft doesn’t change anything about the validation of host routes. * With RTBH, we can auto detect the DDoS and inject an RTBH route. No need to wake up my on-call engineer. If I downgrade a prefix, I’m receiving all the attack traffic still, meaning I will have congested core links. But this would be fine because only the low priority attack traffic is being dropped due to the congestion, not genuine traffic. However, in order not to wake my on-call engineer due to “packet loss on core link” alerts, I need my NMS to alert when any of QoS queues 0,2-7 have packet loss (but not for packet loss in queue 1). This level of detail is not available from all devices via gNMI and not supported by all NMS. This could be tricky. * The barrier to entry for this is quite a bit higher than RTBH. RTBH requires a bit of BGP policy and off you go. Technically, you could say that DOWNGRADE also only requires a bit of BGP policy too, to set the QoS class, but actually, I’ll bet you a round of drinks many networks haven’t got QoS deployed and/or it doesn’t work very well on their devices (we’ve opened 3 separate bug cases with Arista trying to get basic QoS to work). So if you don’t already have QoS working on your network, I’d say the barrier to entry is quite a bit higher. * The main reason for using RTBH is when the attack bandwidth is simply too high (there is collateral damage to other customers because I’m out of capacity, so I need to drop traffic to this prefix to protect my other customers). Because DOWNGRADE doesn’t stop the traffic, the congestion is still present across multiple ASNs. I can de-prioritise the traffic in my network, but my direct peer who doesn’t support DOWNGRADE, their link to me still congests, which affects all customers relying on that link. With RTBH, I can send the RTBH route to my direct peer, they don’t support RTBH, they forward it on to their peer, who does support RTBH, and they drop the traffic, which saves the link capacity between me and my direct peer. With DOWNGRADE, my direct peer doesn’t support it, they forward the DOWNGRADE route to their peer, they de-prioritise the traffic, but the capacity of their physical link to my direct peer is bigger than the capacity of the physical link between me and my direct peer, so the amount of traffic received at the link with my direct peer still causes it to congest and my other customers are still suffering. I would need to fall back to RTBH. Maybe DOWNGRADE becomes an option before RTBH? One could try to use DOWNGRADE, if this doesn’t improve the situation enough we can RTBH. Any thoughts/feedback are appreciated. Cheers, James.
signature.asc
Description: OpenPGP digital signature
_______________________________________________ GROW mailing list -- [email protected] To unsubscribe send an email to [email protected]
