akshaychitneni commented on code in PR #2425:
URL:
https://github.com/apache/datafusion-ballista/pull/2425#discussion_r4112138845
##########
chaos-testing/src/k8s.rs:
##########
@@ -294,6 +302,28 @@ impl K8sCluster {
.map(|_| ())
}
+ /// Total, sustained executor loss *mid-query* (#2029): every executor
dies at
+ /// once and none is rescheduled. Scales the Deployment to 0 first so the
+ /// controller will not recreate the pods, then force-deletes them
+ /// (`--grace-period=0 --force`) so they are SIGKILLed without the graceful
+ /// drain that [`Self::scale_executors`] alone allows — otherwise the
+ /// executors would finish the in-flight query and it would succeed.
+ pub async fn kill_all_executors_hard(&self) -> Result<(), String> {
Review Comment:
Thanks for the feedback. I noticed that the --cascade=orphan deletes the
deployment but leaves the replicaset which reschedules the pods on delete
failing this case. Instead, I have set the `terminationGracePeriodSeconds` to 1
to make sure there is a sigkill following sigterm to replicate this scenario.
Please take another look
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]