On 9/9/26 23:35, Jim Fehlig via Devel wrote: > Back to thinking about this a bit more... > > On 9/1/26 3:35 AM, Daniel P. Berrangé wrote: > > [snip] > >> >> The whole TERM, wait 10 seconds, KILL, wait 30 seconds approach was >> designed from the POV that a normally behaving QEMU will "die" very >> quickly. IOW, any scenario where we reached the KILL stage was almost >> certainly a broken QEMU/kernel in some respect. > > Should we even be sending the KILL to a QEMU undergoing a "graceful" shutdown > (e.g. 'virsh shutdown' or 'systemctl poweroff' within the guest)? It's fine > for > a destroy operation, but sending KILL to a QEMU in the middle of a proper > cleanup/exit seems to open the window for corruption inside the guest > filesystem. > >> >> Clearly this is no longer a valid assumption. When "normal" behaviour >> or QEMU no longer matches libvirt's default mgmt action behaviour >> then I don't think a global qemu.conf setting or a per-VM setting is >> the ideal approach. >> >> We need to ensure libvirt "does the right thing" out of the box, as >> best as we can. >> >> IMHO, this suggests we need to dynamically increase our wait time >> before KILL based on the guest RAM size. eg Add 5 seconds for each >> 100 GB of small page RAM. I pulled that number out of the air, >> you would need to pick something better based on a typical system, >> plus some buffer/fuzz. > > We could attempt to calculate a "safe" wait time before sending KILL, but I > wonder if it's even needed in the graceful shutdown case? Would it be better > to > send TERM only? Doing so could result in behavior change. E.g. processes that > previously needed the KILL to terminate would now remain in a "in shutdown" > state. A subsequent destroy operation would be needed to issue the KILL. > >> >> Also I've noticed that TDX guests are painfully slow to teardown, >> even with tiny RAM sizes. So we might need to increase wait times >> even more when using TDX. > > I hadn't considered CoCo guests. An ill-timed KILL sent during their > cleanup/exit could be more undesirable than non-CoCo guests. > > Regards, > Jim >
I was surprised to learn that libvirt sends SIGKILL in the graceful shutdown case as well after that grace time. I think it should be up to the user (or higher level platform) to decide when it is fed up waiting and issue a virsh destroy separately. So FWIW I would be in favor of not sending SIGKILL to a VM that is gracefully shutting down without user/platform requesting to do so. Ciao, CLaudio
