On 9/9/26 23:35, Jim Fehlig via Devel wrote:
> Back to thinking about this a bit more...
> 
> On 9/1/26 3:35 AM, Daniel P. Berrangé wrote:
> 
> [snip]
> 
>>
>> The whole  TERM, wait 10 seconds, KILL, wait 30 seconds approach was
>> designed from the POV that a normally behaving QEMU will "die" very
>> quickly. IOW, any scenario where we reached the KILL stage was almost
>> certainly a broken QEMU/kernel in some respect.
> 
> Should we even be sending the KILL to a QEMU undergoing a "graceful" shutdown 
> (e.g. 'virsh shutdown' or 'systemctl poweroff' within the guest)? It's fine 
> for 
> a destroy operation, but sending KILL to a QEMU in the middle of a proper 
> cleanup/exit seems to open the window for corruption inside the guest 
> filesystem.
> 
>>
>> Clearly this is no longer a valid assumption. When "normal" behaviour
>> or QEMU no longer matches libvirt's default mgmt action behaviour
>> then I don't think a global qemu.conf setting or a per-VM setting is
>> the ideal approach.
>>
>> We need to ensure libvirt "does the right thing" out of the box, as
>> best as we can.
>>
>> IMHO, this suggests we need to dynamically increase our wait time
>> before KILL based on the guest RAM size. eg Add 5 seconds for each
>> 100 GB of small page RAM. I pulled that number out of the air,
>> you would need to pick something better based on a typical system,
>> plus some buffer/fuzz.
> 
> We could attempt to calculate a "safe" wait time before sending KILL, but I 
> wonder if it's even needed in the graceful shutdown case? Would it be better 
> to 
> send TERM only? Doing so could result in behavior change. E.g. processes that 
> previously needed the KILL to terminate would now remain in a "in shutdown" 
> state. A subsequent destroy operation would be needed to issue the KILL.
> 
>>
>> Also I've noticed that TDX guests are painfully slow to teardown,
>> even with tiny RAM sizes. So we might need to increase wait times
>> even more when using TDX.
> 
> I hadn't considered CoCo guests. An ill-timed KILL sent during their 
> cleanup/exit could be more undesirable than non-CoCo guests.
> 
> Regards,
> Jim
> 

I was surprised to learn that libvirt sends SIGKILL in the graceful shutdown 
case as well after that grace time.
I think it should be up to the user (or higher level platform) to decide when 
it is fed up waiting and issue a virsh destroy separately.

So FWIW I would be in favor of not sending SIGKILL to a VM that is gracefully 
shutting down without user/platform requesting to do so.

Ciao,

CLaudio

Reply via email to