MitchDrage opened a new issue, #14295:
URL: https://github.com/apache/cloudstack/issues/14295

   ### problem
   
   I encountered an issue where a NAS backup of a quiesced database instance 
failed.
   During investigation I have found four issues. Issues 1 and 3 apply to any 
failed backup whatever the cause; issues 2 and 4 are specific to quiescing.
   I've listed these four issues together as they are somewhat related. If 
you'd prefer them broken out into separate issues, just let me know.
   
   The log messages from the failed backup:
   ```
   ERROR [o.a.c.b.NASBackupProvider] (API-Job-Executor-7:[ctx-5f7ccff5, 
job-44506, ctx-f6a4a8d4]) (logid:3a91f756) Failed to take backup for VM 
i-58-3123-VM: Failed to thaw the filesystem for vm i-58-3123-VM: error: guest 
agent command timed out: guest agent didn't respond to command within '5' 
seconds
   ERROR [o.a.c.b.NASBackupProvider] (API-Job-Executor-7:[ctx-5f7ccff5, 
job-44506, ctx-f6a4a8d4]) (logid:3a91f756) Backup cleanup failed for VM 
i-58-3123-VM. Leaving the backup in Error state.
   ```
   
   
   ### 1. A failed scheduled backup is reported as manual
   The failed backup mentioned above is shown in the backups list as MANUAL, 
despite it being a scheduled backup.
   
   <img width="1698" height="105" alt="Image" 
src="https://github.com/user-attachments/assets/93039e29-856a-4081-bce0-5322f7d95875";
 />
   
   From my investigation, I've found that BackupManagerImpl resolves the 
schedule id before the backup is attempted (line 794) and writes it to the row 
only after a successful result (line 835). On failure the throw at 830 runs 
first, so the row keeps backup_schedule_id = NULL. Two effects:
   
   - listBackups defaults to intervaltype: MANUAL.
   - Retention calls listBySchedule, which filters on backup_schedule_id, so 
the row is never counted toward maxbackups and never rotated.
   
   Fix: persist the schedule id when the backup row is created, so the linkage 
survives any outcome.
   
   ### 2. The guest agent timeout is not configurable
   
   The freeze and thaw calls in nasbackup.sh don't pass --timeout to virsh. 
There doesn't appear to be a way to change it: there are timeouts for Veeam and 
Networker only, and there is nothing in agent.properties or libvirt's 
qemu.conf. On a write-heavy instance a thaw exceeds the default (5 seconds) and 
fails the backup. In my case virsh reached timeout but the thaw completed, so a 
successful operation was recorded as a failure.
   
   Fix: a zone-scoped config key, backup.quiesce.agent.timeout, giving a global 
default and a per-zone override (minimum), and ideally a per-VM override.
   
   ### 3. cleanup() tears down the backup destination while the libvirt job is 
still running
   When the thaw failed, the script called cleanup.
   That gave me 17 minutes of errors on the hypervisor after the script had 
already exited:
   
   ```
   Timed out during operation: cannot acquire state change lock (held by 
monitor=remoteDispatchDomainBlockStats)
   Path 
'/tmp/csbackup.zyIrC/i-58-3123-VM/2026.10.01.16.01.20/datadisk.680baabd-859d-4982-9d33-5cdbe1b12e76.qcow2'
 is not accessible: No such file or directory
   End of file while reading data: Input/output error
   ```
   cleanup() does rm -rf $dest and umount with no virsh domjobabort, while the 
job started by backup-begin was still writing there.
   
   Fix: abort the libvirt backup job before removing the destination.
   
   ### 4. A failed freeze is never thawed, and its error is discarded
   I came across this one unexpectedly - I didn't hit this specifically, so 
it's simply a code review and a strong suspicion based on how it reads.
   
   Current code:
   ```
   local thaw=0
   if [[ ${QUIESCE} == "true" ]]; then
     if virsh ... '{"execute":"guest-fsfreeze-freeze"}' > /dev/null 
2>/dev/null; then
       thaw=1
     fi
   fi
   ```
   thaw is set only when the freeze returns success, and the freeze's stderr 
goes to /dev/null. A freeze that fails and then completes inside the guest 
leaves the filesystem frozen with nothing logged.
   I've checked the rest of the path and I couldn't find any code later that 
would unfreeze it.
   
   Fix:
   
   ```
   local thaw=0
   if [[ ${QUIESCE} == "true" ]]; then
     thaw=1
     if ! freeze_err=$(virsh -c qemu:///system qemu-agent-command "$VM" 
'{"execute":"guest-fsfreeze-freeze"}' 2>&1 > /dev/null); then
       echo "Failed to freeze the filesystem for vm $VM: $freeze_err"
     fi
   fi
   ```
   
   
   ### versions
   
   CloudStack 4.22.1.1, KVM on AlmaLinux 9, libvirt/virsh 11.10.0, QEMU 10.1.0, 
Ceph RBD primary storage, NFS backup repository.
   
   ### The steps to reproduce the bug
   
   As above.
   
   ### What to do about it?
   
   As above.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to