poddm opened a new pull request, #13738:
URL: https://github.com/apache/cloudstack/pull/13738

   ### Description
   
   This PR fixes HA-enabled KVM instances getting stuck in the `Running` state 
after an
   out-of-band stop (e.g. the QEMU process is OOM-killed or crashes on the 
host).
   
   **Actual behaviour (before):**
   When the power-state sync detected that a running, HA-enabled KVM instance 
was no longer
   present on the host, the management server called
   `HighAvailabilityManager.scheduleRestart(vm, true)` from
   `handlePowerOffReportWithNoPendingJobsOnVM()`. That path invokes 
`advanceStop()`, which
   submits a VM work job and **blocks waiting for it to complete**. Because 
this handler runs
   on the `AgentManager-Handler` thread — a context that does not dispatch VM 
work jobs — the
   call blocks indefinitely. The instance was left in `Running`, and the 
follow-up HA
   investigation (host still up, domain already gone) concluded "no need to 
restart".
   Additionally, a libvirt `VIR_DOMAIN_CRASHED` state was not mapped to a power 
state, so a
   crashed domain was reported as `PowerUnknown` and only detected later via 
the "missing VM"
   threshold, delaying recovery.
   
   **Expected behaviour (after):**
   An HA-enabled KVM instance that stops out-of-band is transitioned to 
`Stopped` and an HA
   restart is scheduled promptly, without blocking the AgentManager thread.
   
   **Changes:**
   
   1. Management server — 
`engine/orchestration/.../VirtualMachineManagerImpl.java`
      In the HA out-of-band-stop branch of 
`handlePowerOffReportWithNoPendingJobsOnVM()`,
      replace the blocking `scheduleRestart(vm, true)` with a non-blocking path:
      - release the instance's resources,
      - transition the instance to `Stopped` (`FollowAgentPowerOffReport`),
      - insert an HA work item directly into `op_ha_work` with `Step.Scheduled` 
for the HA
        worker to pick up asynchronously.
      A guard ensures a valid `host_id` is present before enqueuing HA work 
(avoids violating
      the FK constraint), and existing pending HA work is not duplicated.
   
   2. KVM agent — `plugins/hypervisors/kvm/.../LibvirtComputingResource.java`
      Map `DomainState.VIR_DOMAIN_CRASHED` to `PowerState.PowerOff` so a 
crashed domain is
      reported as off immediately instead of being omitted from the report 
(which delayed
      detection via the missing-VM threshold).
   
   The power-state report already confirms the instance is off, so 
re-investigation is
   unnecessary in this path.
   
   <!-- Fixes: # -->
   
   ### Types of changes
   
   - [ ] Breaking change (fix or feature that would cause existing 
functionality to change)
   - [ ] New feature (non-breaking change which adds functionality)
   - [x] Bug fix (non-breaking change which fixes an issue)
   - [ ] Enhancement (improves an existing feature and functionality)
   - [ ] Cleanup (Code refactoring and cleanup, that may add test cases)
   - [ ] Build/CI
   - [ ] Test (unit or integration test code)
   
   ### Feature/Enhancement Scale or Bug Severity
   
   #### Feature/Enhancement Scale
   
   - [ ] Major
   - [ ] Minor
   
   #### Bug Severity
   
   - [ ] BLOCKER
   - [x] Critical
   - [ ] Major
   - [ ] Minor
   - [ ] Trivial
   
   ### Screenshots (if appropriate):
   
   N/A
   
   ### How Has This Been Tested?
   
   - Deployed an HA-enabled KVM instance on a multi-host KVM cluster.
   - Simulated an out-of-band stop by killing the QEMU process on the host 
(OOM-kill scenario).
   - Verified the instance transitions from `Running` to `Stopped`, a row is 
created in
     `op_ha_work` with `Step.Scheduled`, and the HA worker restarts the 
instance.
   - Confirmed the `AgentManager-Handler` thread no longer blocks (no stuck 
power-off handler),
     and that `management-server.log` no longer shows the instance stuck in 
`Running` after the
     investigation.
   - Verified a `VIR_DOMAIN_CRASHED` domain is now reported as `PowerOff` and 
triggers the same
     recovery path without waiting for the missing-VM threshold.
   
   #### How did you try to break this feature and the system with this change?
   
   - Killed the QEMU process while the host remained up (host heartbeat 
healthy) to confirm the
     instance is still recovered rather than being deemed "alive" by the 
investigator.
   - Verified behaviour when `host_id`/`last_host_id` is unavailable: HA work 
is not enqueued
     (guarded) to avoid violating the `op_ha_work` FK constraint.
   - Confirmed no duplicate HA work is scheduled when a pending HA task already 
exists for the
     instance.
   - Verified non-KVM hypervisors (VMware, Hyper-V) and non-Running states are 
unaffected, since
     the branch is guarded by `vm.getState() == State.Running` and the 
hypervisor-type checks.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to