nagaboinaramgopal commented on code in PR #14039:
URL: https://github.com/apache/cloudstack/pull/14039#discussion_r3971542597
##########
framework/jobs/src/main/java/org/apache/cloudstack/framework/jobs/impl/AsyncJobManagerImpl.java:
##########
@@ -701,58 +701,58 @@ private int getAndResetPendingSignals(AsyncJob job) {
return signals;
}
- private void executeQueueItem(SyncQueueItemVO item, boolean
fromPreviousSession) {
+ protected void executeQueueItem(SyncQueueItemVO item, boolean
fromPreviousSession) {
AsyncJobVO job = _jobDao.findById(item.getContentId());
- if (job != null) {
+ if (job == null) {
if (logger.isDebugEnabled()) {
- logger.debug("Schedule queued job-" + job.getId());
- }
-
- job.setSyncSource(item);
-
- //
- // TODO: a temporary solution to work-around DB deadlock situation
- //
- // to live with DB deadlocks, we will give a chance for job to be
rescheduled
- // in case of exceptions (most-likely DB deadlock exceptions)
- try {
- job.setExecutingMsid(getMsid());
- _jobDao.update(job.getId(), job);
- } catch (Exception e) {
- logger.warn("Unexpected exception while dispatching job-" +
item.getContentId(), e);
-
- try {
- _queueMgr.returnItem(item.getId());
- } catch (Throwable thr) {
- logger.error("Unexpected exception while returning job-" +
item.getContentId() + " to queue", thr);
- }
+ logger.debug("Unable to find related job for queue item: " +
item.toString());
}
+ _queueMgr.purgeItem(item.getId());
+ return;
+ }
- try {
- scheduleExecution(job);
- } catch (RejectedExecutionException e) {
- logger.warn("Execution for job-" + job.getId() + " is
rejected, return it to the queue for next turn");
+ if (logger.isDebugEnabled()) {
+ logger.debug("Schedule queued job-" + job.getId());
+ }
+ job.setSyncSource(item);
- try {
- _queueMgr.returnItem(item.getId());
- } catch (Exception e2) {
- logger.error("Unexpected exception while returning job-" +
item.getContentId() + " to queue", e2);
- }
+ //
+ // TODO: a temporary solution to work-around DB deadlock situation
+ //
+ // to live with DB deadlocks, we will give a chance for job to be
rescheduled
+ // in case of exceptions (most-likely DB deadlock exceptions)
Review Comment:
Thanks @DaanHoogland, good question.
For context, that TODO predates this PR (it is on main today) and the
restructure only re-indented it, so I kept its intent as is rather than change
it here.
On why it is called temporary: stamping the executing management-server id
via _jobDao.update can hit a DB deadlock, and the workaround gives the job
another chance by returning the queue item to the sync queue to be retried on a
later turn instead of failing it. This PR keeps that behaviour and only fixes
the case where a failed dispatch also fell through and executed the job
immediately, so it ran twice.
For the work forwards, I think the cleaner fix is to address the deadlock at
its source, around the locking and transaction for the sync_queue and async_job
update on dispatch, so the retry is no longer needed. That felt larger than
this bug fix, so I kept it out of scope here, but I would be glad to look into
it as a follow up if you agree that is the right direction.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]