Chi-Hsuan Huang created HDDS-15997:
--------------------------------------
Summary: QuotaRepairTask scan can hang or report partial counts on
worker failure or interrupt
Key: HDDS-15997
URL: https://issues.apache.org/jira/browse/HDDS-15997
Project: Apache Ozone
Issue Type: Bug
Reporter: Chi-Hsuan Huang
QuotaRepairTask scans each table with one producer thread feeding batches to
worker threads through a bounded queue. Two failure paths are broken:# If all
worker threads die with an exception, the producer blocks forever on
queue.put\(\) because nobody drains the queue. The repair task hangs and the
IN\_PROGRESS flag is never cleared, so no further repair can be triggered.
# InterruptedException is swallowed: both producer and workers restore the
interrupt flag and return normally, so repairCount treats partially scanned
counts as a successful result and submits a wrong quota diff through Ratis.
In addition, on failure the scan must still await all worker threads before
returning, because the DB checkpoint is removed in a finally block and no
worker may outlive it.
These paths are currently hard to hit because the per\-record work never
touches the DB, but HDDS\-15835 \(counting incomplete MPU parts\) adds a
consumer that reads RocksDB per record, making worker failure realistic. This
change is a prerequisite for HDDS\-15835.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]