Aryan Gupta created HDDS-16609:
----------------------------------
Summary: Limit DN report queue and send small batches to avoid SCM
re-registration failures
Key: HDDS-16609
URL: https://issues.apache.org/jira/browse/HDDS-16609
Project: Apache Ozone
Issue Type: Improvement
Reporter: Aryan Gupta
Assignee: Aryan Gupta
When an SCM is down or asks DN to re-register, the DN keeps collecting
incremental reports (ICR) in memory.
If this queue becomes too large, heartbeat/register payload can become too big
and cross RPC message size limits.
Then DN fails to register again, and the problem repeats.
We should make DN smarter in this case.
Proposed change:
* Send queued incremental reports in multiple small batches (not one very
large request).
* Add a max payload cap per heartbeat request.
* If queue becomes too old/too large, drop or compact old ICR entries.
* Keep important command status handling safe (do not blindly drop critical
status reports).
* Add metrics/logs for queue size, dropped reports, and batch sends.
Expected result:
* DN can re-register reliably after SCM outage.
* No oversized heartbeat/register RPC due to huge queued reports.
* Better stability when one SCM node is slow/down.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]