Aryan Gupta created HDDS-16609:
----------------------------------

             Summary: Limit DN report queue and send small batches to avoid SCM 
re-registration failures
                 Key: HDDS-16609
                 URL: https://issues.apache.org/jira/browse/HDDS-16609
             Project: Apache Ozone
          Issue Type: Improvement
            Reporter: Aryan Gupta
            Assignee: Aryan Gupta


When an SCM is down or asks DN to re-register, the DN keeps collecting 
incremental reports (ICR) in memory.
If this queue becomes too large, heartbeat/register payload can become too big 
and cross RPC message size limits.
Then DN fails to register again, and the problem repeats.

We should make DN smarter in this case.

Proposed change:
 * Send queued incremental reports in multiple small batches (not one very 
large request).
 * Add a max payload cap per heartbeat request.
 * If queue becomes too old/too large, drop or compact old ICR entries.
 * Keep important command status handling safe (do not blindly drop critical 
status reports).
 * Add metrics/logs for queue size, dropped reports, and batch sends.

Expected result:
 * DN can re-register reliably after SCM outage.
 * No oversized heartbeat/register RPC due to huge queued reports.
 * Better stability when one SCM node is slow/down.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to