Yashaswini G A created HDDS-16122:
-------------------------------------

             Summary: OM becomes unresponsive during sustained FSO snapshot 
creation at ~18.7K snapshots
                 Key: HDDS-16122
                 URL: https://issues.apache.org/jira/browse/HDDS-16122
             Project: Apache Ozone
          Issue Type: Bug
          Components: Ozone Manager
            Reporter: Yashaswini G A


Under sustained snapshot create/delete/snapdiff load - defrag bounday test 
(~18.7K snapshots, ~185K keys, FSO bucket), OM HA cluster stops responding. 
Snapshot create hangs for 900s+, all three OMs fail client failover (50 
attempts), Freon reports Could not determine or connect to OM Leader.



*Steps to reproduce:*
 # Create FSO bucket

 # Loop: write 10×1KB keys + create snapshot

 # Every 600s delete one random snapshot (keep first)

 # Every 5000 snapshots run snapdiff from first to latest

 # Continue toward 65K snapshots

*Observed:* Failure at snapshot {*}18716{*}; OM completely unresponsive 
afterward.


*Configs Applied:*
 * ozone.snapshot.defrag.limit.per.task=65500

 * ozone.snapshot.defrag.limit.per.task: 50

 * ozone.snapshot.defrag.service.timeout: 7200s

 * rlimit_fds:1000000



*Contributing factors :*
 # Sustained snapshot create/delete load (~18.7K snapshots, ~185K keys)

 # Snapdiff from first snapshot to latest every 5,000 snapshots- cost grows 
linearly, at 15K it took 46 min with repeated FAILED/IN_PROGRESS



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to