[
https://issues.apache.org/jira/browse/HDDS-16122?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Siyao Meng resolved HDDS-16122.
-------------------------------
Resolution: Duplicate
> OM becomes unresponsive during sustained FSO snapshot creation at ~18.7K
> snapshots
> ----------------------------------------------------------------------------------
>
> Key: HDDS-16122
> URL: https://issues.apache.org/jira/browse/HDDS-16122
> Project: Apache Ozone
> Issue Type: Bug
> Components: Ozone Manager
> Reporter: Yashaswini G A
> Priority: Major
>
> Under sustained snapshot create/delete/snapdiff load - defrag bounday test
> (~18.7K snapshots, ~185K keys, FSO bucket), OM HA cluster stops responding.
> Snapshot create hangs for 900s+, all three OMs fail client failover (50
> attempts), Freon reports Could not determine or connect to OM Leader.
> *Steps to reproduce:*
> # Create FSO bucket
> # Loop: write 10×1KB keys + create snapshot
> # Every 600s delete one random snapshot (keep first)
> # Every 5000 snapshots run snapdiff from first to latest
> # Continue toward 65K snapshots
> *Observed:* Failure at snapshot {*}18716{*}; OM completely unresponsive
> afterward.
> *Configs Applied:*
> * ozone.snapshot.defrag.limit.per.task=65500
> * ozone.snapshot.defrag.limit.per.task: 50
> * ozone.snapshot.defrag.service.timeout: 7200s
> * rlimit_fds:1000000
> *Contributing factors :*
> # Sustained snapshot create/delete load (~18.7K snapshots, ~185K keys)
> # Snapdiff from first snapshot to latest every 5,000 snapshots- cost grows
> linearly, at 15K it took 46 min with repeated FAILED/IN_PROGRESS
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]