[ 
https://issues.apache.org/jira/browse/HDDS-16122?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

Siyao Meng resolved HDDS-16122.
-------------------------------
    Resolution: Duplicate

> OM becomes unresponsive during sustained FSO snapshot creation at ~18.7K 
> snapshots
> ----------------------------------------------------------------------------------
>
>                 Key: HDDS-16122
>                 URL: https://issues.apache.org/jira/browse/HDDS-16122
>             Project: Apache Ozone
>          Issue Type: Bug
>          Components: Ozone Manager
>            Reporter: Yashaswini G A
>            Priority: Major
>
> Under sustained snapshot create/delete/snapdiff load - defrag bounday test 
> (~18.7K snapshots, ~185K keys, FSO bucket), OM HA cluster stops responding. 
> Snapshot create hangs for 900s+, all three OMs fail client failover (50 
> attempts), Freon reports Could not determine or connect to OM Leader.
> *Steps to reproduce:*
>  # Create FSO bucket
>  # Loop: write 10×1KB keys + create snapshot
>  # Every 600s delete one random snapshot (keep first)
>  # Every 5000 snapshots run snapdiff from first to latest
>  # Continue toward 65K snapshots
> *Observed:* Failure at snapshot {*}18716{*}; OM completely unresponsive 
> afterward.
> *Configs Applied:*
>  * ozone.snapshot.defrag.limit.per.task=65500
>  * ozone.snapshot.defrag.limit.per.task: 50
>  * ozone.snapshot.defrag.service.timeout: 7200s
>  * rlimit_fds:1000000
> *Contributing factors :*
>  # Sustained snapshot create/delete load (~18.7K snapshots, ~185K keys)
>  # Snapdiff from first snapshot to latest every 5,000 snapshots- cost grows 
> linearly, at 15K it took 46 min with repeated FAILED/IN_PROGRESS



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to