[ 
https://issues.apache.org/jira/browse/FLINK-40387?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

Gabor Somogyi reassigned FLINK-40387:
-------------------------------------

    Assignee: Michael

> NativeS3 - configured retry/throttle backoff is ignored due to AWS SDK retry 
> circuit breaker enabled by default - failing checkpoints under S3 503 
> throttling
> -------------------------------------------------------------------------------------------------------------------------------------------------------------
>
>                 Key: FLINK-40387
>                 URL: https://issues.apache.org/jira/browse/FLINK-40387
>             Project: Flink
>          Issue Type: Bug
>          Components: FileSystems
>    Affects Versions: 2.3.0
>            Reporter: Michael
>            Assignee: Michael
>            Priority: Major
>
> h3. Environment
>  * Flink 2.3.0 - flink-s3-fs-native-2.3.0.jar (AWS SDK retries 2.47.5)
>  * Previously: Flink 2.2 - flink-s3-fs-hadoop-2.2.0.jar
>  * ~1 TB RocksDB state, incremental checkpoints
> h3. Problem
>  * Pre-upgrade (S3A plugin): checkpoints succeeded ~99% of the time
>  * Post-upgrade (native plugin): checkpoints fail most of the time on S3 
> throttling:
> {noformat}
> S3Exception: Slow Down (Service: S3, Status Code: 503)
> {noformat}
>  * Entropy is enabled and load *is* spread across prefixes, so hot-prefixes 
> are not the cause
>  * Checkpoints fail in ~1–3s. With max-num-retries=20 and throttle backoff 
> 1s→60s, exhausting retries would take tens of seconds — so retries are being 
> short-circuited, not exhausted.
> h3. Root cause
>  * *{{org.apache.flink.fs.s3native.S3ClientProvider}}* builds the SDK 
> StandardRetryStrategy with maxAttempts + backoff but never sets 
> circuitBreakerEnabled(false)
>  * The breaker is on by default
>  * During a ~1 TB incremental checkpoint, thousands of concurrent 
> getFileStatus/headObject calls throttle at once, drain the shared token 
> bucket, and the SDK stops retrying — regardless of configured backoff.
>  * 
> [https://docs.aws.amazon.com/sdk-for-java/latest/developer-guide/retry-strategy.html]
> h3. Request
>  # Add a config option (e.g. s3.retry.circuit-breaker.enabled) so 
> retry/backoff is honored
>  # Proposal: default it off for the checkpoint filesystem — restores S3A 
> behavior (absorb throttling, keep retrying). At minimum, add the toggle.
> [https://github.com/apache/flink/blob/0a2bad7b052538adecda1f68ba8b558c3aa8f097/flink-filesystems/flink-s3-fs-native/src/main/java/org/apache/flink/fs/s3native/S3ClientProvider.java#L662-L675]
> h3. Stack trace
> {noformat}
> S3Exception: Slow Down (Status Code: 503, Request ID: 9F1GDSDKDKZH5MAM, ...)
>     at ...DefaultS3Client.headObject(DefaultS3Client.java:8133)
>     at o.a.f.f.s.NativeS3FileSystem.getFileStatus(NativeS3FileSystem.java:206)
> {noformat}



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to