[
https://issues.apache.org/jira/browse/HDDS-16315?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
ASF GitHub Bot updated HDDS-16315:
----------------------------------
Labels: pull-request-available (was: )
> Skip MD5 recomputation in CopyObject when source ETag can be reused
> -------------------------------------------------------------------
>
> Key: HDDS-16315
> URL: https://issues.apache.org/jira/browse/HDDS-16315
> Project: Apache Ozone
> Issue Type: Improvement
> Reporter: Chu Cheng Li
> Priority: Major
> Labels: pull-request-available
>
> The S3 Gateway CopyObject path re-hashes every byte of the object during the
> copy: {{ObjectEndpoint#copyObject}} wraps the source {{OzoneInputStream}} in
> a {{java.security.DigestInputStream}}, and the computed MD5 becomes the
> destination key's ETag (in {{ObjectEndpoint#copy}} for the normal write path
> and in {{ObjectEndpointStreaming#copyKeyWithStream}} for the datastream path).
> A whole-object copy produces byte-identical content, and since HDDS-10403 the
> source key's metadata already carries its content-based ETag. When that
> stored ETag is a plain (non-multipart) MD5, it is also the correct content
> MD5 for the destination, so the recomputation is pure wasted CPU.
> Proposed change:
> * On CopyObject (not UploadPartCopy), when the source key metadata holds an
> ETag without the "-N" multipart suffix, skip the {{DigestInputStream}} and
> store the source's ETag on the destination key.
> * Otherwise (missing ETag, or aggregate "-N" ETag from an MPU-created source)
> keep the current digesting behavior, so the destination still gets a proper
> content MD5. For MPU-sourced whole-object copies this matches AWS, which
> returns a plain MD5 for copies under 5GB.
> * Apply the same logic to both write variants: the normal path through
> {{ObjectEndpoint#copy}} / {{S3ObjectWriteGuard}} and the datastream path
> through {{ObjectEndpointStreaming#copyKeyWithStream}}.
> This is a pure CPU optimization. There is no protocol change and no proto
> change, and UploadPartCopy is unaffected because part ETags are digests of
> the part's own bytes and are produced on a separate code path.
> Measured with an off-tree harness driving ObjectEndpoint against the
> in-memory client stub (Apple M4, JDK 21), interleaving the two paths in one
> JVM, medians over >= 12 iterations after warmup:
> || object || digesting (before) || ETag reused (after) || saved ||
> | 1 MiB | 3.78 ms | 1.54 ms | 2.24 ms (59%) |
> | 8 MiB | 14.78 ms | 2.35 ms | 12.43 ms (84%) |
> | 64 MiB | 95.98 ms | 8.76 ms | 87.22 ms (91%) |
> | 256 MiB | 407.32 ms | 46.22 ms | 361.09 ms (89%) |
> The stub holds everything in memory, so those percentages are an upper bound
> rather than end-to-end cluster numbers. The portable result is the absolute
> one: the change removes roughly 1.2-1.4 s of S3 Gateway CPU per GiB copied
> (JDK MD5 measures 0.72-0.81 GiB/s on that machine, and "openssl speed md5"
> independently reports 0.85 GiB/s). Because MD5 is serial and runs on the
> request thread, that is also about 1.2 core-seconds freed per GiB of copy
> throughput on a loaded gateway.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]