lpavanvenkat opened a new pull request, #11369:
URL: https://github.com/apache/ozone/pull/11369
## What changes were proposed in this pull request?
`BasicRootedOzoneClientAdapterImpl#getFileChecksum` issued three OM RPCs per
call: InfoVolume (`objectStore.getVolume`), InfoBucket (`getBucket`) and
lookupKey. Only the lookupKey is needed for the checksum itself.
`getFileChecksumWithCombineMode` uses just `volume.getName()` and
`bucket.getName()`, and the bucket fetch exists only to resolve the (possibly
linked) bucket layout and reject OBS buckets. A checksum over many files in the
same bucket, for example distcp checksum comparison, therefore sends two
redundant RPCs to OM per file.
This PR:
* Adds a client-side Guava `Cache<String, BucketLayout>` keyed by
`volume/bucket` to `BasicRootedOzoneClientAdapterImpl`, bounded by two new
client configs:
* `ozone.client.fs.bucket.layout.cache.expiry` (default `2m`)
* `ozone.client.fs.bucket.layout.cache.size` (default `1000`, `0` disables
caching)
* `getFileChecksum` resolves the layout through the cache using an atomic
get-or-load, so concurrent callers on the same bucket issue a single InfoBucket
RPC. It validates the layout (OBS is still rejected) and builds minimal
`OzoneVolume`/`OzoneBucket` objects from the path, so it no longer issues an
InfoVolume RPC, and issues an InfoBucket RPC only on a cache miss.
* `getBucket` still fetches the bucket and resolves link buckets on every
call, so orphan link detection (`setSourcePathExist(false)`, used by
`listStatus`) is unchanged, and it stores the resolved layout in the cache for
later `getFileChecksum` calls.
* `getFileStatus` is unchanged.
Staleness is bounded by the expiry: if a bucket is deleted and recreated
with a different layout, `getFileChecksum` may use the old layout for up to 2
minutes. If the bucket or volume no longer exists, the lookupKey still fails as
before.
## What is the link to the Apache JIRA
https://issues.apache.org/jira/browse/HDDS-15951
## How was this patch tested?
CI: <link to the workflow run on the fork>
* New unit test `TestBasicRootedOzoneClientAdapterBucketLayoutCache` (9
tests): cache miss/hit, reuse across keys, OBS rejection from the cache, loader
exception unwrapping, `getBucket` populating the cache, empty bucket name,
orphan link detection despite a cached layout, and OBS rejection in `getBucket`
despite a stale cache entry.
* All unit tests in `ozone-filesystem-common` pass (92 tests).
* Existing integration tests pass: `TestOzoneFileChecksum`, `TestOFS`,
`TestOFSWithFSO`, `TestOFSWithFSPaths` (202 tests).
* checkstyle, rat and author checks pass.
* Benchmark (`TestOfsGetFileChecksumCacheBenchmark`, tagged
`@Tag("benchmark")`): one MiniOzoneCluster, 200 FSO buckets each holding one 4
KB file, and 2000 `getFileChecksum` calls (a 90% cache-hit workload), measured
single-threaded and across 10 concurrent client threads sharing one
`FileSystem`. Numbers below are baseline (master) vs this PR, on the same host.
**Single-threaded**
| Metric | Baseline | This PR | Change |
|---|---|---|---|
| InfoVolume RPCs | 2000 (1.00/call) | 0 (0.00/call) | −100% |
| InfoBucket RPCs | 2000 (1.00/call) | 200 (0.10/call) | −90% |
| lookupKey RPCs | 2000 | 2000 | unchanged |
| latency mean | 0.515 ms | 0.330 ms | −36% |
| latency p99 | 1.031 ms | 0.954 ms | −7% |
| throughput | 1,940.1 ops/s | 3,026.8 ops/s | +56% |
**10 concurrent threads**
| Metric | Baseline | This PR | Change |
|---|---|---|---|
| InfoVolume RPCs | 2000 (1.00/call) | 0 (0.00/call) | −100% |
| InfoBucket RPCs | 2000 (1.00/call) | 200 (0.10/call) | −90% |
| lookupKey RPCs | 2000 | 2000 | unchanged |
| latency mean | 2.246 ms | 1.549 ms | −31% |
| latency p99 | 6.726 ms | 4.570 ms | −32% |
| throughput | 4,432.4 ops/s | 6,386.8 ops/s | +44% |
`getFileChecksum` no longer issues an InfoVolume RPC, and issues one
InfoBucket RPC per bucket instead of one per call. The reduction holds
identically at 1 and 10 threads, because the atomic get-or-load keeps
concurrent misses on the same bucket from each calling OM. The lookupKey RPC
count is unchanged, which confirms that every call still reads the key and
computes a real checksum.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]