Devesh Kumar Singh created HDDS-16652:
-----------------------------------------
Summary: Avoid persisting redundant StorageType in datanode block
metadata
Key: HDDS-16652
URL: https://issues.apache.org/jira/browse/HDDS-16652
Project: Apache Ozone
Issue Type: Sub-task
Reporter: Devesh Kumar Singh
Assignee: Devesh Kumar Singh
PR #11223 (HDDS-16388. Datanode putBlock Support StorageType) carries
storageTypeID in ContainerProtos.DatanodeBlockID.
The value is required on client-to-datanode requests to:
- create a missing container on a volume of the requested StorageType; and
- validate block operations against the existing container’s
ContainerData.storageType before mutation.
However, the same value is also persisted in every RocksDB BlockData record
because the database codec and network path use the same serialization method:
PutBlock request
-> BlockData.getFromProtoBuf()
-> Java BlockID.storageType
-> DatanodeStore.putBlockByID(...)
-> BlockData.CODEC
-> BlockData.getProtoBufMessage()
-> BlockID.getDatanodeBlockIDProtobuf()
-> storageTypeID persisted in RocksDB
Storage type is a container-level physical property. All local blocks in a
container reside on the container’s selected volume, and the authoritative
value is already persisted in ContainerData/
container YAML and available from HddsVolume.
Current production validation consumes the request value before the block is
persisted. No production consumer has been identified that requires storage
type to survive a per-block RocksDB round trip.
The known readers are test assertions in
TestContainerPersistence#testPutBlockWithStorageType and the planned patch 16
TestOzoneStoragePolicy integration test.
For current enum values, the protobuf field normally adds approximately two
raw bytes per block record. At one billion blocks, that is roughly 2 GB of
uncompressed protobuf payload, excluding RocksDB
WAL, compaction and write-amplification effects. More importantly, it creates
a duplicate source of truth that could become stale if a container replica is
later imported or moved without rewriting
every block record.
The implementation should separate wire serialization from RocksDB
serialization:
Wire request serialization -> include storageTypeID
RocksDB serialization -> omit storageTypeID
GetBlock response -> omit it or derive it from ContainerData
Existing records containing or omitting the optional field must remain
readable without an eager database migration.
Origin: PR #11223 review discussion
(https://github.com/apache/ozone/pull/11223#pullrequestreview-5354128214).
### Acceptance criteria
- New RocksDB BlockData and lastChunkInfoTable records do not persist
storageTypeID.
- Client-to-datanode requests continue carrying storage type for container
creation and validation.
- Mismatched valid and invalid storage types are rejected before any chunk or
block mutation.
- ContainerData.storageType and the selected HddsVolume remain authoritative.
- Existing block records with or without the optional field remain readable
across restart and mixed-version operation.
- GetBlock behavior is documented; if it must return storage type, the value
is derived from the owning container.
- Tests cover DISK, SSD and ARCHIVE requests, mismatch rejection, incremental
chunk lists, restart compatibility and the chosen GetBlock contract.
- Patch 16 tests are updated so they do not require redundant per-block
persistence.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]