HesandaLiyanage opened a new pull request, #3193:
URL: https://github.com/apache/james-project/pull/3193

   JIRA: https://issues.apache.org/jira/browse/JAMES-4231
   
   ### Context & Problem
   High email ingest volumes in Apache James deployments utilizing S3/MinIO 
object storage result in hundreds of millions of small S3 objects (average 
body/header size ~20–50KB). Operating at this granularity introduces high S3 
API request fees (PUT/GET/LIST), elevated bucket listing latency during storage 
maintenance and garbage collection, and reduced overall object storage 
throughput.
   
   ### Solution Overview
   This PR implements autonomous S3 object compaction for Apache James, packing 
historical standalone blobs into large, immutable chunk files (~100MB target 
size) with slot-level virtual addressing and transparent single-request HTTP 
ranged read support.
   
   ### Key Components & Architecture
   
   1. **Chunk Format (`ChunkFormat`):**
      - Single binary chunk containing concatenated slots and a trailer footer:
        `[Header (16B)][Slot 0]...[Slot N-1][Footer][Footer Size (4B)][End 
Magic 0x454F4643]`
      - **Per-slot independent Zstandard compression:** Each slot has its own 
Zstd codec (`content-encoding=zstd\n`), allowing single-slot ranged reads 
without decompressing or buffering the full 100MB chunk.
      - **Suffix-indexed 64KB footer:** Suffix range read (`readRange(..., 
-65536, -1)`) fetches chunk slot metadata in $O(1)$ without scanning payloads.
      - Fully documented in `src/adr/0076-s3-object-compaction.md`.
   
   2. **Virtual Slot Addressing & Ranged Reads (`ChunkedBlobStoreDAO` & 
`ChunkId`):**
      - Format: `<family>_<generation>_chunk<hash>~<offset>~<limit>`.
      - `ChunkedBlobStoreDAO` transparently intercepts slot references and 
issues HTTP byte-range requests directly against the underlying storage 
connector (`readRange(bucket, chunkId, offset, offset + limit - 1)`).
      - Conforms strictly to `ChunkMarker` regex 
(`^\d+_\d+_chunk[A-Za-z0-9_-]{16,}(~\d+~\d+)?$`) to ensure bloom-filter GC 
never misclassifies chunks as unreferenced garbage.
   
   3. **Compaction Pipeline & Crash Safety (`BlobCompactionAlgorithm`):**
      - **Initial Compaction (`initialCompact`):**
        1. Candidate scanning windowed into batches of 1,000 blobs 
(`DEFAULT_CANDIDATE_BATCH_SIZE = 1000`) so candidate byte arrays are packed and 
freed window-by-window without buffering the generation's payloads in heap.
        2. Chunk assembled and saved to raw S3 storage.
        3. Metadata references updated in Cassandra (`messageV3`, 
`messageIdToImapUid`, `messageIdTable`).
        4. Original standalone blobs deleted from S3 only for candidates whose 
references updated successfully. If an update fails mid-batch, deletion is 
skipped for that candidate, guaranteeing zero dangling references.
      - **GC Compaction (`gcCompact`):**
        - Inspects existing chunks via footer-only ranged reads (metadata-only).
        - Orphan chunks (100% dead slots) deleted with 0 payload bytes read.
        - Chunks exceeding dead ratio rewritten; small adjacent chunks merged. 
Live slots streamed individually via ranged reads, bounding GC heap to 
$O(\text{maxSlotSize})$.
      - **Self-Healing Repairer (`CassandraBlobIdRepairer`):**
        - On missing slot / 404, falls back to original blob ID from 
`blobReferenceMapping` and restores Cassandra references. Also heals any 
transient inconsistency window across Cassandra tables.
   
   4. **Configuration Guards & Scope:**
      - Compaction automatically disables itself with an informative log when 
client-side AES encryption or whole-blob compression is active (AES cipher 
blocks HTTP range slicing).
      - Generation-scoped: triggered via `DELETE 
/blobs?scope=compaction&generation=<gen>&family=<fam>`.
   
   5. **Memory Characteristics & Operational Ceiling:**
      - Candidate payload heap: strictly bounded by $O(\min(N \times 
\text{avgBlobSize}, \text{chunkTargetSize}))$.
      - GC payload heap: strictly bounded by $O(\text{maxSlotSize})$ (~1MB).
      - Reference mapping ceiling: $O(\text{totalLiveGenerationReferences} 
\times \approx 200\text{ bytes})$ (~200MB heap for 1M live references; ~2GB 
heap for 10M). Documented in class javadoc; future follow-up can introduce 
partition-paged lookups.
   
   ### Verification & Tests
   - `server/blob/blob-compaction`: 47 tests passed (100%), including 
allocation-bound proof test and task serialization.
   - `server/blob/blob-storage-strategy`: 78 tests passed (100%).
   - `mailbox/cassandra` (`CassandraBlobId*IntegrationTest`): 3/3 passed with 
real Cassandra 5.0.9 testcontainers.
   - `server/blob/blob-s3`: Ranged read and MinIO S3 end-to-end compaction 
integration tests passed.
   - `server/container/guice/distributed`: 43 tests passed (100%).
   - Checkstyle: 0 errors across all 7 modified modules.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to