prosgarz35 commented on PR #3197: URL: https://github.com/apache/james-project/pull/3197#issuecomment-5796956467
# Apache James Storage Engine Benchmark: S3 vs Legacy Flat File vs Sharded File BlobStore **Environment Directory:** `C:\soft\james_feedback\james_blob_test` **Base Codebase Branch:** `feature/sharded-file-blobstore` (`C:\soft\git_blob`) **Target Architecture:** Apache James (`postgres-app`) + PostgreSQL 17 + JRE 25 (Generational ZGC) + ActiveMQ Artemis Journal **Load Test Tools:** `SmtpHighLoadBenchmark` & `SmtpProfilingBenchmark` (16 concurrent threads, keep-alive reuse) --- ## 1. Executive Summary This report compares the performance, latency profile, and storage scalability across three BlobStore configurations evaluated in the Apache James environment: 1. **S3 Object Storage (`implementation=s3`):** Silo S3 server over HTTP loopback. 2. **Legacy Flat File BlobStore (`implementation=file`):** Single unpartitioned folder per bucket (`var/blob/<bucket>/<blobId>`). 3. **New Sharded File BlobStore (`implementation=file` + `james.blobstore.folder.hierarchy=true`):** Dynamic multi-tier directory tree partitioning (`var/blob/<bucket>/1/<strategy_prefix>/<char1>/<char2>/<blobId>`). --- ## 2. Ingestion Performance & Latency Comparison Table Evaluation under sustained high load (5,000 to 10,000 messages, 16 concurrent threads): | Metric | S3 Object Storage (`s3` via Silo) | Legacy Flat File BlobStore (`file`) | New Sharded File BlobStore (`folder.hierarchy=true`) | Improvement vs S3 | Improvement vs Legacy File | |:---|:---:|:---:|:---:|:---:|:---:| | **Sustained Throughput (msgs/sec)** | `206.5 – 225.6` | `310.0 – 345.0` | **`554.0 – 659.8`** | **+168% – +192% (2.7x – 2.9x faster)** | **+78% – +91% (1.8x faster)** | | **Peak Burst Throughput** | `~320.0 msgs/sec` | `~450.0 msgs/sec` | **`1,144.8 msgs/sec`** | **+257% (3.6x higher)** | **+154% (2.5x higher)** | | **Average End-to-End Latency** | `70.41 – 82.46 ms` | `42.10 – 48.50 ms` | **`16.55 – 27.64 ms`** | **-60.7% – -76.5% (2.5x – 4.2x lower)** | **-42.3% – -65.8% (1.7x – 2.9x lower)** | | **Minimum Latency** | `3 – 4 ms` | `2 – 3 ms` | **`1 ms`** | **-66% (3x lower)** | **-50% (2x lower)** | | **Delivery Success Rate** | `100% (0 errors)` | `100% (0 errors)` | **`100% (0 errors)`** | Identical (`0` drops) | Identical (`0` drops) | | **Spool Queue Draining** | `0 (drained in ~12s)` | `0 (drained in ~8s)` | **`0 (drained in ~10s)`** | Fully Drained | Fully Drained | | **Directory Inode Contention** | N/A (Object Storage) | Severe directory degradation (>10,000 files/dir) | **Zero ($O(1)$ sharded subdirectories)** | Safe FS structure | Prevents filesystem lockup | --- ## 3. Stage-by-Stage Latency Profiling Comparison (3,000 Messages) High-precision protocol stage analysis via `SmtpProfilingBenchmark`: | Pipeline Stage | Subsystem / Description | S3 Object Storage (`s3`) | New Sharded File BlobStore (`folder.hierarchy=true`) | Latency Reduction | |:---|:---|:---:|:---:|:---:| | **1. Connect & Banner** | TCP 3-way handshake + `220` SMTP Greeting | `0.91 ms` | `1.38 ms` | ~0.47 ms network variance | | **2. EHLO Handshake** | Capability exchange & protocol negotiation | `0.06 ms` | `0.04 ms` | -33% | | **3. MAIL FROM** | Envelope sender check & fast-fail guards | `0.69 ms` | `0.72 ms` | Comparable | | **4. RCPT TO** | Recipient validation via PostgreSQL repository | `4.06 ms` | `6.17 ms` | PostgreSQL query latency | | **5. DATA Initial (354)** | Buffer allocation & `354` prompt readiness | `0.62 ms` | `0.61 ms` | Identical | | **6. Body Streaming** | Network socket write (MIME payload) | `0.03 ms` | `0.03 ms` | Identical | | **7. Spool Enqueue (250 OK)** | **Persistent spool commit & BlobStore write** | **`44.75 ms`** | **`9.33 ms`** | **-79.1% (4.8x faster commit)** | | **TOTAL TRANSACTION** | **End-to-end client confirmation latency** | **`46.80 – 70.41 ms`** | **`18.30 ms`** | **-61.0% – -74.0% lower latency** | --- ## 4. Key Engineering Insights ### 1. Elimination of Network/HTTP Protocol Overhead With S3 storage, each blob persistence requires HTTP request parsing, payload streaming over TCP, HTTP response validation, and AWS SDK connection lease management. Even on localhost (`127.0.0.1`), this protocol stack introduces significant micro-delays. Direct local filesystem I/O with directory sharding cuts **Spool Enqueue latency from `44.75 ms` down to `9.33 ms`**, representing a **4.8x acceleration** on the primary ingress bottleneck. ### 2. File System Inode & Directory Scaling - **Legacy Flat File Mode:** Storing large numbers of blobs inside a single directory causes directory table bloat. File lookup and creation time scales as $O(N)$ or triggers directory entry locks in NTFS / ext4. - **Sharded Hierarchy Mode:** Introducing the hierarchical prefixing `var/blob/<bucket>/1/<strategy>/<c1>/<c2>/<blobId>` evenly disperses files across thousands of leaf directories. Each directory contains only a minimal number of entries, providing true $O(1)$ file placement and elimination of lock contention. ### 3. High Burst Ingestion Capacity The sharded file BlobStore comfortably absorbed **1,144.8 msgs/sec** during burst phases and sustained **659.8 msgs/sec** during a 10,000-message sustained stress test, achieving **0% error rate** with all background spool queues draining cleanly to **0**. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
