chibenwa commented on PR #3193: URL: https://github.com/apache/james-project/pull/3193#issuecomment-5792731915
> Also why only S3 or PostgreSQL? Why not filesystem like in Stalwart and Dovecot? We used to have a poorly performing file maildir backend. We deprecated it: bad performance, directory traversal, no maintainer, no known deployment. > No, James itself still does all the compression. S3 is just an object store that receives and stores the bytes James sends it. If options exist to tell s3 to compress and do the job that's nice. I am not aware TBH and we could add support for this in S3BlobStoreDAO ZstdBlobStoreDAO should work out of the box atop file implem. If this is not the case a separated PR is welcomed. > lz4 Bad choice, it won't compress base64 as it doesn't do entropy coding. Most benefit of compression is the base64 overhead on large attachment, which is capped at ~33%. > RustFS S3 for example which can compress by myself with zstd/lz4 If it is standard (ie amazon s3 supports this) then it is welcomed in S3BlobStoreDAO > yeah fair enough, and compression in james isn't forced either way, if the underlying s3 engine you're using already handles it at the storage layer, ceph or rustfs, you can just turn it off on the james side with blobstore.compression.enabled=false and let the backend do its thing. Yes precisely. Also bear in mind that one can implement out-of-band-compression, and solely compress old 3+ month old mail as a tiering strategy... I do not know if RustFS for instance is able to automate that with a S3 policy. > but for plain cloud s3 like aws, wasabi, backblaze b2 etc Event worse if one operate a on prem S3 compatible system like Ceph Rados Gateway. > also only zstd for now ? Lz4 is not fit cf earlier entropy coding point. Zstd had decent java port. More niche algorithm would need to have good perf java implems. I confess I do not have in mind viable java alternative, but with hashing I did bench SHA-256 vs BLAKE3 and fornow JVM implenm of BLAKE3 is worse than SHA-256. I'd be careful with exotic implem and have their viability backed with a JMH micro benchmark. > And deduplication methods its like in ZFS? You are free to drop james dedup and experiment with FS level dedup. My take on it is that FS do not have all the infos James have and would be doing more work to achieve partial results. I'm expecting it not to deduplicate as well as James can do but I do not have solid comparison to answer you and the question indeed is interesting. > Good questions, though i think this is kinda drifting outside the scope of this specific PR Completly, but this is a nice discussion! Cheers, Benoit -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
