clintropolis commented on PR #20382: URL: https://github.com/apache/druid/pull/20382#issuecomment-5765658570
>In this PR(and other object storage like deep storage), the new property druid.storage.zip is used to determine whether segment files are compressed as index.zip or not. However, in https://github.com/apache/druid/pull/18982 , we introduced a new property druid.storage.compressionFormat for HDFS deep storage to indicate which compression format is used for compression(zip and lz4), that PR didn't bring the change to object storage like S3, Azure ang google cloud(because we don't use object storage), but architecturally, that property should apply to all deep storages. So I think we should use the druid.storage.compressionFormat( with a new NONE enum supported ) to replace the druid.storage.zip. My stance is that it would make sense to add `druid.storage.compressionFormat` to `DeepStorageSegmentConfig`, but I didn't personally have much interest in adding support for it to the other deep storage implementations because I am primarily interested in storing V10 segment format files in deep storage uncompressed so that `SegmentRangeReader` can do partial reads to support partial loads on historicals. V10 format stores the segment as a single file (metadata header + a bunch of concatenated internal files so kind similar in spirit to zip 0) so it shouldn't change the overall number of files in deep storage (at least without any extensions to add/attach 'external' files). In terms of compression/encoding, going forward I am much more interested in ways to shrink the internal files stored inside the v10 segment file so that we can preserve the ability to do partial reads, so I'll be spending any time/energy I have to pursue that instead of expanding the types of generic compression we can apply to the whole container >Secondly, do you have any data that shows the segment size without compression as index.zip in the deep storage? As https://github.com/apache/druid/pull/18982#issuecomment-3846035169, zip 0 (which does not comrpess but only pack files together) has a size of 11GiB, while the default zip level currently we are using now shows the size is about 2.48GB, this means that if segments are not compressed, file size in deep storage should be much larger. I'm not sure if such file size increasing is consistent in your case. I don't have any hard data on this, in part because it doesn't matter because of the need to be able to do partial reads and I haven't had much time to do analysis of what we could improve _inside_ the segment (because external compression is basically not an option and size hasn't proven problematic yet). 11GiB is quite a lot larger than most of the segments we deal with, I'm still fairly interested in the composition of the segment in that example to find opportunities to compress the internal files (i theorized at the time in that PR that perhaps a lot of uncompressed complex types were making the difference between compressed and uncompressed so different). -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
