clintropolis commented on PR #20382:
URL: https://github.com/apache/druid/pull/20382#issuecomment-5765658570

   >In this PR(and other object storage like deep storage), the new property 
druid.storage.zip is used to determine whether segment files are compressed as 
index.zip or not. However, in https://github.com/apache/druid/pull/18982 , we 
introduced a new property druid.storage.compressionFormat for HDFS deep storage 
to indicate which compression format is used for compression(zip and lz4), that 
PR didn't bring the change to object storage like S3, Azure ang google 
cloud(because we don't use object storage), but architecturally, that property 
should apply to all deep storages. So I think we should use the 
druid.storage.compressionFormat( with a new NONE enum supported ) to replace 
the druid.storage.zip.
   
   My stance is that it would make sense to add 
`druid.storage.compressionFormat` to `DeepStorageSegmentConfig`, but I didn't 
personally have much interest in adding support for it to the other deep 
storage implementations because I am primarily interested in storing V10 
segment format files in deep storage uncompressed so that `SegmentRangeReader` 
can do partial reads to support partial loads on historicals. V10 format stores 
the segment as a single file (metadata header + a bunch of concatenated 
internal files so kind similar in spirit to zip 0) so it shouldn't change the 
overall number of files in deep storage (at least without any extensions to 
add/attach 'external' files).
   
   In terms of compression/encoding, going forward I am much more interested in 
ways to shrink the internal files stored inside the v10 segment file so that we 
can preserve the ability to do partial reads, so I'll be spending any 
time/energy I have to pursue that instead of expanding the types of generic 
compression we can apply to the whole container
   
   >Secondly, do you have any data that shows the segment size without 
compression as index.zip in the deep storage?
   As https://github.com/apache/druid/pull/18982#issuecomment-3846035169, zip 0 
(which does not comrpess but only pack files together) has a size of 11GiB, 
while the default zip level currently we are using now shows the size is about 
2.48GB, this means that if segments are not compressed, file size in deep 
storage should be much larger. I'm not sure if such file size increasing is 
consistent in your case.
   
   I don't have any hard data on this, in part because it doesn't matter 
because of the need to be able to do partial reads and I haven't had much time 
to do analysis of what we could improve _inside_ the segment (because external 
compression is basically not an option and size hasn't proven problematic yet). 
11GiB is quite a lot larger than most of the segments we deal with, I'm still 
fairly interested in the composition of the segment in that example to find 
opportunities to compress the internal files (i theorized at the time in that 
PR that perhaps a lot of uncompressed complex types were making the difference 
between compressed and uncompressed so different).


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to