[ 
https://issues.apache.org/jira/browse/LUCENE-9211?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=17035624#comment-17035624
 ] 

David Smiley commented on LUCENE-9211:
--------------------------------------

Thanks so much for running the benchmarks [~mharwood]!  When you say you 
modified "this line"; the link did not work.  If you merely changed the default 
spatial.alg to use composite then it's only indexing point data which is not 
realistic for this spatial strategy.  Instead LUCENE-5579 has a spatial.alg 
file that converts those points to random circles and it'll be more 
interesting.  I just did a diff on that spatial.alg with the default one and 
they are pretty similar overall.

> Adding compression to BinaryDocValues storage
> ---------------------------------------------
>
>                 Key: LUCENE-9211
>                 URL: https://issues.apache.org/jira/browse/LUCENE-9211
>             Project: Lucene - Core
>          Issue Type: Improvement
>          Components: core/codecs
>            Reporter: Mark Harwood
>            Assignee: Mark Harwood
>            Priority: Minor
>              Labels: pull-request-available
>
> While SortedSetDocValues can be used today to store identical values in a 
> compact form this is not effective for data with many unique values.
> The proposal is that BinaryDocValues should be stored in LZ4 compressed 
> blocks which can dramatically reduce disk storage costs in many cases. The 
> proposal is blocks of a number of documents are stored as a single compressed 
> blob along with metadata that records offsets where the original document 
> values can be found in the uncompressed content.
> There's a trade-off here between efficient compression (more docs-per-block = 
> better compression) and fast retrieval times (fewer docs-per-block = faster 
> read access for single values). A fixed block size of 32 docs seems like it 
> would be a reasonable compromise for most scenarios.
> A PR is up for review here [https://github.com/apache/lucene-solr/pull/1234]



--
This message was sent by Atlassian Jira
(v8.3.4#803005)

---------------------------------------------------------------------
To unsubscribe, e-mail: issues-unsubscr...@lucene.apache.org
For additional commands, e-mail: issues-h...@lucene.apache.org

Reply via email to