[ https://issues.apache.org/jira/browse/LUCENE-1799?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=13223294#comment-13223294 ]
DM Smith commented on LUCENE-1799: ---------------------------------- Would someone be able to champion this. It appears ready to go. for the last 1.5 years. Looks like it is merely a permission problem. I'd like to see it get in the 3.x series. > Unicode compression > ------------------- > > Key: LUCENE-1799 > URL: https://issues.apache.org/jira/browse/LUCENE-1799 > Project: Lucene - Java > Issue Type: New Feature > Components: core/store > Affects Versions: 2.4.1 > Reporter: DM Smith > Priority: Minor > Attachments: Benchmark.java, Benchmark.java, Benchmark.java, > LUCENE-1779.patch, LUCENE-1799.patch, LUCENE-1799.patch, LUCENE-1799.patch, > LUCENE-1799.patch, LUCENE-1799.patch, LUCENE-1799.patch, LUCENE-1799.patch, > LUCENE-1799.patch, LUCENE-1799.patch, LUCENE-1799.patch, LUCENE-1799.patch, > LUCENE-1799.patch, LUCENE-1799.patch, LUCENE-1799.patch, LUCENE-1799_big.patch > > > In lucene-1793, there is the off-topic suggestion to provide compression of > Unicode data. The motivation was a custom encoding in a Russian analyzer. The > original supposition was that it provided a more compact index. > This led to the comment that a different or compressed encoding would be a > generally useful feature. > BOCU-1 was suggested as a possibility. This is a patented algorithm by IBM > with an implementation in ICU. If Lucene provide it's own implementation a > freely avIlable, royalty-free license would need to be obtained. > SCSU is another Unicode compression algorithm that could be used. > An advantage of these methods is that they work on the whole of Unicode. If > that is not needed an encoding such as iso8859-1 (or whatever covers the > input) could be used. -- This message is automatically generated by JIRA. If you think it was sent incorrectly, please contact your JIRA administrators: https://issues.apache.org/jira/secure/ContactAdministrators!default.jspa For more information on JIRA, see: http://www.atlassian.com/software/jira --------------------------------------------------------------------- To unsubscribe, e-mail: dev-unsubscr...@lucene.apache.org For additional commands, e-mail: dev-h...@lucene.apache.org