[ https://issues.apache.org/jira/browse/LUCENE-3907?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=13652878#comment-13652878 ]
Adrien Grand commented on LUCENE-3907: -------------------------------------- The previous behaviour could trigger highlighting bugs so I think it is important that we fix it in 4.x. In case the broken behaviour is still needed, it can be emulated by providing Version.LUCENE_43 as the Lucene match version. > Improve the Edge/NGramTokenizer/Filters > --------------------------------------- > > Key: LUCENE-3907 > URL: https://issues.apache.org/jira/browse/LUCENE-3907 > Project: Lucene - Core > Issue Type: Improvement > Reporter: Michael McCandless > Assignee: Adrien Grand > Labels: gsoc2013 > Fix For: 4.3 > > Attachments: LUCENE-3907.patch > > > Our ngram tokenizers/filters could use some love. EG, they output ngrams in > multiple passes, instead of "stacked", which messes up offsets/positions and > requires too much buffering (can hit OOME for long tokens). They clip at > 1024 chars (tokenizers) but don't (token filters). The split up surrogate > pairs incorrectly. -- This message is automatically generated by JIRA. If you think it was sent incorrectly, please contact your JIRA administrators For more information on JIRA, see: http://www.atlassian.com/software/jira --------------------------------------------------------------------- To unsubscribe, e-mail: dev-unsubscr...@lucene.apache.org For additional commands, e-mail: dev-h...@lucene.apache.org