[ https://issues.apache.org/jira/browse/LUCENE-3907?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=13596002#comment-13596002 ]
Michael McCandless commented on LUCENE-3907: -------------------------------------------- I think we should remove the Side (BACK/FRONT) enum: an app can always use ReverseStringFilter if it really wants BACK grams (what are BACK grams used for?). > Improve the Edge/NGramTokenizer/Filters > --------------------------------------- > > Key: LUCENE-3907 > URL: https://issues.apache.org/jira/browse/LUCENE-3907 > Project: Lucene - Core > Issue Type: Improvement > Reporter: Michael McCandless > Assignee: Uwe Schindler > Labels: gsoc2012, lucene-gsoc-12 > Fix For: 4.2 > > > Our ngram tokenizers/filters could use some love. EG, they output ngrams in > multiple passes, instead of "stacked", which messes up offsets/positions and > requires too much buffering (can hit OOME for long tokens). They clip at > 1024 chars (tokenizers) but don't (token filters). The split up surrogate > pairs incorrectly. -- This message is automatically generated by JIRA. If you think it was sent incorrectly, please contact your JIRA administrators For more information on JIRA, see: http://www.atlassian.com/software/jira --------------------------------------------------------------------- To unsubscribe, e-mail: dev-unsubscr...@lucene.apache.org For additional commands, e-mail: dev-h...@lucene.apache.org