[jira] Commented: (LUCENE-1522) another highlighter

Mark Miller (JIRA) Mon, 16 Mar 2009 10:52:14 -0700

    [ 
https://issues.apache.org/jira/browse/LUCENE-1522?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=12682387#action_12682387
 ]


Mark Miller commented on LUCENE-1522:
-------------------------------------

I don't think its easy to get a speedy highlighter that works with positions 
for all of the Lucene queries. In the long term, I'd love to see a fast 
highlighter that works with positions for all of Lucene's queries . I'd also 
like it to work if you don't have termvectors stored (though be faster if they 
are perhaps, as it is now).

Essentially we have each of these pieces separately now - the difficulty is 
doing it with one highlighter. 

We have the standard Highlighter with two modes: one that doesn't handle 
positions, and one that handles position sensitive highlighting for Spans and 
almost all of the queries. This framework is great - its customizable, it 
handles a lot of corner cases, it works without termvectors, it gets faster 
with termvectors. Unfortunately, it runs through the source stream one token at 
a time, and doesn't scale well. Getting hit positions for position sensitive 
clauses requires converting the query to a span query and calling getSpans on a 
memory index

We also have the termvector highlighters that can work from offsets in the 
query and avoid running through a token at a time. You need termvectors for 
this approach, and its difficult to handle positions, but it scales.

The difficulty and goal is in merging the qualities of both. 

> another highlighter
> -------------------
>
>                 Key: LUCENE-1522
>                 URL: https://issues.apache.org/jira/browse/LUCENE-1522
>             Project: Lucene - Java
>          Issue Type: Improvement
>          Components: contrib/highlighter
>            Reporter: Koji Sekiguchi
>            Assignee: Michael McCandless
>            Priority: Minor
>             Fix For: 2.9
>
>         Attachments: colored-tag-sample.png, LUCENE-1522.patch, 
> LUCENE-1522.patch
>
>
> I've written this highlighter for my project to support bi-gram token stream 
> (general token stream (e.g. WhitespaceTokenizer) also supported. see test 
> code in patch). The idea was inherited from my previous project with my 
> colleague and LUCENE-644. This approach needs highlight fields to be 
> TermVector.WITH_POSITIONS_OFFSETS, but is fast and can support N-grams. This 
> depends on LUCENE-1448 to get refined term offsets.
> usage:
> {code:java}
> TopDocs docs = searcher.search( query, 10 );
> Highlighter h = new Highlighter();
> FieldQuery fq = h.getFieldQuery( query );
> for( ScoreDoc scoreDoc : docs.scoreDocs ){
>   // fieldName="content", fragCharSize=100, numFragments=3
>   String[] fragments = h.getBestFragments( fq, reader, scoreDoc.doc, 
> "content", 100, 3 );
>   if( fragments != null ){
>     for( String fragment : fragments )
>       System.out.println( fragment );
>   }
> }
> {code}
> features:
> - fast for large docs
> - supports not only whitespace-based token stream, but also "fixed size" 
> N-gram (e.g. (2,2), not (1,3)) (can solve LUCENE-1489)
> - supports PhraseQuery, phrase-unit highlighting with slops
> {noformat}
> q="w1 w2"
> <b>w1 w2</b>
> ---------------
> q="w1 w2"~1
> <b>w1</b> w3 <b>w2</b> w3 <b>w1 w2</b>
> {noformat}
> - highlight fields need to be TermVector.WITH_POSITIONS_OFFSETS
> - easy to apply patch due to independent package (contrib/highlighter2)
> - uses Java 1.5
> - looks query boost to score fragments (currently doesn't see idf, but it 
> should be possible)
> - pluggable FragListBuilder
> - pluggable FragmentsBuilder
> to do:
> - term positions can be unnecessary when phraseHighlight==false
> - collects performance numbers

-- 
This message is automatically generated by JIRA.
-
You can reply to this email to add a comment to the issue online.


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

[jira] Commented: (LUCENE-1522) another highlighter

Reply via email to