[jira] Commented: (SOLR-799) Add support for hash based exact/near duplicate document handling

Thomas Heigl (JIRA) Wed, 24 Mar 2010 00:45:55 -0700

    [ 
https://issues.apache.org/jira/browse/SOLR-799?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=12849091#action_12849091
 ]


Thomas Heigl commented on SOLR-799:
-----------------------------------

Hello,

For my current project I need to implement an index-time mechanism to detect 
(near) duplicate documents. The TextProfileSignature available out-of-the-box 
(http://wiki.apache.org/solr/Deduplication) seems alright but does not use 
global collection statistics in deciding which terms will be used for 
calculating the signature. 
Most state-of-the-art hash-based duplication detection algorithms make use of 
this information to improve precision and recall (e.g. 
http://portal.acm.org/citation.cfm?id=506311&dl=GUIDE&coll=GUIDE&CFID=83187370&CFTOKEN=47052122)

Is it possible to access collection statistics - especially IDF values for all 
non-discarded terms in the current document - from within an implementation of 
the Signature class?

Kind regards,

Thomas


> Add support for hash based exact/near duplicate document handling
> -----------------------------------------------------------------
>
>                 Key: SOLR-799
>                 URL: https://issues.apache.org/jira/browse/SOLR-799
>             Project: Solr
>          Issue Type: New Feature
>          Components: update
>            Reporter: Mark Miller
>            Assignee: Yonik Seeley
>            Priority: Minor
>             Fix For: 1.4
>
>         Attachments: SOLR-799.patch, SOLR-799.patch, SOLR-799.patch, 
> SOLR-799.patch
>
>
> Hash based duplicate document detection is efficient and allows for blocking 
> as well as field collapsing. Lets put it into solr. 
> http://wiki.apache.org/solr/Deduplication

-- 
This message is automatically generated by JIRA.
-
You can reply to this email to add a comment to the issue online.

[jira] Commented: (SOLR-799) Add support for hash based exact/near duplicate document handling

Reply via email to