[
https://issues.apache.org/jira/browse/SOLR-3653?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Lance Norskog updated SOLR-3653:
--------------------------------
Description:
The "Smart" Simplified Chinese toolkit in lucene/analysis/smartcn does not work
in some edge cases. It fails to split certain words which were not part of the
dictionary or training corpus.
This patch supplies a bigramming class to handle these occasional mistakes. The
algorithm creates bigrams out of all "words" longer than two ideograms.
was:
The "Smart" Simplified Chinese toolkit in lucene/analysis/smartcn has no Solr
factories. Also, since it is a statistical algorithm, it is not perfect.
This patch supplies factories and a schema.xml type for the existing Lucene
Smart Chinese implementation, and includes a "fixup" class to handle the
occasional mistake made by the Smart Chinese implementation.
> Custom bigramming filter for to handle Smart Chinese edge cases
> ---------------------------------------------------------------
>
> Key: SOLR-3653
> URL: https://issues.apache.org/jira/browse/SOLR-3653
> Project: Solr
> Issue Type: New Feature
> Components: Schema and Analysis
> Reporter: Lance Norskog
> Attachments: SOLR-3653.patch, SmartChineseType.pdf
>
>
> The "Smart" Simplified Chinese toolkit in lucene/analysis/smartcn does not
> work in some edge cases. It fails to split certain words which were not part
> of the dictionary or training corpus.
> This patch supplies a bigramming class to handle these occasional mistakes.
> The algorithm creates bigrams out of all "words" longer than two ideograms.
--
This message is automatically generated by JIRA.
If you think it was sent incorrectly, please contact your JIRA administrators:
https://issues.apache.org/jira/secure/ContactAdministrators!default.jspa
For more information on JIRA, see: http://www.atlassian.com/software/jira
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]