[jira] [Commented] (LUCENE-8816) Decouple Kuromoji's morphological analyser and its dictionary

Mike Sokolov (JIRA) Tue, 11 Jun 2019 16:34:35 -0700


    [ 
https://issues.apache.org/jira/browse/LUCENE-8816?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=16861609#comment-16861609
 ]


Mike Sokolov commented on LUCENE-8816:
--------------------------------------

I see that in {{BinaryDictionaryWriter}} we restrict incoming leftID (and 
rightID) to be < 4096 because we are going to pack into a 16-bit short with 3 
flag bits. However it seems we have room for one more bit (since 2^(16-3) == 
8192). Am I missing something? Do we use that other bit somewhere? I see eg 
that in {{BinaryDictionary}} when we decode, we >>> 3 to get back the ids, so I 
think it should be OK to allow ids up to 8191. [~rcmuir] do you know why it is 
currenly limited to 4096? Also I think it would make sense to change the 
asserts there to be IllegalArgumentException so they are raised whenever the 
tool is run, since we would get garbage if this limit is exceeded, and (I 
think) nothing else will catch it.

> Decouple Kuromoji's morphological analyser and its dictionary
> -------------------------------------------------------------
>
>                 Key: LUCENE-8816
>                 URL: https://issues.apache.org/jira/browse/LUCENE-8816
>             Project: Lucene - Core
>          Issue Type: Improvement
>          Components: modules/analysis
>            Reporter: Tomoko Uchida
>            Priority: Major
>
> I've inspired by this mail-list thread.
>  
> [http://mail-archives.apache.org/mod_mbox/lucene-java-user/201905.mbox/%3CCAGUSZHA3U_vWpRfxQb4jttT7sAOu%2BuaU8MfvXSYgNP9s9JNsXw%40mail.gmail.com%3E]
> As many Japanese already know, default built-in dictionary bundled with 
> Kuromoji (MeCab IPADIC) is a bit old and no longer maintained for many years. 
> While it has been slowly obsoleted, well-maintained and/or extended 
> dictionaries risen up in recent years (e.g. 
> [mecab-ipadic-neologd|https://github.com/neologd/mecab-ipadic-neologd], 
> [UniDic|https://unidic.ninjal.ac.jp/]). To use them with Kuromoji, some 
> attempts/projects/efforts are made in Japan.
> However current architecture - dictionary bundled jar - is essentially 
> incompatible with the idea "switch the system dictionary", and developers 
> have difficulties to do so.
> Traditionally, the morphological analysis engine (viterbi logic) and the 
> encoded dictionary (language model) had been decoupled (like MeCab, the 
> origin of Kuromoji, or lucene-gosen). So actually decoupling them is a 
> natural idea, and I feel that it's good time to re-think the current 
> architecture.
> Also this would be good for advanced users who have customized/re-trained 
> their own system dictionary.
> Goals of this issue:
>  * Decouple JapaneseTokenizer itself and encoded system dictionary.
>  * Implement dynamic dictionary load mechanism.
>  * Provide developer-oriented dictionary build tool.
> Non-goals:
>   * Provide learner or language model (it's up to users and should be outside 
> the scope).
> I have not dove into the code yet, so have no idea about it's easy or 
> difficult at this moment.



--
This message was sent by Atlassian JIRA
(v7.6.3#76005)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

[jira] [Commented] (LUCENE-8816) Decouple Kuromoji's morphological analyser and its dictionary

Reply via email to