[ https://issues.apache.org/jira/browse/LUCENE-1166?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel ]
Thomas Peuss updated LUCENE-1166: --------------------------------- Attachment: CompoundTokenFilter.patch Updated version: * new dumb decomposition filter ** uses a brute-force approach by generating substrings and checking them against the dictionary ** seems to work better for languages that have no patterns file with a lot of special cases ** Is roughly 3 times slower than the decomposition filter using hyphenation patterns ** No licensing problems because of the hyphenation pattern files * Refactoring to have all methods used by both decomposition filters in one place * Minor performance improvements > A tokenfilter to decompose compound words > ----------------------------------------- > > Key: LUCENE-1166 > URL: https://issues.apache.org/jira/browse/LUCENE-1166 > Project: Lucene - Java > Issue Type: New Feature > Components: Analysis > Reporter: Thomas Peuss > Attachments: CompoundTokenFilter.patch, CompoundTokenFilter.patch, > CompoundTokenFilter.patch, de.xml, hyphenation.dtd > > > A tokenfilter to decompose compound words you find in many germanic languages > (like German, Swedish, ...) into single tokens. > An example: Donaudampfschiff would be decomposed to Donau, dampf, schiff so > that you can find the word even when you only enter "Schiff". > I use the hyphenation code from the Apache XML project FOP > (http://xmlgraphics.apache.org/fop/) to do the first step of decomposition. > Currently I use the FOP jars directly. I only use a handful of classes from > the FOP project. > My question now: > Would it be OK to copy this classes over to the Lucene project (renaming the > packages of course) or should I stick with the dependency to the FOP jars? > The FOP code uses the ASF V2 license as well. > What do you think? -- This message is automatically generated by JIRA. - You can reply to this email to add a comment to the issue online. --------------------------------------------------------------------- To unsubscribe, e-mail: [EMAIL PROTECTED] For additional commands, e-mail: [EMAIL PROTECTED]