[
https://issues.apache.org/jira/browse/TIKA-433?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=12871720#action_12871720
]
Grant Ingersoll commented on TIKA-433:
--------------------------------------
I think it makes sense as a Tika contrib, but that's not for me to determine.
It seems like it is more generally useful than just Behemoth or Mahout and fits
well with what Tika does, along the lines of Tika's command line tool. I don't
have a use for Behemoth and don't wish to inject it into my dep. chain, whereas
I am already using Tika and Hadoop.
> Tika + Hadoop
> -------------
>
> Key: TIKA-433
> URL: https://issues.apache.org/jira/browse/TIKA-433
> Project: Tika
> Issue Type: New Feature
> Components: general
> Reporter: Grant Ingersoll
> Priority: Minor
>
> Would be great to have a Tika contrib that took in an HDFS location with
> "rich" documents on it and an output format (or output processor) and
> converted the docs to XHTML or Solr or whatever. Seems like it should be
> pretty straightforward to do on the Hadoop side of things. Only tricky part,
> I suppose, is the output format and how to make that pluggable.
--
This message is automatically generated by JIRA.
-
You can reply to this email to add a comment to the issue online.