[ 
https://issues.apache.org/jira/browse/NUTCH-1149?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

Markus Jelsma closed NUTCH-1149.
--------------------------------
    Resolution: Won't Fix

Will upload proper patch for NUTCH-1325 soon which already contains numeric 
aggregations for CrawlDB metadata.

> DomainStats should process numeric CrawlDB metadata
> ---------------------------------------------------
>
>                 Key: NUTCH-1149
>                 URL: https://issues.apache.org/jira/browse/NUTCH-1149
>             Project: Nutch
>          Issue Type: Improvement
>            Reporter: Markus Jelsma
>            Assignee: Markus Jelsma
>            Priority: Trivial
>
> Right now the DomainStats program only outputs the sum of fetched records per 
> domain or host. It should also be able to output processed numerics of meta 
> data in order to get the average size (content length) for a given domain or 
> host. This is also useful for generating a metric for adult material (by 
> domain or host) when using a plugin that stores a propability factor of adult 
> material per URL in the Crawl DB.



--
This message was sent by Atlassian JIRA
(v6.3.4#6332)

Reply via email to