Make sure you include: urlnormalizer-regex in your nutch-site.xml to
load the filter
<name>plugin.includes</name>
<value>nutch-extensionpoints|urlnormalizer-regex|protocol-http|urlfilter-(regex|suffix)|parse-(text|html)|index-(basic)|query-(basic|site|url)|summary-basic|scoring-opic|clustering-carrot2|ontology</value>
Also, in your pattern you have an error. "\&|\&" are both the
same thing and the regex requires "&" to be "&", hence they are the same.
<pattern>(\?|\&)jsessionid=[a-zA-Z0-9]{32}(\&)(.*)</pattern>
At 01:43 AM 3/23/2007, you wrote:
hi,
am not able to remove jsessionid while i crawl the web.
I have tried following
<regex>
<pattern>(\?|\&|\&)jsessionid=[a-zA-Z0-9]{32}$</pattern>
<substitution></substitution>
</regex>
<regex>
<pattern>(\?|\&|\&)jsessionid=[a-zA-Z0-9]{32}(\&|\&)(.*)</pattern>
<substitution></substitution>
</regex>
<regex>
am missing something.
Cheers,
Cha
--
View this message in context:
http://www.nabble.com/removing-jsessionid-tf3451965.html#a9629084
Sent from the Nutch - User mailing list archive at Nabble.com.
-------------------------------------------------------------------------
Take Surveys. Earn Cash. Influence the Future of IT
Join SourceForge.net's Techsay panel and you'll get the chance to share your
opinions on IT & business topics through brief surveys-and earn cash
http://www.techsay.com/default.php?page=join.php&p=sourceforge&CID=DEVDEV
_______________________________________________
Nutch-general mailing list
[email protected]
https://lists.sourceforge.net/lists/listinfo/nutch-general