Showing posts with label web crawler. Show all posts
Showing posts with label web crawler. Show all posts

Thursday, 29 December 2011

Common Crawl

http://www.commoncrawl.org/

Common Crawl Foundation is a California 501(c)3 non-profit founded by Gil Elbaz with the goal of democratizing access to web information by producing and maintaining an open repository of web crawl data that is universally accessible.

Saturday, 24 December 2011

WebSPHINX

Link: http://www.cs.cmu.edu/~rcm/websphinx/

Intro: WebSPHINX ( Website-Specific Processors for HTML INformation eXtraction) is a Java class library and interactive development environment for web crawlers. A web crawler (also called a robot or spider) is a program that browses and processes Web pages automatically.