Showing posts with label large-scale. Show all posts
Showing posts with label large-scale. Show all posts

Thursday, 5 February 2015

Available Data from the CommonCrawl

Linkhttp://statmt.org/ngrams/
Intro: Multi-language data used for training large-scale LMs crawled from CommonCrawl.

Tuesday, 29 January 2013

The DBpedia Data Set

Linkhttp://wiki.dbpedia.org/Datasets
Intro: The DBpedia data set uses a large multi-domain ontology which has been derived from Wikipedia. The English version of the DBpedia data set currently describes 3.77 million “things” with 400 million “facts”.