Link: https://github.com/clab/wikipedia-parallel-titles
Intro: Tools for extracting parallel corpora from article titles across languages in Wikipedia
Showing posts with label Wikipedia. Show all posts
Showing posts with label Wikipedia. Show all posts
Sunday, 30 August 2015
Monday, 11 May 2015
Wikipedia and LM
Link: http://trulymadlywordly.blogspot.sg/2011/03/creating-text-corpus-from-wikipedia.html
Intro: How to leverage Wikipedia repository to create huge LM data.
Intro: How to leverage Wikipedia repository to create huge LM data.
Tuesday, 13 May 2014
Wikipedia Extractor
Link: http://medialab.di.unipi.it/wiki/Wikipedia_Extractor
Intro: WikiExtractor.py is a Python script that extracts and cleans text from a Wikipedia database dump. The output is stored in a number of files of similar size in a given directory. Each file contains several documents in the document format.
Tuesday, 29 January 2013
The DBpedia Data Set
Link: http://wiki.dbpedia.org/Datasets
Intro: The DBpedia data set uses a large multi-domain ontology which has been derived from Wikipedia. The English version of the DBpedia data set currently describes 3.77 million “things” with 400 million “facts”.
Labels:
corpus,
DBpedia,
information extraction,
large-scale,
link,
multi-lingual,
NLP,
Wikipedia
Monday, 21 January 2013
Explicit Semantic Analysis - ESA
Intro: ESA is a vector representation of texts based on Wikippedia as external knowledge base.
Link: http://www.cs.technion.ac.il/~gabr/resources/code/esa/esa.html
Link: http://www.cs.technion.ac.il/~gabr/resources/code/esa/esa.html
Labels:
esa,
Explicit Semantic Analysis,
NLP,
research,
semantic relatedness,
Wikipedia
Wednesday, 3 March 2010
Wikipedia issues
All Wikipedia related issues will be posted here:
Dumps of Wikipedia: http://download.wikimedia.org
Extraction of plain text corpus from Wikipedia: http://blog.afterthedeadline.com/2009/12/04/generating-a-plain-text-corpus-from-wikipedia/
--
Cheers,
Vu
Dumps of Wikipedia: http://download.wikimedia.org
Extraction of plain text corpus from Wikipedia: http://blog.afterthedeadline.com/2009/12/04/generating-a-plain-text-corpus-from-wikipedia/
--
Cheers,
Vu
Subscribe to:
Posts (Atom)