Showing posts with label Wikipedia. Show all posts
Showing posts with label Wikipedia. Show all posts

Sunday, 30 August 2015

Wiki Parallel Data Extractor

Link: https://github.com/clab/wikipedia-parallel-titles
Intro: Tools for extracting parallel corpora from article titles across languages in Wikipedia

Monday, 11 May 2015

Tuesday, 13 May 2014

Wikipedia Extractor

Link: http://medialab.di.unipi.it/wiki/Wikipedia_Extractor
Intro: WikiExtractor.py is a Python script that extracts and cleans text from a Wikipedia database dump. The output is stored in a number of files of similar size in a given directory. Each file contains several documents in the document format.

Tuesday, 29 January 2013

The DBpedia Data Set

Linkhttp://wiki.dbpedia.org/Datasets
Intro: The DBpedia data set uses a large multi-domain ontology which has been derived from Wikipedia. The English version of the DBpedia data set currently describes 3.77 million “things” with 400 million “facts”.

Monday, 21 January 2013

Explicit Semantic Analysis - ESA

Intro: ESA is a vector representation of texts based on Wikippedia as external knowledge base.
Linkhttp://www.cs.technion.ac.il/~gabr/resources/code/esa/esa.html

Wednesday, 3 March 2010

Wikipedia issues

All Wikipedia related issues will be posted here:

Dumps of Wikipedia: http://download.wikimedia.org
Extraction of plain text corpus from Wikipedia: http://blog.afterthedeadline.com/2009/12/04/generating-a-plain-text-corpus-from-wikipedia/

--
Cheers,
Vu