Link: http://trulymadlywordly.blogspot.sg/2011/03/creating-text-corpus-from-wikipedia.html
Intro: How to leverage Wikipedia repository to create huge LM data.
Showing posts with label LM. Show all posts
Showing posts with label LM. Show all posts
Monday, 11 May 2015
Sunday, 1 March 2015
Training Google 1T web corpus with IRSTLM
Thanks to http://www44.atwiki.jp/keisks/pages/50.html , here is the way to train enormous LM with Google 1T web corpus using IRSTLM:
I will validate it soon.
build-sublm.pl --size 3 --ngrams "gunzip -c 3gms/*.gz" --sublm LM.000 --witten-bell merge-sublm.pl --size 3 --sublm LM -lm g_3grams_LM.gz compile-lm g_3grams_LM.gz g_3grams_LM.blm (if you get the error: "lt-compile-lm: lmtable.h:247: virtual double lmtable::setlogOOVpenalty(int): Assertion `dub > dict->size()' failed.") compile-lm -dub=100000000 g_3grams_LM g_3grams_LM.blm (make the -dub option bigger)
I will validate it soon.
Thursday, 5 February 2015
Available Data from the CommonCrawl
Link: http://statmt.org/ngrams/
Intro: Multi-language data used for training large-scale LMs crawled from CommonCrawl.
Intro: Multi-language data used for training large-scale LMs crawled from CommonCrawl.
Labels:
data,
large-scale,
LM,
machine translation,
monolingual,
SMT
Monday, 1 December 2014
The ClueWeb09 Dataset
Intro: The ClueWeb09 dataset was created to support research on information retrieval and related human language technologies. It consists of about 1 billion web pages in ten languages that were collected in January and February 2009. The dataset is used by several tracks of the TREC conference.
Note: Huge corpus for LM
Sentence level of parallel texts
Intro: Bleualign is a tool to align parallel texts (i.e. a text and its translation) on a sentence level. Additionally to the source and target text, Bleualign requires an automatic translation of at least one of the texts. The alignment is then performed on the basis of the similarity (modified BLEU score) between the source text sentences (translated into the target language) and the target text sentences
Clustering of Parallel Text
Intro: This program performs sentence-level k-means clustering for parallel texts based on language model similarity.
Subscribe to:
Posts (Atom)