Thursday, 13 January 2011
Wednesday, 12 January 2011
Training SMT incrementally
For word alignments
Force Alignment: http://geek.kyloo.net/software/doku.php/mgiza:forcealignment
MGIZA++: http://geek.kyloo.net/software/doku.php/mgiza:overview and http://sourceforge.net/projects/mgizapp/
For translation models
Stream-based Translation Models for Statistical Machine TranslationAbby Levenberg, Chris Callison-Burch and Miles Osborne, NAACL 2010.
For language models
Abby Levenberg and Miles Osborne, EMNLP 2009.
Sunday, 9 January 2011
Discourse Parsing
Sentence-based level:
SPADE: http://www.isi.edu/licensed-sw/spade/
Text-base level:
HILDA: http://nlp.prendingerlab.net/hilda/
NUS demo: http://wing.comp.nus.edu.sg/~linzihen/parser/demo.html
(to be continued!)
SPADE: http://www.isi.edu/licensed-sw/spade/
Text-base level:
HILDA: http://nlp.prendingerlab.net/hilda/
NUS demo: http://wing.comp.nus.edu.sg/~linzihen/parser/demo.html
(to be continued!)
Thursday, 6 January 2011
Open source toolkits for cloud computing
Eucalyptus: http://www.eucalyptus.com/
OpenNebula: http://www.opennebula.org/doku.php
Wednesday, 5 January 2011
[Book] - Foundations of Computer Science
http://infolab.stanford.edu/~ullman/focs.html
(VNese fresh CS students should consider reading this book!)
Scientext corpus
http://scientext.msh-alpes.fr
Scientext is a new, on-line French and English corpus of scientific texts. The corpus includes 4.8 million running tokens in French, 13 million words of research articles in English (medicine and biology), and an English-language sub-corpus of French undergraduate students’ texts (1,1 million words). The corpus is organized to facilitate the linguistic study of authorial position and reasoning in scientific articles through phraseology and lexico-grammatical markers linked to causality.
Labels:
computational linguistics,
corpus,
links,
NLP,
scientific text
Saturday, 1 January 2011
Subscribe to:
Posts (Atom)