Intro: ESA is a vector representation of texts based on Wikippedia as external knowledge base.
Link: http://www.cs.technion.ac.il/~gabr/resources/code/esa/esa.html
Monday, 21 January 2013
Wednesday, 16 January 2013
Statistical Methods in Language and Linguistic Research
Link: https://www.equinoxpub.com/equinox/books/showbook.asp?bkid=348&keyword=
Monday, 17 December 2012
String Toolkit for Text Processing
Link: http://www.codeproject.com/Articles/23198/C-String-Toolkit-StrTk-Tokenizer
It can be used as lite version of Boost.
It can be used as lite version of Boost.
Tuesday, 13 November 2012
Monday, 5 November 2012
Multeval
Link: https://github.com/jhclark/multeval
Intro: MultEval takes machine translation hypotheses from several runs of an optimizer and provides 3 popular metric scores, as well as, standard deviations (via bootstrap resampling) and p-values (via approximate randomization). This allows researchers to mitigate some of the risk of using unstable optimizers such as MERT, MIRA, and MCMC. It is intended to help in evaluating the impact of in-house experimental variations on translation quality; it is currently not setup to do bake-off style comparisons (bake-offs can't require multiple optimizer runs nor a standard tokenization).
Related: http://www.ark.cs.cmu.edu/MT/ (Code for Statistical Significance Testing for MT Evaluation Metrics)Saturday, 3 November 2012
MacPorts
Link: http://www.macports.org/index.php
Intro: The MacPorts Project is an open-source community initiative to design an easy-to-use system for compiling, installing, and upgrading either command-line, X11 or Aqua based open-source software on the Mac OS X operating system. To that end we provide the command-line driven MacPorts software package under a BSD License, and through it easy access to thousands of ports that greatly simplify the task of compiling and installing open-source software on your Mac.
Tuesday, 16 October 2012
TalkBank
Link: http://talkbank.org/
Intro: The goal of TalkBank is to foster fundamental research in the study of human and animal communication. It will construct sample databases within each of the subfields studying communication. It will use these databases to advance the development of standards and tools for creating, sharing, searching, and commenting upon primary materials via networked computers.
Labels:
child language,
data,
language learning,
NLP,
research,
speech
Subscribe to:
Posts (Atom)