Link: http://phontron.com/japanese-translation-data.php
Showing posts with label SMT. Show all posts
Showing posts with label SMT. Show all posts
Thursday, 23 July 2015
Wednesday, 8 July 2015
NAACL 2015 papers
Here is my subjective list of remarkable papers relating to MT research:
*** Neural Machine Translation
Paul Baltescu and Phil Blunsom. "Pragmatic Neural Language Modelling in Machine Translation"
Adrià de Gispert, Gonzalo Iglesias, Bill Byrne. "Fast and Accurate Preordering for SMT using Neural Networks"
*** Continuous Models for Statistical Machine Translation
Frédéric Blain, Fethi Bougares, Amir Hazem, Loïc Barrault, Holger Schwenk. "Continuous Adaptation to User Feedback for Statistical Machine Translation"
Kai Zhao, Hany Hassan, Michael Auli. "Learning Translation Models from Monolingual Continuous Representations"
*** Multi-language Translation
Raj Dabre, Fabien Cromieres, Sadao Kurohashi, Pushpak Bhattacharyya. "Leveraging Small Multilingual Corpora for SMT Using Many Pivot Languages"
*** Video to Text Translation
Subhashini Venugopalan, Huijuan Xu, Jeff Donahue, Marcus Rohrbach, Raymond Mooney, Kate Saenko. "Translating Videos to Natural Language Using Deep Recurrent Neural Networks"
*** Others
Jonathan H. Clark, Chris Dyer, Alon Lavie. "Locally Non-Linear Learning for Statistical Machine Translation via Discretization and Structured Regularization"
Graham Neubig, Philip Arthur, Kevin Duh. "Multi-Target Machine Translation with Multi-Synchronous Context-free Grammars"
Aurelien Waite and Bill Byrne. "The Geometry of Statistical Machine Translation"
Other papers are also worth reading:
*** News Processing
Areej Alhothali and Jesse Hoey. "Good News or Bad News: Using Affect Control Theory to Analyze Readers' Reaction Towards News Articles"
Other papers are also worth reading:
*** News Processing
Areej Alhothali and Jesse Hoey. "Good News or Bad News: Using Affect Control Theory to Analyze Readers' Reaction Towards News Articles"
Labels:
computational linguistics,
conference,
MT,
NAACL,
NLP,
papers,
research,
SMT
Tuesday, 7 July 2015
Python wrapper for online translators
If you want to use Google Translate and Microsoft Bing Translate for free, you may consider the following Python-based wrappers:
+ Code: Google Translate; Bing Translate
+ Samples:
# Google Translate
import googletrans
gs = googletrans.Googletrans()
import googletrans
gs = googletrans.Googletrans()
languages = gs.get_languages()
print(languages['en'])
print(languages['en'])
print(gs.translate('hello', 'de'))
print(gs.translate('hello', 'zh'))
print(gs.translate('hello', 'vi'))
print(gs.translate('hello', 'zh'))
print(gs.translate('hello', 'vi'))
print(gs.detect('some English words'))
#Bing Translate
from mstranslator import Translator
translator =
Translator('cdvhoang', 'HlUUMftdkETWa8E9/jzD4l1CzC8sOhRSJxH+kk0MDBg=')
Translator('cdvhoang', 'HlUUMftdkETWa8E9/jzD4l1CzC8sOhRSJxH+kk0MDBg=')
print(translator.translate('hello', lang_from='en', lang_to='vi'))
*** Please note that I don't encourage to use the wrapper for Google Translate because you should respect and pay for using its service (simply it's now commercialized ^_^).
Labels:
API,
Bing Translate,
Google Translate,
links,
MT,
online tools,
research,
service,
SMT
Thursday, 5 February 2015
Available Data from the CommonCrawl
Link: http://statmt.org/ngrams/
Intro: Multi-language data used for training large-scale LMs crawled from CommonCrawl.
Intro: Multi-language data used for training large-scale LMs crawled from CommonCrawl.
Labels:
data,
large-scale,
LM,
machine translation,
monolingual,
SMT
Monday, 5 January 2015
MTTK - Machine Translation Toolkit
Intro: MTTK is a collection of software tools for the alignment of parallel text for use in Statistical Machine Translation. With MTTK you can ...
- Align document translation pairs at the sentence or sub-sentence level, sometimes known as chunking. This is a useful pre-processing step to prepare collections of translations for use in estimating the parameters of complex alignment models. Sub-sentence alignment in particular makes it possible to segment long sentences into shorter aligned segments that otherwise would have to be discarded.
- Train statistical models for parallel text alignment. The following models are supported :
- IBM Model-1 and Model-2
- Word-to-Word HMMs
- Word-to-Phrase HMMs , with bigram translation probabilities
- Parallelize your model training procedures. If you have multiple CPUs available, you can partition your translation training texts into subsets, thus speeding up iterative parameter re-estimation procedures and reducing the amount of memory needed in training. This is done under exact EM-based parameter estimation procedures.
- Generate word-to-word and word-to-phrase alignments of parallel text. MTTK can generate Viterbi alignments of parallel text (both training text and other texts) under the supported alignment models.
- Extract word-to-word translation tables from aligned bitext and from the estimated models.
- Extract phrase-to-phrase translation tables (phrase-pair inventories) from aligned parallel text.
- Use the HMM alignment models to induce phrase translations under its statistical models. Phrase-pair induction can generate richer inventories of phrase translations than can be extracted from Viterbi alignments.
- Edit the C++ source code to implement your own estimation and alignment procedures.
Labels:
alignment,
IBM models,
MT,
NLP,
parallel text,
research,
SMT,
toolkit
Thursday, 11 December 2014
Neural Machine Translation
Scientists around the world (especially Google guys) are moving the approaches of Statistical Machine Translation (SMT) (e.g. word-based, statistical with phrase-based or hierarchical, syntax-based) to the next level, namely Neural Machine Translation.
In general, Neural Machine Translation aims to simplify the SMT approaches by taking the source as an input sequence and producing the target as an output sequence via a single, large neural networks.
Here I am trying to catch up the recent progress of Neural Machine Translation.
*** People & Group
1) LISA Lab, University of Montreal 2014 led by Prof. Yoshua Bengio
Latest Demo: http://104.131.78.120/
3) Phil Blunsom's group at Oxford Uni.
4) Dzmitry Bahdanau at Jacobs University Bremen
5) Richard Socher at Stanford Uni.
6) Kyunghyun Cho at NYU?
7) ...
4) Dzmitry Bahdanau at Jacobs University Bremen
5) Richard Socher at Stanford Uni.
6) Kyunghyun Cho at NYU?
7) ...
*** Notable Papers
1a) Sequence to Sequence Learning with Neural Networks (Ilya Sutskever et al., NIPS 2014)
Note:
- The core idea behind neural machine translation.
1b) Generating Sequences With Recurrent Neural Networks (Alex Graves et al., ? 2014)
- TBA
1b) Generating Sequences With Recurrent Neural Networks (Alex Graves et al., ? 2014)
- TBA
2) Neural Machine Translation by Jointly Learning to Align and Translate (Dzmitry Bahdanau et al., EMNLP 2014)
3) On the Properties of Neural Machine Translation: Encoder–Decoder Approaches (Kyunghyun Cho et al., SSST-8 2014)
4) Addressing the Rare Word Problem in Neural Machine Translation (Thang Luong et al., drafted version 2014)
5) On Using Monolingual Corpora in Neural Machine Translation (Caglar Gulcehre et al., arXiv 2015)
6) Ask Me Anything: Dynamic Memory Networks for Natural Language Processing (Richard Socher and co. at MetaMind, arXiv June 2015)
- MT result still not yet released!
7) Effective Approaches to Attention-based Neural Machine Translation (Thang Luong et al., EMNLP'15)
8) (to be updated)
6) Ask Me Anything: Dynamic Memory Networks for Natural Language Processing (Richard Socher and co. at MetaMind, arXiv June 2015)
- MT result still not yet released!
7) Effective Approaches to Attention-based Neural Machine Translation (Thang Luong et al., EMNLP'15)
8) (to be updated)
In addition, some other approaches utilized neural processing to enhance the current state-of-the-art SMT framework, for example:
*** For Language Model:
1) Decoding with large-scale neural language models improves translation (Ashish et al., EMNLP 2013)
Note:
- Resulting toolkit: NPLM ver 0.3 (http://nlg.isi.edu/software/nplm/)
Comments
- It is quite hard to choose the optimized parameters (e.g. hidden layer nodes, input and output embedding dimensions) across data-sets and domains.
- In Moses, NPLM feature will slow down the decoder speed.
- It actually improves the translation performance when being used with n-gram LM features. But I am not sure whether it can completely replace n-gram LM features.
2) OxLM: A Neural Language Modelling Framework for Machine Translation (Paul Baltescu et al., The Prague Bulletin of Mathematical Linguistics 2014)
Note:
- Resulting toolkit: OxLM (https://github.com/pauldb89/oxlm)
- Moses already has this feature.
3) rwthlm - A toolkit for training neural network language models (feed-forward, recurrent, and long short-term memory neural networks). The software was written by Martin Sundermeyer.
4) (to be updated)
4) (to be updated)
*** For Translation Model:
1) Fast and Robust Neural Network Joint Models for Statistical Machine Translation (Devlin et al, ACL 2014)
Note:
- ACL 2014 best paper award.
- Accoding to the paper, they obtained a very impressive performance for Arabic-English Translation; good performance for Chinese-English Translation (datasets: OpenMT 2012, BOLT; domains: news, web forums).
- Moses already has this feature. Basic implementation of this model is already included in Moses under the name "BilingualLM".
- NPLM can be used to train the models for this.
Comments
- Personally, I tried this model with Moses and evaluated with conversational domains (e.g. SMS, Chat, conversational telephone speech) using OpenMT'15 datasets. I obtained good (but not very impressive, 0.7-1.0 BLEU score) performance compared to basic baseline. Using this model together with other strong features did not give significantly better performance as said in the paper :(.
- Optimizing parameters for this model is an exhausted task.
2) Learning Phrase Representations using RNN Encoder–Decoder for Statistical Machine Translation (Kyunghyun Cho et al., EMNLP 2014)
3) (to be updated)
*** For Reordering Model:
*** For Reordering Model:
1) Advancements in Reordering Models for Statistical Machine Translation (Minwei Feng et al., ACL 2013)
2) A Neural Reordering Model for Phrase-based Translation (Peng Li et al., COLING 2014)
3) (to be updated)
3) (to be updated)
Labels:
MT,
neural machine translation,
neural networks,
papers,
research,
SMT
Wednesday, 10 December 2014
Sunday, 10 November 2013
HandAlign
Link: http://www.umiacs.umd.edu/~hal/HandAlign/index.html
Intro: a tool to assist manual alignment in MT, summarization, ...
Intro: a tool to assist manual alignment in MT, summarization, ...
Labels:
MT,
NLP,
phrase alignment,
SMT,
text alignment,
toolkit,
visualization
Monday, 7 October 2013
pialign - Phrasal ITG Aligner
Intro: pialign is a package that allows you to create a phrase table and word alignments from an unaligned parallel corpus. It is unlike other unsupervised word alignment tools in that it is able to create a phrase table using a fully statistical model, no heuristics. As a result, it is able to build phrase tables for phrase-based machine translation that achieve competitive results but are only a fraction of the size of those created with heuristic methods.
*** Note: pialign can extract very compact phrase table directly from unaligned parallel data. This is may be very helpful for SMT system in mobile environment.
Sunday, 22 September 2013
YALIGN
Intro: Statistical Machine Translation relies on parallel corpora for training translation models. However these corpora are limited and take time to create. Yalign is designed to automate this process by finding sentences that are close translation matches from comparable corpora. This opens up avenues for harvesting parallel corpora from sources like translated documents and the web.
Sunday, 15 September 2013
Thursday, 11 April 2013
Nile - Syntax-based Word Alignment Tool
Link: https://code.google.com/p/nile/
Intro: Nile is a supervised, discriminative word alignment package that can make use of arbitrary and overlapping features
Monday, 5 November 2012
Multeval
Link: https://github.com/jhclark/multeval
Intro: MultEval takes machine translation hypotheses from several runs of an optimizer and provides 3 popular metric scores, as well as, standard deviations (via bootstrap resampling) and p-values (via approximate randomization). This allows researchers to mitigate some of the risk of using unstable optimizers such as MERT, MIRA, and MCMC. It is intended to help in evaluating the impact of in-house experimental variations on translation quality; it is currently not setup to do bake-off style comparisons (bake-offs can't require multiple optimizer runs nor a standard tokenization).
Related: http://www.ark.cs.cmu.edu/MT/ (Code for Statistical Significance Testing for MT Evaluation Metrics)Tuesday, 11 September 2012
NiuTrans: A Statistical Machine Translation System
Intro: NiuTrans is an open-source statistical machine translation system developed by the Natural Language Processing Group at Northeastern University, China. The NiuTrans system is fully developed in C++ language. So it runs fast and uses less memory. Currently it has already supported phrase-based, hierarchical phrase-based and syntax-based (string-to-tree, tree-to-string and tree-to-tree) models for research-oriented studies.
Paper (ACL 2012 Demo): http://www.nlplab.com/NiuPlan/paper/ACL2012-NiuTrans.pdf
Labels:
machine translation,
NLP,
research,
SMT,
statistical machine translation,
systems,
toolkits
Wednesday, 5 September 2012
Docent - Document-level SMT
Intro: Docent is a decoder for phrase-based Statistical Machine Translation (SMT). Unlike most existing SMT decoders, it treats complete documents, rather than single sentences, as translation units and permits the inclusion of features with cross-sentence dependencies to facilitate the development of discourse-level models for SMT. Docent implements the local search decoding approach described by Hardmeier et al. (EMNLP 2012).
Paper: https://aclweb.org/anthology-new/D/D12/D12-1108.pdf
Paper: https://aclweb.org/anthology-new/D/D12/D12-1108.pdf
Monday, 27 August 2012
ACCURAT Toolkit
Intro: The ACCURAT project (http://www.accurat-project.eu/) is pleased to announce the release of ACCURAT Toolkit - a collection of tools for comparable corpora collection and multi-level alignment and information extraction from comparable corpora. By using the ACCURAT Toolkit, users may obtain:
- Comparable corpora from the Web (current news corpora, filtered Wikipedia corpora, and narrow domain focussed corpora);
- Comparable document alignments;
- Semi-parallel sentence/phrase mapping from comparable corpora (for SMT training purposes or other tasks);
- Translated terminology extracted and mapped from bilingual comparable corpora;
- Translated named entities extracted and mapped from bilingual comparable corpora.
Labels:
comparable corpora,
machine translation,
SMT,
text alignment,
toolkits
Wednesday, 15 August 2012
Champollion Tool Kit - Text Sentence Aligner
Link: http://champollion.sourceforge.net/
Intro: Built around LDC's champollion sentence aligner kernel, Champollion Tool Kit (CTK) aims to providing ready-to-use parallel text sentence alignment tools for as many language pairs as possible.
Champollion depends heavily on lexical information, but uses sentence length information as well. A translation lexicon is required. Past experiments indicate that champollion's performance improves as the translation lexicon become larger.
Friday, 27 July 2012
PET - Post-Editing Translation Tool
Intro: PET is a stand-alone, open-source (under LGPL) tool written in Java that should help you post-edit and assess machine or human translations while gathering detailed statistics about post-editing time amongst other effort indicators.
Labels:
machine translation,
open source,
post-editing,
SMT,
tools
Tuesday, 24 July 2012
Subtitle Translation
Subtitle corpus: http://opus.lingfil.uu.se/ (more)
*** Google Translation API is no longer freely available. Can we use state-of-the-art SMT techniques to build a subtitle SMT system by ourself? What are challenges???
Labels:
open source,
research,
SMT,
statistical machine translation,
subtitle,
tools
Subscribe to:
Posts (Atom)