Intro: Statistical Machine Translation relies on parallel corpora for training translation models. However these corpora are limited and take time to create. Yalign is designed to automate this process by finding sentences that are close translation matches from comparable corpora. This opens up avenues for harvesting parallel corpora from sources like translated documents and the web.
Showing posts with label comparable corpora. Show all posts
Showing posts with label comparable corpora. Show all posts
Sunday, 22 September 2013
Monday, 27 August 2012
ACCURAT Toolkit
Intro: The ACCURAT project (http://www.accurat-project.eu/) is pleased to announce the release of ACCURAT Toolkit - a collection of tools for comparable corpora collection and multi-level alignment and information extraction from comparable corpora. By using the ACCURAT Toolkit, users may obtain:
- Comparable corpora from the Web (current news corpora, filtered Wikipedia corpora, and narrow domain focussed corpora);
- Comparable document alignments;
- Semi-parallel sentence/phrase mapping from comparable corpora (for SMT training purposes or other tasks);
- Translated terminology extracted and mapped from bilingual comparable corpora;
- Translated named entities extracted and mapped from bilingual comparable corpora.
Labels:
comparable corpora,
machine translation,
SMT,
text alignment,
toolkits
Monday, 19 March 2012
Parallel Text Mining for SMT
Problem: given a relatively large collection of parallel texts and a state-of-the-art SMT system, how to incrementally & automatically mine the parallel texts available on the Web. The newly added texts should ensure to improve the current SMT system.
Papers related:
1) Large Scale Parallel Document Mining for Machine Translation. COLING 2010. Link.2) TBA
Labels:
comparable corpora,
machine translation,
NLP,
parallel corpora,
SMT,
text mining
Subscribe to:
Posts (Atom)