Wednesday, 11 July 2012
Interested Papers at ACL 2012
1) New NLP topic: automatic document dating
P12-1011: Nathanael Chambers
Labeling Documents with Timestamps: Learning from their Time Expressions
2)
P12-1050: Arianna Bisazza; Marcello Federico
Modified Distortion Matrices for Phrase-Based Statistical Machine Translation
3) ...
Sunday, 8 July 2012
Saffron - Extracting the Valuable Threads of Expertise
Link: http://saffron.deri.ie/acl
Intro: Saffron provides insights in a research community or organization by analysing its main topics of investigation and the experts associated with these topics.
Saffron analysis is fully automatic and is based on text mining and linked data principles.
This instance of Saffron analyzes the research community in Natural Language Processing based on the proceedings of the conferences organized by the Association for Computational Linguistics (ACL).
Intro: Saffron provides insights in a research community or organization by analysing its main topics of investigation and the experts associated with these topics.
Saffron analysis is fully automatic and is based on text mining and linked data principles.
This instance of Saffron analyzes the research community in Natural Language Processing based on the proceedings of the conferences organized by the Association for Computational Linguistics (ACL).
Tuesday, 3 July 2012
C++ web development frameworks
1) Witty: http://www.webtoolkit.eu/wt
2) CppCMS: http://cppcms.com/wikipp/en/page/main
3) POCO: http://pocoproject.org/
4) ...
2) CppCMS: http://cppcms.com/wikipp/en/page/main
3) POCO: http://pocoproject.org/
4) ...
Labels:
C programming,
C++,
framework,
links,
web programming
Saturday, 23 June 2012
jusText
http://code.google.com/p/justext/
jusText is a tool for removing boilerplate content, such as navigation
links, headers, and footers from HTML pages. It is designed to preserve
mainly text containing full sentences and it is therefore well suited
for creating linguistic resources such as Web corpora.
Friday, 1 June 2012
WIT3 - Web Inventory of Transcribed and Translated Talks
Link: https://wit3.fbk.eu/
Intro: WIT3 - acronym for Web Inventory of Transcribed and Translated Talks - is a ready-to-use version for research purposes of the multilingual transcriptions of TED talks.
Since 2007, the TED Conference has been posting on its website all video recordings of its talks, English subtitles and their translations in more than 80 languages. In order to make this collection of talks more effectively usable by the research community, the original textual contents are redistributed here, together with MT benchmarks and processing tools.
Labels:
links,
machine translation,
parallel corpora,
TED talks
Tuesday, 29 May 2012
Layout-Aware Text Extraction from Full-text PDF of Scientific Articles
Description: The Portable Document Format (PDF) is the almost universally used file format for online scientific publications. It is also notoriously difficult to read and handle computationally, presenting challenges for developers of biomedical text mining or biocuration informatics systems that use the published literature as an information source. To facilitate the effective use of scientific literature in such systems we introduce Layout-Aware PDF Text Extraction (LA-PDFText). The LA-PDFText system focuses only on the textual content of the research articles and is meant as a baseline for further experiments into more advanced extraction methods that handle multi-modal content, such as images and graphs. The system works in a three-stage process: (1) Detecting contiguous text blocks using spatial layout processing to locate and identify blocks of contiguous text, (2) Classifying text blocks into rhetorical categories using a rule-based method and (3) Stitching classified text blocks together in the correct order resulting in the extraction of text from section-wise grouped blocks.
Monday, 16 April 2012
Subscribe to:
Posts (Atom)