Saturday, 28 May 2011

A USENET corpus (2005-2010)

http://www.psych.ualberta.ca/~westburylab/downloads/usenetcorpus.download.html

This corpus was collected between Oct 2005 and Jan 2011, and covers 47860 English language, non-binary-file news groups:
+ Corpus size: over 30 billion words,
+ Data size: over 34Gb, compressed (delivered as weekly bundles of about 150 Mb each.)

Thursday, 19 May 2011

Mining scientific texts

This post is to collect all papers related to mining scientific texts (entity & relation extraction, summarization, ...).

1) http://www.lrec-conf.org/proceedings/lrec2008/pdf/773_paper.pdf
(Extracting and Querying Relations in Scientific Papers on Language Technology)

2)

Monday, 16 May 2011

hunalign – sentence aligner

http://mokk.bme.hu/resources/hunalign/

Intro

hunalign aligns bilingual text on the sentence level. Its input is tokenized and sentence-segmented text in two languages. In the simplest case, its output is a sequence of bilingual sentence pairs (bisentences).

In the presence of a dictionary, hunalign uses it, combining this information with Gale-Church sentence-length information. In the absence of a dictionary, it first falls back to sentence-length information, and then builds an automatic dictionary based on this alignment. Then it realigns the text in a second pass, using the automatic dictionary.

Like most sentence aligners, hunalign does not deal with changes of sentence order: it is unable to come up with crossing alignments, i.e., segments A and B in one language corresponding to segments B’ A’ in the other language.

There is nothing Hungarian-specific in hunalign, the name simply reflects the fact that it is part of the hun* NLP toolchain.

hunalign was written in portable C++. It can be built under basically any kind of operating system.

YouAlign - Online document alignment solution

http://www.youalign.com/

"Welcome to YouAlign, your online document alignment solution. No software to purchase, no software to install. With YouAlign you can quickly and easily create bitexts from your archived documents. A YouAlign bitext contains a document and its translation aligned at the sentence level. YouAlign generates TMX files that can be loaded into your translation memory. YouAlign can also generate HTML files that you can publish on the Internet, or use with a full-text search engine to search for terminology and phraseology in context.

YouAlign is powered by the AlignFactory engine, which supports all kinds of formats, including Microsoft Word, Excel and PowerPoint, PDF, HTML, XML, Corel WordPerfect, RTF, Lotus WordPro and plain text."

Thursday, 12 May 2011

Google Books Corpus

http://googlebooks.byu.edu/

"This corpus is based on the American English portion of the Google Books data (see http://ngrams.googlelabs.com and especially http://ngrams.googlelabs.com/datasets). It contains 155 *billion* words (155,000,000,000) in more than 1.3 million books from the 1810s-2000s (including 62 billion words from just 1980-2009).

The corpus has most of the functionality of the other corpora from http://corpus.byu.edu (e.g. COCA, COHA, and our interface to the BNC), including: searching by part of speech, wildcards, and lemma (and thus advanced syntactic searches), synonyms, collocate searches, frequency by decade (tables listing each individual string, or charts for total frequency), comparisons of two historical periods (e.g. collocates of "women" or "music" in the 1800s and the 1900s), and more." (From Corpora-List)

Tuesday, 3 May 2011

Interested Papers at SIGIR 2011

SIGIR 2011 accepted papers link: http://www.sigir2011.org/papers.htm

My Interested Papers:
1) Summarizing the Differences in Multilingual News
Xiaojun Wan, Houping Jia

2)
Multifaceted Toponym Recognition for Streaming News
Michael Lieberman, Hanan Samet

3)
Toward Social Context Summarization For Web Documents
Zi Yang, cai keke, Jie Tang, Li Zhang, Zhong Su, Juanzi Li

4)
Evolutionary Timeline Summarization: a Balanced Optimization Framework via Iterative Substitution
Rui Yan, Xiaojun Wan, Jahna Otterbacher, Liang Kong, Yan Zhang, Xiaoming Li

5)
The Economics in Interactive Information Retrieval
Leif Azzopardi

6)
Composite Hashing with Multiple Information Sources
Dan Zhang, Fei Wang, Luo Si

7)
Inverted Indexes for Phrases and Strings
Manish Patil, Sharma V. Thankachan, Rahul Shah, Wing-Kai Hon, Jeffrey Vitter, Sabrina Chandrasekaran

8)
Multimedia Answering: Enriching Text QA with Media Information
Liqiang Nie, Meng Wang, Zha Zhengjun, Li Guangda, Tat Seng Chua

9)
SCENE : A Scalable Two-Stage Personalized News Recommendation System
Lei Li, Dingding Wang, Tao Li

10)
Ranking Related News Predictions
Nattiya Kanhabua, Roi Blanco, Michael Matthews

--
Cheers,
Vu