Intro: The DBpedia data set uses a large multi-domain ontology which has been derived from Wikipedia. The English version of the DBpedia data set currently describes 3.77 million “things” with 400 million “facts”.
Showing posts with label information extraction. Show all posts
Showing posts with label information extraction. Show all posts
Tuesday, 29 January 2013
The DBpedia Data Set
Link: http://wiki.dbpedia.org/Datasets
Labels:
corpus,
DBpedia,
information extraction,
large-scale,
link,
multi-lingual,
NLP,
Wikipedia
Wednesday, 28 March 2012
Calais
Link: http://www.opencalais.com/
Intro: The OpenCalais Web Service automatically creates rich semantic metadata for the content you submit – in well under a second. Using natural language processing (NLP), machine learning and other methods, Calais analyzes your document and finds the entities within it. But, Calais goes well beyond classic entity identification and returns the facts and events hidden within your text as well.
Labels:
information extraction,
machine learning,
NLP,
toolkits
Tuesday, 25 October 2011
AlchemyAPI - Transforming Text into Knowledge
http://www.alchemyapi.com
AlchemyAPI is a series of products of Alchemy company applied for knowledge extraction from text. The name "Alchemy" may make a confusion with Alchemy Open Source AI developed by University of Washington.
It is quite intersting to see how language technologies are used for real applications.
--
Cheers,
Vu
AlchemyAPI is a series of products of Alchemy company applied for knowledge extraction from text. The name "Alchemy" may make a confusion with Alchemy Open Source AI developed by University of Washington.
It is quite intersting to see how language technologies are used for real applications.
--
Cheers,
Vu
Labels:
application,
information extraction,
language technology,
NLP
Wednesday, 20 July 2011
Xtractor
http://www.xtractor.in/ - Really impressive! I am thinking about how to do such a similar service for computational linguistics domain.
Labels:
information extraction,
links,
NLP,
scientific text,
tools,
web service
Thursday, 19 May 2011
Mining scientific texts
This post is to collect all papers related to mining scientific texts (entity & relation extraction, summarization, ...).
1) http://www.lrec-conf.org/proceedings/lrec2008/pdf/773_paper.pdf
(Extracting and Querying Relations in Scientiļ¬c Papers on Language Technology)
2)
1) http://www.lrec-conf.org/proceedings/lrec2008/pdf/773_paper.pdf
(Extracting and Querying Relations in Scientiļ¬c Papers on Language Technology)
2)
Labels:
entity,
information extraction,
links,
NLP,
relation,
research,
scientific text,
summarization,
text mining
Thursday, 10 June 2010
Selected papers at ACL 2010
Machine Translation
Pseudo-word for Phrase-based Machine Translation
Xiangyu Duan, Min Zhang and Haizhou Li
Learning Lexicalized Reordering Models from Reordering Graphs
Jinsong Su, Yang Liu, Yajuan Lv, Haitao Mi and Qun Liu
Filtering Syntactic Constraints for Statistical Machine Translation
Hailong Cao and Eiichiro Sumita
Error Detection for Statistical Machine Translation Using Linguistic Features
Deyi Xiong, Min Zhang and Haizhou Li
Boosting-based System Combination for Machine Translation
Tong Xiao, Jingbo Zhu, Muhua Zhu and Huizhen Wang
Bilingual Sense Similarity for Statistical Machine Translation
Boxing Chen, George Foster and Roland Kuhn
----
Summarization & Generation
A Risk Minimization Framework for Extractive Speech Summarization
Shih-Hsiang Lin and Berlin Chen
Entity-based local coherence modelling using topological fields
Jackie Chi Kit Cheung and Gerald Penn
Automatic Collocation Suggestion in Academic Writing
Jian-Cheng Wu, Yu-Chia Chang, Teruko Mitamura and Jason S. Chang
Identifying Non-Explicit Citing Sentences for Citation-Based Summarization.
Vahed Qazvinian and Dragomir R. Radev
Automatic Generation of Story Highlights
Kristian Woodsend and Mirella Lapata
Plot Induction and Evolutionary Search for Story Generation
Neil McIntyre and Mirella Lapata
Metadata-Aware Measures for Answer Summarization in Community Question Answering
Mattia Tomasoni and Minlie Huang
A Hybrid Hierarchical Model for Multi-Document Summarization
Asli Celikyilmaz and Dilek Hakkani-Tur
A new Approach to Improving Multilingual Summarization using a Genetic Algorithm
Marina Litvak, Mark Last and Menahem Friedman
Cross-Language Document Summarization Based on Machine Translation Quality Prediction
Xiaojun Wan, Huiying Li and Jianguo Xiao
Generating image descriptions using dependency relational patterns
Ahmet Aker and Robert Gaizauskas
----
Information Extraction
Open Information Extraction Using Wikipedia
Fei Wu and Daniel S. Weld
----
Corpus
The Human Language Project: Building a Universal Corpus of the World’s Languages
Steven Abney and Steven Bird
(see full list of papers at http://nlp.csie.ncnu.edu.tw/~shin/acl2010/proceedings/CDROM/ACL/index.html)
Pseudo-word for Phrase-based Machine Translation
Xiangyu Duan, Min Zhang and Haizhou Li
Learning Lexicalized Reordering Models from Reordering Graphs
Jinsong Su, Yang Liu, Yajuan Lv, Haitao Mi and Qun Liu
Filtering Syntactic Constraints for Statistical Machine Translation
Hailong Cao and Eiichiro Sumita
Error Detection for Statistical Machine Translation Using Linguistic Features
Deyi Xiong, Min Zhang and Haizhou Li
Boosting-based System Combination for Machine Translation
Tong Xiao, Jingbo Zhu, Muhua Zhu and Huizhen Wang
Bilingual Sense Similarity for Statistical Machine Translation
Boxing Chen, George Foster and Roland Kuhn
----
Summarization & Generation
A Risk Minimization Framework for Extractive Speech Summarization
Shih-Hsiang Lin and Berlin Chen
Entity-based local coherence modelling using topological fields
Jackie Chi Kit Cheung and Gerald Penn
Automatic Collocation Suggestion in Academic Writing
Jian-Cheng Wu, Yu-Chia Chang, Teruko Mitamura and Jason S. Chang
Identifying Non-Explicit Citing Sentences for Citation-Based Summarization.
Vahed Qazvinian and Dragomir R. Radev
Automatic Generation of Story Highlights
Kristian Woodsend and Mirella Lapata
Plot Induction and Evolutionary Search for Story Generation
Neil McIntyre and Mirella Lapata
Metadata-Aware Measures for Answer Summarization in Community Question Answering
Mattia Tomasoni and Minlie Huang
A Hybrid Hierarchical Model for Multi-Document Summarization
Asli Celikyilmaz and Dilek Hakkani-Tur
A new Approach to Improving Multilingual Summarization using a Genetic Algorithm
Marina Litvak, Mark Last and Menahem Friedman
Cross-Language Document Summarization Based on Machine Translation Quality Prediction
Xiaojun Wan, Huiying Li and Jianguo Xiao
Generating image descriptions using dependency relational patterns
Ahmet Aker and Robert Gaizauskas
----
Information Extraction
Open Information Extraction Using Wikipedia
Fei Wu and Daniel S. Weld
----
Corpus
The Human Language Project: Building a Universal Corpus of the World’s Languages
Steven Abney and Steven Bird
(see full list of papers at http://nlp.csie.ncnu.edu.tw/~shin/acl2010/proceedings/CDROM/ACL/index.html)
Wednesday, 9 June 2010
Topic summarization
Given a scenario in which the system takes the input with a research topic and needs to generate a summary of related works relevant to that topic automatically.
--> I think this research problem is still open and actually very challenging. It requires advanced processing which combines many fields in AI such as: NLP, IR, IE, ...
Some initial works (including mine) as follows:
1) Scientific Paper Summarization Using Citation Summary Networks by Qazvinian V. et al. (COLING 2008).
--> this work only targets single article summarization using a clustering approach based on citation summary networks.
2) Generating surveys of scientific paradigms by Saif Mohammad et al. (NAACL 2009).
--> this work explores the usefulness of citation summary in compared to summary from abstracts or full text of articles.
3) Towards Automated Related Work Summarization by Cong Duy Vu HOANG et al. (COLING 2010)
--> this work does not use citation summary but tries to take advantage of full text of article in generating related work summary.
It makes a strong assumption that each related work summary follows a topic hierarchy tree which is provided as the input of summarization system. The system then proposes two different strategies (general & specific content summarization) based on manual rhetorical analysis on how humans use topic hierarchy tree to generate related work summary.
4) Identifying Non-Explicit Citing Sentences for Citation-Based Summarization by Vahed Qazvinian and Dragomir R. Radev (ACL 2010)
--> TBA
5) Context Identification of Sentences in Related Work Sections using a Conditional Random Field: Towards Intelligent Digital Libraries by Angrosh M. A. et al. (JCDL 2010)
6) Imitating Human Literature Review Writing: An Approach to Multi-document Summarization by Jaidka K. et al. (ICADL 2010)
7) Analysis of the Macro-Level Discourse Structure of Literature Reviews by Jaidka K. et al. (Online Information Review)
8) Ultimate Research Assistant: http://ultimate-research-assistant.com/
9) iResearch Reporter: http://iresearch-reporter.com//
10) TBA
Future works (what I come up in my mind now) includes:
- Given a research topic --> automatically generate a topic hierarchy tree of that topic.
- A systematic comparison of summaries built from citations, abstracts, full text of articles. Which ones are more useful to users?
- An initial add-in component integrated into online ACL anthology system.
- Some other issues improve the summarization performance (i.e. use rhetorical discourse analysis, ...)
- ...
--
Cheers,
Vu
--> I think this research problem is still open and actually very challenging. It requires advanced processing which combines many fields in AI such as: NLP, IR, IE, ...
Some initial works (including mine) as follows:
1) Scientific Paper Summarization Using Citation Summary Networks by Qazvinian V. et al. (COLING 2008).
--> this work only targets single article summarization using a clustering approach based on citation summary networks.
2) Generating surveys of scientific paradigms by Saif Mohammad et al. (NAACL 2009).
--> this work explores the usefulness of citation summary in compared to summary from abstracts or full text of articles.
3) Towards Automated Related Work Summarization by Cong Duy Vu HOANG et al. (COLING 2010)
--> this work does not use citation summary but tries to take advantage of full text of article in generating related work summary.
It makes a strong assumption that each related work summary follows a topic hierarchy tree which is provided as the input of summarization system. The system then proposes two different strategies (general & specific content summarization) based on manual rhetorical analysis on how humans use topic hierarchy tree to generate related work summary.
4) Identifying Non-Explicit Citing Sentences for Citation-Based Summarization by Vahed Qazvinian and Dragomir R. Radev (ACL 2010)
--> TBA
5) Context Identification of Sentences in Related Work Sections using a Conditional Random Field: Towards Intelligent Digital Libraries by Angrosh M. A. et al. (JCDL 2010)
6) Imitating Human Literature Review Writing: An Approach to Multi-document Summarization by Jaidka K. et al. (ICADL 2010)
7) Analysis of the Macro-Level Discourse Structure of Literature Reviews by Jaidka K. et al. (Online Information Review)
8) Ultimate Research Assistant: http://ultimate-research-assistant.com/
9) iResearch Reporter: http://iresearch-reporter.com//
10) TBA
Future works (what I come up in my mind now) includes:
- Given a research topic --> automatically generate a topic hierarchy tree of that topic.
- A systematic comparison of summaries built from citations, abstracts, full text of articles. Which ones are more useful to users?
- An initial add-in component integrated into online ACL anthology system.
- Some other issues improve the summarization performance (i.e. use rhetorical discourse analysis, ...)
- ...
--
Cheers,
Vu
Labels:
idea,
information extraction,
NLP,
research,
scientific text,
topic summarization
Monday, 25 January 2010
Subscribe to:
Posts (Atom)