Intro: This Hansard corpus (or collection of texts) contains nearly every speech given in the British Parliament from 1803-2005, and it allows you to search these speeches (including semantically-based searches) in ways that are not possible with any other resource.
Showing posts with label corpus annotation. Show all posts
Showing posts with label corpus annotation. Show all posts
Tuesday, 21 July 2015
Monday, 19 March 2012
WebAnnotator
WebAnnotator is a new tool for annotating Web pages implemented at LIMSI. Giving it a try will take you no more than 10 minutes.
WebAnnotator is implemented as a Firefox extension, allowing annotation of both offline and inline pages. The HTML rendering is fully preserved and all annotations consist in new HTML spans with specific styles.
WebAnnotator provides an easy and general-purpose framework and is made available under CeCILL free license (close to GNU GPL), so that use and further contributions are made simple.
WebAnnotator can be downloaded on the official Mozilla web page:
https://addons.mozilla.org/en-US/firefox/addon/webannotator/.
https://addons.mozilla.org/en-US/firefox/addon/webannotator/.
A quick manual can be found here:
http://perso.limsi.fr/Individu/xtannier/en/WebAnnotator/
http://perso.limsi.fr/Individu/xtannier/en/WebAnnotator/
All parts of an HTML document can be annotated: text, images, videos, tables, menus, etc. The annotations are created by simply selecting a part of the document and clicking on the relevant type and subtypes. The annotated elements are then highlighted in a specific color. Annotation schemas can be defined by the user by creating a simple DTD representing the types and subtypes that must be highlighted. Finally, annotations can be saved (HTML with highlighted parts of documents) or exported (in a machine-readable format).
WebAnnotator will be presented at LREC conference in May 2012.
Sunday, 19 February 2012
Saturday, 18 June 2011
Detection of Errors and Correction in Corpus Annotation
http://decca.osu.edu/
Intro: The success of data-driven approaches and stochastic modeling in computational linguistic research and applications is rooted in the availability of electronic natural language corpora. Despite the central role that annotated corpora play for computational linguistic research and applications, the question of how errors in the annotation of corpora can be detected and corrected has received only little attention. The DECCA project is designed to address this important gap by exploring an error detection and correction method with potential applicability to a wide range of corpus annotations.
Labels:
corpus annotation,
error detection,
links,
NLP,
research,
tools
Subscribe to:
Posts (Atom)