WikiCorpus - Corpus created from a 2006 dump of the Catalan, Spanish, and English Wikipedias. | TALP :: Language and Speech Technologies and Applications

Authors:
Samuel Reese, Gemma Boleda, Montse Cuadros, Lluís Padró, German Rigau

References:
http://nlp.lsi.upc.edu/wikicorpus

Description:
The Wikicorpus is a freely available trilingual corpus (Catalan, Spanish, English) that contains large portions of the Wikipedia and has been automatically enriched with linguistic information. In its present version, it contains over 750 million words.
The corpora have been annotated with lemma and part of speech information using the open source library FreeLing. Also, they have been sense annotated with the state of the art Word Sense Disambiguation algorithm UKB. As UKB assigns WordNet senses, and WordNet has been aligned across languages via the InterLingual Index, this sort of annotation opens the way to massive explorations in lexical semantics that were not possible before.

Moreover, we also provide an open source Java-based parser for Wikipedia pages developed for the construction of the corpus.

Functionality:

Technology:
Java

Technical Requirements:

Modules:

Innovation:
To our knowledge, the Catalan and Spanish portions of the WikiCorpus are the largest corpora that are freely available to the community. Also, the automatic large-scale WSD annotation is innovative.

Development:

Publications:

Please cite the following publication if you use the corpora:

Samuel Reese, Gemma Boleda, Montse Cuadros, Lluís Padró, German Rigau. Word-Sense Disambiguated Multilingual Wikipedia Corpus. In Proceedings of 7th Language Resources and Evaluation Conference (LREC'10), La Valleta, Malta. May, 2010.

Contact: Gemma Boleda

Search form

WikiCorpus - Corpus created from a 2006 dump of the Catalan, Spanish, and English Wikipedias.

You are here