CMC training corpus Janes-Tag 1.1

Dataset

PID

Janes-Tag is a manually annotated corpus of Slovene Computer-Mediated Communication (CMC). It is meant as a gold-standard training and testing dataset for tokenisation, sentence segmentation, word normalisation, morphosyntactic tagging and lemmatisation of non-standard Slovene. As the corpus has been carefully manually annotated, it is also suitable for detailed linguistic explorations which require highly accurate and reliable annotations.

The corpus is further described in: ERJAVEC, Tomaž, ČIBEJ, Jaka, ARHAR HOLDT, Špela, LJUBEŠIĆ, Nikola, FIŠER, Darja. Gold-standard datasets for annotation of Slovene computer-mediated communication. In Proceedings of RASLAN 2016: Recent Advances in Slavonic Natural Language Processing. Brno: Tribun EU, 2016, pp. 29-40, https://nlp.fi.muni.cz/raslan/raslan16.pdf

Note that a related corpus, Janes-Norm is also available, cf. http://hdl.handle.net/11356/1083.

Identifier
PID	http://hdl.handle.net/11356/1081
Related Identifier	http://hdl.handle.net/11356/1079
Related Identifier	http://hdl.handle.net/11356/1085
Related Identifier	https://nl.ijs.si/janes/
Metadata Access	http://www.clarin.si/repository/oai/request?verb=GetRecord&metadataPrefix=oai_dc&identifier=oai:www.clarin.si:11356/1081

Provenance
Creator	Erjavec, Tomaž; Fišer, Darja; Čibej, Jaka; Arhar Holdt, Špela
Publisher	Jožef Stefan Institute
Publication Year	2016
Rights	Creative Commons - Attribution-ShareAlike 4.0 International (CC BY-SA 4.0); https://creativecommons.org/licenses/by-sa/4.0/; PUB
OpenAccess	true
Contact	info(at)clarin.si

Representation
Language	Slovenian; Slovene
Resource Type	corpus
Format	application/pdf; application/zip; text/plain; charset=utf-8; downloadable_files_count: 4
Discipline	Linguistics