SimLex-999 Slovenian translation SimLex-999-sl 1.0

PID

The resource contains English SimLex-999 (Hill et al. 2015) and their Slovene translations. In the translation process, the word pairs were first translated by two translators independently, and next, for the examples where the translations differed, the final translations were chosen in a consensus meeting. The translators had also access to Croatian Simlex-999 translations (Mrkšić et al. 2017) and received translation guidelines (see next sheet) inspired by guidelines of Multi-SimLex (Vulić et al. 2020). The resources was used for building the CoSimLex resource (Armendariz et al. 2020).

The list contains English original pair of words (Word1 and Word2), their part-of-speech, followed by Slovene translations (Trans1 and Trans2). The last column Comment relates to special cases:

  • "multiword_translation" -> translators were asked to opt for single-word equivalents, in some cases the only appropriate translation was a multi-word expression (for example, "birthday" -> "rojstni dan").
  • "no_translation" -> pairs without a proper translation, i.e. translation pair contains two identical words. Although the translators were asked to find two different translations for the words, in a few examples that was not possible. For example, for the English pair "taxi" and "cab", only "taksi" was considered a good Slovene equivalent.
  • "duplicated_translation" -> in cases where a pair of words is repeated for two different English original pairs, both occurrences are marked as duplicate translations.
  • "duplicated_original" -> in one case, the original word pair was a duplicate, which is also marked.

Cite: If you use the dataset, please cite the Clarin handle and the following paper: Armendariz, Carlos Santos, Purver, Matthew, Ulčar, Matej, Pollak, Senja, Ljubešić, Nikola, Granroth-Wilding, Mark, and Vaik, Kristiina (2020). CoSimLex: A Resource for Evaluating Graded Word Similarity in Context. In Proceedings of the 12th Language Resources and Evaluation Conference, p. 5878--5886. https://www.aclweb.org/anthology/2020.lrec-1.720/

References: Armendariz, Carlos Santos, Purver, Matthew, Ulčar, Matej, Pollak, Senja, Ljubešić, Nikola, Granroth-Wilding, Mark, and Vaik, Kristiina (2020). CoSimLex: A Resource for Evaluating Graded Word Similarity in Context. In Proceedings of the 12th Language Resources and Evaluation Conference, p. 5878--5886. https://www.aclweb.org/anthology/2020.lrec-1.720/ Hill, F., Reichart, R., and Korhonen, A. (2015). Simlex-999: Evaluating semantic models with (genuine) similarity estimation. Computational Linguistics, 41(4):665–695. https://www.aclweb.org/anthology/J15-4004/ Mrkšić, Nikola, Ivan Vulić, Diarmuid Ó Séaghdha, Ira Leviant, Roi Reichart, Milica Gašić, Anna Korhonen, and Steve Young. (2017). Semantic specialisation of distributional word vector spaces using monolingual and cross-lingual constraints. Transactions of the ACL, 5:309–324. https://www.mitpressjournals.org/doi/abs/10.1162/tacl_a_00063 Vulić, Ivan, Baker, Simon, Ponti, Edoardo Maria, Petti, Ulla, Leviant, Ira, Wing, Kelly, Majewska, Olga, Bar, Eden, Malone, Matt, Poibeau, Thierry, Reichart, Roi and Anna Korhonen (2020). Multi-SimLex: A Large-Scale Evaluation of Multilingual and Cross-Lingual Lexical Semantic Similarity. Computational Linguistics. https://doi.org/10.1162/coli_a_00391

Identifier
PID http://hdl.handle.net/11356/1309
Related Identifier https://www.aclweb.org/anthology/2020.lrec-1.720/
Related Identifier http://embeddia.eu/
Metadata Access http://www.clarin.si/repository/oai/request?verb=GetRecord&metadataPrefix=oai_dc&identifier=oai:www.clarin.si:11356/1309
Provenance
Creator Pollak, Senja; Vulić, Ivan; Pelicon, Andraž; Repar, Andraž; Armendariz, Carlos; Matthew, Purver; Ljubešić, Nikola
Publisher University of Ljubljana
Publication Year 2020
Funding Reference info:eu-repo/grantAgreement/EC/H2020/825153
Rights Creative Commons - Attribution 4.0 International (CC BY 4.0); https://creativecommons.org/licenses/by/4.0/; PUB
OpenAccess true
Contact info(at)clarin.si
Representation
Language Slovenian; Slovene; English
Resource Type lexicalConceptualResource
Format text/csv; text/plain; text/plain; charset=utf-8; downloadable_files_count: 3
Discipline Linguistics