Universal Dependencies 2.0 – CoNLL 2017 Shared Task Development and Test Data

PID

Universal Dependencies is a project that seeks to develop cross-linguistically consistent treebank annotation for many languages, with the goal of facilitating multilingual parser development, cross-lingual learning, and parsing research from a language typology perspective. The annotation scheme is based on (universal) Stanford dependencies (de Marneffe et al., 2006, 2008, 2014), Google universal part-of-speech tags (Petrov et al., 2012), and the Interset interlingua for morphosyntactic tagsets (Zeman, 2008).

This release contains the test data used in the CoNLL 2017 shared task on parsing Universal Dependencies. Due to the shared task the test data was held hidden and not released together with the training and development data of UD 2.0. Therefore this release complements the UD 2.0 release (http://hdl.handle.net/11234/1-1983) to a full release of UD treebanks. In addition, the present release contains 18 new parallel test sets and 4 test sets in surprise languages. The present release also includes the development data already released with UD 2.0. Unlike regular UD releases, this one uses the folder-file structure that was visible to the systems participating in the shared task.

Identifier
PID http://hdl.handle.net/11234/1-2184
Related Identifier http://universaldependencies.org/conll17/
Metadata Access http://lindat.mff.cuni.cz/repository/oai/request?verb=GetRecord&metadataPrefix=oai_dc&identifier=oai:lindat.mff.cuni.cz:11234/1-2184
Provenance
Creator Nivre, Joakim; Agić, Željko; Ahrenberg, Lars; Antonsen, Lene; Aranzabe, Maria Jesus; Asahara, Masayuki; Ateyah, Luma; Attia, Mohammed; Atutxa, Aitziber; Badmaeva, Elena; Ballesteros, Miguel; Banerjee, Esha; Bank, Sebastian; Bauer, John; Bengoetxea, Kepa; Bhat, Riyaz Ahmad; Bick, Eckhard; Bosco, Cristina; Bouma, Gosse; Bowman, Sam; Burchardt, Aljoscha; Candito, Marie; Caron, Gauthier; Cebiroğlu Eryiğit, Gülşen; Celano, Giuseppe G. A.; Cetin, Savas; Chalub, Fabricio; Choi, Jinho; Cho, Yongseok; Cinková, Silvie; Çöltekin, Çağrı; Connor, Miriam; de Marneffe, Marie-Catherine; de Paiva, Valeria; Diaz de Ilarraza, Arantza; Dobrovoljc, Kaja; Dozat, Timothy; Droganova, Kira; Eli, Marhaba; Elkahky, Ali; Erjavec, Tomaž; Farkas, Richárd; Fernandez Alcalde, Hector; Foster, Jennifer; Freitas, Cláudia; Gajdošová, Katarína; Galbraith, Daniel; Garcia, Marcos; Ginter, Filip; Goenaga, Iakes; Gojenola, Koldo; Gökırmak, Memduh; Goldberg, Yoav; Gómez Guinovart, Xavier; Gonzáles Saavedra, Berta; Grioni, Matias; Grūzītis, Normunds; Guillaume, Bruno; Habash, Nizar; Hajič, Jan; Hajič jr., Jan; Hà Mỹ, Linh; Harris, Kim; Haug, Dag; Hladká, Barbora; Hlaváčová, Jaroslava; Hohle, Petter; Ion, Radu; Irimia, Elena; Johannsen, Anders; Jørgensen, Fredrik; Kaşıkara, Hüner; Kanayama, Hiroshi; Kanerva, Jenna; Kayadelen, Tolga; Kettnerová, Václava; Kirchner, Jesse; Kotsyba, Natalia; Krek, Simon; Kwak, Sookyoung; Laippala, Veronika; Lambertino, Lorenzo; Lando, Tatiana; Lê Hồng, Phương; Lenci, Alessandro; Lertpradit, Saran; Leung, Herman; Li, Cheuk Ying; Li, Josie; Ljubešić, Nikola; Loginova, Olga; Lyashevskaya, Olga; Lynn, Teresa; Macketanz, Vivien; Makazhanov, Aibek; Mandl, Michael; Manning, Christopher; Manurung, Ruli; Mărănduc, Cătălina; Mareček, David; Marheinecke, Katrin; Martínez Alonso, Héctor; Martins, André; Mašek, Jan; Matsumoto, Yuji; McDonald, Ryan; Mendonça, Gustavo; Missilä, Anna; Mititelu, Verginica; Miyao, Yusuke; Montemagni, Simonetta; More, Amir; Moreno Romero, Laura; Mori, Shunsuke; Moskalevskyi, Bohdan; Muischnek, Kadri; Mustafina, Nina; Müürisep, Kaili; Nainwani, Pinkey; Nedoluzhko, Anna; Nguyễn Thị, Lương; Nguyễn Thị Minh, Huyền; Nikolaev, Vitaly; Nitisaroj, Rattima; Nurmi, Hanna; Ojala, Stina; Osenova, Petya; Øvrelid, Lilja; Pascual, Elena; Passarotti, Marco; Perez, Cenel-Augusto; Perrier, Guy; Petrov, Slav; Piitulainen, Jussi; Pitler, Emily; Plank, Barbara; Popel, Martin; Pretkalniņa, Lauma; Prokopidis, Prokopis; Puolakainen, Tiina; Pyysalo, Sampo; Rademaker, Alexandre; Real, Livy; Reddy, Siva; Rehm, Georg; Rinaldi, Larissa; Rituma, Laura; Rosa, Rudolf; Rovati, Davide; Saleh, Shadi; Sanguinetti, Manuela; Saulīte, Baiba; Sawanakunanon, Yanin; Schuster, Sebastian; Seddah, Djamé; Seeker, Wolfgang; Seraji, Mojgan; Shakurova, Lena; Shen, Mo; Shimada, Atsuko; Shohibussirri, Muh; Silveira, Natalia; Simi, Maria; Simionescu, Radu; Simkó, Katalin; Šimková, Mária; Simov, Kiril; Smith, Aaron; Stella, Antonio; Strnadová, Jana; Suhr, Alane; Sulubacak, Umut; Szántó, Zsolt; Taji, Dima; Tanaka, Takaaki; Trosterud, Trond; Trukhina, Anna; Tsarfaty, Reut; Tyers, Francis; Uematsu, Sumire; Urešová, Zdeňka; Uria, Larraitz; Uszkoreit, Hans; van Noord, Gertjan; Varga, Viktor; Vincze, Veronika; Washington, Jonathan North; Yu, Zhuoran; Žabokrtský, Zdeněk; Zeman, Daniel; Zhu, Hanzhi
Publisher Universal Dependencies Consortium
Publication Year 2017
Rights Licence Universal Dependencies v2.0; https://lindat.mff.cuni.cz/repository/xmlui/page/licence-UD-2.0; PUB
OpenAccess true
Contact lindat-help(at)ufal.mff.cuni.cz
Representation
Language Greek, Ancient (to 1453); Arabic; Basque; Bulgarian; Croatian; Czech; Danish; Dutch; Flemish; English; Estonian; Finnish; French; German; Gothic; Greek, Modern (1453-); Greek; Hebrew; Hindi; Hungarian; Indonesian; Irish; Italian; Japanese; Latin; Norwegian; Church Slavic; Old Slavonic; Church Slavonic; Old Bulgarian; Old Church Slavonic; Persian; Farsi; Polish; Portuguese; Romanian; Moldavian; Moldovan; Slovenian; Slovene; Spanish; Castilian; Swedish; Tamil; Catalan; Valencian; Chinese; Galician; Kazakh; Latvian; Russian; Turkish; Coptic; Sanskrit; Saṁskṛta; Slovak; Ukrainian; Uighur; Uyghur; Vietnamese; Belarusian; Korean; Lithuanian; Urdu; Northern Sami; Upper Sorbian
Resource Type corpus
Format text/plain; charset=utf-8; application/x-gzip; downloadable_files_count: 1
Discipline Linguistics