ParaCrawl Corpus version 1.0

PID

The January 2018 release of the ParaCrawl is the first version of the corpus. It contains parallel corpora for 11 languages paired with English, crawled from a large number of web sites. The selection of websites is based on CommonCrawl, but ParaCrawl is extracted from a brand new crawl which has much higher coverage of these selected websites than CommonCrawl. Since the data is fairly raw, it is released with two quality metrics that can be used for corpus filtering. An official "clean" version of each corpus uses one of the metrics. For more details and raw data download please visit: http://paracrawl.eu/releases.html

Identifier
PID http://hdl.handle.net/11372/LRT-2610
Related Identifier http://paracrawl.eu
Metadata Access http://lindat.mff.cuni.cz/repository/oai/request?verb=GetRecord&metadataPrefix=oai_dc&identifier=oai:lindat.mff.cuni.cz:11372/LRT-2610
Provenance
Creator Koehn, Philipp; Heafield, Kenneth; Forcada, Mikel L.; Esplà-Gomis, Miquel; Ortiz-Rojas, Sergio; Sánchez, Gema Ramírez; Cartagena, Víctor M. Sánchez; Haddow, Barry; Bañón, Marta; Střelec, Marek; Samiotou, Anna; Kamran, Amir
Publisher ParaCrawl
Publication Year 2018
Rights Public Domain Dedication (CC Zero); http://creativecommons.org/publicdomain/zero/1.0/; PUB
OpenAccess true
Contact lindat-help(at)ufal.mff.cuni.cz
Representation
Language English; German; French; Spanish; Castilian; Italian; Portuguese; Dutch; Flemish; Polish; Czech; Romanian; Moldavian; Moldovan; Finnish; Latvian; Russian; Estonian
Resource Type corpus
Format text/plain; charset=utf-8; application/x-gzip; downloadable_files_count: 13
Discipline Linguistics