HinDialect: 26 Hindi-related languages and dialects of the Indic Continuum in North India

PID

HinDialect: 26 Hindi-related languages and dialects of the Indic Continuum in North India

Languages This is a collection of folksongs for 26 languages that form a dialect continuum in North India and nearby regions.

Namely Angika, Awadhi, Baiga, Bengali, Bhadrawahi, Bhili, Bhojpuri, Braj, Bundeli, Chhattisgarhi, Garhwali, Gujarati, Haryanvi, Himachali, Hindi, Kanauji, Khadi Boli, Korku, Kumaoni, Magahi, Malvi, Marathi, Nimadi, Panjabi, Rajasthani, Sanskrit.

This data is originally collected by the Kavita Kosh Project at http://www.kavitakosh.org/ . Here are the main characteristics of the languages in this collection: - They are all Indic languages except for Korku. - The majority of them are closely related to the standard Hindi dialect genealogically (such as Hariyanvi and Bhojpuri), although the collection also contains languages such as Bengali and Gujarati which are more distant relatives. - All except Nepali are primarily spoken in (North) India - All except Sanksrit are alive languages

Data Categorising them by pre-existing available NLP resources, we have: * Band 1 languages : Hindi, Marathi, Punjabi, Sindhi, Gujarati, Bengali, Nepali. These languages already have other large datasets available. Since Kavita Kosh focusses largely on Hindi-related languages, we may have very little data for these other languages in this particular dataset. * Band 2 languages: Bhojpuri, Magahi, Awadhi, Brajbhasha. These languages have growing interest and some datasets of a relatively small size as compared to Band 1 language resources. * Band 3 languages: All other languages in the collection are previously zero-resource languages. These are the languages for which this dataset is the most relevant.

Script This dataset is entirely in Devanagari. Content in the case of languages not written in Devanagari (such as Bengali and Gujarati) has been transliterated by the Kavita Kosh Project.

Format The data is segregated by language, and contains each folksong in a different JSON file.

Identifier
PID http://hdl.handle.net/11234/1-4787
Related Identifier http://hdl.handle.net/11234/1-4839
Related Identifier https://github.com/niyatibafna/north-indian-dialect-modelling
Metadata Access http://lindat.mff.cuni.cz/repository/oai/request?verb=GetRecord&metadataPrefix=oai_dc&identifier=oai:lindat.mff.cuni.cz:11234/1-4787
Provenance
Creator Bafna, Niyati; Žabokrtský, Zdeněk; España-Bonet, Cristina; van Genabith, Josef; Kumar, Lalit "Samyak Lalit"; Suman, Sharda; Shivay, Rahul
Publisher Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL); Kavita Kosh Project
Publication Year 2022
Rights Creative Commons - Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0); http://creativecommons.org/licenses/by-nc-sa/4.0/; PUB
OpenAccess true
Contact lindat-help(at)ufal.mff.cuni.cz
Representation
Language Hindi; Marathi; Marāṭhī; Magahi; Awadhi; Bhojpuri; Braj; Rajasthani; Sanskrit; Saṁskṛta; Angika; Bengali; Bangla; Gujarati; Panjabi; Punjabi; Uncoded languages
Resource Type corpus
Format text/plain; charset=utf-8; application/zip; downloadable_files_count: 1
Discipline Linguistics