Showcase

Project BoLI

Project BoLI, the flagship initiative of [UnReaL-TecE LLP](https://unreal-tece.co.in), is building high-quality datasets to train and benchmark next-generation AI tasks—from speech-to-text to advanced grammatical reasoning—across more than 1,000 underserved Indian languages and dialects. We are pioneering a globally unique data governance framework where all released datasets, derivative assets, and downstream models are permanently co-owned by the contributing communities. Under our groundbreaking BoLI License, commercial data usage is strictly treated as a revocable license rather than a transfer of ownership, ensuring total accountability in the era of LLMs.

—visible projects
—public projects
—workspace projects
—languages

Featured Collections

Project BoLI71 projects
Project BoLI, the flagship initiative of UnReaL-TecE LLP, is building high-quality datasets to train and benchmark next-generation AI tasks—from speech-to-text to advanced grammatical reasoning—across more than 1,000 underserved Indian languages and dialects. We are pioneering a globally unique data governance framework where all released datasets, derivative assets, and downstream models are permanently co-owned by the contributing communities. Under our groundbreaking BoLI License, commercial data usage is strictly treated as a revocable license rather than a transfer of ownership, ensuring total accountability in the era of LLMs.

Project BoLI, the flagship initiative of UnReaL-TecE LLP, is building high-quality datasets to train and benchmark next-generation AI tasks—from speech-to-text to advanced grammatical reasoning—across more than 1,000 underserved Indian languages and dialects. We are pioneering a globally unique data governance framework where all released datasets, derivative assets, and downstream models are permanently co-owned by the contributing communities. Under our groundbreaking BoLI License, commercial data usage is strictly treated as a revocable license rather than a transfer of ownership, ensuring total accountability in the era of LLMs.

124.7
hours
89
speakers
64
languages
223601
entries
SpeeD-TB24 projects
Dataset created under the Speech Datasets and Models for Tibeto-Burman Languages (SpeeD-TB), sponsored under Mission Bhashini by Ministry of Electronics and Information Technology (MEITY), Govt of India. The project aimed to create 1,200 hours of speech dataset and ASR models for 6 underresourced, tribal Tibeto-Burman languages of India speoken in Eastern and North-Eastern parts of India.

Dataset created under the Speech Datasets and Models for Tibeto-Burman Languages (SpeeD-TB), sponsored under Mission Bhashini by Ministry of Electronics and Information Technology (MEITY), Govt of India. The project aimed to create 1,200 hours of speech dataset and ASR models for 6 underresourced, tribal Tibeto-Burman languages of India speoken in Eastern and North-Eastern parts of India.

1,476.8
hours
2066
speakers
6
languages
345468
entries

Discover

Public Projects

Dataset Public preview

BoLI Nyishi Narration

### The Dataset The current data preview of Nyishi is being released as part of the **[Project BoLI](https://boli.unreal-tece.co.in)**. This preview is a reflection of the full dataset and consists of the following - 1. Speech Recordings of 10 narrations in the language. 2. Transcriptions in IPA and native script(s). 3. Translations in English (which also act as prompts for the translation sentences). 4. Detailed speaker metadata, including their demographic, educational and linguistic profile. 5. Prompt in English and Hindi. The full dataset contains the following - 1. Translations of a minimum of 1000 carefully selected sentences. These sentences are selected to represent diverse morphosyntactic categories generally found in Indian languages such as demonstratives, classifiers, TAM morphology, and different sentence structures such as transitives and ditransitives. These sets of sentences are based on the standardised questionnaires built by Linguists for writing the sketch grammar of any human language, thereby, representing almost full range of morphosyntactic properties exhibited in the language. Unlike other benchmarks which focus largely on lexical level evaluation (aka domains), this is the first benchmark dataset that evaluates model's performance on a range of morphosyntactic structures, even the most uncommon ones. 2. More than 600 narrative speeches collected across 8 domains. These are recorded by at least 2 speakers. Narrations, along with their prompts, can be used to evaluate AI models on a range of prompt-based tasks. 3. Along with transcriptions in IPA and other scripts and translation, all data is interlinearly glossed at morphemic level. This gives a word-by-word meaning and morphosyntactic information, thereby, enabling evaluation of models on their deep grammatical knowledge, ability for cross-linguistic comparison and generalisation and reasoning capacity and skills in language-related puzzles. This allows for evaluating the model's capacity on reasoning tasks beyond mathematical reasoning tasks as well as their capacity to arrive at typological generalisations. 4. In addition to the data and prompt itself, as mentioned earlier, we are also making available the detailed speaker metadata and prompt-level metadata viz domain, elicitation method, multilingual prompt, target grammatical category (for translation sentences), etc. 5. In accordance with our data governance policy, all contributors to the dataset, including those recording the dataset and those transcribing it, are named and listed as data contributors. ### Dataset Preparation The speech included as part of this dataset was recorded by native speakers using a data collection application on a mobile device (Karya or in-house app, Atekho), by translating sentences from English or Hindi to the target language. The recorded speech was validated and then transcribed manually by trained linguists, working with the native speakers, following the [guidelines for the project](https://docs.google.com/document/d/1h-EcxxaNiGzdWiwc1-V_S-Xt6n4MwE6F/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) using MATra Lab, a part of the [LiFE Suite Ecosystem](https://life.unreal-tece.co.in), developed by [Unreal Tece LLP](https://unreal-tece.co.in). [Project BoLI Guidelines](https://docs.google.com/document/d/1h-EcxxaNiGzdWiwc1-V_S-Xt6n4MwE6F/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) ### About Nyishi Nyishi belongs to the Tani branch of the Tibeto-Burman language family, specifically categorized within the Western Tani subgroup. ### Speakers According to the [2011 Census of India](https://censusindia.gov.in/nada/index.php/catalog/42561), there are approximately 400,111 speakers of Nyishi. ### Distribution in India The language is spoken primarily in the northeastern state of Arunachal Pradesh, especially across the districts of Papum Pare, Kurung Kumey, Kra Daadi, Lower Subansiri, Kamle, and East Kameng. It is also spoken by a smaller population in the adjacent Lakhimpur and Sonitpur districts of Assam. ### Major grammatical features * **Phonology:** The sound system includes a contrastive vowel length distinction and typically features central vowels (such as /ɨ/ and /ə/) characteristic of Tani languages. Syllable structure is predominantly CV(C), with restrictions on which consonants can occupy the coda position. * **Morphology:** Nyishi is largely agglutinating. It employs an extensive system of numeral classifiers that must agree with the semantic class of the noun being quantified. Nouns inflect for case (including ergative, genitive, accusative, and locative suffixes), while verbs are inflected with complex suffixes indicating tense, aspect, mood, and evidentiality. * **Syntax:** The default word order is Subject-Object-Verb (SOV). The language exhibits postpositional phrase structure and generally adheres to an ergative-absolutive morphosyntactic alignment. ### Further reading For more comprehensive linguistic data, see the [Wikipedia article on Nyishi](https://en.wikipedia.org/wiki/Nyishi_language). ### About Project BoLI [Project BoLI](https://boli.unreal-tece.co.in) is the flagship project of [UnReaL-TecE LLP](https://unreal-tece.co.in), which aims to build high-quality datasets for benchmarking and evaluating different kinds of AI tasks, including speech-to-text, machine translation, grammatical analysis and reasoning tasks and prompt-based evaluation of LLMs. While the project aims to build these datasets for every Indian language and variety, the primary focus is on over 1300 underserved languages and both first and second language varieties of major, scheduled languages. The project's uniqueness is not just limited to the kind of benchmarking tasks it supports (including proposing some novel tasks) and also the kind of languages and communities it supports but also in its contextualisation and implementation of a unique data governance model (not yet implemented anywhere across the globe), which mandates that all datasets released by the project and their derivatives (including the models) are co-owned by the community members and all contributors of the project, and any permission to use the dataset is not a transfer of ownership but a revocable license to use it. The conditions under which the license could be revoked are clearly mentioned as part of the BoLI License. The complete details of the project, languages and communities supported till now, the quantum of data available till now, its data governance model and other relevant documents are all publicly accessible at the [project website](https://boli.unreal-tece.co.in). ### Ethical Considerations, Consent, IPR and Attribution Project BoLI and this repository represents our commitment to not only fair remuneration to the speakers of the language but also to co-ownership and equal IPR to all the contributors who have built the dataset. We believe this is the first step to move away from the extractive data collection and use practices and ensure fairness in our treatment of the community members. As such, we have listed all speakers and transcribers as Contributors to the dataset (we insist that they are co-owners of the dataset, even though the HuggingFace platform does not provide us an explicit way of stating that) and they are further recognised as Speakers and Annotators of the dataset. This dataset is only licensed to other researchers for use in their research projects. More details about licensing and commercial use conditions are given in the License and Commercial Use sections. [Project BoLI - Data Governance Policy](https://docs.google.com/document/d/1DBGZVFJ4RG57j4rFphKIrvgbpfpj8HqZ/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) [BoLI Ethics Principle & Pledge](https://docs.google.com/document/d/1hMQzwsOJF4QDfk_DOSIGLOF3GApLQOZA/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) [Project BoLI - Digital Consent Form](https://docs.google.com/document/d/1BTQUrgVTS-sVszUUDOkbnyUNusAHVYm4/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) [Project BoLI - Field Recording of Oral Consent](https://docs.google.com/document/d/1xoRxXcNES2oQqBajY3Z7HxMwpt0oOyDK/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) [Project BoLI - TnC](https://docs.google.com/document/d/1UuiIAHWXXKF202MYqLrtOy1ftUk__POW/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) ### Dataset Access Full dataset for the language can be browsed and queried on our [app](https://life.unreal-tece.co.in). The results of the evaluation of different models will also be made available on the same link. If you would like to access the full dataset for your own use or would like to work with us in collecting more data for any language or variety, or collaborate with us in this initiative in any other way, please get in touch with us.

Type
transcriptions
Owner
BoLI
Languages
Nyishi
2 speakers 314 audios
Arunachal Pradesh Low-resource Language Nyishi Tibeto-Burman
Dataset Public preview

BoLI Awadhi Translation

### The Dataset The current data preview of Awadhi is being released as part of the **[Project BoLI](https://boli.unreal-tece.co.in)**. This preview is a reflection of the full dataset and consists of the following - 1. Speech Recordings of 10 narrations in the language. 2. Transcriptions in IPA and native script(s). 3. Translations in English (which also act as prompts for the translation sentences). 4. Detailed speaker metadata, including their demographic, educational and linguistic profile. 5. Prompt in English and Hindi. The full dataset contains the following - 1. Translations of a minimum of 1000 carefully selected sentences. These sentences are selected to represent diverse morphosyntactic categories generally found in Indian languages such as demonstratives, classifiers, TAM morphology, and different sentence structures such as transitives and ditransitives. These sets of sentences are based on the standardised questionnaires built by Linguists for writing the sketch grammar of any human language, thereby, representing almost full range of morphosyntactic properties exhibited in the language. Unlike other benchmarks which focus largely on lexical level evaluation (aka domains), this is the first benchmark dataset that evaluates model's performance on a range of morphosyntactic structures, even the most uncommon ones. 2. More than 600 narrative speeches collected across 8 domains. These are recorded by at least 2 speakers. Narrations, along with their prompts, can be used to evaluate AI models on a range of prompt-based tasks. 3. Along with transcriptions in IPA and other scripts and translation, all data is interlinearly glossed at morphemic level. This gives a word-by-word meaning and morphosyntactic information, thereby, enabling evaluation of models on their deep grammatical knowledge, ability for cross-linguistic comparison and generalisation and reasoning capacity and skills in language-related puzzles. This allows for evaluating the model's capacity on reasoning tasks beyond mathematical reasoning tasks as well as their capacity to arrive at typological generalisations. 4. In addition to the data and prompt itself, as mentioned earlier, we are also making available the detailed speaker metadata and prompt-level metadata viz domain, elicitation method, multilingual prompt, target grammatical category (for translation sentences), etc. 5. In accordance with our data governance policy, all contributors to the dataset, including those recording the dataset and those transcribing it, are named and listed as data contributors. ### Dataset Preparation The speech included as part of this dataset was recorded by native speakers using a data collection application on a mobile device (Karya or in-house app, Atekho), by translating sentences from English or Hindi to the target language. The recorded speech was validated and then transcribed manually by trained linguists, working with the native speakers, following the [guidelines for the project](https://docs.google.com/document/d/1h-EcxxaNiGzdWiwc1-V_S-Xt6n4MwE6F/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) using MATra Lab, a part of the [LiFE Suite Ecosystem](https://life.unreal-tece.co.in), developed by [Unreal Tece LLP](https://unreal-tece.co.in). [Project BoLI Guidelines](https://docs.google.com/document/d/1h-EcxxaNiGzdWiwc1-V_S-Xt6n4MwE6F/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) ### About Awadhi Awadhi belongs to the Indo-Aryan language family, specifically classified within the East Central zone of Indo-Aryan languages. According to the [2011 Census of India](https://www.censusindia.gov.in/2011Census/Language-2011/Statement-1.pdf), Awadhi is spoken by approximately 3.85 million people as a mother tongue (often grouped under the broader Hindi category for census purposes). ### Distribution in India Awadhi is primarily spoken in the historic Awadh region of Uttar Pradesh, spanning districts such as Lucknow, Ayodhya, Prayagraj, Gonda, and Bahraich. Its speakers also extend into neighboring areas of Bihar and parts of Nepal. ### Major grammatical features * **Phonology:** Like many Indo-Aryan languages, Awadhi maintains a contrast between aspirated and unaspirated stops, as well as dental and retroflex consonants. Unlike Standard Hindi, Awadhi frequently preserves the short unrounded central vowel (schwa) in word-final positions. * **Morphology:** Nouns are inflected for two genders (masculine and feminine) and two numbers (singular and plural). The language exhibits a rich system of postpositions. A notable feature is its verbal conjugation, which distinguishes itself from Western Hindi through distinct tense and aspect markers, and a weaker or absent ergative marking system in the past tense of transitive verbs, tending towards nominative-accusative alignment. * **Syntax:** The default word order in Awadhi is Subject-Object-Verb (SOV). Auxiliaries and tense-marking clitics generally follow the main lexical verb. ### Further reading * Learn more about the language's literary history and structure on the [Awadhi language Wikipedia page](https://en.wikipedia.org/wiki/Awadhi_language). * Review dialectal classifications and genealogical data on the [Glottolog entry for Awadhi](https://glottolog.org/resource/languoid/id/awad1243). ### About Project BoLI [Project BoLI](https://boli.unreal-tece.co.in) is the flagship project of [UnReaL-TecE LLP](https://unreal-tece.co.in), which aims to build high-quality datasets for benchmarking and evaluating different kinds of AI tasks, including speech-to-text, machine translation, grammatical analysis and reasoning tasks and prompt-based evaluation of LLMs. While the project aims to build these datasets for every Indian language and variety, the primary focus is on over 1300 underserved languages and both first and second language varieties of major, scheduled languages. The project's uniqueness is not just limited to the kind of benchmarking tasks it supports (including proposing some novel tasks) and also the kind of languages and communities it supports but also in its contextualisation and implementation of a unique data governance model (not yet implemented anywhere across the globe), which mandates that all datasets released by the project and their derivatives (including the models) are co-owned by the community members and all contributors of the project, and any permission to use the dataset is not a transfer of ownership but a revocable license to use it. The conditions under which the license could be revoked are clearly mentioned as part of the BoLI License. The complete details of the project, languages and communities supported till now, the quantum of data available till now, its data governance model and other relevant documents are all publicly accessible at the [project website](https://boli.unreal-tece.co.in). ### Ethical Considerations, Consent, IPR and Attribution Project BoLI and this repository represents our commitment to not only fair remuneration to the speakers of the language but also to co-ownership and equal IPR to all the contributors who have built the dataset. We believe this is the first step to move away from the extractive data collection and use practices and ensure fairness in our treatment of the community members. As such, we have listed all speakers and transcribers as Contributors to the dataset (we insist that they are co-owners of the dataset, even though the HuggingFace platform does not provide us an explicit way of stating that) and they are further recognised as Speakers and Annotators of the dataset. This dataset is only licensed to other researchers for use in their research projects. More details about licensing and commercial use conditions are given in the License and Commercial Use sections. [Project BoLI - Data Governance Policy](https://docs.google.com/document/d/1DBGZVFJ4RG57j4rFphKIrvgbpfpj8HqZ/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) [BoLI Ethics Principle & Pledge](https://docs.google.com/document/d/1hMQzwsOJF4QDfk_DOSIGLOF3GApLQOZA/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) [Project BoLI - Digital Consent Form](https://docs.google.com/document/d/1BTQUrgVTS-sVszUUDOkbnyUNusAHVYm4/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) [Project BoLI - Field Recording of Oral Consent](https://docs.google.com/document/d/1xoRxXcNES2oQqBajY3Z7HxMwpt0oOyDK/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) [Project BoLI - TnC](https://docs.google.com/document/d/1UuiIAHWXXKF202MYqLrtOy1ftUk__POW/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) ### Dataset Access Full dataset for the language can be browsed and queried on our [app](https://life.unreal-tece.co.in). The results of the evaluation of different models will also be made available on the same link. If you would like to access the full dataset for your own use or would like to work with us in collecting more data for any language or variety, or collaborate with us in this initiative in any other way, please get in touch with us.

Type
transcriptions
Owner
BoLI
Languages
Awadhi
1 speakers 560 audios
Low-resource Language Unreal Tece Underresourced Language BoLI
Dataset Public preview

BoLI Galo Narration

### The Dataset The current data preview of Galo is being released as part of the **[Project BoLI](https://boli.unreal-tece.co.in)**. This preview is a reflection of the full dataset and consists of the following - 1. Speech Recordings of 10 narrations in the language. 2. Transcriptions in IPA and native script(s). 3. Translations in English (which also act as prompts for the translation sentences). 4. Detailed speaker metadata, including their demographic, educational and linguistic profile. 5. Prompt in English and Hindi. The full dataset contains the following - 1. Translations of a minimum of 1000 carefully selected sentences. These sentences are selected to represent diverse morphosyntactic categories generally found in Indian languages such as demonstratives, classifiers, TAM morphology, and different sentence structures such as transitives and ditransitives. These sets of sentences are based on the standardised questionnaires built by Linguists for writing the sketch grammar of any human language, thereby, representing almost full range of morphosyntactic properties exhibited in the language. Unlike other benchmarks which focus largely on lexical level evaluation (aka domains), this is the first benchmark dataset that evaluates model's performance on a range of morphosyntactic structures, even the most uncommon ones. 2. More than 600 narrative speeches collected across 8 domains. These are recorded by at least 2 speakers. Narrations, along with their prompts, can be used to evaluate AI models on a range of prompt-based tasks. 3. Along with transcriptions in IPA and other scripts and translation, all data is interlinearly glossed at morphemic level. This gives a word-by-word meaning and morphosyntactic information, thereby, enabling evaluation of models on their deep grammatical knowledge, ability for cross-linguistic comparison and generalisation and reasoning capacity and skills in language-related puzzles. This allows for evaluating the model's capacity on reasoning tasks beyond mathematical reasoning tasks as well as their capacity to arrive at typological generalisations. 4. In addition to the data and prompt itself, as mentioned earlier, we are also making available the detailed speaker metadata and prompt-level metadata viz domain, elicitation method, multilingual prompt, target grammatical category (for translation sentences), etc. 5. In accordance with our data governance policy, all contributors to the dataset, including those recording the dataset and those transcribing it, are named and listed as data contributors. ### Dataset Preparation The speech included as part of this dataset was recorded by native speakers using a data collection application on a mobile device (Karya or in-house app, Atekho), by translating sentences from English or Hindi to the target language. The recorded speech was validated and then transcribed manually by trained linguists, working with the native speakers, following the [guidelines for the project](https://docs.google.com/document/d/1h-EcxxaNiGzdWiwc1-V_S-Xt6n4MwE6F/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) using MATra Lab, a part of the [LiFE Suite Ecosystem](https://life.unreal-tece.co.in), developed by [Unreal Tece LLP](https://unreal-tece.co.in). [Project BoLI Guidelines](https://docs.google.com/document/d/1h-EcxxaNiGzdWiwc1-V_S-Xt6n4MwE6F/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) ### About Galo Galo belongs to the Tibeto-Burman language family, specifically classified within the Western Tani subgroup. ### Speakers According to the [2011 Census of India](https://censusindia.gov.in/nada/index.php/catalog/42561), there are approximately 133,499 native speakers of Galo. ### Distribution in India Galo is spoken primarily in the northeastern state of Arunachal Pradesh, with major concentrations in the West Siang, Leparada, Lower Siang, and Upper Subansiri districts. ### Major grammatical features * **Phonology:** Galo features a six-vowel system (/i, e, a, o, u, ɯ/) where vowel length is contrastive. It has a relatively simple consonant inventory and possesses a restricted lexical tone system. * **Morphology:** Highly agglutinative in nature. Grammatical relations are marked predominantly via suffixation. It maintains a complex and rich system of numeral classifiers and extensive verbal suffixation indicating aspect, mood, and evidentiality. * **Syntax:** The language exhibits a rigid Subject-Object-Verb (SOV) default constituent order. It utilizes an ergative-absolutive case-marking alignment, which is subject to differential case marking depending on nominal animacy and topicality. ### Further reading To learn more, explore the [Galo language on Wikipedia](https://en.wikipedia.org/wiki/Special:Search?search=Galo%20language) or refer to the language's structural documentation on [Glottolog](https://glottolog.org/resource/languoid/id/galo1242). ### About Project BoLI [Project BoLI](https://boli.unreal-tece.co.in) is the flagship project of [UnReaL-TecE LLP](https://unreal-tece.co.in), which aims to build high-quality datasets for benchmarking and evaluating different kinds of AI tasks, including speech-to-text, machine translation, grammatical analysis and reasoning tasks and prompt-based evaluation of LLMs. While the project aims to build these datasets for every Indian language and variety, the primary focus is on over 1300 underserved languages and both first and second language varieties of major, scheduled languages. The project's uniqueness is not just limited to the kind of benchmarking tasks it supports (including proposing some novel tasks) and also the kind of languages and communities it supports but also in its contextualisation and implementation of a unique data governance model (not yet implemented anywhere across the globe), which mandates that all datasets released by the project and their derivatives (including the models) are co-owned by the community members and all contributors of the project, and any permission to use the dataset is not a transfer of ownership but a revocable license to use it. The conditions under which the license could be revoked are clearly mentioned as part of the BoLI License. The complete details of the project, languages and communities supported till now, the quantum of data available till now, its data governance model and other relevant documents are all publicly accessible at the [project website](https://boli.unreal-tece.co.in). ### Ethical Considerations, Consent, IPR and Attribution Project BoLI and this repository represents our commitment to not only fair remuneration to the speakers of the language but also to co-ownership and equal IPR to all the contributors who have built the dataset. We believe this is the first step to move away from the extractive data collection and use practices and ensure fairness in our treatment of the community members. As such, we have listed all speakers and transcribers as Contributors to the dataset (we insist that they are co-owners of the dataset, even though the HuggingFace platform does not provide us an explicit way of stating that) and they are further recognised as Speakers and Annotators of the dataset. This dataset is only licensed to other researchers for use in their research projects. More details about licensing and commercial use conditions are given in the License and Commercial Use sections. [Project BoLI - Data Governance Policy](https://docs.google.com/document/d/1DBGZVFJ4RG57j4rFphKIrvgbpfpj8HqZ/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) [BoLI Ethics Principle & Pledge](https://docs.google.com/document/d/1hMQzwsOJF4QDfk_DOSIGLOF3GApLQOZA/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) [Project BoLI - Digital Consent Form](https://docs.google.com/document/d/1BTQUrgVTS-sVszUUDOkbnyUNusAHVYm4/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) [Project BoLI - Field Recording of Oral Consent](https://docs.google.com/document/d/1xoRxXcNES2oQqBajY3Z7HxMwpt0oOyDK/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) [Project BoLI - TnC](https://docs.google.com/document/d/1UuiIAHWXXKF202MYqLrtOy1ftUk__POW/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) ### Dataset Access Full dataset for the language can be browsed and queried on our [app](https://life.unreal-tece.co.in). The results of the evaluation of different models will also be made available on the same link. If you would like to access the full dataset for your own use or would like to work with us in collecting more data for any language or variety, or collaborate with us in this initiative in any other way, please get in touch with us.

Type
transcriptions
Owner
BoLI
Languages
Galo
1 speakers 507 audios
Low-resource Language Unreal Tece Underresourced Language BoLI
Dataset Public preview

BoLI Kokrajhar Bodo Narration

### The Dataset The current data preview of Kokrajhar Bodo is being released as part of the **[Project BoLI](https://boli.unreal-tece.co.in)**. This preview is a reflection of the full dataset and consists of the following - 1. Speech Recordings of 10 narrations in the language. 2. Transcriptions in IPA and native script(s). 3. Translations in English (which also act as prompts for the translation sentences). 4. Detailed speaker metadata, including their demographic, educational and linguistic profile. 5. Prompt in English and Hindi. The full dataset contains the following - 1. Translations of a minimum of 1000 carefully selected sentences. These sentences are selected to represent diverse morphosyntactic categories generally found in Indian languages such as demonstratives, classifiers, TAM morphology, and different sentence structures such as transitives and ditransitives. These sets of sentences are based on the standardised questionnaires built by Linguists for writing the sketch grammar of any human language, thereby, representing almost full range of morphosyntactic properties exhibited in the language. Unlike other benchmarks which focus largely on lexical level evaluation (aka domains), this is the first benchmark dataset that evaluates model's performance on a range of morphosyntactic structures, even the most uncommon ones. 2. More than 600 narrative speeches collected across 8 domains. These are recorded by at least 2 speakers. Narrations, along with their prompts, can be used to evaluate AI models on a range of prompt-based tasks. 3. Along with transcriptions in IPA and other scripts and translation, all data is interlinearly glossed at morphemic level. This gives a word-by-word meaning and morphosyntactic information, thereby, enabling evaluation of models on their deep grammatical knowledge, ability for cross-linguistic comparison and generalisation and reasoning capacity and skills in language-related puzzles. This allows for evaluating the model's capacity on reasoning tasks beyond mathematical reasoning tasks as well as their capacity to arrive at typological generalisations. 4. In addition to the data and prompt itself, as mentioned earlier, we are also making available the detailed speaker metadata and prompt-level metadata viz domain, elicitation method, multilingual prompt, target grammatical category (for translation sentences), etc. 5. In accordance with our data governance policy, all contributors to the dataset, including those recording the dataset and those transcribing it, are named and listed as data contributors. ### Dataset Preparation The speech included as part of this dataset was recorded by native speakers using a data collection application on a mobile device (Karya or in-house app, Atekho), by translating sentences from English or Hindi to the target language. The recorded speech was validated and then transcribed manually by trained linguists, working with the native speakers, following the [guidelines for the project](https://docs.google.com/document/d/1h-EcxxaNiGzdWiwc1-V_S-Xt6n4MwE6F/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) using MATra Lab, a part of the [LiFE Suite Ecosystem](https://life.unreal-tece.co.in), developed by [Unreal Tece LLP](https://unreal-tece.co.in). [Project BoLI Guidelines](https://docs.google.com/document/d/1h-EcxxaNiGzdWiwc1-V_S-Xt6n4MwE6F/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) ### About Kokrajhar Bodo * **Language Family:** Tibeto-Burman * **Subgroup:** Bodo-Garo > Bodo * **ISO 639-3:** brx * **Glottocode:** bodo1269 ### Speakers According to the [2011 Census of India](https://www.censusindia.gov.in/2011Census/Language-2011/Statement-1.pdf), there are approximately 1,482,929 speakers of the Bodo language in India. While this figure represents the total Bodo-speaking population across varieties, the specific "Bodo - Kokrajhar" variety represents the standard northern dialect. ### Distribution in India The Bodo-Kokrajhar variety is primarily spoken in the Bodoland Territorial Region of **Assam**, with the Kokrajhar district acting as its primary linguistic and cultural hub. Speakers are also distributed across adjacent districts of Assam, as well as bordering areas in the states of **West Bengal**, **Meghalaya**, and **Nagaland**. ### Major grammatical features * **Phonology:** Bodo is a tonal language, typically characterized by a contrast between two lexical tones (high and low). Its vowel inventory includes a distinctive high-back unrounded vowel /ɯ/. Consonants feature a contrast between aspirated and unaspirated voiceless stops in word-initial positions. * **Morphology:** The language is predominantly agglutinative and suffixing. Nouns are inflected for number, gender (for animate nouns), and case (including nominative, accusative, dative, genitive, locative, and instrumental). Bodo employs a robust system of numeral classifiers that categorize nouns based on shape, size, and animacy. * **Syntax:** The basic constituent word order is Subject-Object-Verb (SOV). Grammatical relations are marked postpositionally, and modifiers typically precede the head noun. Verbs are highly inflected for tense, aspect, mood, and transitivity/causation. ### Further reading * For general information on this language variety, visit the [Wikipedia Search for Bodo - Kokrajhar language](https://en.wikipedia.org/wiki/Special:Search?search=Bodo%20-%20Kokrajhar%20language). ### About Project BoLI [Project BoLI](https://boli.unreal-tece.co.in) is the flagship project of [UnReaL-TecE LLP](https://unreal-tece.co.in), which aims to build high-quality datasets for benchmarking and evaluating different kinds of AI tasks, including speech-to-text, machine translation, grammatical analysis and reasoning tasks and prompt-based evaluation of LLMs. While the project aims to build these datasets for every Indian language and variety, the primary focus is on over 1300 underserved languages and both first and second language varieties of major, scheduled languages. The project's uniqueness is not just limited to the kind of benchmarking tasks it supports (including proposing some novel tasks) and also the kind of languages and communities it supports but also in its contextualisation and implementation of a unique data governance model (not yet implemented anywhere across the globe), which mandates that all datasets released by the project and their derivatives (including the models) are co-owned by the community members and all contributors of the project, and any permission to use the dataset is not a transfer of ownership but a revocable license to use it. The conditions under which the license could be revoked are clearly mentioned as part of the BoLI License. The complete details of the project, languages and communities supported till now, the quantum of data available till now, its data governance model and other relevant documents are all publicly accessible at the [project website](https://boli.unreal-tece.co.in). ### Ethical Considerations, Consent, IPR and Attribution Project BoLI and this repository represents our commitment to not only fair remuneration to the speakers of the language but also to co-ownership and equal IPR to all the contributors who have built the dataset. We believe this is the first step to move away from the extractive data collection and use practices and ensure fairness in our treatment of the community members. As such, we have listed all speakers and transcribers as Contributors to the dataset (we insist that they are co-owners of the dataset, even though the HuggingFace platform does not provide us an explicit way of stating that) and they are further recognised as Speakers and Annotators of the dataset. This dataset is only licensed to other researchers for use in their research projects. More details about licensing and commercial use conditions are given in the License and Commercial Use sections. [Project BoLI - Data Governance Policy](https://docs.google.com/document/d/1DBGZVFJ4RG57j4rFphKIrvgbpfpj8HqZ/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) [BoLI Ethics Principle & Pledge](https://docs.google.com/document/d/1hMQzwsOJF4QDfk_DOSIGLOF3GApLQOZA/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) [Project BoLI - Digital Consent Form](https://docs.google.com/document/d/1BTQUrgVTS-sVszUUDOkbnyUNusAHVYm4/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) [Project BoLI - Field Recording of Oral Consent](https://docs.google.com/document/d/1xoRxXcNES2oQqBajY3Z7HxMwpt0oOyDK/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) [Project BoLI - TnC](https://docs.google.com/document/d/1UuiIAHWXXKF202MYqLrtOy1ftUk__POW/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) ### Dataset Access Full dataset for the language can be browsed and queried on our [app](https://life.unreal-tece.co.in). The results of the evaluation of different models will also be made available on the same link. If you would like to access the full dataset for your own use or would like to work with us in collecting more data for any language or variety, or collaborate with us in this initiative in any other way, please get in touch with us.

Type
transcriptions
Owner
BoLI
Languages
Bodo
1 speakers 465 audios
Assam Bodo Low-resource Language Tibeto-Burman
Dataset Public preview

BoLI Markodi Translation 2

### The Dataset The current data preview of Markodi is being released as part of the **[Project BoLI](https://boli.unreal-tece.co.in)**. This preview is a reflection of the full dataset and consists of the following - 1. Speech Recordings of 200 sentences in the language. 2. Transcriptions in IPA and native script(s). 3. Translations in English (which also act as prompts for the translation sentences). 4. Detailed speaker metadata, including their demographic, educational and linguistic profile. 5. Prompt in English and Hindi. The full dataset contains the following - 1. Translations of a minimum of 1000 carefully selected sentences. These sentences are selected to represent diverse morphosyntactic categories generally found in Indian languages such as demonstratives, classifiers, TAM morphology, and different sentence structures such as transitives and ditransitives. These sets of sentences are based on the standardised questionnaires built by Linguists for writing the sketch grammar of any human language, thereby, representing almost full range of morphosyntactic properties exhibited in the language. Unlike other benchmarks which focus largely on lexical level evaluation (aka domains), this is the first benchmark dataset that evaluates model's performance on a range of morphosyntactic structures, even the most uncommon ones. 2. More than 600 narrative speeches collected across 8 domains. These are recorded by at least 2 speakers. Narrations, along with their prompts, can be used to evaluate AI models on a range of prompt-based tasks. 3. Along with transcriptions in IPA and other scripts and translation, all data is interlinearly glossed at morphemic level. This gives a word-by-word meaning and morphosyntactic information, thereby, enabling evaluation of models on their deep grammatical knowledge, ability for cross-linguistic comparison and generalisation and reasoning capacity and skills in language-related puzzles. This allows for evaluating the model's capacity on reasoning tasks beyond mathematical reasoning tasks as well as their capacity to arrive at typological generalisations. 4. In addition to the data and prompt itself, as mentioned earlier, we are also making available the detailed speaker metadata and prompt-level metadata viz domain, elicitation method, multilingual prompt, target grammatical category (for translation sentences), etc. 5. In accordance with our data governance policy, all contributors to the dataset, including those recording the dataset and those transcribing it, are named and listed as data contributors. ### Dataset Preparation The speech included as part of this dataset was recorded by native speakers using a data collection application on a mobile device (Karya or in-house app, Atekho), by translating sentences from English or Hindi to the target language. The recorded speech was validated and then transcribed manually by trained linguists, working with the native speakers, following the [guidelines for the project](https://docs.google.com/document/d/1h-EcxxaNiGzdWiwc1-V_S-Xt6n4MwE6F/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) using MATra Lab, a part of the [LiFE Suite Ecosystem](https://life.unreal-tece.co.in), developed by [Unreal Tece LLP](https://unreal-tece.co.in). [Project BoLI Guidelines](https://docs.google.com/document/d/1h-EcxxaNiGzdWiwc1-V_S-Xt6n4MwE6F/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) ### About Markodi Markodi is classified as a member of the **Dravidian** language family. It belongs to the South Dravidian subgroup, showing genetic or areal relationships with major regional languages such as Malayalam, Tulu, or Kannada. ### Speakers The total speaker population of Markodi is currently unknown, as it is not separately enumerated in the [Census of India](https://censusindia.gov.in/) or other major demographic databases. ### Distribution in India Markodi is located in the Kasaragod district of **Kerala**, India, close to the border of Karnataka (approximate coordinates: 12.51° N, 74.99° E). This borderland region is highly multilingual, characterized by intense contact between Malayalam, Kannada, Tulu, and various localized tribal or minority lects. ### Major Grammatical Features Due to the lack of dedicated grammatical descriptions, specific linguistic features of Markodi remain unconfirmed but are expected to align with the typological profile of South Dravidian languages: * **Phonology:** Likely features a contrast between short and long vowels, a rich inventory of Dravidian retroflex consonants (/ʈ/, /ɖ/, /ɳ/, /ɭ/), and a lack of initial consonant clusters. The use of Malayalam script in transcriptions suggests phonological alignment with regional Malayalam or Tulu varieties. * **Morphology:** Highly agglutinating. Nouns are marked for grammatical number (singular/plural) and case (including nominative, accusative, dative, genitive, locative, and sociative) using postpositions or suffixes. Verbs typically inflect for tense (past, present, future), aspect, mood, and person-number-gender (PNG) agreement. * **Syntax:** Strictly head-final with a default Subject-Object-Verb (SOV) word order, extensive use of relative participles instead of relative clauses, and postpositional phrases. ### About Project BoLI [Project BoLI](https://boli.unreal-tece.co.in) is the flagship project of [UnReaL-TecE LLP](https://unreal-tece.co.in), which aims to build high-quality datasets for benchmarking and evaluating different kinds of AI tasks, including speech-to-text, machine translation, grammatical analysis and reasoning tasks and prompt-based evaluation of LLMs. While the project aims to build these datasets for every Indian language and variety, the primary focus is on over 1300 underserved languages and both first and second language varieties of major, scheduled languages. The project's uniqueness is not just limited to the kind of benchmarking tasks it supports (including proposing some novel tasks) and also the kind of languages and communities it supports but also in its contextualisation and implementation of a unique data governance model (not yet implemented anywhere across the globe), which mandates that all datasets released by the project and their derivatives (including the models) are co-owned by the community members and all contributors of the project, and any permission to use the dataset is not a transfer of ownership but a revocable license to use it. The conditions under which the license could be revoked are clearly mentioned as part of the BoLI License. The complete details of the project, languages and communities supported till now, the quantum of data available till now, its data governance model and other relevant documents are all publicly accessible at the [project website](https://boli.unreal-tece.co.in). ### Ethical Considerations, Consent, IPR and Attribution Project BoLI and this repository represents our commitment to not only fair remuneration to the speakers of the language but also to co-ownership and equal IPR to all the contributors who have built the dataset. We believe this is the first step to move away from the extractive data collection and use practices and ensure fairness in our treatment of the community members. As such, we have listed all speakers and transcribers as Contributors to the dataset (we insist that they are co-owners of the dataset, even though the HuggingFace platform does not provide us an explicit way of stating that) and they are further recognised as Speakers and Annotators of the dataset. This dataset is only licensed to other researchers for use in their research projects. More details about licensing and commercial use conditions are given in the License and Commercial Use sections. [Project BoLI - Data Governance Policy](https://docs.google.com/document/d/1DBGZVFJ4RG57j4rFphKIrvgbpfpj8HqZ/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) [BoLI Ethics Principle & Pledge](https://docs.google.com/document/d/1hMQzwsOJF4QDfk_DOSIGLOF3GApLQOZA/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) [Project BoLI - Digital Consent Form](https://docs.google.com/document/d/1BTQUrgVTS-sVszUUDOkbnyUNusAHVYm4/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) [Project BoLI - Field Recording of Oral Consent](https://docs.google.com/document/d/1xoRxXcNES2oQqBajY3Z7HxMwpt0oOyDK/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) [Project BoLI - TnC](https://docs.google.com/document/d/1UuiIAHWXXKF202MYqLrtOy1ftUk__POW/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) ### Dataset Access Full dataset for the language can be browsed and queried on our [app](https://life.unreal-tece.co.in). The results of the evaluation of different models will also be made available on the same link. If you would like to access the full dataset for your own use or would like to work with us in collecting more data for any language or variety, or collaborate with us in this initiative in any other way, please get in touch with us.

Type
transcriptions
Owner
BoLI
Languages
Markodi
1 speakers 29 audios
BoLI Dravidian Low-resource Language Markodi
Dataset Public preview

BoLI Dakhini Translation

### The Dataset The current data preview of Dakkhini is being released as part of the **[Project BoLI](https://boli.unreal-tece.co.in)**. This preview is a reflection of the full dataset and consists of the following - 1. Speech Recordings of 100 sentences in the language. 2. Transcriptions in IPA and native script(s). 3. Translations in English (which also act as prompts for the translation sentences). 4. Detailed speaker metadata, including their demographic, educational and linguistic profile. 5. Prompt in English and Hindi. The full dataset contains the following - 1. Translations of a minimum of 1000 carefully selected sentences. These sentences are selected to represent diverse morphosyntactic categories generally found in Indian languages such as demonstratives, classifiers, TAM morphology, and different sentence structures such as transitives and ditransitives. These sets of sentences are based on the standardised questionnaires built by Linguists for writing the sketch grammar of any human language, thereby, representing almost full range of morphosyntactic properties exhibited in the language. Unlike other benchmarks which focus largely on lexical level evaluation (aka domains), this is the first benchmark dataset that evaluates model's performance on a range of morphosyntactic structures, even the most uncommon ones. 2. More than 600 narrative speeches collected across 8 domains. These are recorded by at least 2 speakers. Narrations, along with their prompts, can be used to evaluate AI models on a range of prompt-based tasks. 3. Along with transcriptions in IPA and other scripts and translation, all data is interlinearly glossed at morphemic level. This gives a word-by-word meaning and morphosyntactic information, thereby, enabling evaluation of models on their deep grammatical knowledge, ability for cross-linguistic comparison and generalisation and reasoning capacity and skills in language-related puzzles. This allows for evaluating the model's capacity on reasoning tasks beyond mathematical reasoning tasks as well as their capacity to arrive at typological generalisations. 4. In addition to the data and prompt itself, as mentioned earlier, we are also making available the detailed speaker metadata and prompt-level metadata viz domain, elicitation method, multilingual prompt, target grammatical category (for translation sentences), etc. 5. In accordance with our data governance policy, all contributors to the dataset, including those recording the dataset and those transcribing it, are named and listed as data contributors. ### Dataset Preparation The speech included as part of this dataset was recorded by native speakers using a data collection application on a mobile device (Karya or in-house app, Atekho), by translating sentences from English or Hindi to the target language. The recorded speech was validated and then transcribed manually by trained linguists, working with the native speakers, following the [guidelines for the project](https://docs.google.com/document/d/1h-EcxxaNiGzdWiwc1-V_S-Xt6n4MwE6F/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) using MATra Lab, a part of the [LiFE Suite Ecosystem](https://life.unreal-tece.co.in), developed by [Unreal Tece LLP](https://unreal-tece.co.in). [Project BoLI Guidelines](https://docs.google.com/document/d/1h-EcxxaNiGzdWiwc1-V_S-Xt6n4MwE6F/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) ### About Dakkhini Dakhini (also known as Deccani) is classified under the Indo-Aryan branch of the Indo-Iranian subfamily of the Indo-European language family. Historically developed as a pluricentric variety of Hindustani, it has been profoundly shaped by long-standing contact with Dravidian languages. ### Speakers Precise population figures specifically for the Hyderabad variety of Dakhini are difficult to isolate because official government surveys generally aggregate Dakhini speakers under the broader category of "Urdu." In the [Census of India 2011](https://censusindia.gov.in/nada/index.php/catalog/42504), Urdu speakers nationwide numbered approximately 50.7 million, with Dakhini speakers constituting a significant, though unquantified, proportion of this population in the southern states. ### Distribution in India Dakhini - Hyderabad is primarily spoken in the metropolitan area of Hyderabad and surrounding districts in the state of Telangana. More broadly, related varieties of Dakhini are distributed across the Deccan plateau, including parts of Andhra Pradesh, northern Karnataka (such as Kalaburagi and Bidar), and central Maharashtra. ### Major grammatical features * **Phonology:** It exhibits a simplified aspirate system compared to Modern Standard Hindi-Urdu, frequently de-aspirating consonants (e.g., standard *samajh* 'understand' becomes *samaj*, and *bhī* 'also' becomes *bī*). It also reflects distinct intonational patterns and phonetic influences from neighboring Dravidian languages. * **Morphology:** Dakhini features unique pronominal forms and case-marking patterns, such as the frequent use of *mere ko* ('to me' / 'me') instead of the standard *mujhe*, and *unon* for the third-person plural/honorific pronoun. Nouns often employ distinct plural markers (like the suffix *-āñ* or *-ā*). * **Syntax:** The language displays strong syntactic convergence with Dravidian structures. Notably, it uses the conjunctive participle *bol ke* (literally 'having said') as a quotative marker and complementizer, mirroring the grammatical function of the Telugu particle *ani*. ### Further reading * Learn more about the history, linguistic structure, and literary heritage of this variety on the [Wikipedia search page for Dakhini - Hyderabad](https://en.wikipedia.org/wiki/Special:Search?search=Dakhini%20-%20Hyderabad%20language). ### About Project BoLI [Project BoLI](https://boli.unreal-tece.co.in) is the flagship project of [UnReaL-TecE LLP](https://unreal-tece.co.in), which aims to build high-quality datasets for benchmarking and evaluating different kinds of AI tasks, including speech-to-text, machine translation, grammatical analysis and reasoning tasks and prompt-based evaluation of LLMs. While the project aims to build these datasets for every Indian language and variety, the primary focus is on over 1300 underserved languages and both first and second language varieties of major, scheduled languages. The project's uniqueness is not just limited to the kind of benchmarking tasks it supports (including proposing some novel tasks) and also the kind of languages and communities it supports but also in its contextualisation and implementation of a unique data governance model (not yet implemented anywhere across the globe), which mandates that all datasets released by the project and their derivatives (including the models) are co-owned by the community members and all contributors of the project, and any permission to use the dataset is not a transfer of ownership but a revocable license to use it. The conditions under which the license could be revoked are clearly mentioned as part of the BoLI License. The complete details of the project, languages and communities supported till now, the quantum of data available till now, its data governance model and other relevant documents are all publicly accessible at the [project website](https://boli.unreal-tece.co.in). ### Ethical Considerations, Consent, IPR and Attribution Project BoLI and this repository represents our commitment to not only fair remuneration to the speakers of the language but also to co-ownership and equal IPR to all the contributors who have built the dataset. We believe this is the first step to move away from the extractive data collection and use practices and ensure fairness in our treatment of the community members. As such, we have listed all speakers and transcribers as Contributors to the dataset (we insist that they are co-owners of the dataset, even though the HuggingFace platform does not provide us an explicit way of stating that) and they are further recognised as Speakers and Annotators of the dataset. This dataset is only licensed to other researchers for use in their research projects. More details about licensing and commercial use conditions are given in the License and Commercial Use sections. [Project BoLI - Data Governance Policy](https://docs.google.com/document/d/1DBGZVFJ4RG57j4rFphKIrvgbpfpj8HqZ/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) [BoLI Ethics Principle & Pledge](https://docs.google.com/document/d/1hMQzwsOJF4QDfk_DOSIGLOF3GApLQOZA/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) [Project BoLI - Digital Consent Form](https://docs.google.com/document/d/1BTQUrgVTS-sVszUUDOkbnyUNusAHVYm4/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) [Project BoLI - Field Recording of Oral Consent](https://docs.google.com/document/d/1xoRxXcNES2oQqBajY3Z7HxMwpt0oOyDK/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) [Project BoLI - TnC](https://docs.google.com/document/d/1UuiIAHWXXKF202MYqLrtOy1ftUk__POW/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) ### Dataset Access Full dataset for the language can be browsed and queried on our [app](https://life.unreal-tece.co.in). The results of the evaluation of different models will also be made available on the same link. If you would like to access the full dataset for your own use or would like to work with us in collecting more data for any language or variety, or collaborate with us in this initiative in any other way, please get in touch with us.

Type
transcriptions
Owner
BoLI
Languages
Dakhini
2 speakers 472 audios
Low-resource Language Unreal Tece Underresourced Language BoLI
Dataset Public preview

BoLI Braj Bhasha Translation 2

### The Dataset The current data preview of Braj Bhasha is being released as part of the **[Project BoLI](https://boli.unreal-tece.co.in)**. This preview is a reflection of the full dataset and consists of the following - 1. Speech Recordings of 100 sentences in the language. 2. Transcriptions in IPA and native script(s). 3. Translations in English (which also act as prompts for the translation sentences). 4. Detailed speaker metadata, including their demographic, educational and linguistic profile. 5. Prompt in English and Hindi. The full dataset contains the following - 1. Translations of a minimum of 1000 carefully selected sentences. These sentences are selected to represent diverse morphosyntactic categories generally found in Indian languages such as demonstratives, classifiers, TAM morphology, and different sentence structures such as transitives and ditransitives. These sets of sentences are based on the standardised questionnaires built by Linguists for writing the sketch grammar of any human language, thereby, representing almost full range of morphosyntactic properties exhibited in the language. Unlike other benchmarks which focus largely on lexical level evaluation (aka domains), this is the first benchmark dataset that evaluates model's performance on a range of morphosyntactic structures, even the most uncommon ones. 2. More than 600 narrative speeches collected across 8 domains. These are recorded by at least 2 speakers. Narrations, along with their prompts, can be used to evaluate AI models on a range of prompt-based tasks. 3. Along with transcriptions in IPA and other scripts and translation, all data is interlinearly glossed at morphemic level. This gives a word-by-word meaning and morphosyntactic information, thereby, enabling evaluation of models on their deep grammatical knowledge, ability for cross-linguistic comparison and generalisation and reasoning capacity and skills in language-related puzzles. This allows for evaluating the model's capacity on reasoning tasks beyond mathematical reasoning tasks as well as their capacity to arrive at typological generalisations. 4. In addition to the data and prompt itself, as mentioned earlier, we are also making available the detailed speaker metadata and prompt-level metadata viz domain, elicitation method, multilingual prompt, target grammatical category (for translation sentences), etc. 5. In accordance with our data governance policy, all contributors to the dataset, including those recording the dataset and those transcribing it, are named and listed as data contributors. ### Dataset Preparation The speech included as part of this dataset was recorded by native speakers using a data collection application on a mobile device (Karya or in-house app, Atekho), by translating sentences from English or Hindi to the target language. The recorded speech was validated and then transcribed manually by trained linguists, working with the native speakers, following the [guidelines for the project](https://docs.google.com/document/d/1h-EcxxaNiGzdWiwc1-V_S-Xt6n4MwE6F/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) using MATra Lab, a part of the [LiFE Suite Ecosystem](https://life.unreal-tece.co.in), developed by [Unreal Tece LLP](https://unreal-tece.co.in). [Project BoLI Guidelines](https://docs.google.com/document/d/1h-EcxxaNiGzdWiwc1-V_S-Xt6n4MwE6F/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) ### About Braj Bhasha Braj Bhasha (ISO 639-3: `bra`, Glottocode: `braj1242`) belongs to the Western group of the Indo-Aryan language family, a branch of the Indo-Iranian subfamily of Indo-European languages. ### Speakers According to the [2011 Census of India (Statement 1)](https://www.censusindia.gov.in/2011Census/Language-2011/Statement-1.pdf), BrajBhasha (grouped under Hindi) has an estimated speaker population of approximately 1.6 million. ### Distribution in India The language is primarily spoken in the historic Braj (Vraj) region of Northern India. Its principal distribution encompasses western Uttar Pradesh (focusing on Mathura, Agra, Aligarh, and Hathras districts), northeastern Rajasthan (principally Bharatpur and Dholpur districts), and the southern fringes of Haryana. ### Major grammatical features * **Phonology:** BrajBhasha distinguishes between aspirated and unaspirated stops, and features a robust contrast between oral and nasalized vowels. A defining phonological and morphological characteristic is the realization of masculine nouns and adjectives ending in a final /-o/ or /-au/ sound, where Standard Hindi typically features /-ā/. * **Morphology:** The language employs postpositions for case-marking, with characteristic genitive markers such as *kau* or *ko*. It features a split-ergative nominal alignment in transitive past-tense constructions. Nouns and adjectives inflect for gender (masculine and feminine) and number. * **Syntax:** BrajBhasha is characterized by a default Subject-Object-Verb (SOV) word order. It is a head-final language where auxiliary verbs follow main verbs and noun modifiers precede the noun they modify. ### Further reading To learn more about the language's historical role as a premier literary medium of medieval Northern India, its grammar, and its modern variants, search the [BrajBhasha Language Wikipedia Entry](https://en.wikipedia.org/wiki/Special:Search?search=BrajBhasha%20language). ### About Project BoLI [Project BoLI](https://boli.unreal-tece.co.in) is the flagship project of [UnReaL-TecE LLP](https://unreal-tece.co.in), which aims to build high-quality datasets for benchmarking and evaluating different kinds of AI tasks, including speech-to-text, machine translation, grammatical analysis and reasoning tasks and prompt-based evaluation of LLMs. While the project aims to build these datasets for every Indian language and variety, the primary focus is on over 1300 underserved languages and both first and second language varieties of major, scheduled languages. The project's uniqueness is not just limited to the kind of benchmarking tasks it supports (including proposing some novel tasks) and also the kind of languages and communities it supports but also in its contextualisation and implementation of a unique data governance model (not yet implemented anywhere across the globe), which mandates that all datasets released by the project and their derivatives (including the models) are co-owned by the community members and all contributors of the project, and any permission to use the dataset is not a transfer of ownership but a revocable license to use it. The conditions under which the license could be revoked are clearly mentioned as part of the BoLI License. The complete details of the project, languages and communities supported till now, the quantum of data available till now, its data governance model and other relevant documents are all publicly accessible at the [project website](https://boli.unreal-tece.co.in). ### Ethical Considerations, Consent, IPR and Attribution Project BoLI and this repository represents our commitment to not only fair remuneration to the speakers of the language but also to co-ownership and equal IPR to all the contributors who have built the dataset. We believe this is the first step to move away from the extractive data collection and use practices and ensure fairness in our treatment of the community members. As such, we have listed all speakers and transcribers as Contributors to the dataset (we insist that they are co-owners of the dataset, even though the HuggingFace platform does not provide us an explicit way of stating that) and they are further recognised as Speakers and Annotators of the dataset. This dataset is only licensed to other researchers for use in their research projects. More details about licensing and commercial use conditions are given in the License and Commercial Use sections. [Project BoLI - Data Governance Policy](https://docs.google.com/document/d/1DBGZVFJ4RG57j4rFphKIrvgbpfpj8HqZ/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) [BoLI Ethics Principle & Pledge](https://docs.google.com/document/d/1hMQzwsOJF4QDfk_DOSIGLOF3GApLQOZA/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) [Project BoLI - Digital Consent Form](https://docs.google.com/document/d/1BTQUrgVTS-sVszUUDOkbnyUNusAHVYm4/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) [Project BoLI - Field Recording of Oral Consent](https://docs.google.com/document/d/1xoRxXcNES2oQqBajY3Z7HxMwpt0oOyDK/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) [Project BoLI - TnC](https://docs.google.com/document/d/1UuiIAHWXXKF202MYqLrtOy1ftUk__POW/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) ### Dataset Access Full dataset for the language can be browsed and queried on our [app](https://life.unreal-tece.co.in). The results of the evaluation of different models will also be made available on the same link. If you would like to access the full dataset for your own use or would like to work with us in collecting more data for any language or variety, or collaborate with us in this initiative in any other way, please get in touch with us.

Type
transcriptions
Owner
BoLI
Languages
Braj
1 speakers 399 audios
BoLI Braj Bhasha Indo-Aryan Low-resource Language
Dataset Public preview

BoLI Braj Bhasha Translation 1

### The Dataset The current data preview of Braj Bhasha is being released as part of the **[Project BoLI](https://boli.unreal-tece.co.in)**. This preview is a reflection of the full dataset and consists of the following - 1. Speech Recordings of 100 sentences in the language. 2. Transcriptions in IPA and native script(s). 3. Translations in English (which also act as prompts for the translation sentences). 4. Detailed speaker metadata, including their demographic, educational and linguistic profile. 5. Prompt in English and Hindi. The full dataset contains the following - 1. Translations of a minimum of 1000 carefully selected sentences. These sentences are selected to represent diverse morphosyntactic categories generally found in Indian languages such as demonstratives, classifiers, TAM morphology, and different sentence structures such as transitives and ditransitives. These sets of sentences are based on the standardised questionnaires built by Linguists for writing the sketch grammar of any human language, thereby, representing almost full range of morphosyntactic properties exhibited in the language. Unlike other benchmarks which focus largely on lexical level evaluation (aka domains), this is the first benchmark dataset that evaluates model's performance on a range of morphosyntactic structures, even the most uncommon ones. 2. More than 600 narrative speeches collected across 8 domains. These are recorded by at least 2 speakers. Narrations, along with their prompts, can be used to evaluate AI models on a range of prompt-based tasks. 3. Along with transcriptions in IPA and other scripts and translation, all data is interlinearly glossed at morphemic level. This gives a word-by-word meaning and morphosyntactic information, thereby, enabling evaluation of models on their deep grammatical knowledge, ability for cross-linguistic comparison and generalisation and reasoning capacity and skills in language-related puzzles. This allows for evaluating the model's capacity on reasoning tasks beyond mathematical reasoning tasks as well as their capacity to arrive at typological generalisations. 4. In addition to the data and prompt itself, as mentioned earlier, we are also making available the detailed speaker metadata and prompt-level metadata viz domain, elicitation method, multilingual prompt, target grammatical category (for translation sentences), etc. 5. In accordance with our data governance policy, all contributors to the dataset, including those recording the dataset and those transcribing it, are named and listed as data contributors. ### Dataset Preparation The speech included as part of this dataset was recorded by native speakers using a data collection application on a mobile device (Karya or in-house app, Atekho), by translating sentences from English or Hindi to the target language. The recorded speech was validated and then transcribed manually by trained linguists, working with the native speakers, following the [guidelines for the project](https://docs.google.com/document/d/1h-EcxxaNiGzdWiwc1-V_S-Xt6n4MwE6F/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) using MATra Lab, a part of the [LiFE Suite Ecosystem](https://life.unreal-tece.co.in), developed by [Unreal Tece LLP](https://unreal-tece.co.in). [Project BoLI Guidelines](https://docs.google.com/document/d/1h-EcxxaNiGzdWiwc1-V_S-Xt6n4MwE6F/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) ### About Braj Bhasha Braj Bhasha (ISO 639-3: `bra`, Glottocode: `braj1242`) belongs to the Western group of the Indo-Aryan language family, a branch of the Indo-Iranian subfamily of Indo-European languages. ### Speakers According to the [2011 Census of India (Statement 1)](https://www.censusindia.gov.in/2011Census/Language-2011/Statement-1.pdf), BrajBhasha (grouped under Hindi) has an estimated speaker population of approximately 1.6 million. ### Distribution in India The language is primarily spoken in the historic Braj (Vraj) region of Northern India. Its principal distribution encompasses western Uttar Pradesh (focusing on Mathura, Agra, Aligarh, and Hathras districts), northeastern Rajasthan (principally Bharatpur and Dholpur districts), and the southern fringes of Haryana. ### Major grammatical features * **Phonology:** BrajBhasha distinguishes between aspirated and unaspirated stops, and features a robust contrast between oral and nasalized vowels. A defining phonological and morphological characteristic is the realization of masculine nouns and adjectives ending in a final /-o/ or /-au/ sound, where Standard Hindi typically features /-ā/. * **Morphology:** The language employs postpositions for case-marking, with characteristic genitive markers such as *kau* or *ko*. It features a split-ergative nominal alignment in transitive past-tense constructions. Nouns and adjectives inflect for gender (masculine and feminine) and number. * **Syntax:** BrajBhasha is characterized by a default Subject-Object-Verb (SOV) word order. It is a head-final language where auxiliary verbs follow main verbs and noun modifiers precede the noun they modify. ### Further reading To learn more about the language's historical role as a premier literary medium of medieval Northern India, its grammar, and its modern variants, search the [BrajBhasha Language Wikipedia Entry](https://en.wikipedia.org/wiki/Special:Search?search=BrajBhasha%20language). ### About Project BoLI [Project BoLI](https://boli.unreal-tece.co.in) is the flagship project of [UnReaL-TecE LLP](https://unreal-tece.co.in), which aims to build high-quality datasets for benchmarking and evaluating different kinds of AI tasks, including speech-to-text, machine translation, grammatical analysis and reasoning tasks and prompt-based evaluation of LLMs. While the project aims to build these datasets for every Indian language and variety, the primary focus is on over 1300 underserved languages and both first and second language varieties of major, scheduled languages. The project's uniqueness is not just limited to the kind of benchmarking tasks it supports (including proposing some novel tasks) and also the kind of languages and communities it supports but also in its contextualisation and implementation of a unique data governance model (not yet implemented anywhere across the globe), which mandates that all datasets released by the project and their derivatives (including the models) are co-owned by the community members and all contributors of the project, and any permission to use the dataset is not a transfer of ownership but a revocable license to use it. The conditions under which the license could be revoked are clearly mentioned as part of the BoLI License. The complete details of the project, languages and communities supported till now, the quantum of data available till now, its data governance model and other relevant documents are all publicly accessible at the [project website](https://boli.unreal-tece.co.in). ### Ethical Considerations, Consent, IPR and Attribution Project BoLI and this repository represents our commitment to not only fair remuneration to the speakers of the language but also to co-ownership and equal IPR to all the contributors who have built the dataset. We believe this is the first step to move away from the extractive data collection and use practices and ensure fairness in our treatment of the community members. As such, we have listed all speakers and transcribers as Contributors to the dataset (we insist that they are co-owners of the dataset, even though the HuggingFace platform does not provide us an explicit way of stating that) and they are further recognised as Speakers and Annotators of the dataset. This dataset is only licensed to other researchers for use in their research projects. More details about licensing and commercial use conditions are given in the License and Commercial Use sections. [Project BoLI - Data Governance Policy](https://docs.google.com/document/d/1DBGZVFJ4RG57j4rFphKIrvgbpfpj8HqZ/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) [BoLI Ethics Principle & Pledge](https://docs.google.com/document/d/1hMQzwsOJF4QDfk_DOSIGLOF3GApLQOZA/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) [Project BoLI - Digital Consent Form](https://docs.google.com/document/d/1BTQUrgVTS-sVszUUDOkbnyUNusAHVYm4/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) [Project BoLI - Field Recording of Oral Consent](https://docs.google.com/document/d/1xoRxXcNES2oQqBajY3Z7HxMwpt0oOyDK/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) [Project BoLI - TnC](https://docs.google.com/document/d/1UuiIAHWXXKF202MYqLrtOy1ftUk__POW/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) ### Dataset Access Full dataset for the language can be browsed and queried on our [app](https://life.unreal-tece.co.in). The results of the evaluation of different models will also be made available on the same link. If you would like to access the full dataset for your own use or would like to work with us in collecting more data for any language or variety, or collaborate with us in this initiative in any other way, please get in touch with us.

Type
transcriptions
Owner
BoLI
Languages
Braj
1 speakers 997 audios
BoLI Braj Bhasha Indo-Aryan Low-resource Language
Dataset Public preview

BoLI Braj Bhasha Narration

### The Dataset The current data preview of Braj Bhasha is being released as part of the **[Project BoLI](https://boli.unreal-tece.co.in)**. This preview is a reflection of the full dataset and consists of the following - 1. Speech Recordings of 10 narrations in the language. 2. Transcriptions in IPA and native script(s). 3. Translations in English (which also act as prompts for the translation sentences). 4. Detailed speaker metadata, including their demographic, educational and linguistic profile. 5. Prompt in English and Hindi. The full dataset contains the following - 1. Translations of a minimum of 1000 carefully selected sentences. These sentences are selected to represent diverse morphosyntactic categories generally found in Indian languages such as demonstratives, classifiers, TAM morphology, and different sentence structures such as transitives and ditransitives. These sets of sentences are based on the standardised questionnaires built by Linguists for writing the sketch grammar of any human language, thereby, representing almost full range of morphosyntactic properties exhibited in the language. Unlike other benchmarks which focus largely on lexical level evaluation (aka domains), this is the first benchmark dataset that evaluates model's performance on a range of morphosyntactic structures, even the most uncommon ones. 2. More than 600 narrative speeches collected across 8 domains. These are recorded by at least 2 speakers. Narrations, along with their prompts, can be used to evaluate AI models on a range of prompt-based tasks. 3. Along with transcriptions in IPA and other scripts and translation, all data is interlinearly glossed at morphemic level. This gives a word-by-word meaning and morphosyntactic information, thereby, enabling evaluation of models on their deep grammatical knowledge, ability for cross-linguistic comparison and generalisation and reasoning capacity and skills in language-related puzzles. This allows for evaluating the model's capacity on reasoning tasks beyond mathematical reasoning tasks as well as their capacity to arrive at typological generalisations. 4. In addition to the data and prompt itself, as mentioned earlier, we are also making available the detailed speaker metadata and prompt-level metadata viz domain, elicitation method, multilingual prompt, target grammatical category (for translation sentences), etc. 5. In accordance with our data governance policy, all contributors to the dataset, including those recording the dataset and those transcribing it, are named and listed as data contributors. ### Dataset Preparation The speech included as part of this dataset was recorded by native speakers using a data collection application on a mobile device (Karya or in-house app, Atekho), by translating sentences from English or Hindi to the target language. The recorded speech was validated and then transcribed manually by trained linguists, working with the native speakers, following the [guidelines for the project](https://docs.google.com/document/d/1h-EcxxaNiGzdWiwc1-V_S-Xt6n4MwE6F/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) using MATra Lab, a part of the [LiFE Suite Ecosystem](https://life.unreal-tece.co.in), developed by [Unreal Tece LLP](https://unreal-tece.co.in). [Project BoLI Guidelines](https://docs.google.com/document/d/1h-EcxxaNiGzdWiwc1-V_S-Xt6n4MwE6F/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) ### About Braj Bhasha Braj Bhasha (ISO 639-3: `bra`, Glottocode: `braj1242`) belongs to the Western group of the Indo-Aryan language family, a branch of the Indo-Iranian subfamily of Indo-European languages. ### Speakers According to the [2011 Census of India (Statement 1)](https://www.censusindia.gov.in/2011Census/Language-2011/Statement-1.pdf), BrajBhasha (grouped under Hindi) has an estimated speaker population of approximately 1.6 million. ### Distribution in India The language is primarily spoken in the historic Braj (Vraj) region of Northern India. Its principal distribution encompasses western Uttar Pradesh (focusing on Mathura, Agra, Aligarh, and Hathras districts), northeastern Rajasthan (principally Bharatpur and Dholpur districts), and the southern fringes of Haryana. ### Major grammatical features * **Phonology:** BrajBhasha distinguishes between aspirated and unaspirated stops, and features a robust contrast between oral and nasalized vowels. A defining phonological and morphological characteristic is the realization of masculine nouns and adjectives ending in a final /-o/ or /-au/ sound, where Standard Hindi typically features /-ā/. * **Morphology:** The language employs postpositions for case-marking, with characteristic genitive markers such as *kau* or *ko*. It features a split-ergative nominal alignment in transitive past-tense constructions. Nouns and adjectives inflect for gender (masculine and feminine) and number. * **Syntax:** BrajBhasha is characterized by a default Subject-Object-Verb (SOV) word order. It is a head-final language where auxiliary verbs follow main verbs and noun modifiers precede the noun they modify. ### Further reading To learn more about the language's historical role as a premier literary medium of medieval Northern India, its grammar, and its modern variants, search the [BrajBhasha Language Wikipedia Entry](https://en.wikipedia.org/wiki/Special:Search?search=BrajBhasha%20language). ### About Project BoLI [Project BoLI](https://boli.unreal-tece.co.in) is the flagship project of [UnReaL-TecE LLP](https://unreal-tece.co.in), which aims to build high-quality datasets for benchmarking and evaluating different kinds of AI tasks, including speech-to-text, machine translation, grammatical analysis and reasoning tasks and prompt-based evaluation of LLMs. While the project aims to build these datasets for every Indian language and variety, the primary focus is on over 1300 underserved languages and both first and second language varieties of major, scheduled languages. The project's uniqueness is not just limited to the kind of benchmarking tasks it supports (including proposing some novel tasks) and also the kind of languages and communities it supports but also in its contextualisation and implementation of a unique data governance model (not yet implemented anywhere across the globe), which mandates that all datasets released by the project and their derivatives (including the models) are co-owned by the community members and all contributors of the project, and any permission to use the dataset is not a transfer of ownership but a revocable license to use it. The conditions under which the license could be revoked are clearly mentioned as part of the BoLI License. The complete details of the project, languages and communities supported till now, the quantum of data available till now, its data governance model and other relevant documents are all publicly accessible at the [project website](https://boli.unreal-tece.co.in). ### Ethical Considerations, Consent, IPR and Attribution Project BoLI and this repository represents our commitment to not only fair remuneration to the speakers of the language but also to co-ownership and equal IPR to all the contributors who have built the dataset. We believe this is the first step to move away from the extractive data collection and use practices and ensure fairness in our treatment of the community members. As such, we have listed all speakers and transcribers as Contributors to the dataset (we insist that they are co-owners of the dataset, even though the HuggingFace platform does not provide us an explicit way of stating that) and they are further recognised as Speakers and Annotators of the dataset. This dataset is only licensed to other researchers for use in their research projects. More details about licensing and commercial use conditions are given in the License and Commercial Use sections. [Project BoLI - Data Governance Policy](https://docs.google.com/document/d/1DBGZVFJ4RG57j4rFphKIrvgbpfpj8HqZ/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) [BoLI Ethics Principle & Pledge](https://docs.google.com/document/d/1hMQzwsOJF4QDfk_DOSIGLOF3GApLQOZA/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) [Project BoLI - Digital Consent Form](https://docs.google.com/document/d/1BTQUrgVTS-sVszUUDOkbnyUNusAHVYm4/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) [Project BoLI - Field Recording of Oral Consent](https://docs.google.com/document/d/1xoRxXcNES2oQqBajY3Z7HxMwpt0oOyDK/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) [Project BoLI - TnC](https://docs.google.com/document/d/1UuiIAHWXXKF202MYqLrtOy1ftUk__POW/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) ### Dataset Access Full dataset for the language can be browsed and queried on our [app](https://life.unreal-tece.co.in). The results of the evaluation of different models will also be made available on the same link. If you would like to access the full dataset for your own use or would like to work with us in collecting more data for any language or variety, or collaborate with us in this initiative in any other way, please get in touch with us.

Type
transcriptions
Owner
SpeeD-TB
Languages
Braj
12 speakers 362 audios
Low-resource Language Unreal Tece Underresourced Language BoLI
Dataset Public preview

BoLI Jhargram Bangla Translation

### The Dataset The current data preview of Jhargram Bangla is being released as part of the **[Project BoLI](https://boli.unreal-tece.co.in)**. This preview is a reflection of the full dataset and consists of the following - 1. Speech Recordings of 100 sentences in the language. 2. Transcriptions in IPA and native script(s). 3. Translations in English (which also act as prompts for the translation sentences). 4. Detailed speaker metadata, including their demographic, educational and linguistic profile. 5. Prompt in English and Hindi. The full dataset contains the following - 1. Translations of a minimum of 1000 carefully selected sentences. These sentences are selected to represent diverse morphosyntactic categories generally found in Indian languages such as demonstratives, classifiers, TAM morphology, and different sentence structures such as transitives and ditransitives. These sets of sentences are based on the standardised questionnaires built by Linguists for writing the sketch grammar of any human language, thereby, representing almost full range of morphosyntactic properties exhibited in the language. Unlike other benchmarks which focus largely on lexical level evaluation (aka domains), this is the first benchmark dataset that evaluates model's performance on a range of morphosyntactic structures, even the most uncommon ones. 2. More than 600 narrative speeches collected across 8 domains. These are recorded by at least 2 speakers. Narrations, along with their prompts, can be used to evaluate AI models on a range of prompt-based tasks. 3. Along with transcriptions in IPA and other scripts and translation, all data is interlinearly glossed at morphemic level. This gives a word-by-word meaning and morphosyntactic information, thereby, enabling evaluation of models on their deep grammatical knowledge, ability for cross-linguistic comparison and generalisation and reasoning capacity and skills in language-related puzzles. This allows for evaluating the model's capacity on reasoning tasks beyond mathematical reasoning tasks as well as their capacity to arrive at typological generalisations. 4. In addition to the data and prompt itself, as mentioned earlier, we are also making available the detailed speaker metadata and prompt-level metadata viz domain, elicitation method, multilingual prompt, target grammatical category (for translation sentences), etc. 5. In accordance with our data governance policy, all contributors to the dataset, including those recording the dataset and those transcribing it, are named and listed as data contributors. ### Dataset Preparation The speech included as part of this dataset was recorded by native speakers using a data collection application on a mobile device (Karya or in-house app, Atekho), by translating sentences from English or Hindi to the target language. The recorded speech was validated and then transcribed manually by trained linguists, working with the native speakers, following the [guidelines for the project](https://docs.google.com/document/d/1h-EcxxaNiGzdWiwc1-V_S-Xt6n4MwE6F/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) using MATra Lab, a part of the [LiFE Suite Ecosystem](https://life.unreal-tece.co.in), developed by [Unreal Tece LLP](https://unreal-tece.co.in). [Project BoLI Guidelines](https://docs.google.com/document/d/1h-EcxxaNiGzdWiwc1-V_S-Xt6n4MwE6F/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) ### About Jhargram Bangla Bangla - Jhargram is an Indo-Aryan language belonging to the Eastern Zone (Magadhan) of the Indo-Aryan branch of the Indo-European language family. It is a regional variety of Bengali, showing transitional linguistic features characteristic of the Southwestern Bengali (Rarhi) and Jharkhandi Bengali dialect groups. ### Speakers The total number of speakers of this specific regional variety is not separately isolated in national censuses. However, the overarching Bengali language is spoken by 97,237,669 people in India, according to the [2011 Census of India](https://censusindia.gov.in/nada/index.php/catalog/42561). The local population of Jhargram district, where this variety is dominant, was recorded as approximately 1,136,548 in the [2011 Census District Handbooks](https://censusindia.gov.in/nada/index.php/catalog/1393), representing the approximate pool of regional variety speakers. ### Distribution in India This variety is primarily spoken in the Jhargram district of West Bengal, India. Its usage extends into the surrounding regions of the Paschim Medinipur district, as well as the border areas of neighboring Jharkhand (especially East Singhbhum district) and Odisha. ### Major grammatical features * **Phonology:** It shares the core inventory of Standard Bengali, featuring a seven-vowel system (/i, e, æ, a, ɔ, o, u/) and a contrast between aspirated and unaspirated stops. However, it displays localized phonetic shifts, such as laxing of high-mid vowels, variations in vowel harmony, and regional intonational contours typical of the borderlands. * **Morphology:** Like other Bengali varieties, it is highly inflectional and suffixing. Nouns are inflected for case (nominative, accusative-dative, genitive, and locative) and animacy. Plurality is often marked by suffixes like *-gula* or *-gulin*. The verb system is complex, with conjugation reflecting tense, aspect, mood, and levels of honorific status (contemptuous, common, and honorific), frequently utilizing the past marker *-l-* and future marker *-b-*. * **Syntax:** The basic word order is Subject-Object-Verb (SOV). Postpositions are used rather than prepositions, and adjectives precede the nouns they modify. Negation is typically placed after the finite verb. ### Further reading To explore more about this linguistic variety and its regional context, visit the [Wikipedia Search for Bangla - Jhargram](https://en.wikipedia.org/wiki/Special:Search?search=Bangla%20-%20Jhargram%20language). ### About Project BoLI [Project BoLI](https://boli.unreal-tece.co.in) is the flagship project of [UnReaL-TecE LLP](https://unreal-tece.co.in), which aims to build high-quality datasets for benchmarking and evaluating different kinds of AI tasks, including speech-to-text, machine translation, grammatical analysis and reasoning tasks and prompt-based evaluation of LLMs. While the project aims to build these datasets for every Indian language and variety, the primary focus is on over 1300 underserved languages and both first and second language varieties of major, scheduled languages. The project's uniqueness is not just limited to the kind of benchmarking tasks it supports (including proposing some novel tasks) and also the kind of languages and communities it supports but also in its contextualisation and implementation of a unique data governance model (not yet implemented anywhere across the globe), which mandates that all datasets released by the project and their derivatives (including the models) are co-owned by the community members and all contributors of the project, and any permission to use the dataset is not a transfer of ownership but a revocable license to use it. The conditions under which the license could be revoked are clearly mentioned as part of the BoLI License. The complete details of the project, languages and communities supported till now, the quantum of data available till now, its data governance model and other relevant documents are all publicly accessible at the [project website](https://boli.unreal-tece.co.in). ### Ethical Considerations, Consent, IPR and Attribution Project BoLI and this repository represents our commitment to not only fair remuneration to the speakers of the language but also to co-ownership and equal IPR to all the contributors who have built the dataset. We believe this is the first step to move away from the extractive data collection and use practices and ensure fairness in our treatment of the community members. As such, we have listed all speakers and transcribers as Contributors to the dataset (we insist that they are co-owners of the dataset, even though the HuggingFace platform does not provide us an explicit way of stating that) and they are further recognised as Speakers and Annotators of the dataset. This dataset is only licensed to other researchers for use in their research projects. More details about licensing and commercial use conditions are given in the License and Commercial Use sections. [Project BoLI - Data Governance Policy](https://docs.google.com/document/d/1DBGZVFJ4RG57j4rFphKIrvgbpfpj8HqZ/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) [BoLI Ethics Principle & Pledge](https://docs.google.com/document/d/1hMQzwsOJF4QDfk_DOSIGLOF3GApLQOZA/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) [Project BoLI - Digital Consent Form](https://docs.google.com/document/d/1BTQUrgVTS-sVszUUDOkbnyUNusAHVYm4/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) [Project BoLI - Field Recording of Oral Consent](https://docs.google.com/document/d/1xoRxXcNES2oQqBajY3Z7HxMwpt0oOyDK/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) [Project BoLI - TnC](https://docs.google.com/document/d/1UuiIAHWXXKF202MYqLrtOy1ftUk__POW/edit?usp=sharing&ouid=106672251267017181111&rtpof=true&sd=true) ### Dataset Access Full dataset for the language can be browsed and queried on our [app](https://life.unreal-tece.co.in). The results of the evaluation of different models will also be made available on the same link. If you would like to access the full dataset for your own use or would like to work with us in collecting more data for any language or variety, or collaborate with us in this initiative in any other way, please get in touch with us.

Type
transcriptions
Owner
BoLI
Languages
Bangla
1 speakers 803 audios
Low-resource Language Unreal Tece Underresourced Language BoLI