Dataset Showcase

SpeeD-TB Meitei Narration (Phase 1)

Whole project public CC BY-NC-SA 4.0 MeiteiSpeeD-TBTribal LanguageUnderresourced LanguageManipurBhashiniNLTMMEITYManipur UniversityIIT-KharagpurLow-resource LanguageUnreal Tece

Research network

Explore this project in context

Follow its languages, people, related projects, and public data collections.

Select a connected item to inspect it and continue navigating.

About this project

Meitei (also known as Manipuri and Meetei; ISO 639-3: mni, Glottocode: meit1246 / mani1292) belongs to the Tibeto-Burman language family. Its precise subgrouping within Tibeto-Burman remains a subject of academic debate, often classified within its own independent branch or grouped tentatively with the Kuki-Chin-Naga languages. According to the 2011 Census of India, there are approximately 1.76 million native speakers of Manipuri in India. The language is primarily spoken in the northeastern state of Manipur, where it serves as the official state language and the lingua franca among diverse ethnic groups. Significant speaker communities also exist in the neighbouring states of Assam, Tripura, and Nagaland. Owing to its official status as one of the scheduled languages of India, a large amount of resource development work has been undertaken for the language. Our current corpus adds to the ever-increasing corpora being collected for the language.

Dataset Description

The Meitei Speech Dataset, developed as part of the Speech Datasets and Models for Tibeto-Burman Languages (Project SpeeD-TB) funded under Mission Bhashini, is a transcribed speech corpus of the language. The full dataset comprises over 200 hours of high-quality audio recordings paired with accurate transcriptions in both IPA and Roman script, making it ** one of the largest speech resources for the language**, which not only enables building and evaluating voice models in low-resource linguistic settings but also enables large-scale linguistic description and documentation of the language. The audio data captures a diverse range of speakers across different age groups, genders, and education levels, ensuring variability in pronunciation, speech patterns, and tone. It includes both spontaneous and read speech collected in naturalistic and semi-controlled environments, thereby reflecting real-world linguistic usage. The transcriptions are carefully prepared and normalised to maintain consistency, supporting robust model training. This dataset also contributes to the preservation and digital documentation of the language and culture by transforming oral knowledge into structured, machine-readable formats.

Almost 60% of the data in the corpus is included from domains of agriculture, education and science & technology. The rest of the data is from varied domains including culture, lifecycle, sports, entertainment, healthcare and oral history, thereby giving extensive coverage. We have also used a variety of elicitation methods for collecting the data, including translations, narrations, lectures, role-play, spontaneous conversations, interviews and picture and video descriptions. The released dataset is meticulously mapped to rich metadata, including demographic and linguistic metadata of the speakers, domains, elicitation methods and individual prompts. The audio included in the current dataset is already sliced at sentence level, thereby ready to be integrated into the model training pipeline out-of-the-box.

The overall dataset of the project is collected over multiple phases and using multiple questionnaires. All the questionnaires and datasets are being released as part of the project.

Ethical Considerations, IPR and Attribution

This repository represents our commitment to not only fair remuneration to the speakers of the language but also to co-ownership and equal IPR to all the contributors who have built the dataset. We believe this is the first step to move away from extractive data collection and use practices and ensure fairness in our treatment of the community members. As such, we have listed all speakers and transcribers as Contributors to the dataset. This dataset is only licensed to other researchers for use in their research projects. More details about licensing and commercial use conditions are given in the License and Commercial Use sections.

Connected research

Relationships and public data

Record collections are summarized by type and remain available through the paginated public browser.

Contributors

Project collaborators

Laxmi LCommunity Collaborator
Nibha KhumanCommunity Collaborator
Dhanabir ThounaojamCommunity Collaborator
Neha SinamCommunity Collaborator
PriyojitCommunity Collaborator
Nameirakpam AmitCommunity Collaborator

Reference

How to cite

A preferred citation has not been supplied.

Licence CC BY-NC-SA 4.0

Contact

Project contacts

  • account_circleProject ownerSpeeD-TB

For data access, use the request or workspace action in the project panel.