Dataset Showcase

SpeeD-TB Nyishi Narration (Phase 1)

## About Nyishi Nyishi is an under-resourced Tibeto-Burman language spoken by the Nyishi people, the largest ethnic group in Arunachal Pradesh, India. According to Census 2011, there are approximately **3 lakh speakers** of the languages. The language belongs to the Tani branch of the Tibeto-Burman language family. Owing to its status as one of the largest languages of Arunachal Pradesh, some resources for the language development such as corpus and a language model have been developed. However, the language lacks any large significant speech or text corpus. ## Dataset Description The Nyishi Speech Dataset, developed as part of the [Speech Datasets and Models for Tibeto-Burman Languages (Project SpeeD-TB)](https://sites.google.com/view/speed-tb/ "SpeeD-TB Project Website") funded under Mission Bhashini, is a transcribed speech corpus of Nyishi,. The full dataset comprises over **200 hours of high-quality audio recordings paired with accurate transcriptions in both IPA and Roman script**, making it the **the largest speech resource for the language** that not only enables building and evaluating voice models in low-resource linguistic settings but also enables large-scale linguistic description and documentation of the language. The audio data captures a diverse range of speakers across different age groups, genders, and education, ensuring variability in pronunciation, speech patterns, and tone. It includes both spontaneous and read speech collected in naturalistic and semi-controlled environments, thereby reflecting real-world linguistic usage. The transcriptions are carefully prepared and normalised to maintain consistency, supporting robust model training. This dataset also contributes to the preservation and digital documentation of the Nyishi language and culture by transforming oral knowledge into structured, machine-readable formats. Almost 60% of the data in the corpus is included from domains of agriculture, education and science & technology. Rest of the data is from varied domains including culture, lifecycle, sports, entertainment, healthcare and oral history, thereby, giving a large coverage. We have also used a variety of elicitation methods for collecting the data including translations, narrations, lectures, role-play, spontaneous conversations, interviews and picture and video descriptions. The released dataset is meticulously mapped to a rich metadats including demographic and linguistic metadata of the speakers, domains, elicitation methods and to individual prompts. The audio included in the current dataset is already sliced at sentence level, thereby, ready to be integrated into the model training pipeline out-of-the-box. The overall dataset of the project is collected over multiple phases and using multiple questionnaires. This repository contains the full dataset for the language collected till now. ## Ethical Considerations, IPR and Attribution This repository represents our committment to not only fair remuneration to the speakers of the language but also to co-ownership and equal IPR to all the contributors who have built the dataset. We believe this is the first step to move away from the extractive data collection and use practices and ensure fairness in our treatment of the community members. As such, we have listed all speakers and transcribers as Contributors to the dataset (we insist that they are co-owners of the dataset even though HuggingFace does not provide us an explicit way of stating that) and they are further recognised as Speakers and Annotators of the dataset. This dataset is only licensed to other researchers for use in their research projects. More details about licensing and commercial use conditions are given in License and Commercial Use sections.

Whole project public 13974 public entries 500 anonymous preview CC BY-NC-SA 4.0 Nyishi SpeeD-TB Tribal Language Underresourced Language Arunachal Pradesh Bhashini NLTM MEITY IIT-Kharagpur Low-resource Language Unreal Tece SpeeD-TB

Browse

Public entries

13974 entries are available to logged-in public viewers. Anonymous visitors can preview 500 entries.

Browse entries

Analysis is available after login for public entries.

Stats

Project signals

101speakers
14479audios
1languages
13974public

Credits

People and terms

Tarak Tatam Community Collaborator
tame Community Collaborator
Bagjam ana Community Collaborator
uku toko Community Collaborator
Likka Aku Community Collaborator
Joram Chobin Community Collaborator
Pel Kamin Community Collaborator
Likha Aku Community Collaborator
yowa Yapin Community Collaborator
Taj Jarbo Community Collaborator
Pochu Community Collaborator
phil tedi Community Collaborator
Biri Yachu Community Collaborator
Byabang Yamin Community Collaborator
Neelam tolu Community Collaborator
Joram Niya Community Collaborator
Tagan Mage Community Collaborator
Joram Puj Community Collaborator
Totosuka Community Collaborator
Neelam Joshua Community Collaborator
Joram Aina and Joram Dui Community Collaborator
Tar Ado Community Collaborator
Toko Yaram Community Collaborator
Hinium Mama Community Collaborator
phill doni Community Collaborator
Joram Maloti Community Collaborator
jorambell Community Collaborator
Meko Singhi Community Collaborator
Byabang sime Community Collaborator
Tania Tayer Community Collaborator
byabang chingsu Community Collaborator
Joram Yano Community Collaborator
Joram Maka Community Collaborator
LIKHA BAI Community Collaborator
Joram Tahe Community Collaborator
Ichiko Pito Community Collaborator
Talomary Community Collaborator
NABAM KHANDU Community Collaborator
Joram Tangam Community Collaborator
Joram Tapu Community Collaborator
Aiya gollom Community Collaborator
Baby Mugli Community Collaborator
Joram Obbi Community Collaborator
Nabam Jakap Community Collaborator
Bengia Ania Community Collaborator
honi japan Community Collaborator
Joram Byani Community Collaborator
Likha Rich Community Collaborator
Nangram Tare Community Collaborator
Toko Ipa Community Collaborator
pachu mello Community Collaborator
Joram I Community Collaborator
Likha Tana Community Collaborator
Joram Taja Community Collaborator
Yana Banang Community Collaborator
Joram Motu Community Collaborator
biki sima Community Collaborator
toko lel Community Collaborator
talo badal Community Collaborator
taho taku Community Collaborator
Likha Ana Community Collaborator
toko laam Community Collaborator
Joram Aku Community Collaborator
lod yayo Community Collaborator
DUI TAKA Community Collaborator
Likha Ramleo Community Collaborator
Shri Joram Taja Community Collaborator
Licha Ribia Community Collaborator
joram paul Community Collaborator
Deadpool Community Collaborator
Lishi Taza Joram Community Collaborator
Likha Ya Joram Community Collaborator
Dobiam Aniya Community Collaborator
Taj Takam Community Collaborator
donu kama Community Collaborator
Pochu Mangu Community Collaborator
Joram Paul and Joram Yama Community Collaborator
Mina Nihnaka Community Collaborator
TaloMartha Community Collaborator
Techi Kujma Community Collaborator
Joram Nama and Nabam Runaya Community Collaborator
Joram Tane Community Collaborator
Dorai Eje Community Collaborator
Joram Tami Community Collaborator
Joram Tuka Community Collaborator
Joram Nikam and Joram Chobin Community Collaborator
Tare Yaya Community Collaborator
Joram Jyoti Community Collaborator
Talo Paga Community Collaborator
Toko Doni Community Collaborator
Nikh Tado Community Collaborator
Taj Jikare Community Collaborator
DOBIAM TAKA Community Collaborator
Lishi Aku Community Collaborator
Joram Nikam Community Collaborator
Pachurinya Community Collaborator
Toko rina Community Collaborator
Techi mary Community Collaborator
_eESc0se98Q Community Collaborator
Adrita Bhattacharya Research Assistant
Ahana Pramanik Research Assistant
AnaghaS Research Assistant
Anindita Research Assistant
AnishaDutta Research Assistant
Ishita Chowdhury Research Assistant
Lishi Aku Research Assistant
ManashiM Research Assistant