← Back to Showcase

Project collection · Ongoing

SpeeD-TB

Speech Datasets and Models for Tibeto-Burman Languages of India

Dataset created under the Speech Datasets and Models for Tibeto-Burman Languages (SpeeD-TB), sponsored under Mission Bhashini by Ministry of Electronics and Information Technology (MEITY), Govt of India. The project aimed to create 1,200 hours of speech dataset and ASR models for 6 underresourced, tribal Tibeto-Burman languages of India speoken in Eastern and North-Eastern parts of India.

Explore the language map →
1,476.7
Hours
6
Languages
345465
Entries
2066
Speakers
Project networkLeadership, partners and funders
8 records

Host / Lead organisations

1

Consortium partners

5

Partners

1

Technology Partner

Unreal Tece LLP

Unreal Tece is a spin-off of the project with the platform being used for data management and transcription as its primary product. They now collaborate in all dsta management and preparation work.

Open partner → Explore relationships →

Funders / Funding agencies

1