← Back to Showcase

Project collection · Ongoing

Project BoLI

Reclaiming Data Sovereignty. Revolutionising AI Development for 1,000+ Indian Languages

Project BoLI, the flagship initiative of UnReaL-TecE LLP, is building high-quality datasets to train and benchmark next-generation AI tasks—from speech-to-text to advanced grammatical reasoning—across more than 1,000 underserved Indian languages and dialects. We are pioneering a globally unique data governance framework where all released datasets, derivative assets, and downstream models are permanently co-owned by the contributing communities. Under our groundbreaking BoLI License, commercial data usage is strictly treated as a revocable license rather than a transfer of ownership, ensuring total accountability in the era of LLMs.

Explore the language map →
268.3
Hours
64
Languages
315163
Entries
170
Speakers
Project networkLeadership, partners and funders

The Project BoLI guidelines are prepared jointly by the Council for Diversity and Innovation and Unreal-Tece LLP in consultation with our community researchers and partners. These serve as a generic set of implementable steps for working with tribal and smaller communities of India. As such, it should be read as more of an implementation guide for the project rather than a traditional set of annotation guidelines. It strongly advocates for community training and implementing community participatory approach in practice, along with some practical advice for the trained researchers and linguists working in the project. You can read the guide below or a friendlier version could be accessed here. A Google Doc link is also give below for downloading or getting a printable version of this document. Please note that it is not a fixed document, so expect changed in it over a period of time, based on our experiences of working with more communities across the country.

Data Collection and Annotation Guidelines