← Back to Showcase

Project collection · Ongoing

Project BoLI

Reclaiming Data Sovereignty. Revolutionising AI Development for 1,000+ Indian Languages

Project BoLI, the flagship initiative of UnReaL-TecE LLP, is building high-quality datasets to train and benchmark next-generation AI tasks—from speech-to-text to advanced grammatical reasoning—across more than 1,000 underserved Indian languages and dialects. We are pioneering a globally unique data governance framework where all released datasets, derivative assets, and downstream models are permanently co-owned by the contributing communities. Under our groundbreaking BoLI License, commercial data usage is strictly treated as a revocable license rather than a transfer of ownership, ensuring total accountability in the era of LLMs.

Explore the language map →
268.3
Hours
64
Languages
315163
Entries
170
Speakers
Project networkLeadership, partners and funders
The BoLI Context

The current language data landscape is built on systemic inequalities that marginalise local communities and exploit the foundational work of linguists. Project BoLI addresses these compounding structural challenges across four critical areas:

Data Inaccessibility & Fragmentation

Language data is frequently locked behind proprietary, poorly documented, and non-standardised formats. Because it is rarely available publicly or through automated internet pipelines, vital linguistic resources remain siloed and unusable for broader research.

Exploitative Labour Practices

Field researchers, documentary linguists, and native community members face severely asymmetrical negotiation power due to a lack of institutional support networks. This leads to unfair compensation and zero visibility for their essential contributions.

Absence of Data Sovereignty

Under the current infrastructure, implementing equitable models like "data royalties" or community ownership of derivative linguistic assets is practically impossible.

Unethical Extraction in the AI Era

The rapid rise of Large Language Models (LLMs) has normalised predatory data-scraping practices. Massive datasets are routinely acquired without explicit permission, community consent, or any regard for Intellectual Property Rights (IPR).

Mission BoLI

Unprecedented Scale: Serving the Underserved

While our ultimate vision spans every Indian language and dialect, we aggressively focus on over 1,300 underserved languages, alongside the first- and second-language varieties of major Scheduled languages. We bring visibility to linguistic communities historically erased by mainstream technology.

Advanced Benchmarking: Moving Beyond Basic AI

We engineer pristine, high-fidelity datasets designed to stress-test and evaluate next-generation AI tasks. Our data powers: * Speech & Translation: Robust speech-to-text pipeline evaluation and complex machine translation. * Cognitive NLP: Advanced grammatical analysis, structural reasoning tasks, and rigorous prompt-based evaluation of Large Language Models (LLMs). * Novel Frameworks: Introducing entirely new benchmarking tasks never before seen in the NLP ecosystem.

A Global First: True Data Sovereignty

Project BoLI introduces a revolutionary, globally unprecedented data governance model. We fundamentally reject the traditional extraction-based data economy:

Permanent Co-Ownership

Every dataset and its derivative assets—including downstream AI models—are legally co-owned by the native community members and contributors who built them.

Revocable Licensing:

Data usage is never a transfer of ownership. Access is granted strictly through a revocable license bound by the pioneering BoLI License.

Enforced Accountability

If the strict ethical and operational conditions of the BoLI License are violated, data access is legally revoked.

Data Governance and Ethics

Project BoLI represents our commitment to not only fair remuneration to the speakers of the language but also to co-ownership and equal IPR to all the contributors who have built the dataset. We believe this is the first step to move away from the extractive data collection and use practices and ensure fairness in our treatment of the community members. As such, we have listed all speakers and transcribers as Contributors to the dataset (we insist that they are co-owners of the dataset, even though the HuggingFace platform does not provide us an explicit way of stating that) and they are further recognised as Speakers and Annotators of the dataset. This dataset is only licensed to other researchers for use in their research projects. More details about licensing and commercial use conditions are given in the License and Commercial Use sections of each dataset.

The BoLI Network

The BoLI Network is the foundational ecosystem powering Project BoLI’s mission. It is a global coalition of forward-thinking organisations, field researchers, native speakers, and linguistic communities who have united to dismantle exploitative data extraction. By contributing to the project, every stakeholder becomes an active co-owner of a truly sovereign, community-centred linguistic infrastructure built specifically for fair AI development.

To honour this collective effort, membership in the network guarantees perpetual, open access to all compiled BoLI datasets for independent research, localized development, and cultural preservation. This ensures that the people who create the value are the ones who permanently benefit from it, fundamentally shifting the power dynamics of the AI era.