Case studies

Project stories built around LiFE.

Case studies describe real projects that used LiFE or its surrounding workflows: what the project needed, how the work was organised, and what resources or systems were produced.

5 project stories in the public directory View datasets and models
01 Language documentation

BoLI

BoLI builds linguistic resources and technologies for Indigenous languages of India, bringing linguists, language experts, and NLP researchers into a shared preservation and revitalization effort.

UnReaL-TecE collaborative network India
Indigenous languages Audio recordings
documentation indigenous languages audio
02 Speech technology

SpeeD-IL

SpeeD-IL creates speech datasets and models for underrepresented Indian languages so speech technology can extend beyond commercially dominant languages.

Mission Bhashini, Government of India India
100 hours per language target 10+ languages per family
speech datasets asr language models
03 Education research

Multilingual Education

This project supports corpus development for multilingual education, teaching, learning, and assessment in India's government schools in the context of the National Education Policy.

University of Cambridge India
50+ hours classroom audio CHAT export
education classroom recordings translanguaging
04 Sociolinguistic corpus

Swiss Tamil Project

The Swiss Tamil project creates a detailed spoken Tamil corpus across South India, Northern Sri Lanka, and Switzerland to study heritage language structure through sociolinguistic cues.

Universite de Lausanne, Switzerland South India, Northern Sri Lanka, and Switzerland
200+ hours speech target Three communities
tamil heritage language speech corpus
05 Computational morphology

MorphGen

MorphGen generates inflectional wordforms and morphological features for more than 10 Indian languages, supporting synthetic dataset creation and system development.

Oxford Languages India
10+ Indian languages Inflectional wordforms
morphology synthetic data nlp