ProgrammeVoice & Language DataIn development

Community Language Data Cooperative.

A community owned model for African language data production, stewardship, training, and benefit sharing.

Overview

The community layer around language data.

The Community Language Data Cooperative is an emerging Neuravox Foundation programme designed to train, pay, and support Ugandan contributors to produce, document, review, and govern local language datasets.

The model connects community storytelling, oral heritage recording, radio derived audio, dataset quality review, and responsible data governance into a benefit sharing structure that supports community income and local language representation in AI systems.

Why it matters

Language knowledge exists. Ownership often doesn't.

African languages remain underrepresented in digital systems and AI tools, yet valuable language knowledge exists across everyday speech, community radio, oral histories, proverbs, folk tales, local expertise, and cultural memory.

Without responsible infrastructure and community ownership models, this knowledge can be extracted without fair benefit, or left out of future AI systems entirely. CLDC is designed to make language data work more ethical, participatory, and economically useful for communities.

How the model works

Three operational components.

Together, these components create a pathway where contributors can participate in language data production, learn AI adjacent digital skills, receive payment for quality reviewed tasks, and help govern how local language resources are documented, stored, licensed, and reused.

01

Community storytelling & oral heritage

Recording oral histories, proverbs, folk tales, and local expertise so cultural memory becomes documented language data.

02

Radio to data pipelines

Connected to VoiceLink Uganda, converting community radio content into structured speech data.

03

Data governance & workforce training

Training contributors and stewards, with consent, metadata, and governance practices built in from the start.

Relationship to VoiceLink Uganda

Voicelink builds the pipeline. CLDC builds the community ownership model around the pipeline.

VoiceLink Uganda provides one of the first infrastructure pathways for CLDC by converting community radio content into structured speech data. CLDC complements this by building the human and governance layer: contributors, reviewers, language stewards, consent processes, metadata practices, and benefit sharing rules.

Explore VoiceLink Uganda
From voices to datasets

Two pathways into the cooperative.

Community storytelling & oral heritage

Contributors record oral histories, proverbs, folk tales, and local expertise. This documents cultural memory and creates community owned audio that would otherwise stay undocumented.

Radio to data pathway

Through VoiceLink Uganda, community radio content is converted into structured speech data. CLDC contributors validate, segment, and document this radio derived audio for responsible reuse.

Training & paid contribution pathways

Skills, then paid, quality reviewed work.

Planned pathways aim to combine training with paid tasks, so participation builds both income and AI adjacent digital skills.

Contributors may be trained in

  • Audio collection and segmentation
  • Transcription support
  • Dataset quality review
  • Metadata tagging
  • Consent management
  • Oral heritage documentation
  • Responsible AI concepts

Paid tasks may include

  • Recording oral heritage content
  • Validating radio derived audio
  • Segmenting speech clips
  • Creating metadata
  • Checking linguistic accuracy
  • Preparing dataset documentation
Governance & benefit sharing

Designed for transparent, cooperative control.

CLDC is designed to support cooperative based benefit sharing. Dataset access, licensing, reuse, and revenue pathways would be governed through transparent rules that protect contributors and communities.

The model aims to support open access datasets where appropriate, and cooperative controlled datasets where content is sensitive or community owned.

  • Transparent licensing and reuse rules
  • Consent and provenance documented per source
  • Open access where appropriate; cooperative control where sensitive
  • Revenue pathways that return value to contributors
12 month pilot targets

Goals for a funded pilot phase.

These are targets for a funded 12 month pilot: goals the model is designed to reach. The pilot has not delivered them yet.

Up to 400

community contributors onboarded and trained

80 to 120 hrs

of oral heritage recordings collected

1,500+

validated radio derived audio clips processed

2 to 3

open access datasets released on public platforms

Micropayments

tied to quality reviewed tasks for contributors

Partnerships

with AI labs, universities, NGOs, media, and cultural bodies

Sustainability model

How the cooperative aims to sustain itself.

CLDC aims to move toward financial sustainability through a mix of services, campaigns, certification, and cooperative revenue sharing to reduce reliance on grants.

Dataset services

Dataset preparation and governance services for institutions needing local language resources.

Advisory support

Advisory to institutions plus curation, translation, transcription, and QA oversight.

Collection campaigns

New language data collection campaigns across additional languages and regions.

Certification programmes

Community Data Stewards, Digital Storytellers, Language Governance Assistants, and Oral Heritage Curators.

Cooperative revenue sharing

Revenue sharing mechanisms that return value to contributors and communities.

Technical infrastructure

A planned low cost stack.

  • Mobile first collection tools
  • Audio segmentation and QA workflows
  • Metadata and governance templates
  • Secure cloud storage for cooperative controlled datasets
  • Open dataset hosting via Mozilla Common Voice, Hugging Face, GitHub, or Zenodo where appropriate
  • Links to VoiceLink infrastructure where relevant
Monitoring & learning

Indicators we plan to track.

  • 01Number of active contributors
  • 02Number of paid tasks completed
  • 03Average earnings per participant
  • 04Number of datasets released
  • 05Hours of validated audio
  • 06Number of institutional partners
  • 07Dataset uptake by AI, education, media, or research partners
  • 08Cooperative fund distributions
  • 09Contributors placed in extended digital roles

Help launch the cooperative pilot.

We're seeking funders, AI labs, universities, NGOs, media, and cultural partners to pilot a community owned model for African language data.