ProgrammeVoice & Language DataIn development

Community Language Data Cooperative.

A community-owned cooperative for language data.

Overview

A cooperative for language data

CLDC trains and pays contributors to record, review, and document language data. Members help set rules for consent, licensing, reuse, and revenue sharing.

Why it matters

Ownership and benefit sharing

Speech, oral histories, radio, and local knowledge can become training data. Communities need a say in how that data is collected, licensed, and used, and they should share in any value it creates.

How the model works

The cooperative model

CLDC combines data collection, contributor training, and governance. Contributors can record material, review clips, create metadata, and take part in decisions about storage, licensing, and reuse.

01

Oral histories

Contributors record stories, proverbs, and local knowledge.

02

Radio archives

VoiceLink Uganda prepares community radio audio for review.

03

Training and governance

Contributors learn data tasks and help set rules for use.

Relationship to VoiceLink Uganda

VoiceLink processes radio audio. CLDC governs how communities contribute and benefit.

VoiceLink provides speech data from community radio. CLDC adds contributors, reviewers, consent records, metadata rules, licensing, and benefit sharing.

Explore VoiceLink Uganda
Data sources

Stories and radio archives

Oral histories

Contributors record stories, proverbs, and local knowledge.

Radio archives

VoiceLink prepares radio audio. CLDC contributors review, segment, and document it.

Participation

Training and paid work

Members can train for paid data tasks.

Contributors may be trained in

  • Audio collection and segmentation
  • Transcription support
  • Dataset quality review
  • Metadata tagging
  • Consent management
  • Oral heritage documentation
  • AI and data rights

Paid tasks may include

  • Recording oral heritage content
  • Validating radio derived audio
  • Segmenting speech clips
  • Creating metadata
  • Checking linguistic accuracy
  • Preparing dataset documentation
Governance

Licensing and benefit sharing

Members set rules for dataset access, licensing, reuse, and revenue sharing.

Open access applies when members approve it. Cooperative control applies to sensitive or community owned material.

  • Licensing and reuse rules
  • Consent and provenance records
  • Access controls for sensitive material
  • Revenue sharing for contributors
Pilot

Pilot targets

These are proposed targets for a funded 12 month pilot.

Up to 400

community contributors onboarded and trained

80 to 120 hrs

of oral heritage recordings collected

1,500+

validated radio derived audio clips processed

2 to 3

open access datasets released on public platforms

Micropayments

for contributor tasks that pass review

Partnerships

with AI labs, universities, NGOs, media, and cultural bodies

Sustainability

Revenue model

The model combines dataset services, collection campaigns, training, and revenue sharing.

Dataset services

Dataset preparation and governance for institutions that need language resources.

Advisory support

Curation, translation, transcription, and quality review for institutions.

Collection campaigns

Language data collection across more languages and regions.

Certification programmes

Community Data Stewards, Digital Storytellers, Language Governance Assistants, and Oral Heritage Curators.

Revenue sharing

A share of revenue returns to contributors and communities.

Infrastructure

Technology

  • Mobile first collection tools
  • Audio segmentation and QA workflows
  • Metadata and governance templates
  • Secure cloud storage for cooperative controlled datasets
  • Open dataset hosting via Mozilla Common Voice, Hugging Face, GitHub, or Zenodo where appropriate
  • Links to VoiceLink infrastructure where relevant
Monitoring

Measures

  • 01Number of active contributors
  • 02Number of paid tasks completed
  • 03Average earnings per participant
  • 04Number of datasets released
  • 05Hours of validated audio
  • 06Number of institutional partners
  • 07Dataset uptake by AI, education, media, or research partners
  • 08Cooperative fund distributions
  • 09Contributors placed in extended digital roles

Partner on the pilot

We are seeking funders and partners for a pilot in Uganda.