Community Language Data Cooperative.
A community-owned cooperative for language data.
A cooperative for language data
CLDC trains and pays contributors to record, review, and document language data. Members help set rules for consent, licensing, reuse, and revenue sharing.
Ownership and benefit sharing
Speech, oral histories, radio, and local knowledge can become training data. Communities need a say in how that data is collected, licensed, and used, and they should share in any value it creates.
The cooperative model
CLDC combines data collection, contributor training, and governance. Contributors can record material, review clips, create metadata, and take part in decisions about storage, licensing, and reuse.
Oral histories
Contributors record stories, proverbs, and local knowledge.
Radio archives
VoiceLink Uganda prepares community radio audio for review.
Training and governance
Contributors learn data tasks and help set rules for use.
VoiceLink processes radio audio. CLDC governs how communities contribute and benefit.
VoiceLink provides speech data from community radio. CLDC adds contributors, reviewers, consent records, metadata rules, licensing, and benefit sharing.
Explore VoiceLink UgandaStories and radio archives
Oral histories
Contributors record stories, proverbs, and local knowledge.
Radio archives
VoiceLink prepares radio audio. CLDC contributors review, segment, and document it.
Training and paid work
Members can train for paid data tasks.
Contributors may be trained in
- Audio collection and segmentation
- Transcription support
- Dataset quality review
- Metadata tagging
- Consent management
- Oral heritage documentation
- AI and data rights
Paid tasks may include
- Recording oral heritage content
- Validating radio derived audio
- Segmenting speech clips
- Creating metadata
- Checking linguistic accuracy
- Preparing dataset documentation
Licensing and benefit sharing
Members set rules for dataset access, licensing, reuse, and revenue sharing.
Open access applies when members approve it. Cooperative control applies to sensitive or community owned material.
- Licensing and reuse rules
- Consent and provenance records
- Access controls for sensitive material
- Revenue sharing for contributors
Pilot targets
These are proposed targets for a funded 12 month pilot.
Up to 400
community contributors onboarded and trained
80 to 120 hrs
of oral heritage recordings collected
1,500+
validated radio derived audio clips processed
2 to 3
open access datasets released on public platforms
Micropayments
for contributor tasks that pass review
Partnerships
with AI labs, universities, NGOs, media, and cultural bodies
Revenue model
The model combines dataset services, collection campaigns, training, and revenue sharing.
Dataset services
Dataset preparation and governance for institutions that need language resources.
Advisory support
Curation, translation, transcription, and quality review for institutions.
Collection campaigns
Language data collection across more languages and regions.
Certification programmes
Community Data Stewards, Digital Storytellers, Language Governance Assistants, and Oral Heritage Curators.
Revenue sharing
A share of revenue returns to contributors and communities.
Technology
- Mobile first collection tools
- Audio segmentation and QA workflows
- Metadata and governance templates
- Secure cloud storage for cooperative controlled datasets
- Open dataset hosting via Mozilla Common Voice, Hugging Face, GitHub, or Zenodo where appropriate
- Links to VoiceLink infrastructure where relevant
Measures
- 01Number of active contributors
- 02Number of paid tasks completed
- 03Average earnings per participant
- 04Number of datasets released
- 05Hours of validated audio
- 06Number of institutional partners
- 07Dataset uptake by AI, education, media, or research partners
- 08Cooperative fund distributions
- 09Contributors placed in extended digital roles
Partner on the pilot
We are seeking funders and partners for a pilot in Uganda.