Voicelink Uganda.
Community radio infrastructure for open African speech data.
Hardware verification with radio partner Voice of Teso, Soroti, UgandaTurning community radio into governed African speech data.
Voicelink Uganda is a Neuravox Foundation initiative building the infrastructure needed to transform community radio content into governed, reusable African speech datasets.
The project works with radio stations and local partners to process spontaneous speech, segment long form recordings, document language and consent contexts, and prepare speech data for responsible use in language AI systems. It is infrastructure, governance, and data stewardship work.
Valuable speech data is locked out of reach.
African languages are significantly underrepresented in modern speech and language technologies. At the same time, valuable speech data already exists across community radio stations, call in programmes, local discussions, interviews, and public interest broadcasts.
Much of this material is not structured, documented, or technically prepared for responsible reuse. Without better infrastructure, African communities risk remaining data poor in the systems that will shape future access to information, services, and AI tools.
Data infrastructure, end to end.
Voicelink builds data infrastructure for collecting, processing, and governing African speech data from community radio sources, combining station partnerships, ingestion workflows, segmentation, documentation, monitoring, and governance.
Station partnerships
Working with community radio stations and local media partners to identify suitable audio sources.
Audio ingestion
Ingesting long form radio recordings into the Voicelink processing pipeline.
Speech segmentation
Processing long recordings into shorter, usable speech clips suitable for reuse.
Metadata documentation
Structuring metadata around language, source, station, programme context, and processing status.
Pipeline monitoring
Operating a live processing pipeline and dashboard that track data as it moves through each stage.
Governance processes
Reviewing consent, data quality, and publication pathways that support responsible reuse.
From broadcast to governed dataset.
Each recording moves through a documented pipeline so that data quality, consent, and provenance are clear before anything is reused.
- 01Radio stations and community media partners provide suitable audio sources.
- 02Long form recordings are ingested into the Voicelink pipeline.
- 03Audio is processed into shorter speech clips.
- 04Metadata is structured around language, source, station, programme context, and processing status.
- 05Data quality, consent, governance, and publication pathways are reviewed before reuse.
- 06Outputs can support open speech datasets, language technology, and public interest AI research.
Building momentum, responsibly.
The project has already processed hundreds of hours of radio audio and generated thousands of speech clips through the Voicelink pipeline. Live figures are available on the project's pipeline dashboard.
Hundreds of hours
of radio audio processed
Thousands
of speech clips generated
Multiple
African languages documented
Community
radio partners in Uganda
A governed data pathway.
Voicelink is designed around community benefit, transparency, and accountable data stewardship. The project creates a governed pathway where speech resources are documented, reviewed, and reused in ways that benefit African language communities.
Community benefit
Speech resources are reused in ways that benefit African language communities.
Transparency
Sources, languages, and processing status are documented so data provenance is clear.
Consent & documentation
Consent and programme context are recorded and reviewed before any publication or reuse.
Accountable stewardship
A governed pathway reviews data quality before anything is published.
A live pipeline you can watch work.
Voicelink runs a live processing pipeline and dashboard that track audio as it is ingested, segmented, documented, and prepared for reuse. The dashboard is the project's operational proof of work: explore it directly.
Built with communities and supporters.
Voicelink Uganda has been supported by Mozilla Common Voice and is implemented by Neuravox Foundation with community radio partners in Uganda.
The full project lives at voicelink.cloud.
Support open African speech data.
We work with funders, radio stations, researchers, and language communities building responsible speech data infrastructure for African languages.