ProjectVoice & Language DataActive

Voicelink Uganda.

Community radio infrastructure for open speech data.

Neuravox team and station engineers beside the FM transmitter at Voice of Teso in Soroti, UgandaHardware verification with radio partner Voice of Teso, Soroti, Uganda
Project overview

Community radio for open speech data

VoiceLink Uganda turns community radio archives into open speech data. The project processes recordings, segments speech, records source and language metadata, reviews consent, and prepares datasets for reuse.

The problem

The data gap

Many languages lack the speech data needed for speech and language technology. Community radio stations already hold recordings of interviews, discussions, call in programmes, and public interest broadcasts.

These archives are often unstructured. VoiceLink provides a process for documenting, preparing, and governing them for reuse.

What we are building

The VoiceLink pipeline

The pipeline covers station agreements, audio intake, segmentation, metadata, monitoring, and release review.

01

Station partnerships

Community radio stations provide audio that meets the project requirements.

02

Audio ingestion

Long recordings enter the processing pipeline.

03

Speech segmentation

Recordings are divided into speech clips.

04

Metadata

Each clip records its language, source, station, programme, and processing status.

05

Monitoring

A dashboard tracks audio through each stage.

06

Release review

Consent, quality, and publication requirements are checked before release.

How it works

How audio becomes data

Each recording follows the same process before release.

  1. 01Radio stations and community media partners provide audio.
  2. 02Long form recordings are ingested into the Voicelink pipeline.
  3. 03Audio is processed into shorter speech clips.
  4. 04Metadata is structured around language, source, station, programme context, and processing status.
  5. 05Data quality, consent, governance, and publication pathways are reviewed before reuse.
  6. 06Outputs can support open speech datasets, language technology, and public interest AI research.
Progress & outputs

Current outputs

The project has processed hundreds of hours of radio audio and produced thousands of speech clips. The dashboard reports the current figures.

Hundreds of hours

of radio audio processed

Thousands

of speech clips generated

Multiple

languages documented

Community

radio partners in Uganda

Governance & data stewardship

Data governance

VoiceLink records consent, source, processing status, and reuse terms before release.

Community benefit

Data use should return value to the communities that provided it.

Transparency

Each dataset records its sources and processing history.

Consent

Consent and programme context are recorded before release.

Data review

Data quality is checked before publication.

Dashboard

Pipeline dashboard

The dashboard tracks audio from intake to release.

Partners & support

Project partners

Voicelink Uganda has been supported by Mozilla Common Voice and is implemented by Neuravox Foundation with community radio partners in Uganda.

Collaborate

Support open speech data

We work with funders, radio stations, researchers, and language communities on open speech data.