Voice & Language Data · Services

Verified speech data for African languages.

Speech data verification takes native speakers listening to clips, checking the language, judging audio quality, and confirming or correcting transcripts. Neuravox runs that work as a service, part of our Voice & Language Data focus area.

Send us audio and a language brief.

We recruit and vet native speaker reviewers, run the campaign through Luper, our speech verification platform, and return a dataset with review records, QA results, and export ready files.

Who this is for

If you have audio and need human review before the data can be used for model training, publication, evaluation, or release, this is for you.

Language technology teamsResearch labsCommon Voice communitiesNLP groupsNGOsArchive projectsASR and TTS builders
What you receive

A cleaned dataset in CSV, with JSONL and Hugging Face formats on request, delivered with a checksummed manifest.

Each reviewed clip carries its decision, transcript status, rejection reason where relevant, reviewer code, timestamp, and QA status. With it comes a delivery report covering acceptance rate, rejection breakdown, reviewer agreement, gold clip accuracy where used, time statistics, verified audio hours, and payment counts. Every figure is traceable to timestamped review records, so the dataset can be defended.

Three review modes
01

Verify

Native reviewers confirm the language, speech content, audio quality, and rejection reason for each clip. Best for corpus cleaning, archive filtering, and dataset readiness checks.

02

Correct

Reviewers fix machine generated draft transcripts. Where a usable ASR model exists, correction cuts review time and yields consistent transcript quality.

03

Transcribe

Reviewers write transcripts from scratch for languages where no usable ASR draft exists.

Service tiers

Basic

Single review per clip, standard rejection taxonomy, and a summary report.

Dataset ready

Review plus QA sampling, reviewer agreement checks, rejection breakdown, and a full delivery report.

Gold

Language lead review, gold seed accuracy checks, adjudication of disputed clips, publication grade documentation, and a dataset manifest.

Why Neuravox

The reviewers exist. The rails have been missing.

Native speakers with the right language skills are reachable through local networks, universities, radio communities, and civil society groups. What has been missing is a system for task delivery, reviewer vetting, QA, attribution, and fair payment accounting. We built Luper to run that workflow.

Reviewers are paid per submitted item at rates shown to them before work begins, from Luper's review ledger, within seven days of campaign close.

Proof from VoiceLink Uganda

We built this service from our own field work. VoiceLink Uganda ran on the workflow behind Luper: human audit, ASR draft correction, reviewer attribution, and preparation of verified audio and transcripts for open dataset release, subject to rights and platform requirements.

514
hours of Luganda community radio
170,000
clips processed
What we need from you

Audio files

The language, and any dialect that matters

Rights confirmation

Your preferred output format

ASR drafts, if you have them

Your timeline

Start a campaign.

Write to contact@neuravox.org. We reply with a quote within two business days. Campaigns start after rights confirmation and deposit. We review audio only when you confirm you hold the right to share it.

Start a campaign