Verified speech data for African languages.
Speech data verification takes native speakers listening to clips, checking the language, judging audio quality, and confirming or correcting transcripts. Neuravox runs that work as a service, part of our Voice & Language Data focus area.
Send us audio and a language brief.
We recruit and vet native speaker reviewers, run the campaign through Luper, our speech verification platform, and return a dataset with review records, QA results, and export ready files.
If you have audio and need human review before the data can be used for model training, publication, evaluation, or release, this is for you.
A cleaned dataset in CSV, with JSONL and Hugging Face formats on request, delivered with a checksummed manifest.
Each reviewed clip carries its decision, transcript status, rejection reason where relevant, reviewer code, timestamp, and QA status. With it comes a delivery report covering acceptance rate, rejection breakdown, reviewer agreement, gold clip accuracy where used, time statistics, verified audio hours, and payment counts. Every figure is traceable to timestamped review records, so the dataset can be defended.
Verify
Native reviewers confirm the language, speech content, audio quality, and rejection reason for each clip. Best for corpus cleaning, archive filtering, and dataset readiness checks.
Correct
Reviewers fix machine generated draft transcripts. Where a usable ASR model exists, correction cuts review time and yields consistent transcript quality.
Transcribe
Reviewers write transcripts from scratch for languages where no usable ASR draft exists.
Basic
Single review per clip, standard rejection taxonomy, and a summary report.
Dataset ready
Review plus QA sampling, reviewer agreement checks, rejection breakdown, and a full delivery report.
Gold
Language lead review, gold seed accuracy checks, adjudication of disputed clips, publication grade documentation, and a dataset manifest.
The reviewers exist. The rails have been missing.
Native speakers with the right language skills are reachable through local networks, universities, radio communities, and civil society groups. What has been missing is a system for task delivery, reviewer vetting, QA, attribution, and fair payment accounting. We built Luper to run that workflow.
Reviewers are paid per submitted item at rates shown to them before work begins, from Luper's review ledger, within seven days of campaign close.
We built this service from our own field work. VoiceLink Uganda ran on the workflow behind Luper: human audit, ASR draft correction, reviewer attribution, and preparation of verified audio and transcripts for open dataset release, subject to rights and platform requirements.
Audio files
The language, and any dialect that matters
Rights confirmation
Your preferred output format
ASR drafts, if you have them
Your timeline
Start a campaign.
Write to contact@neuravox.org. We reply with a quote within two business days. Campaigns start after rights confirmation and deposit. We review audio only when you confirm you hold the right to share it.