Voice & Language Data · Services

Verified speech data for languages.

Neuravox recruits native speakers to review audio, confirm the language, check audio quality, and correct transcripts. We manage the work through our Voice & Language Data service.

Send us audio and a language brief.

We recruit and vet native speaker reviewers, run the campaign through Luper, our speech verification platform, and return a dataset with review records, QA results, and export ready files.

Who this is for

If you have audio and need human review before the data can be used for model training, publication, evaluation, or release, this is for you.

Language technology teamsResearch labsCommon Voice communitiesNLP groupsNGOsArchive projectsASR and TTS builders
What you receive

A cleaned dataset in CSV, with JSONL and Hugging Face formats on request, delivered with a checksummed manifest.

Each reviewed clip carries its decision, transcript status, rejection reason where relevant, reviewer code, timestamp, and QA status. With it comes a delivery report covering acceptance rate, rejection breakdown, reviewer agreement, gold clip accuracy where used, time statistics, verified audio hours, and payment counts. Every figure is traceable to timestamped review records, so the dataset can be defended.

How we review speech data
01

Verify

Native reviewers confirm the language, speech content, audio quality, and rejection reason for each clip. Best for corpus cleaning, archive filtering, and dataset readiness checks.

02

Correct

Reviewers fix machine generated draft transcripts. Where an ASR draft exists, correction reduces review time.

03

Transcribe

Reviewers write transcripts from scratch for languages where no usable ASR draft exists.

Service tiers

Basic

Single review per clip, standard rejection taxonomy, and a summary report.

Dataset ready

Review plus QA sampling, reviewer agreement checks, rejection breakdown, and a full delivery report.

Gold

Language lead review, seed accuracy checks, adjudication of disputed clips, documentation for publication, and a dataset manifest.

Why Neuravox

We connect skilled reviewers to accountable workflows.

Native speakers with the right language skills are reachable through local networks, universities, radio communities, and civil society groups. Luper brings task delivery, reviewer vetting, quality checks, attribution, and payment records into one workflow.

Reviewers are paid per submitted item at rates shown to them before work begins, from Luper's review ledger, within seven days of campaign close.

Proof from VoiceLink Uganda

We built this service from our own field work. VoiceLink Uganda ran on the workflow behind Luper, combining human audit, ASR draft correction, reviewer attribution, and the preparation of verified audio and transcripts for open dataset release, subject to rights and platform requirements.

514
hours of Luganda community radio
170,000
clips processed
What we need from you

Audio files

The language, and any dialect that matters

Rights confirmation

Your preferred output format

ASR drafts, if you have them

Your timeline

Start a campaign.

Write to contact [at] neuravox.org. We reply with a quote within two business days. Campaigns start after rights confirmation and deposit. We review audio only when you confirm you hold the right to share it.

Start a campaign