Labari Voice

Voice infrastructure · low-resource languages

The voice data your models have never heard.

Eight languages. Around 200 million speakers. Native-speaker recordings, transcription, 14 annotation layers, commercial rights included — produced in studio with native speakers under contract, delivered via API.

labari.dev · hello@labari.dev

Why this data doesn't exist

Speech models were trained on the languages that were easiest to collect. Hundreds of millions of speakers fall outside that set.

This is not a demand problem. It is a production problem — and it is why the gap has stayed open.

You cannot scrape this data. It has to be recorded: in country, in studio, with native speakers under contract, on terms that clear commercial rights at the source. Then transcribed, in languages whose written standards are still consolidating and whose competent transcribers are few.

That barrier is the reason the data isn't available. It is also what we built.

Languages

Language Primary countries
Fon Benin
Ewe Togo, Ghana
Yoruba Nigeria, Benin
Hausa Nigeria, Niger
Wolof Senegal, Gambia
Pulaar Senegal, Mauritania, Guinea
Bambara Mali
Maninka Guinea, Mali

Need a language that isn't listed? Our sourcing network extends beyond the published catalog.

What a Labari dataset contains

How we produce

End-to-end, in-house: speaker sourcing, studio recording, segmentation, transcription, orthographic validation and QA.

Every corpus carries its own audit trail — segmentation thresholds, per-segment quality measurements, SHA-256 checksums for sources and segments, and the full production parameters. Where a recording chain applies processing that affects a measurement, we say so and flag what it affects. We would rather publish a caveat than let a buyer find it after delivery.

That traceability is the product as much as the audio: it is what lets you filter a corpus on quality before training, and what lets you defend its provenance afterwards.

Licensing

Describe your needs →

Public samples

The corpora published on this page are pilot releases: small, unlabeled, and intended for method review rather than training. They exist so the segmentation, measurement and documentation behind the catalog can be audited before you evaluate the annotated product.

The commercial catalog — transcribed, annotated, licensed — is not distributed here.

Browse the catalog → · Request a sample →