Sable Foundry — The training-data platform for African languages
sable/shona-duplex-conversation v1.0.0 is live — the first commercial full-duplex Shona speech corpus  Browse →
COLLECTING NOW — HARARE · BULAWAYO · MUTARE

The training-data platform for languages AI left behind

License production-grade Shona and Ndebele speech corpora. Benchmark your models on the only public leaderboards for Zimbabwean languages. Or get paid to build the dataset yourself.

12.4 hrs validated speech
3,000 translation pairs
78 paid contributors
10/10 provinces covered
sablefoundry.co.zw/datasets/shona-duplex-conversation
sable / shona-duplex-conversation
Shona Full-Duplex Conversational Corpus
v1.0.0 License ↓
Sessions Schema Quality report Bias register Consent
{{ s.id }}
{{ s.meta }} {{ s.dur }} {{ s.qa }}
47 sessions · 96 speaker channels · WAV PCM 48 kHz manifest.csv · SHA-256 verified
BUILT FOR THE PIPELINE: ASR training/ TTS & voice modelling/ machine translation/ benchmarking & evals/ voice-banking & health lines
DATASET CATALOG

Speech data that exists nowhere else

dual release: commercial licence + 20% open subset (CC BY 4.0)
DATASETMODALITYSIZELICENCESTATUS
sable/{{ d.slug }} {{ d.desc }}
{{ d.modality }} {{ d.size }} {{ d.licence }} {{ d.status }}
AI-READY, PROVABLY

If it doesn't load in Python untouched, it doesn't ship

Every release is a complete handover package: audio, word-aligned JSONL transcripts, machine-readable schema, validation reports, bias register and consent documentation. The loading script is included — running it is our release criterion, not your integration problem.

SHA-256 manifest for every file, semantic versioning, changelogs
Dual-review transcripts — ≤5% word-error disagreement to pass
Automated gates: clipping, SNR, silence ratio, duplicates, schema
Coverage gaps disclosed in the bias register, never concealed
load_corpus.py ships with every release
# pandas + librosa. No manual edits. That's the point.
import pandas as pd, librosa

utts = pd.read_json("processed/utterances.jsonl", lines=True)
wav, sr = librosa.load(utts.audio_file[0], sr=None)

assert (utts.consent_ref.notna()).all()   # always true
assert (utts.qa_status == "pass").all()  # always true

# 47 sessions · 96 channels · word-level alignment
# dialects: zezuru · karanga · manyika · korekore
Shona ASR Benchmark — Season 1 metric: word error rate ↓ · held-out test set · 214 entries
LIVE · 18 days left
{{ row.rank }}
{{ row.initials }} {{ row.team }} {{ row.affil }}
{{ row.wer }}
baseline: whisper-large-v3 · 41.2 WER full leaderboard →
THE OPEN DATA COMMONS

Zimbabwean talent, finally training on Zimbabwean data

Every corpus releases a free CC BY 4.0 subset to the commons — with competitions, held-out test sets and public leaderboards. Students and developers stop practising on foreign data about foreign problems. Institutions bring real problems and sponsor the next season.

Submitted solutions, evaluation results and benchmarks flow back into the commons — the competition itself becomes a dataset.

Enter the benchmark
CONTRIBUTE & EARN

Your language is an export. Get paid like it.

Record conversations, read scripts, review translations — from an entry-level Android phone, on constrained networks, with resumable uploads. Published per-task rates, foreign-currency payouts over mobile money, and a ledger you can audit yourself. Contributor payments are the first call on revenue, at a fixed published share.

TASK 01 Conversation 10–15 min, paired, full-duplex
TASK 02 Scripted reading phonetically balanced sets
TASK 03 Translation sn ⇄ en, peer-reviewed
Join the contributor network
My earnings ● payout SLA on track
$86.50 this month · 23 validated tasks
{{ l.task }} {{ l.ref }}
{{ l.amount }} {{ l.state }}
Withdraw to EcoCash
GOVERNANCE

Consent is a schema field, not a promise

A recording without a valid consent link isn't repaired — it's rejected, automatically, by the pipeline. Zimbabwe's Data Protection Act [Chapter 12:07] is the floor, not the ceiling.

{{ g.tag }} {{ g.title }} {{ g.body }}
INZWI REDU · ILIZWI LETHU · OUR VOICE

The next decade of AI will speak Shona and Ndebele. Because we recorded it.

License the corpus Contribute & earn