Whisper-Small-Shona πΏπΌ
The first open Shona (chiShona) speech recognition model trained on real Shona speech data.
Part of Project Nyaradzai β open language infrastructure for chiShona and isiNdebele.
Model description
This model is a fine-tuned version of openai/whisper-small for automatic speech recognition (ASR) in chiShona, the primary language of Zimbabwe spoken by over 10 million people.
It was trained on the Google WAXAL dataset (google/WaxalNLP, config sna_asr) β a large-scale, CC-BY-4.0 licensed multilingual African speech corpus. This is currently the largest openly available Shona speech dataset.
| Property | Detail |
|---|---|
| Base model | openai/whisper-small |
| Language | chiShona (sn) |
| Task | Automatic Speech Recognition |
| Training data | google/WaxalNLP (sna_asr split) |
| Training examples | ~13,000 labelled utterances |
| Training steps | 4,000 |
| Word Error Rate | 36.42% (WAXAL test split) |
| License | Apache-2.0 |
Why this model matters
ChiShona has been effectively absent from speech technology. Speakers cannot:
- Caption or transcribe Shona TV and video
- Use voice-to-text in Shona on phones or computers
- Access accessibility tools (subtitles, screen readers) in their language
This model is a first step toward fixing that. A 36.42% WER on a first fine-tune with no Zimbabwean broadcast-domain data is a solid foundation β improving with more data is the next step.
How to use
from transformers import pipeline
asr = pipeline(
"automatic-speech-recognition",
model="Starsm91/whisper-small-shona"
)
# Transcribe an audio file
result = asr(
"your_shona_audio.mp3",
generate_kwargs={"language": "shona", "task": "transcribe"},
return_timestamps=True
)
print(result["text"])
For audio clips shorter than 30 seconds you can omit return_timestamps:
result = asr(
"short_clip.mp3",
generate_kwargs={"language": "shona", "task": "transcribe"}
)
print(result["text"])
Test on WAXAL data directly
from transformers import pipeline
from datasets import load_dataset
asr = pipeline("automatic-speech-recognition",
model="Starsm91/whisper-small-shona")
ds = load_dataset("google/WaxalNLP", "sna_asr",
split="test", streaming=True)
for i, item in enumerate(ds):
if i >= 3:
break
result = asr(item["audio"]["array"],
generate_kwargs={"language": "shona", "task": "transcribe"},
return_timestamps=True)
print(f"Model: {result['text'].strip()}")
print(f"Reference: {item['transcription'].strip()}\n")
Training details
| Parameter | Value |
|---|---|
| Learning rate | 1e-5 |
| Batch size | 4 (with gradient accumulation of 4) |
| Warmup steps | 200 |
| Max steps | 4,000 |
| FP16 | Yes |
| Framework | HuggingFace Transformers |
| Hardware | Kaggle T4 GPU |
| Training time | ~5 hours |
Data preprocessing
- Loaded
trainsplit only (~14,109 examples) - Filtered examples with missing transcripts or audio shorter than 0.1 seconds
- 95/5 train/eval split (seed=42)
- Audio resampled to 16kHz
Evaluation
Word Error Rate (WER): 36.42% on the WAXAL test split.
This means the model correctly transcribes roughly 2 in every 3 Shona words. For context:
- This is a first fine-tune with no Zimbabwean broadcast-domain data
- The WAXAL data covers diverse Shona speakers from multiple regions
- Expected improvement with: more training steps, Zimbabwean broadcast audio, Common Voice data
Limitations
- Trained exclusively on WAXAL speech β may struggle with strong regional Zimbabwean accents (Harare urban, rural dialects)
- No broadcast/news domain data β radio and TV speech may score higher WER
- 36.42% WER means roughly 1 in 3 words will need correction for formal use
- Best suited for: voice memos, conversational transcription, accessibility tools, research
Roadmap
The next planned improvements (via Project Nyaradzai):
- v0.6 β Fine-tune on Zimbabwean broadcast audio (ZBC radio archives, community radio)
- Common Voice campaign β crowdsource Shona recordings across Zimbabwe for a locally-representative dataset
- whisper-medium-shona β larger model for higher accuracy once more data is available
- Live captioning pipeline β deploy for Zimbabwean broadcasters
Citation
If you use this model in your research or work, please cite:
@misc{whisper-small-shona-2026,
title={Whisper-Small-Shona: First Open Shona ASR Model},
author={Mateta, Stanley},
year={2026},
publisher={Hugging Face},
url={https://huggingface.co/Starsm91/whisper-small-shona},
note={Part of Project Nyaradzai: https://github.com/stanleymateta-tech/Project-Nyaradzai}
}
Data attribution
Training data: Google WAXAL (West African and eXtra-continental African Languages) dataset. License: CC-BY-4.0. Citation: WAXAL: A Multilingual Speech Dataset for African Languages, Google Research, 2026.
Project Nyaradzai
This model is part of a larger open-source project to bring chiShona and isiNdebele into the digital era:
- π€ Spellchecker: 214,000+ word forms for Word, LibreOffice, Firefox
- π£ ASR model: This model
- β¨οΈ Android keyboard: In development
- π Validated lexicon: Community-driven, in progress
GitHub: github.com/stanleymateta-tech/Project-Nyaradzai
Contributions welcome β especially from Shona and Ndebele native speakers. No coding required to contribute: report a word.
Mutauro wedu, panyika yose β Our language, everywhere in the world.
- Downloads last month
- 58
Model tree for Starsm91/whisper-small-shona
Base model
openai/whisper-small