Whisper-Small-Shona πŸ‡ΏπŸ‡Ό

The first open Shona (chiShona) speech recognition model trained on real Shona speech data.

Part of Project Nyaradzai β€” open language infrastructure for chiShona and isiNdebele.

Model description

This model is a fine-tuned version of openai/whisper-small for automatic speech recognition (ASR) in chiShona, the primary language of Zimbabwe spoken by over 10 million people.

It was trained on the Google WAXAL dataset (google/WaxalNLP, config sna_asr) β€” a large-scale, CC-BY-4.0 licensed multilingual African speech corpus. This is currently the largest openly available Shona speech dataset.

Property Detail
Base model openai/whisper-small
Language chiShona (sn)
Task Automatic Speech Recognition
Training data google/WaxalNLP (sna_asr split)
Training examples ~13,000 labelled utterances
Training steps 4,000
Word Error Rate 36.42% (WAXAL test split)
License Apache-2.0

Why this model matters

ChiShona has been effectively absent from speech technology. Speakers cannot:

  • Caption or transcribe Shona TV and video
  • Use voice-to-text in Shona on phones or computers
  • Access accessibility tools (subtitles, screen readers) in their language

This model is a first step toward fixing that. A 36.42% WER on a first fine-tune with no Zimbabwean broadcast-domain data is a solid foundation β€” improving with more data is the next step.

How to use

from transformers import pipeline

asr = pipeline(
    "automatic-speech-recognition",
    model="Starsm91/whisper-small-shona"
)

# Transcribe an audio file
result = asr(
    "your_shona_audio.mp3",
    generate_kwargs={"language": "shona", "task": "transcribe"},
    return_timestamps=True
)
print(result["text"])

For audio clips shorter than 30 seconds you can omit return_timestamps:

result = asr(
    "short_clip.mp3",
    generate_kwargs={"language": "shona", "task": "transcribe"}
)
print(result["text"])

Test on WAXAL data directly

from transformers import pipeline
from datasets import load_dataset

asr = pipeline("automatic-speech-recognition",
               model="Starsm91/whisper-small-shona")

ds = load_dataset("google/WaxalNLP", "sna_asr",
                  split="test", streaming=True)

for i, item in enumerate(ds):
    if i >= 3:
        break
    result = asr(item["audio"]["array"],
                 generate_kwargs={"language": "shona", "task": "transcribe"},
                 return_timestamps=True)
    print(f"Model:     {result['text'].strip()}")
    print(f"Reference: {item['transcription'].strip()}\n")

Training details

Parameter Value
Learning rate 1e-5
Batch size 4 (with gradient accumulation of 4)
Warmup steps 200
Max steps 4,000
FP16 Yes
Framework HuggingFace Transformers
Hardware Kaggle T4 GPU
Training time ~5 hours

Data preprocessing

  • Loaded train split only (~14,109 examples)
  • Filtered examples with missing transcripts or audio shorter than 0.1 seconds
  • 95/5 train/eval split (seed=42)
  • Audio resampled to 16kHz

Evaluation

Word Error Rate (WER): 36.42% on the WAXAL test split.

This means the model correctly transcribes roughly 2 in every 3 Shona words. For context:

  • This is a first fine-tune with no Zimbabwean broadcast-domain data
  • The WAXAL data covers diverse Shona speakers from multiple regions
  • Expected improvement with: more training steps, Zimbabwean broadcast audio, Common Voice data

Limitations

  • Trained exclusively on WAXAL speech β€” may struggle with strong regional Zimbabwean accents (Harare urban, rural dialects)
  • No broadcast/news domain data β€” radio and TV speech may score higher WER
  • 36.42% WER means roughly 1 in 3 words will need correction for formal use
  • Best suited for: voice memos, conversational transcription, accessibility tools, research

Roadmap

The next planned improvements (via Project Nyaradzai):

  1. v0.6 β€” Fine-tune on Zimbabwean broadcast audio (ZBC radio archives, community radio)
  2. Common Voice campaign β€” crowdsource Shona recordings across Zimbabwe for a locally-representative dataset
  3. whisper-medium-shona β€” larger model for higher accuracy once more data is available
  4. Live captioning pipeline β€” deploy for Zimbabwean broadcasters

Citation

If you use this model in your research or work, please cite:

@misc{whisper-small-shona-2026,
  title={Whisper-Small-Shona: First Open Shona ASR Model},
  author={Mateta, Stanley},
  year={2026},
  publisher={Hugging Face},
  url={https://huggingface.co/Starsm91/whisper-small-shona},
  note={Part of Project Nyaradzai: https://github.com/stanleymateta-tech/Project-Nyaradzai}
}

Data attribution

Training data: Google WAXAL (West African and eXtra-continental African Languages) dataset. License: CC-BY-4.0. Citation: WAXAL: A Multilingual Speech Dataset for African Languages, Google Research, 2026.

Project Nyaradzai

This model is part of a larger open-source project to bring chiShona and isiNdebele into the digital era:

  • πŸ”€ Spellchecker: 214,000+ word forms for Word, LibreOffice, Firefox
  • πŸ—£ ASR model: This model
  • ⌨️ Android keyboard: In development
  • πŸ“š Validated lexicon: Community-driven, in progress

GitHub: github.com/stanleymateta-tech/Project-Nyaradzai

Contributions welcome β€” especially from Shona and Ndebele native speakers. No coding required to contribute: report a word.


Mutauro wedu, panyika yose β€” Our language, everywhere in the world.

Downloads last month
58
Safetensors
Model size
0.2B params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for Starsm91/whisper-small-shona

Finetuned
(3793)
this model

Dataset used to train Starsm91/whisper-small-shona

Space using Starsm91/whisper-small-shona 1