beethogedeon/fongbe-speech
Viewer • Updated • 8.16k • 146 • 1
How to use whettenr/asr-fon-streaming-conformer-without-diacritics with speechbrain:
# interface not specified in config.json
from speechbrain.inference.ASR import StreamingASR
from speechbrain.utils.dynamic_chunk_training import DynChunkTrainConfig
asr_model = StreamingASR.from_hparams(
source="whettenr/asr-fon-streaming-conformer-without-diacritics",
savedir="pretrained_models/asr-fon-streaming-conformer-without-diacritics"
)
asr_model.transcribe_file(
"whettenr/asr-fon-without-diacritics/example.wav",
# select a chunk size of ~960ms with 4 chunks of left context
DynChunkTrainConfig(24, 8),
# disable torchaudio streaming to allow fetching from HuggingFace
# set this to True for your own files or streams to allow for streaming file decoding
use_torchaudio_streaming=False,
)
# expected output:
# huzuhuzu gɔngɔn ɖe ɖo dandan
~100M parameters, 12 layer conformer encoder, Transducer (LSTM) decoder
pretrained using BEST-RQ on 700 hours for 400k steps
finetuned with Transducer decoder loss on training sets of
# other citation coming soon
# dataset citation
@inproceedings{kponou25_interspeech,
title = {{Extending the Fongbe to French Speech Translation Corpus: resources, models and benchmark}},
author = {D. Fortuné Kponou and Salima Mdhaffar and Fréjus A. A. Laleye and Eugène C. Ezin and Yannick Estève},
year = {2025},
booktitle = {{Interspeech 2025}},
pages = {4533--4537},
doi = {10.21437/Interspeech.2025-1801},
issn = {2958-1796},
}