Instructions to use rookes/cantocaptions-cantonese-asr with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use rookes/cantocaptions-cantonese-asr with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("automatic-speech-recognition", model="rookes/cantocaptions-cantonese-asr")# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("rookes/cantocaptions-cantonese-asr") model = AutoModelForMultimodalLM.from_pretrained("rookes/cantocaptions-cantonese-asr", device_map="auto") - Notebooks
- Google Colab
- Kaggle
CantoCaptions Cantonese ASR
- A fine-tuning of Qwen3-ASR on a subset of the CantoCaptions dataset
Details
This model is designed for high-quality written Cantonese transcription using the CantoCaptions standards, which mostly follow acceptable usages and variants documented in the words.hk dictionary. Sentence final particles (SFPs) such as 啦 laa1 / 喇 laa3 / 嗱 laa4 are separated out by tone as described on the CantoCaptions website.
The current model was trained using LoRa rank=128 and trained over a single epoch on ~75h of training audio, with an additional 4h dev and 4h test audio derived from the same dataset. At the time of writing, the model is in an early stage of development and may be updated.
- Downloads last month
- 1,652
Model tree for rookes/cantocaptions-cantonese-asr
Base model
Qwen/Qwen3-ASR-1.7B-hf