Raon-OpenTTS 0.3B โ€” GGUF for CrispASR

GGUF conversion of KRAFTON/Raon-OpenTTS-0.3B โ€” an English F5-TTS DiT (flow-matching, zero-shot voice cloning) paired with a 16 kHz HiFi-GAN vocoder (speechbrain/tts-hifigan-libritts-16kHz lineage).

License: CC-BY-NC-4.0 โ€” non-commercial use only (upstream model license). Attribution: Raon-OpenTTS by KRAFTON. The slaney mel filterbank + Hann window are computed with torchaudio and shipped inside the GGUF.

Single self-contained GGUF: F5-TTS DiT (dim=1024, 22 layers) + HiFi-GAN vocoder + shipped mel filterbank/window + 5555-char vocab.

Usage (CrispASR)

crispasr --backend raon -m auto \
    --voice reference.wav --ref-text "transcript of the reference" \
    --tts "Text to synthesize." --tts-output out.wav --i-have-rights

-m auto downloads this GGUF (with the CC-BY-NC-4.0 acceptance notice).

Validation: the full CrispASR pipeline (DiT + sbhifigan mel + HiFi-GAN) passes a TTSโ†’ASR roundtrip โ€” synthesized speech transcribes back to the input text at 0.90 word overlap. The vocoder was verified in isolation against the reference (cosine 0.998). Converted with models/convert-raon-opentts-to-gguf.py.

Performance note: the HiFi-GAN vocoder currently runs on CPU (~40 s per short utterance); the DiT is fast on a GPU build. A ggml vocoder path is planned.

Downloads last month
372
GGUF
Model size
0.4B params
Architecture
f5-tts
Hardware compatibility
Log In to add your hardware

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for cstr/raon-opentts-0.3b-GGUF

Quantized
(1)
this model