qwen3-asr-all_devices: Qwen3-ASR 1.7B packaged for the A14 generation (M1 Macs) too

#1
by takeargmax - opened
Argmax org

Adds qwen3-asr-all_devices/, the Qwen3-ASR 1.7B tree packaged to run on the
A14-generation Neural Engine (M1 Macs) as well as on newer chips. Argmax SDK
serves it as Qwen3ASRVariant.qwen3ASR_1_7B_AllDevices (argmax-sdk-swift#480);
the default variant keeps downloading qwen3-asr/, which this PR does not touch.

Same layout as qwen3-asr/ (audio_encoder/1.7b + text_decoder/1.7b, 44 files,
1.79 GB), copied file for file from argmaxinc/qwenasrkit-pro-internal/qwen3-asr-kvcombined
(Hub commit f12b353), where it was staged and verified against its sources:

  • audio_encoder/1.7b/*, Speculator.mlmodelc, the tokenizer files,
    decoder_manifest.json, LICENSE_NOTICE.txt: identical to the shipped qwen3-asr/ files.
  • TextDecoderC0.mlmodelc / TextDecoderC1.mlmodelc: argmaxinc/axon-coreml
    qwen3_asr/text_decoder_c{0,1}/1.7b-hf/W6A16-l1024-kvcombined-multifunction β€”
    the same weights with one combined KV state per layer (14 per chunk instead of 28),
    which is what fits the A14 generation's 26-state budget. Measured to load and run
    on an M1 Mac mini and an iPhone 12 mini; no measured cost on newer chips.
  • TextDecoderEmbedHead.mlmodelc: argmaxinc/axon-coreml
    qwen3_asr/text_decoder_embed_head/1.7b-hf/W6A16-multifunction. It has to be this
    head: C0 bakes a tether term derived from it and the shipped head's row differs
    (a mismatch decodes a lone EOS). In the SDK A/B, transcripts match the shipped
    folder on jfk and LibriSpeech and differ by one token on three of five clips.

Existing SDK releases are unaffected: they download qwen3-asr/ through the glob
*qwen3-asr/*, which does not match this folder.

EduardoPacheco changed pull request status to open
EduardoPacheco changed pull request status to merged
EduardoPacheco deleted the refs/pr/1 ref

Sign up or log in to comment