qwen3-asr-all_devices: Qwen3-ASR 1.7B packaged for the A14 generation (M1 Macs) too
#1
by takeargmax - opened
Adds qwen3-asr-all_devices/, the Qwen3-ASR 1.7B tree packaged to run on the
A14-generation Neural Engine (M1 Macs) as well as on newer chips. Argmax SDK
serves it as Qwen3ASRVariant.qwen3ASR_1_7B_AllDevices (argmax-sdk-swift#480);
the default variant keeps downloading qwen3-asr/, which this PR does not touch.
Same layout as qwen3-asr/ (audio_encoder/1.7b + text_decoder/1.7b, 44 files,
1.79 GB), copied file for file from argmaxinc/qwenasrkit-pro-internal/qwen3-asr-kvcombined
(Hub commit f12b353), where it was staged and verified against its sources:
audio_encoder/1.7b/*,Speculator.mlmodelc, the tokenizer files,decoder_manifest.json,LICENSE_NOTICE.txt: identical to the shippedqwen3-asr/files.TextDecoderC0.mlmodelc/TextDecoderC1.mlmodelc:argmaxinc/axon-coremlqwen3_asr/text_decoder_c{0,1}/1.7b-hf/W6A16-l1024-kvcombined-multifunctionβ
the same weights with one combined KV state per layer (14 per chunk instead of 28),
which is what fits the A14 generation's 26-state budget. Measured to load and run
on an M1 Mac mini and an iPhone 12 mini; no measured cost on newer chips.TextDecoderEmbedHead.mlmodelc:argmaxinc/axon-coremlqwen3_asr/text_decoder_embed_head/1.7b-hf/W6A16-multifunction. It has to be this
head: C0 bakes a tether term derived from it and the shipped head's row differs
(a mismatch decodes a lone EOS). In the SDK A/B, transcripts match the shipped
folder on jfk and LibriSpeech and differ by one token on three of five clips.
Existing SDK releases are unaffected: they download qwen3-asr/ through the glob*qwen3-asr/*, which does not match this folder.
EduardoPacheco changed pull request status to open
EduardoPacheco changed pull request status to merged
EduardoPacheco deleted the
refs/pr/1 ref