hancheolp commited on
Commit
119ea32
·
verified ·
1 Parent(s): 5c04c4e

add model card

Browse files
Files changed (1) hide show
  1. README.md +92 -0
README.md ADDED
@@ -0,0 +1,92 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ base_model: Qwen/Qwen3.8-27B
4
+ tags: [mlx, quantized, mixed-precision, 2bit, 4bit, auto-round]
5
+ pipeline_tag: image-text-to-text
6
+ ---
7
+
8
+ # Qwen3.8-27B MLX 4-bit with 100% of FFN at 2-bit
9
+
10
+ One rung of a five-model ladder built to measure **decode throughput vs. FFN bit-width**
11
+ on Apple Silicon. Accuracy was deliberately not tuned — these exist to answer one question:
12
+
13
+ > Is MLX's 2-bit `qmv` kernel as efficient as the 4-bit one?
14
+
15
+ ## This model
16
+
17
+ Base scheme is uniform **4-bit, group_size 64, affine**, matching
18
+ [`mlx-community/Qwen3.8-27B-4bit`](https://huggingface.co/mlx-community/Qwen3.8-27B-4bit)
19
+ (498 quantized modules, vision tower bf16, MTP dropped). On top of that, **192 of the
20
+ 192 FFN tensors are dropped to 2-bit** (still g64/affine). Attention, `lm_head` and
21
+ `embed_tokens` stay at 4-bit.
22
+
23
+ - Layers with all three FFN projections at 2-bit: **64 (L0-63)**
24
+ - Layers only partially converted: **0**
25
+
26
+ | | |
27
+ |---|---|
28
+ | 2-bit FFN tensors | 192 / 192 |
29
+ | On disk | 11.78 GB |
30
+ | Weights read per decoded token | 10.13 GB |
31
+ | Bandwidth-bound ceiling vs. baseline | **1.422x** |
32
+
33
+ **That last number is arithmetic, not a measurement.** It is
34
+ `baseline_bytes / this_model_bytes`, assuming batch-1 decode is purely memory-bandwidth
35
+ bound. `embed_tokens` is excluded from the read figure because decoding gathers a single
36
+ row rather than streaming the matrix. **No tok/s has been measured on any hardware.**
37
+ Real measurements, when they exist, belong below this line.
38
+
39
+ ## The full ladder
40
+
41
+ | model | 2-bit FFN tensors | disk | ceiling |
42
+ |---|---|---|---|
43
+ | [`test-4bit`](https://huggingface.co/hancheolp/test-4bit) | 0 / 192 | 16.05 GB | 1.000x |
44
+ | [`test-4bit25`](https://huggingface.co/hancheolp/test-4bit25) | 48 / 192 | 14.98 GB | 1.080x |
45
+ | [`test-4bit50`](https://huggingface.co/hancheolp/test-4bit50) | 96 / 192 | 13.92 GB | 1.174x |
46
+ | [`test-4bit75`](https://huggingface.co/hancheolp/test-4bit75) | 144 / 192 | 12.85 GB | 1.286x |
47
+ | [`test-4bit100`](https://huggingface.co/hancheolp/test-4bit100) | 192 / 192 | 11.78 GB | 1.422x |
48
+
49
+ Even at 100% FFN coverage the ceiling is 1.42x, and quantizing *everything* to 2-bit would
50
+ only reach 1.80x. The g64 metadata (fp16 scale + fp16 bias = 0.5 bpw) does not shrink with
51
+ bit-width, so 4-bit is really 4.5 bpw and 2-bit is 2.5 bpw.
52
+
53
+ ## Which tensors go to 2-bit
54
+
55
+ Selected by ascending KL sensitivity, using the per-tensor sweep published in
56
+ [`mlx-community/Qwen3.8-27B-OptiQ-4bit`](https://huggingface.co/mlx-community/Qwen3.8-27B-OptiQ-4bit)
57
+ (`optiq/sensitivity.json`). FFN sensitivity in this model falls monotonically with depth —
58
+ mean KL is 0.01584 for L0-7 and 0.00072 for L56-63, a 22x spread — so the least-sensitive
59
+ tensors all sit near the output. All three FFN projections hold the same parameter count,
60
+ so coverage alone fixes size and speed; the ranking only decides which tensors take the
61
+ damage.
62
+
63
+ ## Build
64
+
65
+ Quantized with [AutoRound](https://github.com/intel/auto-round) 0.15.0 in plain RTN mode
66
+ (`iters=0`, `disable_opt_rtn=True`, data-free), exported via `--format mlx`, then repaired.
67
+
68
+ The repair step is not optional. AutoRound's MLX exporter leaves 97 layers unquantized:
69
+
70
+ - `embed_tokens` — `SUPPORTED_LAYER_TYPES` is `(Linear, Conv2d, Conv1D)`; `nn.Embedding`
71
+ entries are dropped by the layer-config resolver.
72
+ - `linear_attn.in_proj_a` / `in_proj_b` (96 tensors, shape `[48, 5120]`) —
73
+ `_is_mlx_quantizable()` requires `out_dim % 64 == 0`. MLX imposes no such rule, and
74
+ `mlx-community/Qwen3.8-27B-4bit` quantizes all 96.
75
+
76
+ Those 97 are filled in afterwards with `mx.quantize` at 4-bit/g64, so the packing is
77
+ bit-exact MLX rather than a reimplementation of the affine formula. The exporter also omits
78
+ `"mode": "affine"` and emits ~57 stray `false` entries for vision layers; both are fixed.
79
+
80
+ ## Usage
81
+
82
+ ```python
83
+ from mlx_vlm import load, generate
84
+ model, processor = load("hancheolp/test-4bit100")
85
+ ```
86
+
87
+ ## Caveats
88
+
89
+ - Accuracy is unmeasured. This rung is not recommended for real use.
90
+ - Plain RTN, no calibration. Not representative of AutoRound's tuned modes.
91
+ - MTP is dropped, as in the mlx-community conversion. For speculative decoding see
92
+ [`mlx-community/Qwen3.8-27B-MTP-4bit`](https://huggingface.co/mlx-community/Qwen3.8-27B-MTP-4bit).