GLM-5.3-Flash MLX
Apple Silicon (MLX) builds of GLM-5.3-Flash with a validated, fixed glm5_next runtime: github.com/PipeNetwork/glm53-flash-mlx
Image-Text-to-Text • 314B • Updated • 3.55k • 4Note 334 GB. The anchor (bf16 does not fit a 512 GB Mac): ppl 3.4607 on 289k tokens of wikitext-2.
pipenetwork/GLM-5.3-Flash-MLX-6bit
Image-Text-to-Text • 314B • Updated • 2.18k • 3Note 256 GB. Indistinguishable from 8-bit: ppl 3.4646, paired ratio 0.9989 [0.9962, 1.0017], worse on 89/141 windows.
pipenetwork/GLM-5.3-Flash-MLX-mixed-4_8bit
Image-Text-to-Text • 314B • Updated • 3.01k • 5Note 182 GB. ppl 3.5705 (+3.2%): 4-bit experts, 8-bit for the ~9B non-expert weights. Fits a 256 GB Mac.
pipenetwork/GLM-5.3-Flash-MLX-4bit
Image-Text-to-Text • 314B • Updated • 3.62k • 2Note 178 GB. ppl 3.7549 (+8.5%): 4.4 GB smaller than mixed-4_8bit for 5 more points of perplexity.
pipenetwork/GLM-5.3-Flash-REAP25-MLX-mixed-4_8bit
Image-Text-to-Text • 238B • Updated • 434Note 139 GB. 216 of 288 experts per layer. ppl 4.2249 (+18% vs the unpruned mixed build).
pipenetwork/GLM-5.3-Flash-REAP37-MLX-mixed-4_8bit
Image-Text-to-Text • 201B • Updated • 418Note 118 GB, 128 GB Mac. 181 of 288 experts. ppl 4.8752 (+37%).
pipenetwork/GLM-5.3-Flash-REAP50-MLX-mixed-4_8bit
Image-Text-to-Text • 162B • Updated • 603Note 96 GB, 128 GB Mac with room. 144 of 288 experts. ppl 6.0757 (+70%) — coherent, but pruning costs this family a lot; use the unpruned builds when they fit.