ddalcu/Muse-Glimmer-30B-MLX-Serve-4bit

🤗 Hugging Face 来源text-generationapache-2.029.8B 参数60 GBsafetensors✓ 6 个校验和今天更新
一条命令提交

在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo ddalcu/Muse-Glimmer-30B-MLX-Serve-4bit ./model-folder
需要做种者 →

Muse-Glimmer-30B, MLX 4-bit for mlx-serve

4-bit MLX conversion of meta-models/Muse-Glimmer-30B, built for mlx-serve — the native Zig MLX server for Apple Silicon (mlxserve.com).

Half the memory of the 8-bit conversion, same features. Pick this one to fit a 32 GB machine; pick 8-bit if you have the headroom.

Conversion

  • Affine 4-bit, group size 64, on every matmul-read weight: text attention (incl. the sigmoid output gate projection), MLPs, lm_head, and the vision tower.
  • Kept bf16: embed_tokens and the vision position embedding table (gather-read), the patch embedder (input dim not divisible by 64), all norms and biases.
  • Weights, config and tokenizer are otherwise verbatim from the base repo. Trunk 18.7 GB, plus the 2.7 GB drafter below — 21.4 GB total.

DFlash drafter included

drafter/ holds a DFlash block-drafter for this trunk, so speculative decoding works with no flags and no second download. mlx-serve probes that subdirectory when the model loads, which also means a hot model switch brings the drafter with it. --no-drafter opts out; an explicit --drafter <dir> still wins.

It is a 5-layer assistant (block_size 16, mask_token_id 201818, reading the trunk at layers 1/13/25/37/49), shipped pre-packed at 8-bit so the server serves it as-is. One assistant pass drafts a whole block and the trunk verifies it in a single forward.

Speed on mlx-serve

All numbers below are mlx-serve (mlxserve.com) on an M4 Max (128 GB), temperature 0, median of 3, taken from the server's own reported decode rate:

coding / agent workload mlx-serve + drafter mlx-serve, drafter off
write a class from scratch 60.9 tok/s 28.3 tok/s 2.15x
edit a file and re-emit it 75.5 tok/s 54.3 tok/s 1.39x
emit a tool call 58.9 tok/s 31.6 tok/s 1.86x

62-84% of drafted tokens accepted. Code is where this pays: indentation, closing brackets, repeated identifiers and schema keys are all predictable enough to draft several ahead. Free-form prose accepts closer to 30% and gains proportionally less.

The effective block is a hardware property, not a checkpoint constant: the verify width a machine can run depends on its qmm lanes, so on Apple silicon without the wide lane the server caps the block (it logs capped (no wide verify lane) at load) and uses the full 16 where the lane exists. Greedy output is byte-identical to decoding without the drafter — the drafter changes speed, never the tokens.

Serving with mlx-serve

Install and run — see mlxserve.com:

mlx-serve --model <this-repo> --serve

Text, tool calling and thinking are served. The model always reasons in a to=self channel; mlx-serve returns it as reasoning_content and parses the ATEM (<atem:invoke>) tool-call format natively, on both the OpenAI and Anthropic APIs. Vision weights are included in this conversion but image input is not served yet.

Needs about 24 GB of memory with the drafter resident.