ddalcu/Muse-Glimmer-30B-MLX-Serve-8bit

🤗 Hugging Face sourcetext-generationapache-2.029.8B params60 GBsafetensors✓ 9 checksumsupdated today
Submit in one command

Run it next to your model folder. It makes the torrent, checks your files against Hugging Face, and submits it. You just start seeding and paste your key from your account. It only reads your files and never changes them. Read the script first if you like.

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo ddalcu/Muse-Glimmer-30B-MLX-Serve-8bit ./model-folder
Needs a seeder →

Muse-Glimmer-30B, MLX 8-bit for mlx-serve

8-bit MLX conversion of meta-models/Muse-Glimmer-30B, built for mlx-serve — the native Zig MLX server for Apple Silicon (mlxserve.com).

Conversion

  • Affine 8-bit, group size 64, on every matmul-read weight: text attention (incl. the sigmoid output gate projection), MLPs, lm_head, and the vision tower.
  • Kept bf16: embed_tokens and the vision position embedding table (gather-read), the patch embedder (input dim not divisible by 64), all norms and biases.
  • Weights, config and tokenizer are otherwise verbatim from the base repo. Trunk 32.9 GB, plus the 2.7 GB drafter below — 35.7 GB total.

DFlash drafter included

drafter/ holds a DFlash block-drafter for this trunk, so speculative decoding works with no flags and no second download. mlx-serve probes that subdirectory when the model loads, which also means a hot model switch brings the drafter with it. --no-drafter opts out; an explicit --drafter <dir> still wins.

It is a 5-layer assistant (block_size 16, mask_token_id 201818, reading the trunk at layers 1/13/25/37/49), shipped pre-packed at 8-bit so the server serves it as-is. One assistant pass drafts a whole block and the trunk verifies it in a single forward.

Speed on mlx-serve

All numbers below are mlx-serve (mlxserve.com) on an M4 Max (128 GB), temperature 0, median of 3, taken from the server's own reported decode rate:

coding / agent workload mlx-serve + drafter mlx-serve, drafter off
write a class from scratch 41.4 tok/s 15.9 tok/s 2.61x
edit a file and re-emit it 51.7 tok/s 26.9 tok/s 1.92x
emit a tool call 42.1 tok/s 18.6 tok/s 2.26x

62-86% of drafted tokens accepted. Code is where this pays: indentation, closing brackets, repeated identifiers and schema keys are all predictable enough to draft several ahead. Free-form prose accepts closer to 35% and gains proportionally less.

The effective block is a hardware property, not a checkpoint constant: the verify width a machine can run depends on its qmm lanes, so on Apple silicon without the wide lane the server caps the block (it logs capped (no wide verify lane) at load) and uses the full 16 where the lane exists. Greedy output is byte-identical to decoding without the drafter — the drafter changes speed, never the tokens.

Serving with mlx-serve

Install and run — see mlxserve.com:

mlx-serve --model <this-repo> --serve

Text, tool calling and thinking are served. The model always reasons in a to=self channel; mlx-serve returns it as reasoning_content and parses the ATEM (<atem:invoke>) tool-call format natively, on both the OpenAI and Anthropic APIs. Vision weights are included in this conversion but image input is not served yet.

Needs about 38 GB of memory with the drafter resident.