Muse-Glimmer-30B-Heretic-VL · NVFP4 + DFlash (MXFP8 draft) · with vision
Download one folder and it runs. NVFP4 main model, the vision tower's MLP quantized to NVFP4 as well, and three DFlash draft heads bundled in.
⚠️ This variant requires a small vLLM patch (shipped in
vllm_patch/). See "Why the vision tower must be compressed" below. If you only need text → useMuse-Glimmer-30B-Heretic-NVFP4-DFlash-MXFP8(no patch needed).
Measured (2×RTX 5070Ti 16GB · sm_120 · TP2 · vLLM 0.26.1)
| Metric | This variant (with vision) | Text-only variant |
|---|---|---|
| Decode speed (single stream, 7 prompts, mean) | 92.8 tok/s | 93.4 tok/s |
| KV pool | 503,761 tokens | 700,880 tokens |
| Concurrency at 131,072 context | 3.84x | 5.35x |
| DFlash accept_len (K=15) | 3.31 | 3.28 |
★ The KV pool number is meaningful here. A 16 GB card is a constrained system: whatever remains after weights, vision tower, draft head and activations is the KV pool, and this is already near the practical ceiling. It reflects what this hardware can do, not what I configured. (Contrast: the GB10 section below deliberately does not quote one — see why there.)
Carrying vision costs almost no decode speed (92.8 vs 93.4) — the text decode path never reads the vision weights. The cost shows up purely in the KV pool (the vision tower takes 0.93 GiB/card). Both variants clear 3.8x concurrency at 131,072 context.
Vision correctness check: given a synthetic image, the model correctly reported a red
rectangle outline on the left, a blue circle outline on the right, an off-white background and
the position of two text lines; it OCR'd CANARY-7492 correctly, and spontaneously pointed out
that the caption saying "red square" actually depicts a rectangle.
⚠️ Scope of these numbers — please read before quoting them
Everything on this page is a short, single-domain measurement. It is not a long-running or broad benchmark:
- 7 prompts per configuration, single stream, one run each. No repeated runs for variance, no multi-hour stability testing.
- One prompt domain only: programming / technical Q&A. Chinese prose, long-form reports, agent / tool-calling traces, multi-turn chat and very-long-context continuation are all not covered. Vision was verified for correctness, not benchmarked for throughput.
- This matters especially for speculative decoding: DFlash acceptance is essentially a function of how redundant the text is, and the same draft head can differ by more than 2× in accept_len across domains (in practice: tool-calling > code > Chinese reports > Chinese prose). The ranking between the three draft heads below only holds for this one domain.
- Concurrency figures are the theoretical value vLLM prints at startup (KV pool ÷ context), not throughput measured under real concurrent load.
If you are making a decision based on this, re-measure on your own prompt distribution.
★ Why the vision tower must be compressed (it's the entry ticket, not an optimization)
The vision tower is 3.58 GiB, so under TP2 you would expect 1.79 GiB per card. That is wrong.
MuseGlimmerVisionMLP is built from plain nn.Linear (muse_glimmer.py:650-651) — and
plain nn.Linear has no tensor-parallel logic, so every card holds a full copy:
| Component | Total | Actually resident per card (TP2) |
|---|---|---|
MLP c_fc/c_proj |
2.564 GiB (71.6%) | 2.564 — fully replicated |
| attn q/k/v/o | 0.879 GiB | 0.440 (QKVParallelLinear does shard) |
| adapter / projection / patch_embed / norms | 0.135 GiB | 0.135 — fully replicated |
| 3.58 GiB | ≈ 3.14 GiB/card |
On 2×16 GB that leaves only 0.12 GiB/card for KV — not even one 131k sequence fits.
After the patch, the MLP becomes ColumnParallelLinear/RowParallelLinear, which
buys quantizability and sharding at the same time: the vision tower drops from
3.14 to 0.93 GiB/card.
Text backbone quantization recipe
| Component | Precision | Size |
|---|---|---|
language_model mlp |
W4A16_NVFP4 · group 16 | 10.86 GiB |
language_model self_attn |
W4A16_NVFP4 · group 16 | 2.32 GiB |
lm_head |
MXFP8 · group 32 | 1.29 GiB |
embed_tokens |
BF16 | 2.51 GiB |
| norms | BF16 | ~0 |
How this recipe was decided
Not "quantize everything to 4-bit". Each component was decided separately, by one rule:
open vLLM's model file and check whether that layer is constructed with a quant_config.
If it isn't, it becomes a BF16 parameter — quantizing it guarantees a load failure.
| Component | Decision | Basis |
|---|---|---|
language_model mlp / self_attn |
W4A16_NVFP4 g16 | They are MergedColumnParallelLinear / QKVParallelLinear and do take a quant_config. ★Fused layers (gate+up, q+k+v) must all share one precision, otherwise the weight_scale_2 values disagree and the fused kernel breaks |
lm_head |
MXFP8 g32 | ParallelLMHead(..., quant_config=quant_config) does take one. It is read in full for every single token, so shrinking it directly buys speed |
embed_tokens |
BF16 | VocabParallelEmbedding(vocab, hidden) takes no quant_config; and decoding only reads a single row |
| norms | BF16 | Should not be quantized |
Why NVFP4 and not FP8 for the backbone: vLLM's ModelOptFp8LinearMethod is a W8A8 path.
It needs an input_scale, i.e. calibration. There is no "weight-only FP8 without calibration"
here — feed it uncalibrated FP8 and the model loads fine and then emits garbage. Skipping
calibration leaves you NVFP4 (weight-only) or BF16.
Vision quantization recipe
Vision tower spec: 50 layers · hidden 1536 · intermediate 8960 · 16 heads.
100 modules were quantized (50 layers × mlp.fc1 + mlp.fc2) with
W4A16_NVFP4 · group_size 16. Each module emits three tensors: weight (packed uint8),
weight_scale (FP8 E4M3, one per 16 elements) and weight_scale_2 (F32, global). Biases stay BF16.
| Component | Tensors | Precision | Size | Why |
|---|---|---|---|---|
mlp.fc1 / mlp.fc2 |
100 | W4A16_NVFP4 g16 | 2.564 → 0.722 GiB | Independent, never fused — clean to quantize. Also 71.6% of the tower, so the only block worth touching |
attn.q/k/v/proj |
200 | BF16 (untouched) | 0.879 GiB | ★q/k/v get fused into qkv_proj by vLLM, and NVFP4's weight_scale_2 is a per-module global scale — quantized separately the three scales disagree and fusion breaks. Compressing these requires MXFP8 instead (per-32-block E8M0 scales, no global scale → immune to fusion) |
norm1/norm2/ln_pre/ln_post |
204 | BF16 | ~0.003 GiB | LayerNorms should not be quantized |
patch_embedder / positional_embedding |
2 | BF16 | 0.006 GiB | Input side, too small to matter |
vision_adapter / vision_projection |
3 | BF16 | 0.129 GiB | Plain nn.Linear, no quant_config on the vLLM side — quantized weights would not load |
Result: vision shard 3.580 → 1.737 GiB; 3.14 → 0.93 GiB per card under TP2.
Quantization procedure (in case you want to reproduce it):
- Container-style manual export, bypassing
export_hf_checkpoint— it traces a forward pass, which weight-only quantization never needs. Instead: wrap each weight in an isolatednn.Linear, runmtq.quantize(..., forward_loop=None), then call_export_quantized_weight()per module. - ★Gate on a real scan before export: the number of modules actually quantized must equal
100, and so must the number of
uint8tensors — abort otherwise. When modelopt'squant_cfgis given as a list, later rules silently override earlier ones, so exclusions placed first get swallowed. Only counting the real result catches this. - Emit the new
model-vision-nvfp4.safetensorsand updatequantized_layersinconfig.json(★using vLLM-side names — see "Two easy traps" below).
Running it
⚠️ The image must have
muse_glimmersupport —vllm/vllm-openai:latestwill not start. Upstream vLLM has nomuse_glimmer.pyin main, v0.26.0, v0.26.1 or v0.27.0 (verified 2026-08-13; all four return 404).Use the public build on Docker Hub,
vllm/vllm-openai:muse-glimmer— vLLM0.26.1rc1.dev608+g99a10304d· 2026-08-11 · multi-arch (amd64 + arm64):vllm/vllm-openai@sha256:413c8fbecb1204a218117c77a4ea4b3a211d5686ff99c31d82ba6dd0cec8c5a6Every command on this page and
run_example.shpin that digest.
huggingface-cli download pottokao/Muse-Glimmer-30B-Heretic-VL-NVFP4-DFlash-MXFP8 \
--local-dir ./muse-vl
1. Prepare the patch (required)
vllm_patch/muse_glimmer.py is the result of 8 edits applied to vLLM 0.26.1's muse_glimmer.py.
If your vLLM version differs, do not use it as-is — regenerate it from your own image with
the patcher in the same directory:
docker run --rm -v $PWD:/host --entrypoint python3 <your vllm image> \
/host/muse-vl/vllm_patch/patch_muse_vision_quant.py /host/muse_glimmer_patched.py
The patcher asserts that each of its edits matches exactly once, and calls sys.exit
otherwise — so a version change fails loudly instead of silently patching the wrong place.
2. Start it
docker run -d --name muse-vl --gpus all --ipc host \
-e VLLM_ALLOW_LONG_MAX_MODEL_LEN=1 -e TORCH_CUDA_ARCH_LIST=12.0 \
-e PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True \
-v $PWD/muse-vl:/model \
-v $PWD/muse-vl/vllm_patch/muse_glimmer.py:/usr/local/lib/python3.12/dist-packages/vllm/model_executor/models/muse_glimmer.py:ro \
-p 8000:8000 --entrypoint vllm \
vllm/vllm-openai@sha256:413c8fbecb1204a218117c77a4ea4b3a211d5686ff99c31d82ba6dd0cec8c5a6 \
serve /model --tensor-parallel-size 2 --generation-config auto \
--quantization compressed-tensors --attention-backend TRITON_ATTN --kv-cache-dtype fp8 \
--max-model-len 131072 --gpu-memory-utilization 0.95 --kv-cache-memory 2300000000 \
--max-num-batched-tokens 2048 --max-num-seqs 4 --enable-chunked-prefill \
--chat-template /model/chat_template.jinja \
--reasoning-parser muse_glimmer --tool-call-parser muse_glimmer --enable-auto-tool-choice \
--speculative-config '{"method":"dflash","model":"/model/draft-mxfp8","num_speculative_tokens":15,"attention_backend":"TRITON_ATTN"}' \
--trust-remote-code
Three draft heads — pick one
DRAFT=draft-nvfp4 bash run_example.sh switches between them:
| Directory | Precision | Size | Per card (TP2) | Hit rate | accept_len | tok/s |
|---|---|---|---|---|---|---|
draft-mxfp8/ |
MXFP8 g32 | 2.76 GiB | 1.38 GiB | 15.5% | 3.32 | 31.2 |
draft-nvfp4/ |
W4A16_NVFP4 g16 | 1.80 GiB | 0.90 GiB | 10.5% | 2.58 | 25.6 |
draft-bf16/ |
BF16 (upstream) | 4.76 GiB | 2.38 GiB | 14.7% | 3.21 | 27.8 |
| (no speculative decoding) | — | — | — | — | 1.00 | 13.2 |
The two columns mean different things: hit rate = accepted ÷ proposed draft tokens (K=15);
accept_len = tokens emitted per round = 1 + hit_rate × K, and that is the one that scales
with speed.
tok/s was measured on a GB10 single card (same target, same K=15, same 7 prompts) and is only for ranking; absolute numbers for 2×5070Ti are at the top, and the last row is the same machine with speculative decoding off.
★ The BF16 draft head is not the best one — bigger and slower than MXFP8, with slightly
lower acceptance (3.21 vs 3.32).
★ This variant is already tighter on memory; switching to draft-bf16/ needs
--kv-cache-memory lowered by roughly 2×10⁹. See the text-only README for the full trade-off.
Contents
model-0000{1,2}-of-00002.safetensors 16.98 GiB main model (text)
model-vision-nvfp4.safetensors 1.74 GiB vision tower (MLP quantized to NVFP4)
model.safetensors.index.json 2,469 tensors / 3 files
config.json 517 quantization entries (100 of them vision)
draft-mxfp8/ 2.76 GiB DFlash draft head (default)
draft-nvfp4/ 1.80 GiB DFlash draft head (smallest)
draft-bf16/ 4.76 GiB DFlash draft head (upstream, unquantized)
vllm_patch/ patched file + the patcher that generates it
Dual-card configuration (2×16GB, as measured)
| Parameter | Value | Reason |
|---|---|---|
--tensor-parallel-size |
2 |
Weights, KV and vision attention shard; plain nn.Linear does not (that is what the patch fixes) |
--attention-backend |
TRITON_ATTN |
★On sm_120, FLASH_ATTN does not support FP8 KV (requires FA3 on SM90 or FA4 on SM100). The only way in |
--kv-cache-dtype |
fp8 |
Prerequisite for 128k on 16 GB cards |
--kv-cache-memory |
2300000000 |
★Explicit, and not filled to the brim — startup does not reserve for activations, so too large a value starts fine and OOMs on the first request |
--gpu-memory-utilization |
0.95 |
Used with the explicit KV value |
--max-num-seqs |
4 |
★KV pool is set by --kv-cache-memory; raising seqs is nearly free |
--max-num-batched-tokens |
2048 |
Pairs with chunked prefill |
num_speculative_tokens |
15 |
Raising K costs CUDA-graph memory only (K=4→15 ≈ +1.17 GiB), not linear slowdown — DFlash emits a whole block per forward. K=8 keeps ~95% of the benefit |
attention_backend (inside speculative-config) |
TRITON_ATTN |
★The drafter needs its own setting; flash_attn asserts at K≥16 |
TORCH_CUDA_ARCH_LIST |
12.0 |
RTX 50-series is sm_120 |
NCCL_P2P_DISABLE=1 |
optional | Avoids NCCL hangs on consumer dual-card setups without working P2P |
Single large-memory card (NVIDIA GB10 / DGX Spark, as measured)
The same folder on a GB10 (sm_121, 121.6 GiB unified memory) single card.
★A single card has no "MLP fully replicated" problem, but the patch is still required —
the weights are already packed NVFP4, and an unpatched nn.Linear cannot accept them
(shape error at load). Three changes:
docker run -d --name muse --gpus all --ipc host \
-e VLLM_ALLOW_LONG_MAX_MODEL_LEN=1 \
-e PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True \
-v $PWD/muse-vl/vllm_patch/muse_glimmer.py:/usr/local/lib/python3.12/dist-packages/vllm/model_executor/models/muse_glimmer.py:ro \
-v $PWD/muse-vl:/model -p 8000:8000 --entrypoint vllm \
vllm/vllm-openai@sha256:413c8fbecb1204a218117c77a4ea4b3a211d5686ff99c31d82ba6dd0cec8c5a6 \
serve /model --tensor-parallel-size 1 --generation-config auto \
--quantization compressed-tensors --attention-backend TRITON_ATTN --kv-cache-dtype fp8 \
--max-model-len 131072 --gpu-memory-utilization 0.75 \
--max-num-batched-tokens 4096 --max-num-seqs 16 --enable-chunked-prefill \
--chat-template /model/chat_template.jinja \
--reasoning-parser muse_glimmer --tool-call-parser muse_glimmer --enable-auto-tool-choice \
--speculative-config '{"method":"dflash","model":"/model/draft-mxfp8","num_speculative_tokens":15,"attention_backend":"TRITON_ATTN"}' \
--trust-remote-code
| Dual 16GB | GB10 single card | Reason | |
|---|---|---|---|
--tensor-parallel-size |
2 | 1 | Single card |
--kv-cache-memory |
explicit | omit, leave it to util 0.75 |
Memory is plentiful |
--max-num-seqs |
4 | 16 | Much larger KV pool |
--max-num-batched-tokens |
2048 | 4096 | Same |
TORCH_CUDA_ARCH_LIST |
12.0 |
unset | GB10 is sm_121, not sm_120 |
Measured on GB10 (figures below are for the text-only variant; the vision variant was not separately benchmarked on GB10, but the vision tower only adds 1.74 GiB of weights, which is negligible against 121 GiB):
| Metric | GB10 single card | 2×5070Ti |
|---|---|---|
| tok/s | 34.5 | 93.4 |
| accept_len | 3.26 | 3.28 |
| Weights + non-torch | 22.41 GiB | same weights, split across two cards |
| Peak activation | 2.87 GiB | — |
| CUDA graphs (K=15) | 1.4 GiB | — |
No KV pool figure is quoted here on purpose. On GB10 the KV size is entirely determined by
how much memory you hand it (util 0.75 above is just a test value), and such machines usually
host other services too — publishing a number would invite reading it as a property of the model.
Compute it yourself: subtract the three rows above (~26.7 GiB) and everything left can be KV;
this model needs roughly 9,140 bytes/token with FP8 KV.
⇒ The two setups are opposites: dual 5070Ti buys speed (2.7×); GB10 buys "make the context as large as you like". TRITON_ATTN is equally mandatory on GB10 (sm_121 is neither SM90 nor SM100).
★ Two easy traps
1. Vision cannot be switched off with has_vision: false alone.
AutoWeightsLoader(self) in muse_glimmer.py is constructed without skip_prefixes, so the
model.vision_tower.* tensors in the checkpoint find no home and it fails outright:
ValueError: There is no module or parameter named 'vision_encoder' in MuseGlimmerForCausalLM
Configuration can disable the compute path, not the weight loading. To actually save the memory, use the text-only repo — it keeps the vision tower in a dedicated shard that its index does not list, and since vLLM filters at file granularity that file is never even opened (the HF equivalent of GGUF's external mmproj).
2. quantized_layers keys use vLLM-side names, not checkpoint names.
ModelOptMixedPrecisionConfig.apply_vllm_mapper() runs the keys through the weight mapper, but
the mapper's substring rules require a dot on both sides (".mlp.fc1." → ".mlp.c_fc."),
and a module name has no trailing dot. Write checkpoint names and they will not be translated —
the layer silently falls back to Unquantized and then explodes on shape mismatch at load.
So the vision entries here are written as vision_encoder.transformer.N.mlp.c_fc, while the
tensor names keep their original checkpoint form (they end in .weight, which the rules do match).
Other pitfalls (why TRITON_ATTN is mandatory, the drafter backend, kv-cache-memory, cost of K)
are identical to the text-only variant — see that README.
Provenance and credits
- Base model:
meta-models/Muse-Glimmer-30B(Meta Superintelligence Lab, Apache 2.0) - DFlash draft head: the upstream Muse Glimmer DFlash assistant
(
draft-*/README_upstream.mdis the original upstream model card) - Abliteration source:
mlasli/Muse-Glimmer-30B-Heretic-Abliterated-BF16, pinned at commitdabdf7eb48a770e4fda6a4bfc485ea21fae5fc84— refusal-direction removal with Heretic over 500 Optuna trials, 6.5% refusals / KL 0.076. This repo only quantizes; it makes no capability changes. - Quantization: NVIDIA TensorRT Model Optimizer 0.45.0
Use of this model is subject to the upstream Usage Policy.