sakamakismile/Ling-3.0-flash-W4A4-NVFP4

🤗 On Hugging Facetext-generationmit74.6B params134 GBsafetensors✓ Checksum-verifiedupdated 0d ago
Magnet

Ling-3.0-flash — NVFP4 W4A4 (calibrated, MTP verified working)

The first W4A4 NVFP4 quantization of Ling-3.0-flash

(124B total / 5.1B active, hybrid KDA+MLA MoE) — weights and activations in NVFP4,

calibrated with 128 samples via llm-compressor.

| | |

|---|---|

| Size | 81.4 GB (from 255 GB BF16) |

| Format | compressed-tensors nvfp4-pack-quantized, W4A4 (E2M1 + FP8-E4M3 group-16 scales, static input global scales) |

| Measured | 110 tok/s single / 493 tok/s @ 8-way generation, ~7,000 tok/s prefill — 8×RTX PRO 2000 Blackwell (16 GB), TP=8, 32K context |

| Long context | 131,072 tokens verified (eager mode; prefill 5,900–8,000 tok/s) |

| MTP | Works: 93% acceptance (k=1). Off by default — see below |

| Sanity | 6/6 (math / logic / code / JP idiom / instruction-following / long-form JP); tool-calling 4/4 |

Recommended: fp8 KV cache — 131K context at full speed

With the bundled vllm_patch/triton_decode_attention.py

(1-line shared-memory fix for sm_120), --kv-cache-dtype fp8 works and everything fits at once:

| | context | KV budget | generation | prefill |

|---|---|---|---|---|

| fp8 KV + patch (recommended) | 131,072 | 458,752 tok | 112.6 / 482–493 tok/s (1 / 8-way) | 1.9k single / 6.5k @ 4-way |

| bf16 KV, 32K | 32,768 | ~187K tok | 110 / 493 tok/s | up to 7.6k single (4096 chunks) |

| bf16 KV, 131K (--enforce-eager) | 131,072 | ~187K tok | 16.5 / 31.7 tok/s | 5.9k–8.0k |

458K tokens of KV = three concurrent 131K streams, or eight at 57K each.

Quality battery under fp8 KV: 6/6 (math / logic / code / JP idiom / instruction-following /

long-form JP) — no measured degradation vs bf16 KV.

Single-stream prefill in the recommended config is chunking-bound (`--max-num-batched-tokens

2048`, the largest stable value at 131K on 16 GB); parallel prefill reaches ~6,500 tok/s.

Why context is cheap here but headroom is not. Only 7 of 42 layers are MLA (the other

35 are KDA linear attention and carry no KV), so the KV cost is 8.2 KB/token — roughly

9× cheaper per token than an all-MLA model like DeepSeek-V3. The constraint is the card:

weights + non-torch overhead take 12.66 GB of a 16 GB GPU, leaving ~1.5 GB for KV and

KDA state — a total budget of about 187K tokens to split between length and concurrency

(131K × 1 stream, or 23K × 8 streams, etc.). On 80 GB cards this limit essentially vanishes.

How this differs from olka-fi/Ling-3.0-flash-NVFP4

(weight-only W4A16): this build quantizes activations too (W4A4) with real calibration data,

and ships a working MTP path.

Quantization boundary

Same proven boundary as the W4A16 build: routed experts only (layers 2–41,

mlp.experts.*.{gate,up,down}_proj, 120.8B of 127.4B params) → NVFP4 W4A4.

Kept in BF16: shared experts, KDA & gated-MLA attention, router gate + expert_bias,

dense layers 0–1, the MTP layer 42, embeddings, lm_head, norms.

Serving (vLLM fork required)

Upstream vLLM does not support bailing_hybrid v3. Use the prebuilt image

olkafi/vllm-bailing-v3 (inclusionAI/vllm @ ling_3_0 + SwiGLU-clamp patch).

This exact command reproduces every number in this card (131K context, fp8 KV,

458,752-token KV budget, ~112 tok/s single-stream). Both bundled patch files are

mounted over the image's own copies — no rebuild needed, they are pure Python:

REPO=/path/to/this-repo   # e.g. $(huggingface-cli download sakamakismile/Ling-3.0-flash-W4A4-NVFP4)

docker run --gpus all -d --name ling3 --ipc=host -p 8000:8000 \
  -e NCCL_P2P_DISABLE=1 -e NCCL_CUMEM_ENABLE=0 \
  -v $REPO:/models/ling3:ro \
  -v $REPO/vllm_patch/unquantized.py:/opt/vllm-ling3/vllm/model_executor/layers/fused_moe/oracle/unquantized.py:ro \
  -v $REPO/vllm_patch/triton_decode_attention.py:/opt/vllm-ling3/vllm/v1/attention/ops/triton_decode_attention.py:ro \
  olkafi/vllm-bailing-v3:latest /models/ling3 \
  --served-model-name ling3 --trust-remote-code --host 0.0.0.0 --port 8000 \
  --tensor-parallel-size 8 --disable-custom-all-reduce \
  --kernel-config '{"moe_backend":"marlin"}' --kv-cache-dtype fp8 \
  --compilation-config '{"max_cudagraph_capture_size":16}' \
  --gpu-memory-utilization 0.95 --max-model-len 131072 --max-num-seqs 8 \
  --max-num-batched-tokens 2048 \
  --enable-prefix-caching --mamba-cache-mode align \
  --enable-auto-tool-choice --tool-call-parser ling3 --reasoning-parser ling3

Notes for reproducers:

  • --max-num-batched-tokens 2048: 4096 OOMs at 131K on 16 GB cards (chunked-prefill

activations); 2048 is stable.

  • On cards with more than 16 GB, drop the two patch mounts if you don't need them

(flashinfer_cutlass may still be broken on sm_120 — keep moe_backend=marlin there),

and raise --max-num-batched-tokens / --max-num-seqs freely.

  • NCCL flags are for machines without GPU P2P; harmless otherwise.

Sampling (from the base card): temperature=0.6, top_p=0.95, top_k=20, thinking enabled.

⚠️ Known landmines (sm_120 / consumer Blackwell)

1. moe_backend=marlin is mandatory. The auto-selected flashinfer_cutlass

backend silently corrupts output (endless !!!!) on sm_120 with EP, and plain

TP=8 hits NotImplementedError (per-rank intermediate 768/8=96 needs unsupported

padding). cutlass rejects EP. Marlin works and is the source of the numbers above.

2. KV headroom is tight — on 16 GB cards use --gpu-memory-utilization 0.95,

--max-num-seqs 8 and --compilation-config '{"max_cudagraph_capture_size":16}'.

32K fits comfortably that way; 131K needs --enforce-eager (see table above).

3. --kv-cache-dtype fp8 needs the bundled kernel patch on sm_120. The TRITON_MLA

fp8 decode kernel asks for 102,400 bytes of shared memory at num_stages=2; consumer

Blackwell caps at 101,376 — 1 KB short. vllm_patch/triton_decode_attention.py

drops fp8-KV MLA to num_stages=1, which fits with no measured speed loss

(112.6 tok/s single-stream, same as bf16 KV).

4. The bundled vllm_patch/unquantized.py makes the

unquantized MTP-layer MoE fall back to triton when moe_backend=marlin is forced

globally — without it, serving with --speculative-config fails at startup.

MTP / speculative decoding: verified working, off by default

The W4A16 release reported 0% acceptance and shipped with MTP disabled. With this

build + the bundled patch, MTP works:

| config | single-stream | acceptance |

|---|---|---|

| no MTP (recommended) | 113.2 tok/s | — |

| MTP k=1 | 88.3 tok/s | 93.0% |

| MTP k=2 | 110.0 tok/s | 68%/tok |

The draft (BF16 MTP layer on the triton path) currently costs more than the accepted

tokens buy back on this hardware, so MTP is a proof-of-life, not a speed win — leave

--speculative-config off for throughput. To try it:

--speculative-config '{"method":"mtp","num_speculative_tokens":1}' (add

--compilation-config '{"max_cudagraph_capture_size":32}' --max-num-seqs 32 for memory).

Bake provenance

llm-compressor 0.12 (lna-lab/llmc-bake:v0.24.0), 128 calibration samples

(neuralmagic/calibration, chat-templated) @ 2048 tokens, basic pipeline with

CPU-resident weights and single-GPU onload calibration (fla/KDA kernels are

triton/cuda-only, so pure-CPU forward is impossible). Recipe quirks that will bite

reproducers: transformers 5.x needs three shims in the custom modeling file

(is_torch_fx_available, ROPE_INIT_FUNCTIONS['default'], _tied_weights_keys);

rope_scaling must be forced back to None after config normalization;

llm-compressor drops model_type from the saved config (breaking vLLM's MLA

detection — restore "model_type": "bailing_hybrid"); and post-calibration weight

observation must be moved to CPU (set_onload_device) or 61,440 expert observers

accumulate on one GPU and OOM.

Acknowledgements

Built on inclusionAI's Ling-3.0-flash (MIT). Quantization boundary and serving

groundwork follow olka-fi's pioneering W4A16 release and vendor-fork patches.

MIT license, inherited from the base model.