sakamakismile/Ling-3.0-flash-W4A4-NVFP4

🤗 Hugging Face sourcetext-generationmit74.6B params134 GBsafetensors✓ 11 checksumsupdated today
Submit in one command

Run it next to your model folder. It makes the torrent, checks your files against Hugging Face, and submits it. You just start seeding and paste your key from your account. It only reads your files and never changes them. Read the script first if you like.

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo sakamakismile/Ling-3.0-flash-W4A4-NVFP4 ./model-folder
Needs a seeder →

Ling-3.0-flash — NVFP4 W4A4 (calibrated, MTP verified working)

The first W4A4 NVFP4 quantization of Ling-3.0-flash (124B total / 5.1B active, hybrid KDA+MLA MoE) — weights and activations in NVFP4, calibrated with 128 samples via llm-compressor.

Size 81.4 GB (from 255 GB BF16)
Format compressed-tensors nvfp4-pack-quantized, W4A4 (E2M1 + FP8-E4M3 group-16 scales, static input global scales)
Measured 110 tok/s single / 493 tok/s @ 8-way generation, ~7,000 tok/s prefill — 8×RTX PRO 2000 Blackwell (16 GB), TP=8, 32K context
Long context 131,072 tokens verified (eager mode; prefill 5,900–8,000 tok/s)
MTP Works: 93% acceptance (k=1). Off by default — see below
Sanity 6/6 (math / logic / code / JP idiom / instruction-following / long-form JP); tool-calling 4/4

Recommended: fp8 KV cache — 131K context at full speed

With the bundled vllm_patch/triton_decode_attention.py (1-line shared-memory fix for sm_120), --kv-cache-dtype fp8 works and everything fits at once:

context KV budget generation prefill
fp8 KV + patch (recommended) 131,072 458,752 tok 112.6 / 482–493 tok/s (1 / 8-way) 1.9k single / 6.5k @ 4-way
bf16 KV, 32K 32,768 ~187K tok 110 / 493 tok/s up to 7.6k single (4096 chunks)
bf16 KV, 131K (--enforce-eager) 131,072 ~187K tok 16.5 / 31.7 tok/s 5.9k–8.0k

458K tokens of KV = three concurrent 131K streams, or eight at 57K each. Quality battery under fp8 KV: 6/6 (math / logic / code / JP idiom / instruction-following / long-form JP) — no measured degradation vs bf16 KV. Single-stream prefill in the recommended config is chunking-bound (--max-num-batched-tokens 2048, the largest stable value at 131K on 16 GB); parallel prefill reaches ~6,500 tok/s.

Why context is cheap here but headroom is not. Only 7 of 42 layers are MLA (the other 35 are KDA linear attention and carry no KV), so the KV cost is 8.2 KB/token — roughly 9× cheaper per token than an all-MLA model like DeepSeek-V3. The constraint is the card: weights + non-torch overhead take 12.66 GB of a 16 GB GPU, leaving ~1.5 GB for KV and KDA state — a total budget of about 187K tokens to split between length and concurrency (131K × 1 stream, or 23K × 8 streams, etc.). On 80 GB cards this limit essentially vanishes.

How this differs from olka-fi/Ling-3.0-flash-NVFP4 (weight-only W4A16): this build quantizes activations too (W4A4) with real calibration data, and ships a working MTP path.

Quantization boundary

Same proven boundary as the W4A16 build: routed experts only (layers 2–41, mlp.experts.*.{gate,up,down}_proj, 120.8B of 127.4B params) → NVFP4 W4A4. Kept in BF16: shared experts, KDA & gated-MLA attention, router gate + expert_bias, dense layers 0–1, the MTP layer 42, embeddings, lm_head, norms.

Serving (vLLM fork required)

Upstream vLLM does not support bailing_hybrid v3. Use the prebuilt image olkafi/vllm-bailing-v3 (inclusionAI/vllm @ ling_3_0 + SwiGLU-clamp patch).

This exact command reproduces every number in this card (131K context, fp8 KV, 458,752-token KV budget, ~112 tok/s single-stream). Both bundled patch files are mounted over the image's own copies — no rebuild needed, they are pure Python:

REPO=/path/to/this-repo   # e.g. $(huggingface-cli download sakamakismile/Ling-3.0-flash-W4A4-NVFP4)

docker run --gpus all -d --name ling3 --ipc=host -p 8000:8000 \
  -e NCCL_P2P_DISABLE=1 -e NCCL_CUMEM_ENABLE=0 \
  -v $REPO:/models/ling3:ro \
  -v $REPO/vllm_patch/unquantized.py:/opt/vllm-ling3/vllm/model_executor/layers/fused_moe/oracle/unquantized.py:ro \
  -v $REPO/vllm_patch/triton_decode_attention.py:/opt/vllm-ling3/vllm/v1/attention/ops/triton_decode_attention.py:ro \
  olkafi/vllm-bailing-v3:latest /models/ling3 \
  --served-model-name ling3 --trust-remote-code --host 0.0.0.0 --port 8000 \
  --tensor-parallel-size 8 --disable-custom-all-reduce \
  --kernel-config '{"moe_backend":"marlin"}' --kv-cache-dtype fp8 \
  --compilation-config '{"max_cudagraph_capture_size":16}' \
  --gpu-memory-utilization 0.95 --max-model-len 131072 --max-num-seqs 8 \
  --max-num-batched-tokens 2048 \
  --enable-prefix-caching --mamba-cache-mode align \
  --enable-auto-tool-choice --tool-call-parser ling3 --reasoning-parser ling3

Notes for reproducers:

  • --max-num-batched-tokens 2048: 4096 OOMs at 131K on 16 GB cards (chunked-prefill activations); 2048 is stable.
  • On cards with more than 16 GB, drop the two patch mounts if you don't need them (flashinfer_cutlass may still be broken on sm_120 — keep moe_backend=marlin there), and raise --max-num-batched-tokens / --max-num-seqs freely.
  • NCCL flags are for machines without GPU P2P; harmless otherwise.

Sampling (from the base card): temperature=0.6, top_p=0.95, top_k=20, thinking enabled.

⚠️ Known landmines (sm_120 / consumer Blackwell)

  1. moe_backend=marlin is mandatory. The auto-selected flashinfer_cutlass backend silently corrupts output (endless !!!!) on sm_120 with EP, and plain TP=8 hits NotImplementedError (per-rank intermediate 768/8=96 needs unsupported padding). cutlass rejects EP. Marlin works and is the source of the numbers above.
  2. KV headroom is tight — on 16 GB cards use --gpu-memory-utilization 0.95, --max-num-seqs 8 and --compilation-config '{"max_cudagraph_capture_size":16}'. 32K fits comfortably that way; 131K needs --enforce-eager (see table above).
  3. --kv-cache-dtype fp8 needs the bundled kernel patch on sm_120. The TRITON_MLA fp8 decode kernel asks for 102,400 bytes of shared memory at num_stages=2; consumer Blackwell caps at 101,376 — 1 KB short. vllm_patch/triton_decode_attention.py drops fp8-KV MLA to num_stages=1, which fits with no measured speed loss (112.6 tok/s single-stream, same as bf16 KV).
  4. The bundled vllm_patch/unquantized.py makes the unquantized MTP-layer MoE fall back to triton when moe_backend=marlin is forced globally — without it, serving with --speculative-config fails at startup.

MTP / speculative decoding: verified working, off by default

The W4A16 release reported 0% acceptance and shipped with MTP disabled. With this build + the bundled patch, MTP works:

config single-stream acceptance
no MTP (recommended) 113.2 tok/s —
MTP k=1 88.3 tok/s 93.0%
MTP k=2 110.0 tok/s 68%/tok

The draft (BF16 MTP layer on the triton path) currently costs more than the accepted tokens buy back on this hardware, so MTP is a proof-of-life, not a speed win — leave --speculative-config off for throughput. To try it: --speculative-config '{"method":"mtp","num_speculative_tokens":1}' (add --compilation-config '{"max_cudagraph_capture_size":32}' --max-num-seqs 32 for memory).

Bake provenance

llm-compressor 0.12 (lna-lab/llmc-bake:v0.24.0), 128 calibration samples (neuralmagic/calibration, chat-templated) @ 2048 tokens, basic pipeline with CPU-resident weights and single-GPU onload calibration (fla/KDA kernels are triton/cuda-only, so pure-CPU forward is impossible). Recipe quirks that will bite reproducers: transformers 5.x needs three shims in the custom modeling file (is_torch_fx_available, ROPE_INIT_FUNCTIONS['default'], _tied_weights_keys); rope_scaling must be forced back to None after config normalization; llm-compressor drops model_type from the saved config (breaking vLLM's MLA detection — restore "model_type": "bailing_hybrid"); and post-calibration weight observation must be moved to CPU (set_onload_device) or 61,440 expert observers accumulate on one GPU and OOM.

Acknowledgements

Built on inclusionAI's Ling-3.0-flash (MIT). Quantization boundary and serving groundwork follow olka-fi's pioneering W4A16 release and vendor-fork patches.

MIT license, inherited from the base model.