Qwen3.8-Flash-Next-Uncensored-NVFP4-FP8PLE
NVFP4 abliterated build of Qwen/Qwen3.8-Flash-Next in which the PLE n-gram embedding table is quantized to FP8 e4m3 (single global scale) instead of bf16. The result is byte-for-byte the same model as orcarouter/Qwen3.8-Flash-Next-Uncensored-NVFP4 with one structural change: the 102.4 GB bf16 PLE shard is replaced by 10 FP8 shards, shrinking the checkpoint from 170.93 GiB to 123.25 GiB on disk. The PLE layout is deliberately matched to RadixArk/Qwen3.8-Flash-Next-NVFP4 so the checkpoint slots into the MiaAI-Lab Dual-DGX-Sparks recipe unchanged.
Provenance chain:
Qwen/Qwen3.8-Flash-Next (BF16)
└─ orcarouter/Qwen3.8-Flash-Next-Uncensored-NVFP4 (abliteration + NVFP4 experts / FP8 attention; PLE kept bf16; 170.93 GiB)
└─ this repo (PLE n-gram table → FP8 e4m3, global scale; 123.25 GiB)
What changed vs orcarouter (everything else identical)
| Item | orcarouter upstream | This repo |
|---|---|---|
PLE n-gram table (128 ngram_embedding.shard_N.weight tensors, [2500012, 160] each) |
bf16, in model-00002-of-00017.safetensors (102.4 GB) |
FP8 e4m3, in model-plefp8-00000…00009.safetensors (51.2 GB) |
model-00002-of-00017.safetensors |
shipped | removed |
model-plefp8-00000…00009.safetensors (10 files, 13 tensors each, last has 12) |
n/a | added |
model.safetensors.index.json |
130 n-gram keys, 223,046 total | 131 n-gram keys (128 shards + ngram_heads_offsets + ngram_heads_vocab_sizes + 1 global weight_scale BF16 [1] in model-plefp8-00009; the two small metadata tensors stay bf16 in model-00003), 223,047 total; total_size updated to 132,279,699,754 bytes |
config.json → text_config |
no ple_embedding_dtype field |
ple_embedding_dtype: "float8_e4m3fn" added |
| Checkpoint size on disk | 170.93 GiB | 123.25 GiB (−47.7 GiB) |
| All other 37 files | — | byte-identical (same hash-verified content as the serving copy) |
Note: the two small PLE metadata tensors (ngram_heads_offsets, ngram_heads_vocab_sizes) stay bf16/int in model-00003-of-00017.safetensors, exactly as upstream.
Quantization details: single global scale per table, scale = amax/448 rounded to bf16 (global amax = 0.0893555, scale = 0.0001993179321) — the same law RadixArk published for its FP8 PLE, and this value matches their shipped scale to the displayed digit. Worst-case measured single-value dequantization error: 0.00317 in bf16 units. The table was quantized from orca's own abliterated bf16 PLE table (streamed 1:1, no resampling), not transplanted from any other checkpoint; see Similarity to RadixArk.
Similarity to RadixArk/Qwen3.8-Flash-Next-NVFP4
This checkpoint reproduces RadixArk's FP8-PLE serialization so the two are drop-in interchangeable at the layout level:
| Aspect | RadixArk | This repo |
|---|---|---|
| PLE file naming | model-plefp8-00000…00009.safetensors |
identical |
| Tensors per PLE file | 13 (last: 12) | identical |
| Shard key naming | …ple.ple_embedding.ngram_embedding.shard_N.weight (single-level) |
identical |
| Global scale | 1× BF16 weight_scale in model-plefp8-00009 |
identical |
ple_embedding_dtype |
float8_e4m3fn |
identical |
| Global-scale law | amax/448, rounded to bf16 |
identical |
| PLE data on disk | 47.68 GiB (128 shards + scale) | 47.63 GiB (51.2 GB) — same 13/13/…/12-tensor shard split, file sizes identical to within ~5 MB |
What is NOT the same: the PLE weights. RadixArk quantized from the PLE tables shipped in the official Qwen/Qwen3.8-Flash-Next-FP8 revision; this repo quantizes from orca's abliterated bf16 table, and independent byte-comparison shows the PLE value_proj matrices differ (~44% of bytes) between the two sources. Do not mix PLE files across these two checkpoints: the loader layout is compatible, the data is not. Architecture-level config (48 layers, 36 Gated-DeltaNet + 12 full-attention at full_attention_interval: 4, 512 experts, MTP) is the same family in both.
config.json layer_types uses full_attention ×12 / linear_attention ×36 (no qwen_sparse_attention entries), matching RadixArk and required by the ModelConfig allowlist of the pinned vLLM image.
Compatibility with the dual-DGX-Spark recipe
Primary recipe: a fork of the MiaAI-Lab recipe with recipe-switching support, delivered as PR #32 ("Uncensored mode with custom-quant PLE from orcarouter"). MODEL_ID points at this repo by default in the fork's .env.sample; one config line switches the served build:
Modified recipe available here: Forked Recipe
cp .env.sample .env # MODEL_ID is already set to this repo
# switch to RadixArk or orca's bf16-PLE build by changing the single MODEL_ID line
./start.sh
The fork's start.sh auto-derives PLE_QUANT_OVERRIDE from the model's config.json (ple_embedding_dtype: float8_e4m3fn → fp8), so no other line needs to change. This repo is public, so no HF_TOKEN is needed to pull it (only the orca bf16-PLE build is gated).
Running the original MiaAI-Lab recipe instead works the same: point its MODEL_ID at this repo and keep PLE_QUANT_OVERRIDE=fp8 (the recipe hardcodes it, which is what this checkpoint needs).
Details (both recipes):
- Image:
vllm/vllm-openai:qwen38-flash-next(vLLMv0.1.dev20073+g8e685d198) — the recipe's pinned day-0 image. - PLE patch: both recipes bind-mount the resolver shim (
files/ple_layer_patched.py→vllm/models/qwen3_8_flash_next/nvidia/ple_layer.py). WithPLE_QUANT_OVERRIDE=fp8(the fork's auto-derived value for this checkpoint; the original recipe hardcodes it) the PLE embedding routes to the image'sQwen3_8FlashNextPLEFp8EmbeddingMethod, which handles "FP8 shards + one global weight_scale" — this layout. Without it, the engine tries to build a ~102 GB BF16 PLE and OOMs on boot. - Launch flags: the recipe's full flag set applies unchanged (
--tensor-parallel-size 2 --nnodes 2 --enable-expert-parallel --all2all-backend allgather_reducescatter --speculative-config '{"method":"mtp","num_speculative_tokens":3}' --gpu-memory-utilization 0.835 --max-num-seqs 8 --max-num-batched-tokens 8192 --kv-cache-dtype auto --load-format safetensors --safetensors-load-strategy lazy --reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder --compilation-config '{"mode":0,"cudagraph_mode":"FULL_DECODE_ONLY"}'plus the YaRN--hf-overridesfor 1M context andVLLM_ALLOW_LONG_MAX_MODEL_LEN=1). - Local-dir alternative: the fork also accepts an absolute-path
MODEL_IDfor a flat repo dir staged on both nodes (skips download + rsync, bind-mounts read-only). - KV cache: the recipe mandates
KV_CACHE_DTYPE=auto(bf16). Do not set FP8 KV: on sm_121 it crashes QSA prefill in this image (documented by the recipe and confirmed by multiple deployments).
End-user notes (read before deploying)
- Smaller PLE footprint, slightly tighter KV pool than RadixArk. The FP8 PLE halves the runtime embedding memory vs the bf16-PLE upstream (102.4 GB → 51.2 GB on disk). Against RadixArk specifically, this PLE carries ~3.5 GiB of more data (51.20 vs 47.68 GiB), so in an otherwise identical launch the KV pool lands slightly below the reference build's measured 2,481,424 tokens (35.11 GiB at
GPU_MEMORY_UTILIZATION=0.835, 1M YaRN, MTP3, as measured in the recipe repo). The recipe's own measured ladder for that configuration: 0.835 → 2.48M tokens; 0.85 → 2.57M;--kv-cache-memory=45502283776(42.38 GiB) → ≈3.0M. No 2.7M pool figure is documented in the recipe or in either upstream repo; treat any such number as unverified. Measure on your own boot and treat that log as ground truth:docker logs vllm-fn 2>&1 | grep -E "Available KV cache memory|GPU KV cache size" PLE_QUANT_OVERRIDE=fp8is mandatory on the pinned image (see above). Theconfig.jsonple_embedding_dtypefield is informational for this image's loader and runtime-neutral.- Boot gotchas already fixed in this checkpoint: (a)
layer_typescontains onlyfull_attention/linear_attention, which the pinned image's ModelConfig allowlist accepts; (b) n-gram shard keys are single-levelngram_embedding.shard_N.weight, which is what the image'sQwen3_8FlashNextNGramEmbeddingloader expects. If you regenerate the FP8 PLE yourself, keep both invariants. - Accuracy impact of FP8 PLE: bounded by one E4M3 ULP per element on the n-gram prefix-lookup table. The PLE table is not a refusal/behavior-carrying tensor in this model and abliterated behavior is unchanged relative to the bf16-PLE upstream; no eval deltas were observed in smoke testing, but treat any PLE-quantization-sensitive task (heavy prefix/n-gram retrieval) as something to spot-check for your use case.
- Hardware: Blackwell with FP4 tensor cores (DGX Spark GB10, B100/B200/GB200, RTX 50-series) for the NVFP4 experts;
qwen4_expruntime required (the pinned image above). - Check integrity on download: every file is listed in
SHA256SUMS.txt(39 entries); verify withsha256sum -c SHA256SUMS.txt.
Requirements
vllm/vllm-openai:qwen38-flash-next(or any build withqwen4_exp+ the PLE FP8 resolver shim).transformers>=5.16for config/tokenizer handling if loading outside vLLM.- ~126 GiB of local storage per node when using the dual-Spark recipe.
Upstream provenance & license
Derived from orcarouter/Qwen3.8-Flash-Next-Uncensored-NVFP4 (Apache 2.0), itself an abliteration + NVFP4 quantization of Qwen/Qwen3.8-Flash-Next. License: Apache 2.0 (inherited; see upstream for full text and terms). This model has had its safety alignment substantially removed; it is released strictly for legitimate research, interpretability, AI-safety and red-teaming work. You assume full responsibility for its use; add your own safety and moderation layers before any deployment. The recipe.yaml included here is the upstream ModelOpt quantization recipe, carried over for provenance.