lychee888/Qwen3.8-Flash-Next-Uncensored-NVFP4-FP8PLE

🤗 Hugging Face 来源image-text-to-textapache-2.0180B 参数185 GBsafetensors✓ 29 个校验和今天更新
一条命令提交

在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo lychee888/Qwen3.8-Flash-Next-Uncensored-NVFP4-FP8PLE ./model-folder
需要做种者 →

Qwen3.8-Flash-Next-Uncensored-NVFP4-FP8PLE

NVFP4 abliterated build of Qwen/Qwen3.8-Flash-Next in which the PLE n-gram embedding table is quantized to FP8 e4m3 (single global scale) instead of bf16. The result is byte-for-byte the same model as orcarouter/Qwen3.8-Flash-Next-Uncensored-NVFP4 with one structural change: the 102.4 GB bf16 PLE shard is replaced by 10 FP8 shards, shrinking the checkpoint from 170.93 GiB to 123.25 GiB on disk. The PLE layout is deliberately matched to RadixArk/Qwen3.8-Flash-Next-NVFP4 so the checkpoint slots into the MiaAI-Lab Dual-DGX-Sparks recipe unchanged.

Provenance chain:

Qwen/Qwen3.8-Flash-Next (BF16)
  └─ orcarouter/Qwen3.8-Flash-Next-Uncensored-NVFP4   (abliteration + NVFP4 experts / FP8 attention; PLE kept bf16; 170.93 GiB)
       └─ this repo                                    (PLE n-gram table → FP8 e4m3, global scale; 123.25 GiB)

What changed vs orcarouter (everything else identical)

Item orcarouter upstream This repo
PLE n-gram table (128 ngram_embedding.shard_N.weight tensors, [2500012, 160] each) bf16, in model-00002-of-00017.safetensors (102.4 GB) FP8 e4m3, in model-plefp8-00000…00009.safetensors (51.2 GB)
model-00002-of-00017.safetensors shipped removed
model-plefp8-00000…00009.safetensors (10 files, 13 tensors each, last has 12) n/a added
model.safetensors.index.json 130 n-gram keys, 223,046 total 131 n-gram keys (128 shards + ngram_heads_offsets + ngram_heads_vocab_sizes + 1 global weight_scale BF16 [1] in model-plefp8-00009; the two small metadata tensors stay bf16 in model-00003), 223,047 total; total_size updated to 132,279,699,754 bytes
config.json → text_config no ple_embedding_dtype field ple_embedding_dtype: "float8_e4m3fn" added
Checkpoint size on disk 170.93 GiB 123.25 GiB (−47.7 GiB)
All other 37 files — byte-identical (same hash-verified content as the serving copy)

Note: the two small PLE metadata tensors (ngram_heads_offsets, ngram_heads_vocab_sizes) stay bf16/int in model-00003-of-00017.safetensors, exactly as upstream.

Quantization details: single global scale per table, scale = amax/448 rounded to bf16 (global amax = 0.0893555, scale = 0.0001993179321) — the same law RadixArk published for its FP8 PLE, and this value matches their shipped scale to the displayed digit. Worst-case measured single-value dequantization error: 0.00317 in bf16 units. The table was quantized from orca's own abliterated bf16 PLE table (streamed 1:1, no resampling), not transplanted from any other checkpoint; see Similarity to RadixArk.

Similarity to RadixArk/Qwen3.8-Flash-Next-NVFP4

This checkpoint reproduces RadixArk's FP8-PLE serialization so the two are drop-in interchangeable at the layout level:

Aspect RadixArk This repo
PLE file naming model-plefp8-00000…00009.safetensors identical
Tensors per PLE file 13 (last: 12) identical
Shard key naming …ple.ple_embedding.ngram_embedding.shard_N.weight (single-level) identical
Global scale 1× BF16 weight_scale in model-plefp8-00009 identical
ple_embedding_dtype float8_e4m3fn identical
Global-scale law amax/448, rounded to bf16 identical
PLE data on disk 47.68 GiB (128 shards + scale) 47.63 GiB (51.2 GB) — same 13/13/…/12-tensor shard split, file sizes identical to within ~5 MB

What is NOT the same: the PLE weights. RadixArk quantized from the PLE tables shipped in the official Qwen/Qwen3.8-Flash-Next-FP8 revision; this repo quantizes from orca's abliterated bf16 table, and independent byte-comparison shows the PLE value_proj matrices differ (~44% of bytes) between the two sources. Do not mix PLE files across these two checkpoints: the loader layout is compatible, the data is not. Architecture-level config (48 layers, 36 Gated-DeltaNet + 12 full-attention at full_attention_interval: 4, 512 experts, MTP) is the same family in both.

config.json layer_types uses full_attention ×12 / linear_attention ×36 (no qwen_sparse_attention entries), matching RadixArk and required by the ModelConfig allowlist of the pinned vLLM image.

Compatibility with the dual-DGX-Spark recipe

Primary recipe: a fork of the MiaAI-Lab recipe with recipe-switching support, delivered as PR #32 ("Uncensored mode with custom-quant PLE from orcarouter"). MODEL_ID points at this repo by default in the fork's .env.sample; one config line switches the served build: Modified recipe available here: Forked Recipe

cp .env.sample .env   # MODEL_ID is already set to this repo
# switch to RadixArk or orca's bf16-PLE build by changing the single MODEL_ID line
./start.sh

The fork's start.sh auto-derives PLE_QUANT_OVERRIDE from the model's config.json (ple_embedding_dtype: float8_e4m3fn → fp8), so no other line needs to change. This repo is public, so no HF_TOKEN is needed to pull it (only the orca bf16-PLE build is gated).

Running the original MiaAI-Lab recipe instead works the same: point its MODEL_ID at this repo and keep PLE_QUANT_OVERRIDE=fp8 (the recipe hardcodes it, which is what this checkpoint needs).

Details (both recipes):

  • Image: vllm/vllm-openai:qwen38-flash-next (vLLM v0.1.dev20073+g8e685d198) — the recipe's pinned day-0 image.
  • PLE patch: both recipes bind-mount the resolver shim (files/ple_layer_patched.py → vllm/models/qwen3_8_flash_next/nvidia/ple_layer.py). With PLE_QUANT_OVERRIDE=fp8 (the fork's auto-derived value for this checkpoint; the original recipe hardcodes it) the PLE embedding routes to the image's Qwen3_8FlashNextPLEFp8EmbeddingMethod, which handles "FP8 shards + one global weight_scale" — this layout. Without it, the engine tries to build a ~102 GB BF16 PLE and OOMs on boot.
  • Launch flags: the recipe's full flag set applies unchanged (--tensor-parallel-size 2 --nnodes 2 --enable-expert-parallel --all2all-backend allgather_reducescatter --speculative-config '{"method":"mtp","num_speculative_tokens":3}' --gpu-memory-utilization 0.835 --max-num-seqs 8 --max-num-batched-tokens 8192 --kv-cache-dtype auto --load-format safetensors --safetensors-load-strategy lazy --reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder --compilation-config '{"mode":0,"cudagraph_mode":"FULL_DECODE_ONLY"}' plus the YaRN --hf-overrides for 1M context and VLLM_ALLOW_LONG_MAX_MODEL_LEN=1).
  • Local-dir alternative: the fork also accepts an absolute-path MODEL_ID for a flat repo dir staged on both nodes (skips download + rsync, bind-mounts read-only).
  • KV cache: the recipe mandates KV_CACHE_DTYPE=auto (bf16). Do not set FP8 KV: on sm_121 it crashes QSA prefill in this image (documented by the recipe and confirmed by multiple deployments).

End-user notes (read before deploying)

  1. Smaller PLE footprint, slightly tighter KV pool than RadixArk. The FP8 PLE halves the runtime embedding memory vs the bf16-PLE upstream (102.4 GB → 51.2 GB on disk). Against RadixArk specifically, this PLE carries ~3.5 GiB of more data (51.20 vs 47.68 GiB), so in an otherwise identical launch the KV pool lands slightly below the reference build's measured 2,481,424 tokens (35.11 GiB at GPU_MEMORY_UTILIZATION=0.835, 1M YaRN, MTP3, as measured in the recipe repo). The recipe's own measured ladder for that configuration: 0.835 → 2.48M tokens; 0.85 → 2.57M; --kv-cache-memory=45502283776 (42.38 GiB) → ≈3.0M. No 2.7M pool figure is documented in the recipe or in either upstream repo; treat any such number as unverified. Measure on your own boot and treat that log as ground truth:
    docker logs vllm-fn 2>&1 | grep -E "Available KV cache memory|GPU KV cache size"
    
  2. PLE_QUANT_OVERRIDE=fp8 is mandatory on the pinned image (see above). The config.json ple_embedding_dtype field is informational for this image's loader and runtime-neutral.
  3. Boot gotchas already fixed in this checkpoint: (a) layer_types contains only full_attention/linear_attention, which the pinned image's ModelConfig allowlist accepts; (b) n-gram shard keys are single-level ngram_embedding.shard_N.weight, which is what the image's Qwen3_8FlashNextNGramEmbedding loader expects. If you regenerate the FP8 PLE yourself, keep both invariants.
  4. Accuracy impact of FP8 PLE: bounded by one E4M3 ULP per element on the n-gram prefix-lookup table. The PLE table is not a refusal/behavior-carrying tensor in this model and abliterated behavior is unchanged relative to the bf16-PLE upstream; no eval deltas were observed in smoke testing, but treat any PLE-quantization-sensitive task (heavy prefix/n-gram retrieval) as something to spot-check for your use case.
  5. Hardware: Blackwell with FP4 tensor cores (DGX Spark GB10, B100/B200/GB200, RTX 50-series) for the NVFP4 experts; qwen4_exp runtime required (the pinned image above).
  6. Check integrity on download: every file is listed in SHA256SUMS.txt (39 entries); verify with sha256sum -c SHA256SUMS.txt.

Requirements

  • vllm/vllm-openai:qwen38-flash-next (or any build with qwen4_exp + the PLE FP8 resolver shim).
  • transformers>=5.16 for config/tokenizer handling if loading outside vLLM.
  • ~126 GiB of local storage per node when using the dual-Spark recipe.

Upstream provenance & license

Derived from orcarouter/Qwen3.8-Flash-Next-Uncensored-NVFP4 (Apache 2.0), itself an abliteration + NVFP4 quantization of Qwen/Qwen3.8-Flash-Next. License: Apache 2.0 (inherited; see upstream for full text and terms). This model has had its safety alignment substantially removed; it is released strictly for legitimate research, interpretability, AI-safety and red-teaming work. You assume full responsibility for its use; add your own safety and moderation layers before any deployment. The recipe.yaml included here is the upstream ModelOpt quantization recipe, carried over for provenance.