jan1k/Qwen3.8-27B-Uncensored-Genesis-NVFP4

🤗 Hugging Face 来源text-generationapache-2.0激活 27B31 GBGGUF✓ 2 个校验和今天更新
需要做种者 →

Qwen3.8-27B-Uncensored-Genesis — NVFP4 GGUF

NVFP4 quantisation of LuffyTheFox/Qwen3.8-27B-Uncensored-Genesis-GGUF, built with advanced-gguf-quantizer (a llama.cpp fork focused on NVFP4 quantization).

This is the Genesis-repaired version of HauhauCS/Qwen3.8-27B-Uncensored-HauhauCS-Aggressive — Luffy's tensor repair applied to the HauhauCS uncensored base.

Files

v4 — Recommended (inline scales, LM Studio compatible)

File Calibration MTP Size Tensors
Qwen3.8-27B-Uncensored-Genesis-NVFP4-v4.gguf none (data-free) yes (blk.64 preserved) 15.76 GB 866
Qwen3.8-27B-Uncensored-Genesis-NVFP4-v4-noMTP.gguf none (data-free) no (blk.64 stripped) 15.53 GB 851

v4 uses --nvfp4-inline-scales-only to emit native inline UE4M3 scales without separate .scale/.input_scale tensors. Required for LM Studio / Pelican and other runtimes that do not support the extended NVFP4 scale tensor contract.

Imatrix variants will follow after the plain v4 is confirmed working in target runtimes.

Source

Source GGUF Qwen3.8-27B-Uncensored-Genesis-Q8_K_P.gguf (31.5 GB)
Architecture qwen35 (dense), 64 layers + 1 MTP/NextN block
Layers 48 Gated DeltaNet + 16 gated-attention
general.file_type 39 (LLAMA_FTYPE_MOSTLY_NVFP4)
MTP/NextN qwen35.nextn_predict_layers=1 (native, preserved)

Single-step Q8_K_P → NVFP4 with per-tensor-type protections.

Tensor mix (v4)

type count notes
F32 360 norms, ssm scalars (a/dt/conv1d), nextn norms
F16 4 blk.0 attn_gate/attn_qkv, blk.0/blk.13 ffn_down
NVFP4 502 bulk weights incl. output.weight, token_embd
total 866
separate .scale/.input_scale 0 inline UE4M3 only

Tensor protection policy

F16 singular-collapse protection:

tensor type
blk.0.attn_gate.weight F16
blk.0.attn_qkv.weight F16
blk.0.ffn_down.weight F16
blk.13.ffn_down.weight F16

F32 architecture-specific protection:

  • blk.*.attn_norm.weight, blk.*.post_attention_norm.weight
  • blk.*.attn_q_norm.weight, blk.*.attn_k_norm.weight
  • blk.*.ssm_norm.weight, blk.*.nextn.*.norm.weight
  • output_norm.weight
  • blk.*.ssm_conv1d.weight, blk.*.ssm_dt.bias, blk.*.ssm_a

Forced NVFP4 (do not push lower):

tensor type
blk.0.ssm_out.weight NVFP4
blk.1.attn_gate.weight NVFP4
blk.1.attn_qkv.weight NVFP4

Usage

llama-cli -m Qwen3.8-27B-Uncensored-Genesis-NVFP4-v4.gguf \
  --jinja -c 131072 -ngl 99

For runtimes without FastMTP support, use Qwen3.8-27B-Uncensored-Genesis-NVFP4-v4-noMTP.gguf. The MTP variant enables speculative decoding (--spec-type draft-mtp where supported).

  • Set K cache and V cache quantization to F16
  • Vision support requires the mmproj file from the source repository

Hardware

  • Blackwell (RTX 50xx): native FP4 path, fastest
  • Ampere (RTX 30xx): NVFP4 inference works via fallback kernels
  • Quantisation was done CPU-only (Ampere CUDA NVFP4 encoder is unreliable)

Credits