Qwen3.8-27B-Uncensored-Genesis — NVFP4 GGUF
NVFP4 quantisation of LuffyTheFox/Qwen3.8-27B-Uncensored-Genesis-GGUF, built with advanced-gguf-quantizer (a llama.cpp fork focused on NVFP4 quantization).
This is the Genesis-repaired version of HauhauCS/Qwen3.8-27B-Uncensored-HauhauCS-Aggressive — Luffy's tensor repair applied to the HauhauCS uncensored base.
Files
v4 — Recommended (inline scales, LM Studio compatible)
| File | Calibration | MTP | Size | Tensors |
|---|---|---|---|---|
Qwen3.8-27B-Uncensored-Genesis-NVFP4-v4.gguf |
none (data-free) | yes (blk.64 preserved) | 15.76 GB | 866 |
Qwen3.8-27B-Uncensored-Genesis-NVFP4-v4-noMTP.gguf |
none (data-free) | no (blk.64 stripped) | 15.53 GB | 851 |
v4 uses --nvfp4-inline-scales-only to emit native inline UE4M3 scales
without separate .scale/.input_scale tensors. Required for LM Studio /
Pelican and other runtimes that do not support the extended NVFP4 scale
tensor contract.
Imatrix variants will follow after the plain v4 is confirmed working in target runtimes.
Source
| Source GGUF | Qwen3.8-27B-Uncensored-Genesis-Q8_K_P.gguf (31.5 GB) |
| Architecture | qwen35 (dense), 64 layers + 1 MTP/NextN block |
| Layers | 48 Gated DeltaNet + 16 gated-attention |
general.file_type |
39 (LLAMA_FTYPE_MOSTLY_NVFP4) |
| MTP/NextN | qwen35.nextn_predict_layers=1 (native, preserved) |
Single-step Q8_K_P → NVFP4 with per-tensor-type protections.
Tensor mix (v4)
| type | count | notes |
|---|---|---|
| F32 | 360 | norms, ssm scalars (a/dt/conv1d), nextn norms |
| F16 | 4 | blk.0 attn_gate/attn_qkv, blk.0/blk.13 ffn_down |
| NVFP4 | 502 | bulk weights incl. output.weight, token_embd |
| total | 866 | |
separate .scale/.input_scale |
0 | inline UE4M3 only |
Tensor protection policy
F16 singular-collapse protection:
| tensor | type |
|---|---|
blk.0.attn_gate.weight |
F16 |
blk.0.attn_qkv.weight |
F16 |
blk.0.ffn_down.weight |
F16 |
blk.13.ffn_down.weight |
F16 |
F32 architecture-specific protection:
blk.*.attn_norm.weight,blk.*.post_attention_norm.weightblk.*.attn_q_norm.weight,blk.*.attn_k_norm.weightblk.*.ssm_norm.weight,blk.*.nextn.*.norm.weightoutput_norm.weightblk.*.ssm_conv1d.weight,blk.*.ssm_dt.bias,blk.*.ssm_a
Forced NVFP4 (do not push lower):
| tensor | type |
|---|---|
blk.0.ssm_out.weight |
NVFP4 |
blk.1.attn_gate.weight |
NVFP4 |
blk.1.attn_qkv.weight |
NVFP4 |
Usage
llama-cli -m Qwen3.8-27B-Uncensored-Genesis-NVFP4-v4.gguf \
--jinja -c 131072 -ngl 99
For runtimes without FastMTP support, use
Qwen3.8-27B-Uncensored-Genesis-NVFP4-v4-noMTP.gguf. The MTP variant enables
speculative decoding (--spec-type draft-mtp where supported).
- Set K cache and V cache quantization to F16
- Vision support requires the mmproj file from the source repository
Hardware
- Blackwell (RTX 50xx): native FP4 path, fastest
- Ampere (RTX 30xx): NVFP4 inference works via fallback kernels
- Quantisation was done CPU-only (Ampere CUDA NVFP4 encoder is unreliable)
Credits
- Base model: HauhauCS (uncensored fine-tune)
- Genesis tensor repair + GGUF: LuffyTheFox
- NVFP4 quantisation: jan1k
- Quantiser: advanced-gguf-quantizer