Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-Final — NVFP4 GGUF
NVFP4 quantisation of LuffyTheFox/Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-Final-GGUF, built with advanced-gguf-quantizer (a llama.cpp fork focused on NVFP4 quantization).
- Original Model & Genesis Tensor Repair: LuffyTheFox
- Base Model: HauhauCS/Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive (0/465 refusals)
- Hermes Finetune: DJLougen/hermes-qwen3.5-35b-a3b-GGUF (transferred Hermes data onto the uncensored base)
- Architecture: Mixture of Experts (MoE) — 35B total / ~3B active per token (8 routed + 1 shared)
- Multimodal (Vision): Supported via
mmproj-Hermes3.6-35B-A3B-Uncensored-Genesis.gguf
Files
v4 — Recommended (inline scales, LM Studio compatible)
| File | Size | Tensors |
|---|---|---|
Hermes3.6-35B-A3B-Uncensored-Genesis-Final-NVFP4-v4.gguf |
~20 GB | 733 |
mmproj-Hermes3.6-35B-A3B-Uncensored-Genesis.gguf |
~899 MB | vision projector |
chat_template.jinja |
16 KB | chat template |
System_Prompt.txt |
6 KB | recommended system prompt |
System_Prompt_Agent.txt |
1 KB | agentic / tool-calling prompt |
System_Prompt_Creative.txt |
6 KB | creative prompt |
v4 uses --nvfp4-inline-scales-only to emit native inline UE4M3 scales
without separate .scale/.input_scale tensors. Required for LM Studio /
Pelican and other runtimes that do not support the extended NVFP4 scale
tensor contract.
Imatrix variants will follow after the plain v4 is confirmed working in target runtimes.
Source
| Source GGUF | Hermes3.6-35B-A3B-Uncensored-Genesis-Final-Q8_K_P.gguf (43.6 GB, 10.06 BPW) |
| Architecture | qwen35moe (MoE), 40 layers |
| Context | 262144 |
general.file_type |
39 (LLAMA_FTYPE_MOSTLY_NVFP4) |
| MTP/NextN | none in source |
Two-step pipeline: Q8_K_P → F16 intermediate → NVFP4. The F16 step gives the NVFP4 encoder clean data (no Q8 quantization noise). Quantised CPU-only (Ampere CUDA NVFP4 encoder hangs on MoE tensors).
Tensor mix (v4)
| type | count | notes |
|---|---|---|
| F32 | 331 | norms, ssm scalars, gate inputs |
| F16 | 175 | sensitive weights (blk.0 attn, ssm, ffn_down_exps) |
| Q6_K | 1 | output.weight |
| NVFP4 | 226 | bulk weights |
| total | 733 | |
separate .scale/.input_scale |
0 | inline UE4M3 only |
Tensor protection policy
F16 singular-collapse protection:
| tensor | type |
|---|---|
blk.0.attn_gate.weight |
F16 |
blk.0.attn_qkv.weight |
F16 |
blk.0.ffn_down_exps.weight |
F16 |
blk.13.ffn_down_exps.weight |
F16 |
F32 architecture-specific protection:
- all norm weights (attn/post_attention/q/k/ssm norms)
blk.*.ssm_conv1d.weight,blk.*.ssm_dt.bias,blk.*.ssm_ablk.*.ffn_gate_inp_shexp.weight(shared expert gate input)token_embd.weight→ F16
Forced NVFP4 (do not push lower):
| tensor | type |
|---|---|
blk.0.ssm_out.weight |
NVFP4 |
blk.1.attn_gate.weight |
NVFP4 |
blk.1.attn_qkv.weight |
NVFP4 |
Usage
llama-cli -m Hermes3.6-35B-A3B-Uncensored-Genesis-Final-NVFP4-v4.gguf \
--mmproj mmproj-Hermes3.6-35B-A3B-Uncensored-Genesis.gguf \
--jinja -c 131072 -ngl 99
- Set K cache and V cache quantization to F16
- GPU offload maximum, active experts 8
- Use
chat_template.jinjawith--jinja; system prompts included in-repo
Hardware
- Blackwell (RTX 50xx): native FP4 path, fastest
- Ampere (RTX 30xx): NVFP4 inference works via fallback kernels
Reproducibility
# Step 1: Q8_K_P -> F16 intermediate
llama-quantize --allow-requantize \
Hermes3.6-35B-A3B-Uncensored-Genesis-Final-Q8_K_P.gguf temp_f16.gguf F16 6
# Step 2: F16 -> NVFP4 v4 (inline scales only, CPU-only)
CUDA_VISIBLE_DEVICES="" llama-quantize \
--allow-requantize \
--nvfp4-inline-scales-only \
--tensor-type-file tensor_types_protection.txt \
temp_f16.gguf \
Hermes3.6-35B-A3B-Uncensored-Genesis-Final-NVFP4-v4.gguf \
NVFP4 6
Credits
- Genesis algorithm & repair: LuffyTheFox
- Base model: HauhauCS
- Hermes dataset: NousResearch / DJLougen
- NVFP4 quantisation: jan1k
- Quantiser: advanced-gguf-quantizer