jan1k/Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-Final-NVFP4

🤗 Hugging Face sourceimage-text-to-textapache-2.03B activated22 GBGGUF✓ 2 checksumsupdated today
Needs seeder →

Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-Final — NVFP4 GGUF

NVFP4 quantisation of LuffyTheFox/Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-Final-GGUF, built with advanced-gguf-quantizer (a llama.cpp fork focused on NVFP4 quantization).

Files

v4 — Recommended (inline scales, LM Studio compatible)

File Size Tensors
Hermes3.6-35B-A3B-Uncensored-Genesis-Final-NVFP4-v4.gguf ~20 GB 733
mmproj-Hermes3.6-35B-A3B-Uncensored-Genesis.gguf ~899 MB vision projector
chat_template.jinja 16 KB chat template
System_Prompt.txt 6 KB recommended system prompt
System_Prompt_Agent.txt 1 KB agentic / tool-calling prompt
System_Prompt_Creative.txt 6 KB creative prompt

v4 uses --nvfp4-inline-scales-only to emit native inline UE4M3 scales without separate .scale/.input_scale tensors. Required for LM Studio / Pelican and other runtimes that do not support the extended NVFP4 scale tensor contract.

Imatrix variants will follow after the plain v4 is confirmed working in target runtimes.

Source

Source GGUF Hermes3.6-35B-A3B-Uncensored-Genesis-Final-Q8_K_P.gguf (43.6 GB, 10.06 BPW)
Architecture qwen35moe (MoE), 40 layers
Context 262144
general.file_type 39 (LLAMA_FTYPE_MOSTLY_NVFP4)
MTP/NextN none in source

Two-step pipeline: Q8_K_P → F16 intermediate → NVFP4. The F16 step gives the NVFP4 encoder clean data (no Q8 quantization noise). Quantised CPU-only (Ampere CUDA NVFP4 encoder hangs on MoE tensors).

Tensor mix (v4)

type count notes
F32 331 norms, ssm scalars, gate inputs
F16 175 sensitive weights (blk.0 attn, ssm, ffn_down_exps)
Q6_K 1 output.weight
NVFP4 226 bulk weights
total 733
separate .scale/.input_scale 0 inline UE4M3 only

Tensor protection policy

F16 singular-collapse protection:

tensor type
blk.0.attn_gate.weight F16
blk.0.attn_qkv.weight F16
blk.0.ffn_down_exps.weight F16
blk.13.ffn_down_exps.weight F16

F32 architecture-specific protection:

  • all norm weights (attn/post_attention/q/k/ssm norms)
  • blk.*.ssm_conv1d.weight, blk.*.ssm_dt.bias, blk.*.ssm_a
  • blk.*.ffn_gate_inp_shexp.weight (shared expert gate input)
  • token_embd.weight → F16

Forced NVFP4 (do not push lower):

tensor type
blk.0.ssm_out.weight NVFP4
blk.1.attn_gate.weight NVFP4
blk.1.attn_qkv.weight NVFP4

Usage

llama-cli -m Hermes3.6-35B-A3B-Uncensored-Genesis-Final-NVFP4-v4.gguf \
  --mmproj mmproj-Hermes3.6-35B-A3B-Uncensored-Genesis.gguf \
  --jinja -c 131072 -ngl 99
  • Set K cache and V cache quantization to F16
  • GPU offload maximum, active experts 8
  • Use chat_template.jinja with --jinja; system prompts included in-repo

Hardware

  • Blackwell (RTX 50xx): native FP4 path, fastest
  • Ampere (RTX 30xx): NVFP4 inference works via fallback kernels

Reproducibility

# Step 1: Q8_K_P -> F16 intermediate
llama-quantize --allow-requantize \
  Hermes3.6-35B-A3B-Uncensored-Genesis-Final-Q8_K_P.gguf temp_f16.gguf F16 6

# Step 2: F16 -> NVFP4 v4 (inline scales only, CPU-only)
CUDA_VISIBLE_DEVICES="" llama-quantize \
  --allow-requantize \
  --nvfp4-inline-scales-only \
  --tensor-type-file tensor_types_protection.txt \
  temp_f16.gguf \
  Hermes3.6-35B-A3B-Uncensored-Genesis-Final-NVFP4-v4.gguf \
  NVFP4 6

Credits