jan1k/Swift-Qwen3.8-27B-Genesis-NVFP4-GGUF

🤗 Hugging Face sourcetext-generationapache-2.027B activated31 GBGGUF✓ 2 checksumsupdated today
Needs seeder →

Swift-Qwen3.8-27B-Genesis — NVFP4 GGUF

NVFP4 quantisation of LuffyTheFox/Swift-Qwen3.8-27B-Genesis-F16-GGUF, built with advanced-gguf-quantizer (a llama.cpp fork focused on NVFP4 quantization).

Base model: ukisai/Swift-Qwen3.8-27B-GGUF with Genesis tensor repair by LuffyTheFox.

Files

v4 — Recommended (inline scales, LM Studio compatible)

File Calibration MTP Size Tensors
Swift-Qwen3.8-27B-Genesis-NVFP4-v4.gguf none (data-free) yes (blk.64 preserved) 15.76 GB 866
Swift-Qwen3.8-27B-Genesis-NVFP4-v4-noMTP.gguf none (data-free) no (blk.64 stripped) 15.53 GB 851

v4 uses --nvfp4-inline-scales-only to emit native inline UE4M3 scales without separate .scale/.input_scale tensors. Required for LM Studio / Pelican and other runtimes that do not support the extended NVFP4 scale tensor contract.

Imatrix variants will follow after the plain v4 is confirmed working in target runtimes.

Source

Source GGUF Swift-Qwen3.8-27B-F16-0000{1..3}-of-00003.gguf (55.6 GB total, split)
Architecture qwen35 (dense), 64 layers + 1 MTP/NextN block
Layers 48 Gated DeltaNet + 16 gated-attention
general.file_type 39 (LLAMA_FTYPE_MOSTLY_NVFP4)
MTP/NextN qwen35.nextn_predict_layers=1 (native, preserved)

Single-step F16 → NVFP4 — the F16 source gives the NVFP4 encoder clean data with no intermediate quantization noise. llama-quantize reads the split source directly via ...-00001-of-00003.gguf.

Tensor mix (v4)

type count notes
F32 360 norms, ssm scalars (a/dt/conv1d), nextn norms
F16 4 blk.0 attn_gate/attn_qkv, blk.0/blk.13 ffn_down
NVFP4 502 bulk weights incl. output.weight, token_embd
total 866
separate .scale/.input_scale 0 inline UE4M3 only

Tensor protection policy

F16 singular-collapse protection:

tensor type
blk.0.attn_gate.weight F16
blk.0.attn_qkv.weight F16
blk.0.ffn_down.weight F16
blk.13.ffn_down.weight F16

F32 architecture-specific protection:

  • blk.*.attn_norm.weight, blk.*.post_attention_norm.weight
  • blk.*.attn_q_norm.weight, blk.*.attn_k_norm.weight
  • blk.*.ssm_norm.weight, blk.*.nextn.*.norm.weight
  • output_norm.weight
  • blk.*.ssm_conv1d.weight, blk.*.ssm_dt.bias, blk.*.ssm_a

Forced NVFP4 (do not push lower):

tensor type
blk.0.ssm_out.weight NVFP4
blk.1.attn_gate.weight NVFP4
blk.1.attn_qkv.weight NVFP4

Usage

llama-cli -m Swift-Qwen3.8-27B-Genesis-NVFP4-v4.gguf \
  --mmproj mmproj-Swift-Qwen3.8-27B-F16.gguf \
  --jinja -c 131072 -ngl 99

For the MTP variant, use Swift-Qwen3.8-27B-Genesis-NVFP4-v4-noMTP.gguf when the runtime lacks FastMTP support, or the MTP file for speculative decoding.

  • Set K cache and V cache quantization to F16
  • Vision support requires the mmproj-Swift-Qwen3.8-27B-F16.gguf from the source repository

Hardware

  • Blackwell (RTX 50xx): native FP4 path, fastest
  • Ampere (RTX 30xx): NVFP4 inference works via fallback kernels
  • Quantisation was done CPU-only (Ampere CUDA NVFP4 encoder is unreliable)

Reproducibility

# F16 split source -> NVFP4 v4 (inline scales only, CPU-only)
CUDA_VISIBLE_DEVICES="" llama-quantize \
  --allow-requantize --mode fast --nvfp4-inline-scales-only \
  --tensor-type '.*=nvfp4' \
  --tensor-type '^blk\..*\.attn_norm\.weight$=f32' \
  --tensor-type '^blk\..*\.post_attention_norm\.weight$=f32' \
  --tensor-type '^blk\..*\.attn_q_norm\.weight$=f32' \
  --tensor-type '^blk\..*\.attn_k_norm\.weight$=f32' \
  --tensor-type '^blk\..*\.ssm_norm\.weight$=f32' \
  --tensor-type '^blk\..*\.nextn\..*norm\.weight$=f32' \
  --tensor-type '^output_norm\.weight$=f32' \
  --tensor-type '^blk\..*\.ssm_conv1d\.weight$=f32' \
  --tensor-type '^blk\..*\.ssm_dt\.bias$=f32' \
  --tensor-type '^blk\..*\.ssm_a$=f32' \
  --tensor-type '^blk.0.ssm_out.weight$=nvfp4' \
  --tensor-type '^blk.1.attn_gate.weight$=nvfp4' \
  --tensor-type '^blk.1.attn_qkv.weight$=nvfp4' \
  --tensor-type '^blk.0.attn_gate.weight$=f16' \
  --tensor-type '^blk.0.attn_qkv.weight$=f16' \
  --tensor-type '^blk.0.ffn_down.weight$=f16' \
  --tensor-type '^blk.13.ffn_down.weight$=f16' \
  Swift-Qwen3.8-27B-F16-00001-of-00003.gguf \
  Swift-Qwen3.8-27B-Genesis-NVFP4-v4.gguf NVFP4 6

Credits