esatapedico/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4-GGUF

🤗 Hugging Face sourcetext-generationapache-2.0138 GBGGUFHF checksums availableupdated today
No torrent yet

Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4-GGUF

HF repos: esatapedico/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4 (safetensors, compressed-tensors NVFP4) and esatapedico/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4-GGUF (this GGUF family, 8 tiers). Budget no-MTP variants (BUDGET 14.72, STARVED 14.59 GB, Q3_K/Q2_K heads, 1,107 tensors, block_count 64) live in the companion repo esatapedico/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4-BUDGET-GGUF.

A family of eight GGUF files of DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU, DavidAU's TURBO 735-882 tune (Fable plus Cold Fusion plus Heretic/Uncensored, 735/882 Heretic merge and DPO layers) on Qwen3.8-27B: a 27B dense hybrid model (Gated DeltaNet plus Gated Attention every fourth layer, 262K native context, embedded MTP speculative head, native vision tower). NVFP4 here means W4A16 with FP8 scales, group size 16, weight only. Linear layers are NVFP4, vision tower, linear attention path, lm_head, embeddings and MTP head are kept in BF16 at the quantization source. The MTP head is baked into every file in this repo, no separate drafter is needed (--spec-type draft-mtp).

My part here is only the numerics: I converted the NVFP4 checkpoint to GGUF and built a size and precision ladder for the tensors that most affect output quality and decode speed. All credit for the model itself belongs upstream (full chain below).

Follow along & support

I post updates on new conversions, benchmarks, and what I'm working on over on Ko-fi. Follow along there to keep up with new releases and the work in progress. If you'd like to support more of it, a coffee is always welcome. I do this on consumer hardware and like seeing how far it goes. More is on the way.

ko-fi.com/esatapedico. Updates, work-in-progress, and an optional coffee.

The eight files

Eight tiers in this repo share the MTP head (block_count 65, nextn_predict_layers 1, 1,122 tensors). Seven tiers share a byte-identical 448-tensor native NVFP4 backbone and differ only in lm_head / token_embd / MTP precision; HIGHEST is the fidelity-max tier that preserves more of the source precision (256 NVFP4 plus Q8_0 attention/SSM, BF16 token_embd and MTP) and is closest to ORIG:

File Size (decimal GB) lm_head (output.weight) token_embd MTP head (blk.64) Backbone
Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4-VERY-LOW.gguf 15.15 GB Q3_K Q2_K Q2_K NVFP4 448
Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4-COMPACT-LOW.gguf 15.16 GB Q4_K Q3_K Q2_K NVFP4 448
Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4-LOW.gguf 15.75 GB Q5_0 IQ4_XS IQ4_XS NVFP4 448
Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4-MEDIUM.gguf 16.60 GB Q8_0 Q6_K IQ4_XS NVFP4 448
Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4-MID-HIGH.gguf 16.91 GB Q8_0 Q8_0 Q8_0 NVFP4 448
Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4-HIGH.gguf 17.83 GB BF16 Q6_K IQ4_XS NVFP4 448
Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4-VERY-HIGH.gguf 19.33 GB BF16 BF16 BF16 NVFP4 448
Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4-HIGHEST.gguf 21.28 GB Q8_0 BF16 BF16 256 NVFP4 + Q8_0

Sizes are decimal GB as shown in the HF file browser. Picking a tier: MID-HIGH is the highest precision compact option among the 448-backbone tiers (all three head groups at Q8_0) and our expected fastest compact decode on dual GPU split; LOW and VERY-LOW trade some head precision for about 2 GB less VRAM; HIGH and VERY-HIGH restore BF16 heads where VRAM allows. VERY-HIGH needs about 18 GiB VRAM plus KV, so single 16 GB cards will fall back to CPU offload. HIGHEST is the fidelity-max tier (21.28 GB, 19.82 GiB reported, Q8_0 attention/SSM plus BF16 token_embd/MTP, closest to ORIG) and needs about 19.8 GiB VRAM plus KV; use dual-GPU split or 24 GB cards for full GPU.

Budget companions in the sibling repo: Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4-BUDGET.gguf (14.72 GB, Q3_K head, Q2_K emb) and Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4-STARVED.gguf (14.59 GB, Q2_K head and emb), both 1,107 tensors, block_count 64, no MTP, same 448 NVFP4 backbone with the draft block removed. Both fit a single 16 GB card.

Tensor layout

At a high level, the GGUFs in this repo contain 1,122 tensors. Seven tiers contribute a uniform 448-tensor native NVFP4 block shared byte identical across those seven tiers; HIGHEST contributes 256 NVFP4 plus Q8_0 attention/SSM extras (193 Q8_0, 9 BF16) and is not byte-identical to the 448 block but is the closest to ORIG. The NVFP4 format is W4A16 with group size 16 and FP8 E4M3 scales. Remaining tensors are F32 norms and scales, with lm_head, token embeddings and MTP head varying per tier as in the table above. Budget no-MTP variants in the companion repo have 1,107 tensors (block_count 64, nextn_predict_layers 0) and share the same 448-tensor NVFP4 backbone as the seven MTP tiers with the MTP block removed.

Field Value
Tensors per file 1,122
NVFP4 per tier 448 (HIGHEST 256)
qwen35.block_count 65 (64 plus 1 MTP blk.64, 15 tensors)
qwen35.context_length 262144
qwen35.nextn_predict_layers 1
tokenizer.chat_template 8953 chars, intact
general.file_type advisory only, per tensor type is authoritative
general.quantization_version 2
general.name Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4-<TIER>
general.description NVFP4 <TIER> backbone. Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU.
general.license apache-2.0 in every file

Vision

The TURBO 735-882 tune leaves the original Qwen3.8 vision tower untouched. Pair any tier with the matching vision projector mmproj-BF16.gguf via --mmproj. No mmproj is bundled in this repo, use the projector from the base model.

How this was made

At a high level, the steps were:

  1. Converted the BF16 base model to compressed tensors NVFP4 (W4A16, vision, linear attention, lm_head and MTP kept in BF16, no calibration).
  2. Verified NVFP4 metadata and ran a short vLLM smoke check.
  3. Converted the NVFP4 checkpoint to GGUF.
  4. Built each tier over a shared 448 tensor NVFP4 backbone, varying only lm_head, token embedding and MTP head precision per tier.
  5. Verified per tier NVFP4 tensor count and backbone byte identity across all tiers, and patched GGUF KV for name, description and license without changing tensor data.

Benchmarks (naive, single run, not comparable across setups)

These are rough sanity checks to confirm the files load and generate, not a formal benchmark. Method and hardware are noted so you can interpret them in context. Results will vary with hardware, sampling and context length.

All runs used direct GPU via llama.cpp tools, with a diverse synthetic English payload (about 75 kB, 28 chunks times 512 context). LocalAI gateway numbers are end to end and include gateway overhead, don't compare them to llama-bench decode.

Smoke: 8 of 8 coherent

Prompt What is a black hole? via single turn sampling, each tier produces a natural coherent completion with no repetition or truncation:

Tier Result
VERY-LOW ok, coherent, 710 chars
COMPACT-LOW ok, coherent, 892 chars
LOW ok, coherent, 981 chars
MEDIUM ok, coherent, 838 chars
MID-HIGH ok, coherent, 978 chars
HIGH ok, coherent, 829 chars
VERY-HIGH ok, coherent, 1.1k chars
HIGHEST ok, coherent, 1.0k chars

Budget companions also smoke 2 of 2 pass (BUDGET and STARVED coherent on the same prompt).

Perplexity (PPL)

Direct llama-perplexity on the 75 kB diverse payload, single GPU, same chunks for all tiers:

Tier PPL (Final estimate)
VERY-LOW 3.3410 +/- 0.07288
COMPACT-LOW 3.2808 +/- 0.07074
LOW 3.2761 +/- 0.07056
MEDIUM 3.2858 +/- 0.07109
MID-HIGH 3.2903 +/- 0.07107
HIGH 3.2851 +/- 0.07107
VERY-HIGH 3.2835 +/- 0.07089
HIGHEST 3.2367 +/- 0.07020

HIGHEST is best of the whole family at 3.2367, about 0.10 better than VERY-LOW and about 0.04 better than LOW. The six 448-backbone tiers span only 0.065 from best to worst, so the quantization costs almost nothing even at the smallest tier. VERY-LOW is highest as expected (Q3_K and Q2_K heads least precise). For reference, the sibling BUDGET repo measures BUDGET 3.3410 (identical to VERY-LOW, both Q3_K/Q2_K heads) and STARVED 3.4560 (+0.11, Q2_K everywhere).

Speed (llama-bench, pp512 prompt processing, tg128 generation, tok/s)

Single 5070 Ti (16 GB class) with full CUDA offload, three runs per tier:

Tier VRAM (llama-bench) pp512 tok/s tg128 tok/s Notes
VERY-LOW 14.10 GiB 2748.05 +/- 263.70 47.30 +/- 0.04 fits, full GPU
COMPACT-LOW 14.11 GiB 2437.61 +/- 343.88 47.49 +/- 0.08 fits, full GPU
LOW 14.66 GiB 2728.03 +/- 264.88 46.77 +/- 0.08 fits
MEDIUM 15.45 GiB 2704.70 +/- 300.87 45.55 +/- 0.04 fits
MID-HIGH 15.74 GiB 2714.21 +/- 277.94 45.55 +/- 0.10 fits
HIGH 16.59 GiB 34.89 +/- 0.05 6.32 +/- 0.01 exceeds single 16 GB, CPU fallback, about 30 times slower on this bench
VERY-HIGH 17.99 GiB 34.44 +/- 0.05 6.12 +/- 0.12 exceeds, CPU fallback
HIGHEST 19.81 GiB 35.40 +/- 0.03 4.73 +/- 0.00 exceeds single 16.3 GiB, CPU fallback on single GPU, needs dual split or 24 GB

The four compact tiers show flat prefill around 2700 tok/s and decode around 46 tok/s with less than 4 percent variance, consistent with an identical backbone and heads that are decode bound. HIGH, VERY-HIGH and HIGHEST exceed single 16 GB VRAM and fall back to partial CPU offload on this single GPU bench, on dual GPU split or 24 GB cards they run at full GPU speed. For reference, sibling BUDGET 13.70 GiB at 2753.10 +/- 249.27 pp512 and 47.51 +/- 0.06 tg128 fits single 5070 Ti, STARVED 13.58 GiB at 2762.68 +/- 235.81 and 48.05 +/- 0.13 also fits.

vLLM NVFP4 smoke

A short vLLM smoke check with tensor parallel 2 and 2048 context on the NVFP4 checkpoint passed with coherent output, confirming the checkpoint loads and generates before GGUF tier building.

LocalAI gateway check (naive, one-at-a-time)

Each tier was sent one large request with max_tokens=20000 and finish_reason=stop for all tiers. These are end to end gateway timings, not pure decode, so treat as sanity and responsiveness, not a formal benchmark.

Tier request_seconds tokens per sec (20000 / s) finish_reason reasoning chars content chars
VERY-LOW 661.94 30.21 stop 1325 4941
COMPACT-LOW 392.98 50.89 stop 1392 9382
LOW 370.89 53.92 stop 795 2806
MEDIUM 374.34 53.43 stop 1464 2474
MID-HIGH 422.28 47.36 stop 3320 8399
HIGH 507.82 39.38 stop 2115 8685
VERY-HIGH 505.75 39.55 stop 1993 7551
HIGHEST 416.45 48.02 stop 3557 5056

Budget companions in the sibling repo on the same gateway method: BUDGET 178.84s, 111.83 tok/s, stop; STARVED 303.68s, 65.86 tok/s, stop. Both finished with finish_reason stop, no repetition collapse. Gateway overhead included, single run only, don't compare to llama-bench decode.

Tokens per sec is 20000 divided by request_seconds. It reflects the whole gateway round trip for this specific large prompt, not llama-bench decode. All tiers had distinct 5-gram ratios near 1.0 and no adjacent duplication, so no repetition collapse.

What to take away, naive framing

  • Single run only, one large synthetic prompt, gateway overhead included, no warmup average. Don't compare gateway tok/s to direct llama-bench numbers and don't compare across hardware.
  • LOW and MEDIUM were the fastest gateway round trips for this prompt at about 54 tok/s, MID-HIGH at 47 tok/s, HIGHEST at 48 tok/s, HIGH and VERY-HIGH slower at about 39 tok/s. VERY-LOW was the outlier at 30 tok/s with the longest wall time but still finished cleanly with stop. All produced coherent output with no truncation. BUDGET was fastest at 111 tok/s on its cycle for the same gateway prompt, STARVED at 65 tok/s.

SHA-256

c58f0a932957825cea5e6b73966ddf451f33b287668f670e1293a8ef132bf561  Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4-VERY-LOW.gguf
56f8d6f5c656ac96da20086c4e9e546c1b30d723e185386d50219b08b5e47f83  Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4-COMPACT-LOW.gguf
6b2d5f6d5795ceef5a9dcc18a444ef9d03dd47b1d3bffaf80dd392f5a6cc425b  Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4-LOW.gguf
1c9e7f3e77c6938fb0f1218ec0da9de93dcd7a83357e244d76e4ebff99e36058  Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4-MEDIUM.gguf
4e1b87edfc2b7f58a78c657afe965344dea2415ae71264f93ecea1f94d4ac24f  Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4-MID-HIGH.gguf
10f60fd0f0597e9d97641e6c77b637fc247e343fb39b84b145f17254dd588bc8  Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4-HIGH.gguf
b08f4dcd1c5c48463c185746dd6ba594cae466c24c5e2a26628351dbadabc702  Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4-VERY-HIGH.gguf
c18d844039ccbd5cf5195c12937bf7043c26324a03de3c5b6ff3a821a6c68473  Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4-HIGHEST.gguf

Attribution & provenance

This is a derivative work built entirely from existing Apache 2.0 artifacts. Nothing here was trained or fine tuned. Credit belongs to:

  1. Alibaba and Qwen team for the base model, Qwen/Qwen3.8-27B (Apache 2.0): 27B dense, 64 blocks, Gated DeltaNet plus Gated Attention hybrid, native vision language, 262,144 token context, MTP head.
  2. DavidAU for the tune itself, Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU (Apache 2.0): the TURBO 735-882 Fable plus Cold Fusion plus Heretic/Uncensored DPO stack that this family converts.
  3. Unsloth, whose trainers and systems power the underlying training methods.
  4. This repo's author for the GGUF conversion and the tier ladder only.

Repository contents

  • Eight tier GGUFs (table above, 15.15 to 21.28 GB) plus 8 override maps (overrides-*.txt, 1,122 lines each: per-tensor target types, includes output.weight/token_embd.weight per tier). HIGHEST override pins attention/SSM to Q8_0 and keeps token_embd/MTP at BF16.
  • Budget no-MTP companion repo: esatapedico/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4-BUDGET-GGUF (BUDGET 14.72 GB Q3_K/Q2_K, STARVED 14.59 GB Q2_K/Q2_K, 1,107 tensors, no MTP)
  • Corresponding safetensors NVFP4 checkpoint (single model.safetensors, compressed tensors nvfp4-pack-quantized): esatapedico/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4

License

apache-2.0 (inherits from Qwen base and DavidAU tune). general.license = apache-2.0 is set inside every GGUF. general.name equals filename without extension, general.description mentions full base Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU, qwen35.block_count 65, qwen35.context_length 262144.

Card written by AI assistance at the request of the repository author, who did the engineering. As with any AI-generated text, there may be errors; please verify anything important (hashes, sizes, commands) against the file itself before relying on it.