Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4
NVFP4 checkpoint of DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU. NVFP4 here means W4A16 with FP8 scales, group size 16, weight only. Linear layers are NVFP4, vision tower, linear attention path, lm_head and MTP head are kept in BF16 at the source. No calibration data was used. This matches the pattern used for other Qwen3.8 Cold Fusion NVFP4 builds.
HF repo: esatapedico/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4
GGUF family: esatapedico/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4-GGUF (6 tiers, see the GGUF model card)
At a glance
| Field | Value |
|---|---|
| Format | compressed-tensors nvfp4-pack-quantized |
| Quantization | W4A16, group size 16, FP8 E4M3 scales, weight only |
| Kept in BF16 | vision tower, linear attention path, lm_head, embeddings, MTP head |
| Calibration | none |
| Base | DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU (Qwen3.8-27B, Apache 2.0) |
Quantized set covers MLP on all 64 layers plus Q, K, V, O on the 16 full attention layers. The DeltaNet path stays BF16 in this checkpoint and is normalized to NVFP4 in the GGUF tier builds (448 tensor backbone).
Checkpoint
- Single
model.safetensorsabout 26 GB config.jsonquantization_config.format=nvfp4-pack-quantized,quant_method=compressed-tensors,Qwen3_5ForCausalLM, 64 layers, hybrid GatedDeltaNet, 262144 context, MTP headtokenizer.jsonintact, chat template intact- NVFP4 tensors verified via metadata checks
Provenance
Derivative of DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU (Apache 2.0) which itself derives from Qwen/Qwen3.8-27B. See the GGUF model card for full attribution. This checkpoint is the source for the GGUF family, which contains 1,122 tensors per file with a shared 448 tensor NVFP4 backbone.
License
apache-2.0
Card written with AI assistance.