aday777/Qwen3.8-27B-ARA-abliterated-NVFP4-MTP

🤗 Hugging Face 来源image-text-to-textapache-2.016.7B 参数31 GBsafetensors✓ 6 个校验和今天更新
一条命令提交

在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo aday777/Qwen3.8-27B-ARA-abliterated-NVFP4-MTP ./model-folder
需要做种者 →

Qwen3.8-27B ARA Abliterated / Uncensored NVFP4 + MTP

This is a multimodal, refusal-ablated (commonly described as "uncensored") W4A4 NVFP4 derivative of the reproducible trohrbaugh/Qwen3.8-27B-heretic-ara BF16 checkpoint. It is intended for native Blackwell NVFP4 inference.

What is preserved

  • The vision tower, recurrent convolutions, language head, and all 15 native MTP tensors remain BF16.
  • The MTP tensors were grafted from the hash-verified source after Transformers serialization and verified bit-exact.
  • All 333 vision tensors were verified bit-exact against the BF16 source.
  • The language-model linear layers use compressed-tensors NVFP4 W4A4 group-16 quantization.

Validation

This artifact passed its text-capability, image-vision, benign refusal-surface, native MTP-acceptance, integrity, and clean-load gates on vLLM 0.23. See BUILD_MANIFEST.json, VALIDATION_REPORT.json, and SHA256SUMS for exact provenance and results. Video tensors/processors are preserved, but video input was not part of the live runtime gate.

The live gate used Qwen3_5ForConditionalGeneration, native three-token MTP, the FlashInfer CUTLASS NVFP4 kernel, an 8,192-token context, and an RTX PRO 6000 Blackwell GPU. Loaded model memory was approximately 19.53 GiB.

vLLM example

vllm serve aday777/Qwen3.8-27B-ARA-abliterated-NVFP4-MTP \
  --served-model-name Qwen3.8-27B-ARA-NVFP4-MTP \
  --max-model-len 8192 \
  --speculative-config '{"method":"mtp","num_speculative_tokens":3}'

Use a recent vLLM build with Qwen3.5 multimodal and compressed-tensors NVFP4 support. Native NVFP4 execution requires compatible Blackwell hardware and CUDA runtime support.

Notes

"Abliterated" or "uncensored" describes the source checkpoint's refusal-ablation process; it is not a guarantee that every prompt will receive a particular answer. Users remain responsible for evaluating outputs and applying safeguards appropriate to their deployment.


Support

If this model is useful to you, Bitcoin donations are welcome:

bc1q5ayht3fxhj0v95fk0z8l2f6900g3awdsw5842p