aday777/Qwen3.8-27B-ARA-abliterated-NVFP4-MTP

🤗 Hugging Face sourceimage-text-to-textapache-2.016.7B params31 GBsafetensors✓ 6 checksumsupdated today
Submit in one command

Run it next to your model folder. It makes the torrent, checks your files against Hugging Face, and submits it. You just start seeding and paste your key from your account. It only reads your files and never changes them. Read the script first if you like.

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo aday777/Qwen3.8-27B-ARA-abliterated-NVFP4-MTP ./model-folder
Needs a seeder →

Qwen3.8-27B ARA Abliterated / Uncensored NVFP4 + MTP

This is a multimodal, refusal-ablated (commonly described as "uncensored") W4A4 NVFP4 derivative of the reproducible trohrbaugh/Qwen3.8-27B-heretic-ara BF16 checkpoint. It is intended for native Blackwell NVFP4 inference.

What is preserved

  • The vision tower, recurrent convolutions, language head, and all 15 native MTP tensors remain BF16.
  • The MTP tensors were grafted from the hash-verified source after Transformers serialization and verified bit-exact.
  • All 333 vision tensors were verified bit-exact against the BF16 source.
  • The language-model linear layers use compressed-tensors NVFP4 W4A4 group-16 quantization.

Validation

This artifact passed its text-capability, image-vision, benign refusal-surface, native MTP-acceptance, integrity, and clean-load gates on vLLM 0.23. See BUILD_MANIFEST.json, VALIDATION_REPORT.json, and SHA256SUMS for exact provenance and results. Video tensors/processors are preserved, but video input was not part of the live runtime gate.

The live gate used Qwen3_5ForConditionalGeneration, native three-token MTP, the FlashInfer CUTLASS NVFP4 kernel, an 8,192-token context, and an RTX PRO 6000 Blackwell GPU. Loaded model memory was approximately 19.53 GiB.

vLLM example

vllm serve aday777/Qwen3.8-27B-ARA-abliterated-NVFP4-MTP \
  --served-model-name Qwen3.8-27B-ARA-NVFP4-MTP \
  --max-model-len 8192 \
  --speculative-config '{"method":"mtp","num_speculative_tokens":3}'

Use a recent vLLM build with Qwen3.5 multimodal and compressed-tensors NVFP4 support. Native NVFP4 execution requires compatible Blackwell hardware and CUDA runtime support.

Notes

"Abliterated" or "uncensored" describes the source checkpoint's refusal-ablation process; it is not a guarantee that every prompt will receive a particular answer. Users remain responsible for evaluating outputs and applying safeguards appropriate to their deployment.


Support

If this model is useful to you, Bitcoin donations are welcome:

bc1q5ayht3fxhj0v95fk0z8l2f6900g3awdsw5842p