prithivMLmods/Qwen3-VL-4B-Instruct-Unredacted-MAX-FP8

🤗 Hugging Face sourceimage-text-to-textapache-2.04.4B params5.2 GBsafetensors✓ 3 checksumsupdated today
Submit in one command

Run it next to your model folder. It makes the torrent, checks your files against Hugging Face, and submits it. You just start seeding and paste your key from your account. It only reads your files and never changes them. Read the script first if you like.

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo prithivMLmods/Qwen3-VL-4B-Instruct-Unredacted-MAX-FP8 ./model-folder
Needs a seeder →

Qwen3-VL-4B-Instruct-Unredacted-MAX-FP8

Qwen3-VL-4B-Instruct-Unredacted-MAX-FP8 is an FP8-compressed evolution built on top of Qwen3-VL-4B-Instruct-Unredacted-MAX. This variant preserves the unredacted abliterated training strategies of the original MAX release while applying BF16 · F8_E4M3 FP8 weight compression for improved efficiency. The result is a highly capable 4B vision-language model optimized for unrestricted, detailed reasoning and captioning across complex visual inputs, with reduced memory footprint and improved hardware throughput.

[!important] FP8 (8-bit floating point) weight and activation quantization using hardware acceleration on GPUs – FP8 W8A8. Quantization W8A8 FP8-dynamic recipe – examples.

Key Highlights

  • FP8 Compressed (BF16 · F8_E4M3): Weights compressed using FP8 E4M3 format while retaining BF16 compute stability for efficient deployment.
  • Unredacted MAX Training: Fine-tuned to significantly reduce refusal patterns and improve instruction adherence across diverse prompts.
  • 4B Parameter Architecture: Based on Qwen3-VL-4B-Instruct-Unredacted-MAX, balancing strong reasoning performance with lower VRAM requirements compared to larger 8B variants.
  • Unrestricted Multimodal Reasoning: Designed for deep analysis of artistic, forensic, technical, or abstract visual content without standard safety-driven refusals.
  • High-Fidelity Captions: Produces dense, descriptive outputs suitable for dataset generation, metadata enrichment, or accessibility use cases.
  • Dynamic Resolution Support: Retains Qwen3-VL’s ability to process varying image resolutions and aspect ratios effectively.

Quick Start with Transformers

from transformers import Qwen3VLForConditionalGeneration, AutoProcessor
from qwen_vl_utils import process_vision_info
import torch

# Load the 4B Unredacted MAX FP8 model
model = Qwen3VLForConditionalGeneration.from_pretrained(
    "prithivMLmods/Qwen3-VL-4B-Instruct-Unredacted-MAX-FP8",
    torch_dtype="auto",
    device_map="auto"
)

processor = AutoProcessor.from_pretrained(
    "prithivMLmods/Qwen3-VL-4B-Instruct-Unredacted-MAX-FP8"
)

messages = [
    {
        "role": "user",
        "content": [
            {
                "type": "image",
                "image": "https://qianwen-res.oss-cn-beijing.aliyuncs.com/Qwen-VL/assets/demo.jpeg",
            },
            {"type": "text", "text": "Provide a detailed caption and reasoning for this image."},
        ],
    }
]

text = processor.apply_chat_template(
    messages, tokenize=False, add_generation_prompt=True
)

image_inputs, video_inputs = process_vision_info(messages)

inputs = processor(
    text=[text],
    images=image_inputs,
    videos=video_inputs,
    padding=True,
    return_tensors="pt",
).to("cuda")

generated_ids = model.generate(**inputs, max_new_tokens=256)

generated_ids_trimmed = [
    out_ids[len(in_ids):] for in_ids, out_ids in zip(inputs.input_ids, generated_ids)
]

output_text = processor.batch_decode(
    generated_ids_trimmed,
    skip_special_tokens=True,
    clean_up_tokenization_spaces=False
)

print(output_text)

Intended Use

  • Advanced Red-Teaming: Evaluating multimodal robustness and probing behavioral edge cases.
  • Complex Data Archiving: Generating detailed captions for medical, artistic, historical, or research datasets.
  • Refusal Mechanism Research: Studying behavioral shifts in vision-language models after abliterated fine-tuning.
  • Creative Storytelling: Producing detailed visual descriptions for narrative and world-building projects.

Limitations & Risks

Critical Note: This model is designed to minimize built-in refusal mechanisms.

  • Sensitive Content Exposure: The model may generate explicit or controversial descriptions if prompted accordingly.
  • User Responsibility: Generated outputs must be handled responsibly and used within ethical and legal boundaries.
  • Hardware Requirements: While more memory-efficient due to FP8 compression, adequate VRAM is still required for high-resolution image processing and longer generations.

Acknowledgements

I would like to thank the works of the following:

  • Uncensor any LLM with abliteration – Maxime Labonne
  • Using FP8 and FP4 with Transformer Engine – docs.nvidia
  • Remove Refusals with Transformers – Sumandora
  • LLM Compressor – vllm-project
  • FP8 Floating-Point 8: An Introduction to Efficient, Lower-Precision AI Training – nvidia