catplusplus/Qwen3-VL-30B-A3B-Instruct-Heretic-NVFP4

Verified creator catplusplus verified
🤗 Hugging Face sourceapache-2.018B params3B activated19 GBsafetensors✓ 5 checksumsupdated today
Submit in one command

Run it next to your model folder. It makes the torrent, checks your files against Hugging Face, and submits it. You just start seeding and paste your key from your account. It only reads your files and never changes them. Read the script first if you like.

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo catplusplus/Qwen3-VL-30B-A3B-Instruct-Heretic-NVFP4 ./model-folder
Needs a seeder →

This is an uncensored (not perfect but doesn't refuse much in practice) model compressed with LLM compressor using the following script: import torch import sys from datasets import load_dataset from transformers import AutoProcessor, Qwen3VLMoeForConditionalGeneration

from llmcompressor import oneshot from llmcompressor.modeling.moe_context import moe_calibration_context from llmcompressor.modifiers.quantization import QuantizationModifier from llmcompressor.utils import dispatch_for_generation

NOTE: Requires a minimum of transformers 4.57.0

#MODEL_ID = "Qwen/Qwen3-VL-235B-A22B-Instruct" MODEL_ID = sys.argv[1]

Load model.

model = Qwen3VLMoeForConditionalGeneration.from_pretrained(MODEL_ID, torch_dtype="auto") processor = AutoProcessor.from_pretrained(MODEL_ID)

DATASET_ID = "neuralmagic/calibration" NUM_CALIBRATION_SAMPLES = 20 MAX_SEQUENCE_LENGTH = 8192

ds = load_dataset(DATASET_ID, name="LLM", split=f"train[:{NUM_CALIBRATION_SAMPLES}]")

def preprocess_function(example): messgages = [] for message in example["messages"]: messgages.append( { "role": message["role"], "content": [{"type": "text", "text": message["content"]}], } )

return processor.apply_chat_template(
    messgages,
    return_tensors="pt",
    padding=False,
    truncation=True,
    max_length=MAX_SEQUENCE_LENGTH,
    tokenize=True,
    add_special_tokens=False,
    return_dict=True,
    add_generation_prompt=False,
)

ds = ds.map(preprocess_function, batched=False, remove_columns=ds.column_names)

def data_collator(batch): assert len(batch) == 1 return { key: ( torch.tensor(value) if key != "pixel_values" else torch.tensor(value, dtype=torch.bfloat16).squeeze(0) ) for key, value in batch[0].items() }

Configure the quantization algorithm and scheme.

In this case, we:

* quantize the weights to fp4 with group-wise quantization

* quantize the activations to fp4 with dynamic group activations

recipe = QuantizationModifier( targets="Linear", scheme="NVFP4", ignore=[ "re:.lm_head", "re:visual.", "re:model.visual.*", "re:.*mlp.gate$", ], )

Apply quantization.

with moe_calibration_context(model): oneshot( model=model, recipe=recipe, max_seq_length=MAX_SEQUENCE_LENGTH, num_calibration_samples=NUM_CALIBRATION_SAMPLES, dataset=ds, data_collator=data_collator, )

print("========== SAMPLE GENERATION ==============") dispatch_for_generation(model) input_ids = processor(text="Hello my name is", return_tensors="pt").input_ids.to("cuda") output = model.generate(input_ids, max_new_tokens=20) print(processor.decode(output[0])) print("==========================================")

Save to disk in compressed-tensors format.

SAVE_DIR = MODEL_ID.rstrip("/").split("/")[-1] + "-NVFP4" model.save_pretrained(SAVE_DIR) processor.save_pretrained(SAVE_DIR)