catplusplus/Qwen3-VL-30B-A3B-Instruct-Heretic-NVFP4

认证创作者 catplusplus 已认证
🤗 Hugging Face 来源apache-2.018B 参数激活 3B19 GBsafetensors✓ 5 个校验和今天更新
一条命令提交

在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo catplusplus/Qwen3-VL-30B-A3B-Instruct-Heretic-NVFP4 ./model-folder
需要做种者 →

This is an uncensored (not perfect but doesn't refuse much in practice) model compressed with LLM compressor using the following script: import torch import sys from datasets import load_dataset from transformers import AutoProcessor, Qwen3VLMoeForConditionalGeneration

from llmcompressor import oneshot from llmcompressor.modeling.moe_context import moe_calibration_context from llmcompressor.modifiers.quantization import QuantizationModifier from llmcompressor.utils import dispatch_for_generation

NOTE: Requires a minimum of transformers 4.57.0

#MODEL_ID = "Qwen/Qwen3-VL-235B-A22B-Instruct" MODEL_ID = sys.argv[1]

Load model.

model = Qwen3VLMoeForConditionalGeneration.from_pretrained(MODEL_ID, torch_dtype="auto") processor = AutoProcessor.from_pretrained(MODEL_ID)

DATASET_ID = "neuralmagic/calibration" NUM_CALIBRATION_SAMPLES = 20 MAX_SEQUENCE_LENGTH = 8192

ds = load_dataset(DATASET_ID, name="LLM", split=f"train[:{NUM_CALIBRATION_SAMPLES}]")

def preprocess_function(example): messgages = [] for message in example["messages"]: messgages.append( { "role": message["role"], "content": [{"type": "text", "text": message["content"]}], } )

return processor.apply_chat_template(
    messgages,
    return_tensors="pt",
    padding=False,
    truncation=True,
    max_length=MAX_SEQUENCE_LENGTH,
    tokenize=True,
    add_special_tokens=False,
    return_dict=True,
    add_generation_prompt=False,
)

ds = ds.map(preprocess_function, batched=False, remove_columns=ds.column_names)

def data_collator(batch): assert len(batch) == 1 return { key: ( torch.tensor(value) if key != "pixel_values" else torch.tensor(value, dtype=torch.bfloat16).squeeze(0) ) for key, value in batch[0].items() }

Configure the quantization algorithm and scheme.

In this case, we:

* quantize the weights to fp4 with group-wise quantization

* quantize the activations to fp4 with dynamic group activations

recipe = QuantizationModifier( targets="Linear", scheme="NVFP4", ignore=[ "re:.lm_head", "re:visual.", "re:model.visual.*", "re:.*mlp.gate$", ], )

Apply quantization.

with moe_calibration_context(model): oneshot( model=model, recipe=recipe, max_seq_length=MAX_SEQUENCE_LENGTH, num_calibration_samples=NUM_CALIBRATION_SAMPLES, dataset=ds, data_collator=data_collator, )

print("========== SAMPLE GENERATION ==============") dispatch_for_generation(model) input_ids = processor(text="Hello my name is", return_tensors="pt").input_ids.to("cuda") output = model.generate(input_ids, max_new_tokens=20) print(processor.decode(output[0])) print("==========================================")

Save to disk in compressed-tensors format.

SAVE_DIR = MODEL_ID.rstrip("/").split("/")[-1] + "-NVFP4" model.save_pretrained(SAVE_DIR) processor.save_pretrained(SAVE_DIR)