inference-optimization/Kimi-K3-0.40B-MXFP4

🤗 Hugging Face 来源mit396M 参数1.5 GBsafetensors✓ 2 个校验和今天更新
一条命令提交

在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo inference-optimization/Kimi-K3-0.40B-MXFP4 ./model-folder
需要做种者 →

Kimi-K3-0.40B-MXFP4

This is an MXFP4-quantized version of inference-optimization/Kimi-K3-0.40B, a tiny model derived from moonshotai/Kimi-K3. Created for testing and development.

Model Details

  • Base Model: inference-optimization/Kimi-K3-0.40B
  • Architecture: kimi_k3
  • Total Parameters: 0.40B
  • Activated Parameters: ~0.22B (MoE: 2 of 8 experts active per token, plus 1 shared expert)
  • Quantization: W4A16 MXFP4 (mxfp4-pack-quantized), group size 32

Quantization Config

Matches the quantization scheme used in moonshotai/Kimi-K3:

Field Value
Format mxfp4-pack-quantized
Weights 4-bit float, group_size=32, minmax observer
Scale dtype torch.uint8
Activations unquantized (W4A16)
Ignored layers self_attn, shared_experts, lm_head, vision_tower

Usage

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
from compressed_tensors.offload import dispatch_model

model = AutoModelForCausalLM.from_pretrained(
    "inference-optimization/Kimi-K3-0.40B-MXFP4",
    trust_remote_code=True,
    dtype=torch.bfloat16,
)
tokenizer = AutoTokenizer.from_pretrained(
    "inference-optimization/Kimi-K3-0.40B-MXFP4",
    trust_remote_code=True,
)
dispatch_model(model)

sample = tokenizer("Hello my name is", return_tensors="pt")
sample = {k: v.to(model.device) for k, v in sample.items()}
output = model.generate(
    **sample,
    max_new_tokens=100,
    eos_token_id=tokenizer.eos_token_id,
    pad_token_id=tokenizer.pad_token_id,
)
print(tokenizer.decode(output[0].tolist(), skip_special_tokens=True))

Creation Process

Quantized using llm-compressor:

from transformers import AutoModelForCausalLM, AutoTokenizer
from llmcompressor import oneshot
from llmcompressor.modifiers.quantization import QuantizationModifier

MODEL_ID = "inference-optimization/Kimi-K3-0.40B"
model = AutoModelForCausalLM.from_pretrained(MODEL_ID, trust_remote_code=True)
tokenizer = AutoTokenizer.from_pretrained(MODEL_ID, trust_remote_code=True)

recipe = QuantizationModifier(
    targets="Linear",
    scheme="MXFP4A16",
    ignore=[
        "re:.*self_attn.*",
        "re:.*shared_experts.*",
        "re:.*lm_head.*",
        "re:.*vision_tower.*",
    ],
)
oneshot(model=model, recipe=recipe)
model.save_pretrained(SAVE_DIR, save_compressed=True)
tokenizer.save_pretrained(SAVE_DIR)

Notes

  • trust_remote_code=True is required to load the custom modeling files.
  • Load with dtype=torch.bfloat16 to match the decompressed weight dtype.