inference-optimization/Kimi-K3-0.40B-MXFP4

🤗 Hugging Face sourcemit396M params1.5 GBsafetensors✓ 2 checksumsupdated today
Submit in one command

Run it next to your model folder. It makes the torrent, checks your files against Hugging Face, and submits it. You just start seeding and paste your key from your account. It only reads your files and never changes them. Read the script first if you like.

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo inference-optimization/Kimi-K3-0.40B-MXFP4 ./model-folder
Needs a seeder →

Kimi-K3-0.40B-MXFP4

This is an MXFP4-quantized version of inference-optimization/Kimi-K3-0.40B, a tiny model derived from moonshotai/Kimi-K3. Created for testing and development.

Model Details

  • Base Model: inference-optimization/Kimi-K3-0.40B
  • Architecture: kimi_k3
  • Total Parameters: 0.40B
  • Activated Parameters: ~0.22B (MoE: 2 of 8 experts active per token, plus 1 shared expert)
  • Quantization: W4A16 MXFP4 (mxfp4-pack-quantized), group size 32

Quantization Config

Matches the quantization scheme used in moonshotai/Kimi-K3:

Field Value
Format mxfp4-pack-quantized
Weights 4-bit float, group_size=32, minmax observer
Scale dtype torch.uint8
Activations unquantized (W4A16)
Ignored layers self_attn, shared_experts, lm_head, vision_tower

Usage

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
from compressed_tensors.offload import dispatch_model

model = AutoModelForCausalLM.from_pretrained(
    "inference-optimization/Kimi-K3-0.40B-MXFP4",
    trust_remote_code=True,
    dtype=torch.bfloat16,
)
tokenizer = AutoTokenizer.from_pretrained(
    "inference-optimization/Kimi-K3-0.40B-MXFP4",
    trust_remote_code=True,
)
dispatch_model(model)

sample = tokenizer("Hello my name is", return_tensors="pt")
sample = {k: v.to(model.device) for k, v in sample.items()}
output = model.generate(
    **sample,
    max_new_tokens=100,
    eos_token_id=tokenizer.eos_token_id,
    pad_token_id=tokenizer.pad_token_id,
)
print(tokenizer.decode(output[0].tolist(), skip_special_tokens=True))

Creation Process

Quantized using llm-compressor:

from transformers import AutoModelForCausalLM, AutoTokenizer
from llmcompressor import oneshot
from llmcompressor.modifiers.quantization import QuantizationModifier

MODEL_ID = "inference-optimization/Kimi-K3-0.40B"
model = AutoModelForCausalLM.from_pretrained(MODEL_ID, trust_remote_code=True)
tokenizer = AutoTokenizer.from_pretrained(MODEL_ID, trust_remote_code=True)

recipe = QuantizationModifier(
    targets="Linear",
    scheme="MXFP4A16",
    ignore=[
        "re:.*self_attn.*",
        "re:.*shared_experts.*",
        "re:.*lm_head.*",
        "re:.*vision_tower.*",
    ],
)
oneshot(model=model, recipe=recipe)
model.save_pretrained(SAVE_DIR, save_compressed=True)
tokenizer.save_pretrained(SAVE_DIR)

Notes

  • trust_remote_code=True is required to load the custom modeling files.
  • Load with dtype=torch.bfloat16 to match the decompressed weight dtype.