pipenetwork/Inkling-Small-MLX-4bit

🤗 Hugging Face sourceimage-text-to-textapache-2.0264B params527 GBsafetensors✓ 30 checksumsupdated today
Submit in one command

Run it next to your model folder. It makes the torrent, checks your files against Hugging Face, and submits it. You just start seeding and paste your key from your account. It only reads your files and never changes them. Read the script first if you like.

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo pipenetwork/Inkling-Small-MLX-4bit ./model-folder
Needs a seeder →

Inkling-Small-MLX-4bit

Built with Inkling (Thinking Machines Lab).

MLX (Apple Silicon) conversion of thinkingmachines/Inkling-Small, quantized to 4-bit (affine group quant, group size 64).

Code / loader: github.com/PipeNetwork/inkling-mlx

Inkling Small is a 276B-total / 12B-active sparse-MoE, natively multimodal model (text + image/video + audio → text). This is the full multimodal conversion: all three towers (text backbone, HMLP vision, dMel audio) are ported; the multi-token-prediction head is dropped (inference-irrelevant).

Builds

Variant Size Text ppl Notes
8bit ~280 GB 5.569 near-lossless
6bit ~214 GB 5.569 high quality
4bit ~148 GB 5.452 balanced default
3bit ~115 GB 6.706 ⚠️ experimental — visibly degraded

Perplexity is teacher-forcing over one fixed held-out set (prose / code / reasoning / multilingual) — identical inputs across builds, so the columns compare directly. 4-bit shows no measurable loss vs 8-bit.

There is also a REAP-pruned build: REAP25-4bit keeps 4-bit precision with 192 of 256 routed experts, fitting a 128 GB Mac at ~112 GB for no measurable perplexity cost, with vision and speech intact.

No bf16 build is published. The MLX bf16 conversion is bit-identical to the upstream checkpoint (name-mapping and layout only — the dtype cast is a no-op), so it would carry nothing thinkingmachines/Inkling-Small does not already have, and at ~527 GB it does not fit a 512 GB Mac. If you want it as a requant source, scripts/convert_all.sh regenerates it from the upstream weights in about three minutes.

Quantization scheme: affine int4 (not NVFP4 / MXFP4)

MLX supports FP4 modes and Thinking Machines ships an Inkling-NVFP4 checkpoint — so for the record, we benchmarked round-trip reconstruction error (‖W − Ŵ‖ / ‖W‖ vs bf16) on real Inkling expert weights:

Scheme bits/weight reconstruction error
affine int4 (group 64) 4.50 ~9.1%
nvfp4 (group 16) 4.50 ~10.2%
mxfp4 (group 32) 4.25 ~12.3%

Affine int4 is the most faithful: it is asymmetric (per-group scale and zero-point, 16 uniform levels), which centers on Inkling's near-Gaussian expert weights better than symmetric FP4's fixed non-uniform levels. FP4's real payoff is heavy-tailed activations and native Blackwell FP4 tensor cores — neither helps weight fidelity on Apple Silicon, where MLX would dequantize FP4 anyway. So these builds use affine int4.

⚠️ Loading requires the bundled inkling_mlx loader

The inkling_mm_model architecture is not in stock mlx-lm / mlx-vlm, so this repo bundles a minimal, numerically-validated MLX implementation under inkling_mlx/.

pip install mlx mlx-lm transformers
from inkling_mlx.load import load
from inkling_mlx.generate import greedy_generate
from transformers import AutoTokenizer

model, config = load("/path/to/this/repo")
tok = AutoTokenizer.from_pretrained("/path/to/this/repo", trust_remote_code=True)
ids = tok("The capital of France is")["input_ids"]
print(tok.decode(greedy_generate(model, config, ids, max_new_tokens=64)))

Needs an Apple-Silicon Mac with enough unified memory to hold the weights (≈ the size above).

Status & caveats

  • Text generation works end-to-end via an incremental KV + short-convolution cache.
  • Multimodal is supported end-to-end: the vision/audio towers and their preprocessing (InklingProcessor — image patchify/normalize, audio log-mel→dMel, validated ~1e-7 vs the reference) are included. Pass images/audio via the processor.
  • Quantized: attention / MLP / expert projections, token embed+unembed, and the vision/audio matmuls. Kept in higher precision: the MoE router, RMSNorms, the four short-convolutions per layer, and the relative-position bias.

Conversion is streaming (tensor-by-tensor; the ~527 GB bf16 model never fully loads into RAM) and was validated with fp32 numerical parity against transformers PR #47347. License: Apache-2.0 (inherits the base model).