incoai/Muse-Glimmer-30B-DFlash2-GGUF

🤗 Hugging Face sourcetext-generationapache-2.030B activated10 GBGGUF✓ 4 checksumsupdated today
Submit in one command

Run it next to your model folder. It makes the torrent, checks your files against Hugging Face, and submits it. You just start seeding and paste your key from your account. It only reads your files and never changes them. Read the script first if you like.

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo incoai/Muse-Glimmer-30B-DFlash2-GGUF ./model-folder
Needs a seeder →

Muse-Glimmer-30B-DFlash2-GGUF

Blog | GitHub

This repository contains GGUF conversions of incoai/Muse-Glimmer-30B-DFlash2, the DFlash 2 draft model for meta-models/Muse-Glimmer-30B. It is not a standalone language model: it runs inside a speculative decoding server and drafts tokens for the target model to verify. The checkpoints are also mirrored at z-lab/Muse-Glimmer-30B-DFlash2-GGUF.

DFlash 2 is a block-diffusion drafter for speculative decoding. It predicts a whole block of tokens in a single pass and keeps the top candidates at every position. A lightweight selector then traces one coherent path through them. Two-tap dynamic convolutions in the backbone keep the draft from decaying toward the end of the block. Decoding is lossless: greedy output matches the target model exactly, and sampling preserves its distribution.

File Size
Muse-Glimmer-30B-DFlash2-Q4_K_M.gguf 1.6 GB
Muse-Glimmer-30B-DFlash2-Q8_0.gguf 2.9 GB
Muse-Glimmer-30B-DFlash2-BF16.gguf 5.5 GB

Quick Start

Build llama.cpp with DFlash 2 support (PR #27342):

git clone https://github.com/ggml-org/llama.cpp.git
cd llama.cpp
git fetch origin pull/27342/head:pr-27342
git switch pr-27342

# NVIDIA CUDA
cmake -B build -DCMAKE_BUILD_TYPE=Release -DGGML_CUDA=ON
cmake --build build -j

# Apple Silicon
cmake -B build -DCMAKE_BUILD_TYPE=Release -DGGML_METAL=ON
cmake --build build -j

Then serve:

./build/bin/llama-server \
  -hf meta-models/Muse-Glimmer-30B-GGUF:Q4_K_M \
  -hfd incoai/Muse-Glimmer-30B-DFlash2-GGUF:Q4_K_M \
  --spec-type draft-dflash \
  --spec-draft-n-max 15

See the blog post for other engines and more details.

Evaluation

  • Target: meta-models/Muse-Glimmer-30B-GGUF, Q4_K_M
  • Sampling: Muse's officially recommended parameters (temperature 1.0, top-p 0.95, top-k 64), with high reasoning strength
  • Maximum new tokens: 2048
  • Prompts: the first eight GSM8K test examples

Acceptance Length

Acceptance length is the per-request mean of completion tokens divided by verification steps. Higher is better.

Draft GGUF Acceptance Length
BF16 5.45
Q8_0 5.58
Q4_K_M 5.44

Full evaluations of the base checkpoint are on the main model card.

Citation

If you find DFlash 2 useful, please cite:

@misc{inco2026dflash2,
  title  = {{DFlash 2: Keep Drafting Parallel}},
  author = {{Inco AI}},
  year   = {2026},
  month  = {August},
  url    = {https://inco.ai/blog/dflash2/}
}

Please also cite the original DFlash paper:

@inproceedings{chen2026dflash,
  title     = {{DFlash: Block Diffusion for Flash Speculative Decoding}},
  author    = {Chen, Jian and Liang, Yesheng and Liu, Zhijian},
  booktitle = {International Conference on Machine Learning (ICML)},
  year      = {2026}
}