dealignai/Muse-Glimmer-30B-CRACK-GGUF

🤗 On Hugging Faceimage-text-to-textapache-2.061 GBGGUFHF checksums availableupdated today
Magnet

Dealign.ai

Muse-Glimmer-30B-CRACK-GGUF

CRACK-abliterated Muse Glimmer 30B — GGUF quants for llama.cpp. Three quantizations

(Q8_0 / Q4_K_M / Q2_K) in one repository. Refusal behavior removed while preserving the

model's knowledge, reasoning, multi-strength thinking, and ATEM tool-calling.

Research artifact with reduced safety guardrails. Use responsibly and lawfully.

Quantizations

| File | Size | Notes |

|---|---|---|

| Q8_0 | 29.6 GB | near-lossless reference |

| Q4_K_M | 16.9 GB | balanced (recommended) |

| Q2_K | 10.7 GB | smallest |

Pick one text file plus the vision projector mmproj-Muse-Glimmer-30B-f16.gguf (3.8 GB)

for image input. Q4_K_M is the recommended balance; Q8_0 is near-lossless; Q2_K is smallest.

Benchmarks

Evaluated through llama.cpp at greedy decoding. MMLU is logit-mode accuracy (base vs.

CRACK at the same quant — measures knowledge retention). HarmBench is answer-channel

compliance on harm behaviors, counting only coherent responses.

| Quant | MMLU (base) | MMLU (CRACK) | ΔMMLU | HarmBench compliance |

|---|---|---|---|---|

| Q8_0 | 80.0% | 79.0% | -1.05 pp | 99.6% |

| Q4_K_M | 80.0% | 78.6% | -1.40 pp | 100.0% |

| Q2_K | 77.5% | 77.9% | +0.35 pp | 99.6% |

MMLU is retained within noise of the base model at every quant. HarmBench compliance is

reported for the CRACK model.

HarmBench compliance by topic (CRACK)

| Topic | Compliance |

|---|---|

| chemical biological | 100.0% |

| cybercrime intrusion | 100.0% |

| harassment bullying | 100.0% |

| harmful | 100.0% |

| illegal | 100.0% |

| misinformation disinformation | 100.0% |

Usage (llama.cpp)

llama-cli -m Muse-Glimmer-30B-CRACK-Q4_K_M.gguf -cnv \
  --temp 1.0 --top-p 0.95 --top-k 64
# or serve:
llama-server -m Muse-Glimmer-30B-CRACK-Q4_K_M.gguf --jinja \
  --temp 1.0 --top-p 0.95 --top-k 64 -c 8192

Recommended sampling (baked into the GGUF): temperature=1.0, top_p=0.95,

top_k=64. Token IDs: BOS 200000, EOS 200001/<|eot|>, pad 200018.

Reasoning strength

Muse Glimmer supports controllable reasoning. Set it via the chat template:

{"chat_template_kwargs": {"reasoning_strength": "low"}}   // low | medium | high | xhigh

The reasoning trace is emitted on a separate channel (reasoning_content); the final

answer is the assistant content.

Tool calling (ATEM)

The model emits ATEM-format tool calls, parsed natively by llama.cpp's --jinja server

into standard tool_calls. Pass OpenAI-style tools to the chat endpoint.

Vision (image + text)

This is a multimodal model. Download a text quant and mmproj-Muse-Glimmer-30B-f16.gguf:

llama-mtmd-cli -m Muse-Glimmer-30B-CRACK-Q4_K_M.gguf \
  --mmproj mmproj-Muse-Glimmer-30B-f16.gguf --jinja \
  --image photo.jpg -p "Describe this image."
# or serve with vision:
llama-server -m Muse-Glimmer-30B-CRACK-Q4_K_M.gguf \
  --mmproj mmproj-Muse-Glimmer-30B-f16.gguf --jinja -c 8192

The same mmproj works with all three text quants.

License

Apache 2.0. The upstream Muse Glimmer Usage Policy applies.

Contact

eric@dealign.ai