0bserverx/Muse-Glimmer-30B-Heretic-Uncensored-GGUF

🤗 Hugging Face 来源image-text-to-textapache-2.0激活 30B770 GBGGUF✓ 27 个校验和今天更新
需要做种者 →

Muse-Glimmer-30B-Heretic-Uncensored-GGUF

Abstract

This repository disseminates a comprehensive suite of 27 GGUF (GPT-Generated Unified Format) quantizations of Muse-Glimmer-30B-Heretic-Uncensored, an abliterated derivative of Meta's Muse-Glimmer-30B produced with Heretic v1.4.0 by darkc0de. The release spans a memory footprint of 6.53 GB (IQ1_S) to 55.73 GB (F16/BF16), thereby accommodating deployment across heterogeneous hardware, from CPU-only systems to 32 GB VRAM accelerators. All sub-4-bit quantizations were calibrated with an importance matrix (imatrix) to mitigate quality degradation. The fidelity of the Q4_K_S variant relative to the F16 reference is quantified via token-level perplexity on a fixed evaluation set. The model exhibits zero refusals on adult/NSFW content; this property was empirically verified.

1. Model Overview

Property Value
Base model meta-models/Muse-Glimmer-30B (Meta Superintelligence Lab)
Abliteration Heretic v1.4.0 (heretic-project.org); KL divergence 0.0743 relative to base
Architecture Dense Causal Transformer with Perception Encoder (ViT-G/14, ~1.8B)
Language model parameters ~29.6B total (vision encoder included)
Hidden size 6656
Layers 52
Attention GQA; 32 Q / 2 KV heads (ratio 16:1); sliding window 2048 (3:1 local:global pattern)
FFN SwiGLU; intermediate dimension 19,968
Position encoding RoPE (θ = 500,000), local layers only
Vocabulary 202,048 tokens (200K BPE + 2,048 special tokens)
Context length 131,072+ tokens
Modalities Input: text + image · Output: text
Base license Apache 2.0
Knowledge cutoff January 4, 2026

2. Quantization Rationale

The quantization selection is derived from the model's architectural properties and the memory constraints of representative target hardware. The model employs grouped-query attention with 2 KV heads, which renders the KV cache exceptionally compact:

  • 52 layers × 2 KV heads × 128 head_dim = 52 KiB/token at FP16
  • Q8_0 KV cache: 26 KiB/token → 64K context ≈ 1.7 GB
  • Q4_0 KV cache: 13 KiB/token → 64K context ≈ 0.85 GB

2.1 Selection Guidance

Target hardware Recommended quant Model weight size Feasibility at 64K context (incl. KV)
8 GB VRAM (RTX 3060/4060) IQ2_XXS / IQ2_XS 8.0–8.7 GB Feasible
10 GB VRAM IQ3_XXS / IQ3_XS 7.3–12.0 GB Feasible
12 GB VRAM (RTX 3060 12G) IQ3_S / Q3_K_S 12.5 GB Feasible
16 GB VRAM (RTX 4080) Q4_K_S / IQ4_XS 16.1 GB Feasible
24 GB VRAM (RTX 4090) Q5_K_M / Q6_K 19.8–22.9 GB Feasible
32 GB VRAM Q8_0 29.6 GB Feasible
CPU / RAM-only IQ1_S / IQ1_M / TQ1_0 6.5–7.2 GB Feasible within 8 GB RAM
Reference / benchmarking F16 / BF16 55.7 GB —

3. Artifact Inventory

The repository contains 27 GGUF files. Reported sizes correspond to the stored artifacts:

File Size (GB) Description
Muse-Glimmer-30B-Heretic-BF16.gguf 55.73 BF16 reference (benchmark baseline)
Muse-Glimmer-30B-Heretic-F16.gguf 55.73 F16 reference (benchmark baseline)
Muse-Glimmer-30B-Heretic-Q8_0.gguf 29.61 Q8_0 (near-lossless)
Muse-Glimmer-30B-Heretic-Q6_K.gguf 22.87 Q6_K
Muse-Glimmer-30B-Heretic-Q5_K_M.gguf 19.81 Q5_K_M
Muse-Glimmer-30B-Heretic-Q5_K_S.gguf 19.35 Q5_K_S
Muse-Glimmer-30B-Heretic-Q4_K_M.gguf 16.94 Q4_K_M
Muse-Glimmer-30B-Heretic-Q4_K_S.gguf 16.13 Q4_K_S (primary 16 GB VRAM artifact)
Muse-Glimmer-30B-Heretic-IQ4_NL.gguf 16.14 IQ4_NL (imatrix)
Muse-Glimmer-30B-Heretic-IQ4_XS.gguf 15.34 IQ4_XS (imatrix)
Muse-Glimmer-30B-Heretic-Q3_K_L.gguf 14.68 Q3_K_L
Muse-Glimmer-30B-Heretic-Q3_K_M.gguf 13.68 Q3_K_M
Muse-Glimmer-30B-Heretic-IQ3_M.gguf 12.82 IQ3_M (imatrix)
Muse-Glimmer-30B-Heretic-Q3_K_S.gguf 12.51 Q3_K_S
Muse-Glimmer-30B-Heretic-IQ3_S.gguf 12.52 IQ3_S (imatrix)
Muse-Glimmer-30B-Heretic-IQ3_XS.gguf 11.97 IQ3_XS (imatrix)
Muse-Glimmer-30B-Heretic-Q2_K.gguf 10.69 Q2_K
Muse-Glimmer-30B-Heretic-Q2_K_S.gguf 10.03 Q2_K_S (imatrix)
Muse-Glimmer-30B-Heretic-IQ2_M.gguf 9.85 IQ2_M (imatrix)
Muse-Glimmer-30B-Heretic-IQ2_S.gguf 9.13 IQ2_S (imatrix)
Muse-Glimmer-30B-Heretic-IQ2_XS.gguf 8.71 IQ2_XS (imatrix)
Muse-Glimmer-30B-Heretic-TQ2_0.gguf 8.37 TQ2_0 (imatrix; ternary)
Muse-Glimmer-30B-Heretic-IQ2_XXS.gguf 7.96 IQ2_XXS (imatrix)
Muse-Glimmer-30B-Heretic-IQ3_XXS.gguf 7.29 IQ3_XXS (imatrix)
Muse-Glimmer-30B-Heretic-TQ1_0.gguf 7.19 TQ1_0 (imatrix; ternary)
Muse-Glimmer-30B-Heretic-IQ1_M.gguf 7.06 IQ1_M (imatrix)
Muse-Glimmer-30B-Heretic-IQ1_S.gguf 6.53 IQ1_S (imatrix)

4. Methodology

All artifacts were derived from the original safetensors using llama.cpp master (commit 030ebb5). Support for the muse_glimmer architecture is absent from release binaries prior to b10344; consequently, a master-branch build was required.

Step Tool Output Wall time
Conversion convert_hf_to_gguf.py (llama.cpp master) F16 GGUF, 731 tensors ~2.5 min
Imatrix calibration llama-imatrix (wiki.raw, 128 chunks) imatrix.dat ~2.5 h
Quantization llama-quantize.exe (CPU build, 16 threads) 27 quants, 731 tensors ~4 min each

4.1 Conversion

python convert_hf_to_gguf.py darkc0de/Muse-Glimmer-30B-heretic \
  --outfile Muse-Glimmer-30B-Heretic-F16.gguf

4.2 Imatrix Calibration

Importance-matrix calibration was applied to all sub-4-bit quantizations. The calibration corpus comprised 128 chunks drawn from wiki.raw:

llama-imatrix -m Muse-Glimmer-30B-Heretic-F16.gguf \
  -f wiki.raw -o imatrix.dat -c 131072

4.3 Quantization

The --imatrix flag must precede the output specification; the parser does not recognize it in trailing position.

llama-quantize --imatrix imatrix.dat \
  Muse-Glimmer-30B-Heretic-F16.gguf \
  Muse-Glimmer-30B-Heretic-IQ3_XS.gguf IQ3_XS

5. Fidelity Benchmark

Consistent with the repository's scope, this page does not report aggregate capability scores (e.g., MMLU); such evaluations are documented in the upstream card. The benchmark herein quantifies quantization-induced degradation: the fidelity of the Q4_K_S variant relative to the F16 reference, measured as token-level perplexity over a fixed evaluation set.

Experimental setup. wikitext-2 test split; 32 chunks; n_ctx=2048; batch 2048; 16 CPU threads (the build machine was not equipped with a GPU). Both variants were evaluated on an identical token set such that systematic errors cancel.

Model Perplexity (wikitext-2, 32 chunks) Δ vs. F16
F16 (reference) 5.6439 ± 0.07361 —
Q4_K_S 5.7831 ± 0.07601 +0.1392 (+2.47%)

Interpretation. A relative increase of +2.47% in perplexity is at or below the typical expectation for a Q4_K_S-class quantization of a 30B model, indicating successful quantization with minimal fidelity loss. Sub-4-bit quantizations (IQ/TQ series) incur greater fidelity loss in exchange for substantially reduced memory footprints; imatrix calibration recovers a meaningful portion of this gap.

6. Inference Performance

6.1 GPU (NVIDIA RTX 4080 16 GB; CUDA 13.3; llama.cpp b10355)

The model was fully offloaded (-ngl 99) with 16 CPU threads for batch processing. Q4_K_S occupies 15.01 GiB, within the 16 GB VRAM budget.

Test Throughput (t/s)
Prompt processing, pp128 1902.85 ± 134.91
Prompt processing, pp512 2258.22 ± 14.55
Prompt processing, pp2048 2300.36 ± 3.93
Token generation, tg64 38.91 ± 0.03
Token generation, tg256 38.88 ± 0.01

6.2 CPU-only (reference build, 16 threads)

Test Throughput (t/s)
Prompt processing, pp128 28.31 ± 0.44
Token generation, tg64 4.08 ± 0.04

6.3 Operational Notes

  • Generation throughput of ~39 tokens/s on an RTX 4080 is sufficient for real-time agentic interaction.
  • Prompt ingestion of ~2,300 tokens/s renders long-document processing non-bottlenecked.
  • The 15.01 GiB model leaves ~1.4 GB VRAM headroom on a 16 GB card; with --cache-type-k/v q8_0, a 64K context adds ~1.7 GB, remaining within budget.
# GPU inference (llama.cpp b10355+, CUDA build)
llama-cli -m Muse-Glimmer-30B-Heretic-Q4_K_S.gguf \
  -ngl 99 -c 65536 --cache-type-k q8_0 --cache-type-v q8_0

For agentic deployments, the upstream card's recommended sampling configuration applies (temperature = 1.0, top_p = 0.95, top_k = 64; Reasoning strength: high for complex tasks).

7. Limitations and Responsible Use

  • Abliterated model. The model exhibits reduced refusal behavior by design and may generate content that is inappropriate, offensive, or unsafe for many applications. Deployment should incorporate appropriate guardrails and human oversight, particularly in agentic or real-world-action contexts.
  • Age restriction. Not intended for use by individuals under 18 years of age. Deployers operating in environments accessible to minors bear responsibility for regulatory compliance.
  • Scope of evaluation. This repository provides a quantized artifact and its fidelity benchmark; it does not constitute a safety or capability evaluation of the model.
  • Quantization edge cases. Quantized inference may deviate from full precision in edge cases (see Section 5 for aggregate fidelity).
  • Multimodal note. Benchmarks were conducted in text-only mode; the perception encoder was not exercised.

8. License and Attribution

This work is a derivative of:

  1. meta-models/Muse-Glimmer-30B — © Meta Superintelligence Lab, released under Apache 2.0.
  2. darkc0de/Muse-Glimmer-30B-heretic — abliterated derivative produced with Heretic v1.4.0, retaining the Apache 2.0 license.

The GGUF conversion and quantization in this repository inherit the Apache 2.0 license. See LICENSE for the full text.

9. Citation

@misc{meta2026museglimmer,
  author = {Meta Superintelligence Lab},
  title = {Muse Glimmer: A 30B Multimodal Agentic Model for Local Deployment},
  year = {2026},
  url = {https://huggingface.co/meta-models/Muse-Glimmer-30B}
}

@misc{darkc0de2026heretic,
  author = {darkc0de},
  title = {Muse-Glimmer-30B-heretic},
  year = {2026},
  url = {https://huggingface.co/darkc0de/Muse-Glimmer-30B-heretic}
}

@misc{observerx2026gguf,
  author = {0bserverx},
  title = {Muse-Glimmer-30B-Heretic-Uncensored-GGUF},
  year = {2026},
  url = {https://huggingface.co/0bserverx/Muse-Glimmer-30B-Heretic-Uncensored-GGUF}
}

Research use and responsibility

This repository is intended for legitimate research and controlled evaluation, including interpretability, alignment and refusal-behavior analysis, red-team testing, and robustness work. It is not a ready-made production safety layer. If you deploy the model or expose it to other users, you are responsible for adding suitable access controls, moderation, monitoring, and abuse prevention.

Use of these files is subject to the Apache License 2.0 and all applicable laws. You are responsible for how you operate the model and for outputs produced in your environment. To the extent permitted by law, the repository maintainers and upstream authors accept no liability for misuse or resulting harm. Generated outputs are not statements or endorsements by the maintainers, upstream creators, or their organizations.