Abiray/Muse-Glimmer-30B-GGUF

🤗 Hugging Face sourceimage-text-to-textapache-2.030B activated178 GBGGUF✓ 12 checksumsupdated today
Submit in one command

Run it next to your model folder. It makes the torrent, checks your files against Hugging Face, and submits it. You just start seeding and paste your key from your account. It only reads your files and never changes them. Read the script first if you like.

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo Abiray/Muse-Glimmer-30B-GGUF ./model-folder
Needs a seeder →

Muse-Glimmer-30B - GGUF Quants

This repository contains GGUF quants for meta-models/Muse-Glimmer-30B, created using llama.cpp. Muse Glimmer is a 30-billion-parameter multimodal model distilled from Muse Spark, purpose-built for autonomous agentic workflows, long-horizon multi-step reasoning, SWE-bench coding tasks, and reliable function calling on consumer hardware.


File Availability & Recommended Hardware

To run vision inputs, download one core model file (.gguf) along with one multimodal vision projector (mmproj-*.gguf).

Core Model Files

File Name Size Quantization Rec. Memory / VRAM Description
Muse-Glimmer-30B-Q8_0.gguf 29.6 GB Q8_0 32 GB - 48 GB Maximum precision. Virtually lossless retention compared to BF16.
Muse-Glimmer-30B-Q6_K.gguf 22.9 GB Q6_K 28 GB - 32 GB Near-lossless output precision. Great for 32GB system/VRAM setup.
Muse-Glimmer-30B-Q5_K_M.gguf 19.8 GB Q5_K_M 24 GB High Quality balance. Ideal fit for GPUs with 24GB VRAM (e.g., RTX 3090/4090/5090).
Muse-Glimmer-30B-Q4_K_M.gguf 16.9 GB Q4_K_M 20 GB - 24 GB Recommended Sweet Spot. Optimal trade-off between speed, memory, and reasoning capacity.
Muse-Glimmer-30B-IQ4_NL.gguf 16.1 GB IQ4_NL 20 GB Non-Linear 4-bit importance matrix quantization. Strong performance under 17GB.
Muse-Glimmer-30B-IQ4_XS.gguf 15.3 GB IQ4_XS 18 GB - 20 GB Extra-small 4-bit iQuant for constrained VRAM environments.
Muse-Glimmer-30B-Q3_K_M.gguf 14.0 GB Q3_K_M 16 GB - 18 GB Standard 3-bit K-quant. Good option for 16GB VRAM cards.
Muse-Glimmer-30B-IQ3_M.gguf 13.1 GB IQ3_M 16 GB 3-bit medium importance quant with better reasoning recovery than baseline Q3.
Muse-Glimmer-30B-IQ3_XS.gguf 12.3 GB IQ3_XS 14 GB - 16 GB 3-bit extra-small iQuant for lower memory targets.
Muse-Glimmer-30B-IQ3_XXS.gguf 11.5 GB IQ3_XXS 12 GB - 16 GB Highly compressed 3-bit iQuant. Fits tight memory budgets.

Vision Projector Files (mmproj)

File Name Size Precision Usage
mmproj-Muse-Glimmer-30B-BF16.gguf 3.85 GB BF16 Full-precision ~1.8B ViT perception projector for maximum image fidelity.
mmproj-Muse-Glimmer-30B-Q8_0.gguf 2.05 GB Q8_0 Recommended. 8-bit quantized vision projector preserving high image understanding at nearly half the RAM.

Quickstart & Usage

1. llama.cpp CLI (With Vision Support)

To run the model with multimodal vision capability:

# Start server with vision support
llama-server \
  -m Muse-Glimmer-30B-Q4_K_M.gguf \
  --mmproj mmproj-Muse-Glimmer-30B-Q8_0.gguf \
  -c 131072 \
  --port 8080