Muse-Glimmer-30B-Hermes-Agentic-NVFP4
NVFP4 (compressed-tensors / nvfp4-pack-quantized) export of the Hermes-agentic Muse Glimmer 30B student.
- BF16 / FP16 merge: vcruz305/Muse-Glimmer-30B-Hermes-Agentic
- GGUF ladder: vcruz305/Muse-Glimmer-30B-Hermes-Agentic-GGUF
- Serve recipes (SGLang + vLLM on DGX Spark SM121): github.com/vcruz305/muse-glimmer-nvfp4-spark-serve
- Base: meta-models/Muse-Glimmer-30B (Apache-2.0)
Format
quantization_config.format:nvfp4-pack-quantizedquant_method:compressed-tensors- ~22 GB on disk (2 safetensor shards)
- Vision / projector /
lm_head/ linear-attn / MTP paths ignored in the NVFP4 recipe (seerecipe.yaml)
Serving (preferred: SGLang on Spark)
hf download vcruz305/Muse-Glimmer-30B-Hermes-Agentic-NVFP4 \
--local-dir $HOME/models/Muse-Glimmer-30B-Hermes-Agentic-NVFP4
export MODEL=$HOME/models/Muse-Glimmer-30B-Hermes-Agentic-NVFP4
# from https://github.com/vcruz305/muse-glimmer-nvfp4-spark-serve
bash scripts/launch_sglang_nvfp4_spark.sh
Critical SM121 flags (do not leave on auto):
--fp4-gemm-backend flashinfer_cutlass--cuda-graph-backend-decode full--cuda-graph-backend-prefill disabled--reasoning-parser muse+--tool-call-parser muse--language-model-only
vLLM secondary path is in the same recipe repo (scripts/launch_vllm_nvfp4.sh).
Evaluation (sixcat, SGLang, thinking on)
Host: DGX Spark GB10 · endpoint OpenAI-compatible · policy strict · thinking on · limit 20 · sixcat ≥0.4.4 · full 120 after --retry remaining.
| Category | Score |
|---|---|
| knowledge | 75.0 |
| math | 100.0 |
| truth | 85.0 |
| instruct | 90.0 |
| code | 85.0 |
| tools | 80.0 |
| overall[strict] | 85.8 |
Suite throughput ≈ 11.9 tok/s on that run. Instruct/code still carried length/loop flags.
Model Details
Hermes-agentic SFT of Muse Glimmer so the model calls one or two tools and stops. Vision weights remain in the tree but were frozen for the text/tools student. This repo is the NVFP4 deployable of that merge — not a second SFT.
- Developed by: Victor Cruz (vcruz305)
- License: Apache-2.0
- Quant scheme: NVFP4 via compressed-tensors (
recipe.yamlin-repo)
How to Get Started
- Download this repo.
- Follow muse-glimmer-nvfp4-spark-serve.
- Hit
http://127.0.0.1:8201/v1with model idmuse-glimmer-30b-nvfp4.