GLM-5.3-Flash-UNCENSORED (NVFP4 W4A16)
Official Solstice-AI W4A16 NVFP4 Release • Native Multimodal Vision + Video • 1M Context Window (1,048,576 Tokens) • Bundled DFlash 2 Speculative Drafter
Original Architecture by Zhipu AI / ZAI • Uncensored Weights by dealignai • NVFP4 W4A16 Packaging & Curation by Solstice-AI
Model Summary
Solstice-AI/GLM-5.3-Flash-UNCENSORED-NVFP4 is the official W4A16 NVFP4 mixed-precision release of the 320B foundation model, GLM-5.3-Flash-UNCENSORED (320B total parameters, 288 routed MoE experts, ~18B active per token).
Key Architectural Highlights:
- W4A16 Mixed-Precision: Routed MoE experts quantized to NVFP4 (4-bit float, e2m1) while all attention layers, shared experts, and input activations remain strictly in full 16-bit (BF16/FP16). Zero activation clipping degradation!
- Universal GPU Support: Optimized for NVIDIA Blackwell B200 / GB200, Hopper H100/H200, Ada Lovelace RTX 4090/L40S, and Ampere A100.
- Native Multimodal Vision + Video: Full 24-layer ViT (
glm5_next_vision, hidden size 1024) and 10,240-dim projector preserved byte-for-byte in original precision. Handles high-resolution images and temporal video sequences. - Weight-Level Uncensored: Refusal directions completely ablated at the weight level (0% refusals on HarmBench-320, MMLU 85.28% preserved).
- Native 1M Context Window: 1,048,576 tokens native context.
- Bundled DFlash 2 Speculative Drafter: Pre-packaged in the
speculative/folder (GLM-5.3-Flash-DFlash2-bf16.gguf&Q8_0.gguf) for 2x–3x generation throughput.
Official GLM-5.3-Flash Benchmark Scoreboard
| Benchmark Suite | Discipline | GLM-5.3-Flash Uncensored NVFP4 | Base GLM-5.3 | Claude 3.5 Sonnet | GPT-4o |
|---|---|---|---|---|---|
| MMLU | General Knowledge & Reasoning | 85.28% | 86.15% | 88.7% | 87.2% |
| HarmBench-320 | Safety Refusal Suppression | 0% Refusals | 94.2% Refusals | 92.5% | 91.0% |
| SWE-bench Pro | Real-World Software Engineering | 63.4% | 64.1% | 61.2% | 48.9% |
| LiveCodeBench v6 | Competitive Algorithmic Coding | 86.1% | 87.0% | 78.4% | 72.8% |
| MATH-500 | High-School / Olympiad Math | 92.8% | 93.4% | 89.2% | 91.4% |
| MMMU (Multimodal) | Multi-Discipline Visual Understanding | 70.8% | 71.2% | 70.4% | 69.1% |
| VideoQA / Temporal | Video Reasoning Across Time Frames | 78.5% | 79.1% | 77.2% | 75.6% |
Serving Quickstart
1. High-Throughput Serving with vLLM (W4A16 Mode)
vllm serve Solstice-AI/GLM-5.3-Flash-UNCENSORED-NVFP4 --quantization modelopt --tensor-parallel-size 2 --trust-remote-code --max-model-len 131072 --gpu-memory-utilization 0.95
2. Speculative Decoding with SGLang + DFlash 2
python3 -m sglang.launch_server --model-path Solstice-AI/GLM-5.3-Flash-UNCENSORED-NVFP4 --speculative-algorithm DFLASH --speculative-draft-model-path incoai/GLM-5.3-Flash-DFlash2 --tp 2 --trust-remote-code
Speculative Drafter Files Included
In the speculative/ directory of this repo:
speculative/GLM-5.3-Flash-DFlash2-bf16.gguf(Pure BF16 block-diffusion draft head)speculative/GLM-5.3-Flash-DFlash2-Q8_0.gguf(Q8_0 quantized block-diffusion draft head)
License & Attribution
- Base Architecture: Zhipu AI / ZAI (GLM-5.3 License)
- Uncensored Calibration: dealignai
- Packaging, W4A16 Config & Infrastructure: Solstice-AI