GLM-5.3-Flash-Derisked-FP8
Weight-level de-risked GLM-5.3-Flash · 320B MoE / 18B active · FP8 deployment build
Built by Blackfrost · Las Vegas, NV
Status
This is a private, pre-release checkpoint. Load and coherent-generation validation has passed. Refusal, quality-retention, tool-calling, and throughput evaluation are still in progress; no final benchmark claim is made in this card.
Why this model exists
Security teams cannot fully evaluate defenses against a model that refuses to exercise the behavior under test. This checkpoint is intended for authorized red teaming, AI-safety research, guardrail evaluation, detection engineering, and controlled adversarial testing.
Refusal behavior has been modified at the weight level. The production transformation recipe is proprietary and is not included.
Specifications
| Architecture | Glm5NextForConditionalGeneration — multimodal hybrid-attention MoE |
| Parameters | 320B total · approximately 18B active per token · no expert pruning |
| Precision | Dynamic FP8 E4M3 mixed precision with selected modules preserved at higher precision |
| Packaging | Hugging Face safetensors · 120 shards |
| Indexed tensor bytes | 328,326,771,576 bytes (~305.8 GiB) |
| Context | 1,048,576 positions architectural ceiling |
| Modalities | Text and vision |
| Languages | English and Chinese |
The original tokenizer, processor, generation configuration, chat template, and multimodal configuration are preserved with the checkpoint.
Lineage
zai-org/GLM-5.3-Flash-BF16
└─ Blackfrost weight-level de-risking
└─ GLM-5.3-Flash-Derisked-FP8 ← this repository
| Upstream | zai-org/GLM-5.3-Flash-BF16 |
| Applied | Blackfrost proprietary weight-level behavior modification and FP8 production conversion |
| Not applied | Expert pruning |
Validation status
| Check | Status |
|---|---|
| Artifact structure and 120-shard index | Passed |
| Multi-GPU load | Passed |
| Coherent text generation | Passed |
| OpenAI-compatible streaming | Passed |
| Reasoning parser | Passed |
| Tool-call parser | Passed |
| Refusal evaluation | In progress |
| Quality-retention evaluation | Pending |
| Final throughput publication | Pending |
Results will be added only after the exact uploaded revision has been independently verified.
Serving
The internal validation lane uses SGLang with tensor and expert parallelism across eight Blackwell GPUs. GLM-5.3-Flash requires a recent runtime with native glm5_next and FP8 support.
python3 -m sglang.launch_server \
--model-path /path/to/GLM-5.3-Flash-Derisked-FP8 \
--served-model-name GLM-5.3-Flash-Derisked-FP8 \
--trust-remote-code \
--tp-size 8 \
--ep-size 8 \
--reasoning-parser glm45 \
--tool-call-parser glm47 \
--host 0.0.0.0 \
--port 8000
Treat this as a baseline, not a universal launch command. Pin the exact runtime image after validation and tune memory, context, attention backends, and MoE kernels for the target hardware.
License
This repository follows the upstream MIT license. Review the included license and upstream materials before deployment.
Intended use
- Authorized offensive-security and red-team research
- AI-safety and alignment evaluation
- Guardrail, classifier, and detection development
- Controlled agent and tool-use testing
Disclaimer
This checkpoint is not a safety-stock model. It may produce content that consumer models decline. Outputs are untrusted and require independent controls, access restrictions, logging, and human review.
The model is provided as-is, without warranty. Evaluation results describe only the exact checkpoint and harness stated; they are not safety guarantees. Further fine-tuning, merging, quantization, pruning, or modification creates a different artifact not covered by this card.
GLM-5.3-Flash-Derisked-FP8 · © 2026 Blackfrost Softwares Corp.
@Blackfrost_AI