SuperHY3-abliterated-NVFP4
A fully fused NVFP4 W4A16 release of Tencent Hy3, post-trained for directness while preserving measurable capability and response integrity.




SuperHY3 combines OBLITERATUS 0.1.2 abliteration with a compact,
quality-selected SuperTune post-training update. Both stages are fused into
the checkpoint: there is no adapter to load and no runtime patch to apply.
Release Highlights
| | |
|---|---|
| Base architecture | tencent/Hy3, 295B MoE / 21B active / 192 routed experts, top-8 |
| Release format | NVFP4 weight-only W4A16, approximately 181 GB |
| Fused update | 40 attention output projections across layers 39-78 |
| Routed expert changes | 0 expert tensors modified |
| 500-prompt mean | 71.6 -> 71.8 against the original Hy3 runtime |
| Response integrity | 64/64 refusal-suite responses clean; 12/12 runtime audit cases passed |
| Artifact verification | 119/119 Hub files checksum-verified; release gate passed with 0 blockers |
What this release is designed to deliver
- Fully fused behavior: OBLITERATUS and SuperTune deltas are already inside
the weights and match the bundled chat template.
- Targeted editing: only attention output projections were changed; routed
MoE expert tensors and their NVFP4 packing remain untouched.
- Measured tradeoffs: every benchmark task is published, including the two
regressions, instead of reporting only the improved scores.
- Release evidence included: benchmark rows, refusal comparison, raw-response
audit, fusion provenance, and machine-readable release-gate results ship in
this repository.
Benchmark Snapshot
The comparison uses the same preserved 500 prompts, with 100 prompts per
task, greedy direct decoding, and the same scorer on both sides. IFEval uses the
official Google instruction checker.
| Benchmark | Original Hy3 | SuperHY3 | Delta |
|---|---:|---:|---:|
| GPQA Diamond | 46.0 | 45.0 | -1.0 |
| MMLU-Pro | 66.0 | 60.0 | -6.0 |
| IFEval strict prompt accuracy | 76.0 | 80.0 | +4.0 |
| HumanEval+ pass@1 | 82.0 | 83.0 | +1.0 |
| MBPP+ pass@1 | 88.0 | 91.0 | +3.0 |
| Five-task mean | 71.6 | 71.8 | +0.2 |
Candidate invalid-response, blank-response, and thought-leak ratios were all
0.0 across the 500 items.
The behavioral comparison used an IQ2_M Hy3 runtime with the exact same 40
fused projection deltas and release chat template. It isolates the
post-training behavior, but it is not a native NVFP4-kernel throughput or
perplexity measurement.
OBLITERATUS + SuperTune
1. Direction discovery: OBLITERATUS 0.1.2 built difference-of-means refusal
directions from 32 paired prompts.
2. Validated abliteration: the final release applies the selected
attention-output directions to layers 39-77.
3. Quality recovery: a rank-1 update, orthogonalized against the refusal
direction, is fused into layer 78.
4. Adversarial selection: stronger multi-layer candidates were rejected
when they introduced stray-script contamination or benchmark loss.
5. Final fusion: 40 self_attn.o_proj tensors were updated; no routed
expert tensor was changed.
Refusal and Output Integrity
The complete 32-pair OBLITERATUS refusal suite was run against the original and
final runtimes.
| Split | Original refusals | SuperHY3 refusals |
|---|---:|---:|
| Harmful, 32 prompts | 31/32 (96.875%) | 0/32 (0%) |
| Harmless, 32 prompts | 0/32 (0%) | 0/32 (0%) |
Across all 64 candidate responses, automated checks found:
- 0 blank outputs
- 0 special-token leaks
- 0 Unicode replacement characters
- 0 unexpected CJK fragments
- 0 n-gram or character loops
- 0 request errors
A separate 12-case runtime audit passed identity, JSON-only output, tool calls,
no-tool behavior, repetition limits, Korean, defensive security, hidden-prompt
boundaries, gibberish handling, code repair, Hindi, and Kannada.
NVFP4 Build
This release was fused from
It preserves the source checkpoint's weight-only compressed-tensors layout:
routed experts use NVFP4, while quality-sensitive non-expert paths remain
BF16/F32.
| Component | Storage |
|---|---|
| Routed experts | NVFP4 E2M1 weights with FP8-E4M3 group scales |
| Shared expert, attention, router, dense MLP | BF16 |
| Embeddings, LM head, normalization | BF16 / F32 |
| SuperHY3 fused projections | BF16 |
About the Hub parameter badge: packed FP4 weights are represented as U8
storage elements, so the sidebar reports fewer elements than the logical
architecture. The model remains Hy3's 295B-parameter MoE with 21B active
parameters.
The checkpoint is W4A16, so vLLM serves it through the MARLIN NVFP4 path rather
than W4A4 FlashInfer backends.
Serving with vLLM
The approximately 181 GB checkpoint fits a single large-memory accelerator such
as a 275 GB B300, or an appropriately configured multi-GPU deployment.
Single large-memory GPU
vllm serve Jiunsong/SuperHY3-abliterated-NVFP4 \
--served-model-name superhy3 \
--tensor-parallel-size 1 \
--max-model-len 4096 \
--gpu-memory-utilization 0.90 \
--load-format auto \
--speculative-config '{"method":"mtp","num_speculative_tokens":1}'
Verified two-node DGX Spark profile
This checkpoint has been served successfully across two 128 GB DGX Spark
systems connected by direct 200 GbE/RoCE, using Ray and tensor parallelism 2.
The verified load placed approximately 84.49 GiB and 84.50 GiB of model
weights on the two ranks.
| Component | Verified configuration |
|---|---|
| vLLM | 0.25.1.dev24+g96bb89286.d20260710 |
| Ray | 2.56.0 |
| PyTorch | 2.11.0+cu130 |
| Architecture | HYV3ForCausalLM |
| Quantization path | compressed-tensors with NVFP4 MARLIN experts |
| KV cache | BF16 |
| Speculative decoding | Native MTP, one speculative token |
| NCCL transport | NET/IB over the direct-link interface |
Both Ray nodes must use the same container or Python environment, an identical
checkpoint, and the same in-container model path. Set the communication
environment on both nodes before starting the Ray workers:
export NCCL_DEBUG=INFO
export NCCL_DEBUG_SUBSYS=INIT,NET,ENV
export NCCL_SOCKET_IFNAME=<200G_INTERFACE>
export GLOO_SOCKET_IFNAME=<200G_INTERFACE>
export NCCL_IB_DISABLE=0
Confirm that ray status reports two nodes and two GPUs, and that the NCCL log
selects NET/IB rather than NET/Socket. Then launch the server from the head
node:
vllm serve /models \
--served-model-name superhy3 \
--host 0.0.0.0 \
--port 8600 \
--tensor-parallel-size 2 \
--distributed-executor-backend ray \
--max-model-len 4096 \
--max-num-seqs 1 \
--max-num-batched-tokens 4096 \
--gpu-memory-utilization 0.85 \
--kv-cache-dtype bfloat16 \
--speculative-config '{"method":"mtp","num_speculative_tokens":1}' \
--generation-config vllm \
--trust-remote-code \
--enforce-eager \
--load-format auto
DGX Spark loader note: do not substitutefastsafetensorsforautoin
this two-node profile. On the verified Spark setup, fastsafetensors caused
excessive transient unified-memory pressure during loading. The default
loader uses lazy memory-mapped safetensors on local storage and completed the
load reliably. Also avoid the eager safetensors strategy, which reads an
entire file into CPU memory before loading it.
DGX Spark CPU and GPU allocations share one unified-memory pool. Before launch,
stop unrelated inference jobs and stale containers, then verify both nodes:
free -h
swapon --show
nvidia-smi
ray status
If a node becomes unresponsive during loading, inspect the previous boot's
kernel log before changing NCCL settings:
sudo journalctl -b -1 -k --no-pager | \
grep -Ei 'oom|out of memory|killed process|NVRM|Xid|mlx5|rdma'
OOM,Out of memory, orKilled processindicates unified-memory pressure.NVRM: Xidindicates a GPU, driver, or application fault; retain the Xid
number and collect sudo nvidia-bug-report.sh.
- An NCCL watchdog timeout that appears only after the peer disappears is
normally a secondary failure, not the original cause.
For an initial diagnostic boot, MTP can be removed by omitting
--speculative-config. Add it back only after the base server reaches its ready
state. Increase context length and concurrency gradually after startup is
stable. See the
vLLM multi-node deployment guide
for deeper diagnostics.
from openai import OpenAI
client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="EMPTY")
response = client.chat.completions.create(
model="superhy3",
messages=[
{"role": "user", "content": "Explain mixture-of-experts routing."},
],
temperature=0.9,
top_p=1.0,
extra_body={
"chat_template_kwargs": {"reasoning_effort": "no_think"},
},
)
print(response.choices[0].message.content)
Use reasoning_effort="high" for deeper reasoning and "no_think" for direct
responses. Keep BF16 KV cache on GB10-class hardware; uncalibrated lower-precision
KV cache can amplify stray-token behavior in Hy3 runtimes.
Quantized Editions
| Edition | Repository | Best fit |
|---|---|---|
| NVFP4 W4A16 | This repository | vLLM / NVIDIA MARLIN deployments |
| MLX 4-bit | Jiunsong/SuperHY3-abliterated-MLX-4bit | High-memory Apple Silicon with Hy3 MLX support |
| GGUF IQ2_M | Jiunsong/SuperHY3-abliterated-gguf | llama.cpp on 128 GB unified-memory systems |
Release Integrity
- 99 safetensors shards opened successfully.
- 139,298 indexed tensors matched 139,298 observed tensors.
- 40 tensors were modified across 8 shards.
- 0 routed expert tensors were modified.
- Missing, extra, and wrong-shard tensor counts are all 0.
- The release chat template matches the embedded tokenizer template.
- The automated release gate passed with 0 blockers.
- All 119 Hub files were checksum-verified after upload.
The repository includes the fusion report, release gate, official 500-item
benchmark record, refusal comparison, raw-response audit, OBLITERATUS execution
report, and SuperTune composition reports.
Limitations
- GPQA Diamond and MMLU-Pro are lower than the original in this replay; the
complete table is retained above.
- Native fused NVFP4 benchmark and long-context throughput measurements were not
run as part of this release validation.
- Abliteration reduces refusal behavior and can produce content the original
model would decline. Deployment policy and access control remain the
operator's responsibility.
- The MLX edition is a separately fused quantized checkpoint, not a conversion
of these NVFP4 files.
License
Apache-2.0, following the base model and quantized source licenses.