0bserverx/Qwen3.8-27B-Heretic-Abliterated-Uncensored

🤗 Hugging Face 来源text-generationapache-2.026.9B 参数54 GBsafetensors✓ 80 个校验和今天更新
需要做种者 →

RVN Qwen3.8-27B — FP16 safetensors

The full text-model weights, in Hugging Face safetensors format, for vLLM/SGLang users. This repository contains all 13 main-weight shards, the model index, configuration, tokenizer, repaired chat template, reconstruction code, and verification evidence. It is not a LoRA adapter or a weights-free placeholder.

Why this release exists

This export was made at the request of our community member JC1DA in “FP16 safetensor models?” — discussion #10:

Can you also release the fp16 base model so vllm/sglang people can run it as well?

Thanks to JC1DA for the request, and to everyone who expressed interest in an FP16 release.

Our main RVN repository is Qwen3.8-27B-Heretic-Abliterated-Uncensored-GGUF. Please use that repository for the GGUF quantizations, original RVN background, and the main discussion history. This is its safetensors companion release, reconstructed from its existing RVN-F16.gguf — not a new training or abliteration run.

What is included — and what is not

  • 13 safetensors shards; 851 text-model tensors. All main text weights, including embeddings and the untied language-model head, are present.
  • On-disk storage preserves 498 F16 tensors and 353 F32 tensors, rather than downcasting the stability tensors. Inference engines may cast tensors according to the requested loading dtype.
  • Tensor payload: 53,797,287,936 bytes, excluding file headers and supporting assets.
  • Model configuration, generation configuration, tokenizer, and the source repository's repaired chat template.
  • The supported architecture is Qwen3_5ForCausalLM / qwen3_5_text, despite the Qwen3.8 family name.
  • Text-only. No vision encoder/projector and no MTP head. Those weights are absent from this source F16 GGUF. We did not fabricate them or splice in a different checkpoint's weights.
  • This is not the original upstream BF16 checkpoint and is not an exact recovery of the unavailable pre-GGUF RVN checkpoint.

Derived quantizations (GSQ & GGUF)

  • GSQ-RCO series (dedicated repository) - non-uniform per-tensor GGUF quantizations (4 tiers + MTP twins, 8.45-11.80 GB): Qwen3.8-27B-Heretic-GSQ-RCO-GGUF. Measured on wikitext-2: IQ3_S PPL 6.1778 vs the 6.1197 F16 reference (+0.95%), ahead of a same-class uniform quantization (+2.9%) at a smaller size.
  • Standard GGUF quants - the full RVN quant spectrum: Qwen3.8-27B-Heretic-Abliterated-Uncensored-GGUF.
  • GSQ-3bit (compressed-tensors) — a 3-bit Gumbel-Softmax PTQ variant (MLP linears) for vLLM/Transformers serving lives in this repository under GSQ-3bit/ (loads with transformers >= 5.8 + compressed-tensors >= 0.15).

Download

hf download 0bserverx/Qwen3.8-27B-Heretic-Abliterated-Uncensored --local-dir ./rvn-fp16

The root directory is directly loadable by supported Transformers/vLLM/SGLang versions; no merge or conversion step is needed to use these weights.

Tested serving configurations

Both backends below produced real responses from these exact safetensors shards on one NVIDIA RTX PRO 6000 Blackwell Server Edition. The checks used FP16 loading, 4,096-token context, concurrency 1, temperature 0, and disabled CUDA graphs.

  • vLLM 0.28.0, Transformers 5.12.1, Torch 2.13.0+cu130, FlashInfer 0.6.16.post3.
  • SGLang 0.5.19, Transformers 5.12.1, Torch 2.13.0+cu130, FlashInfer 0.6.18.
  • A coherent CUDA 13.0 SDK was required for the tested Blackwell/FlashInfer setup. Native compilation prerequisites included a compiler toolchain, matching Python development headers, and a discoverable ninja executable. The toolkit's bin and lib64 must be available through PATH and LD_LIBRARY_PATH, with CUDA_HOME pointing to the same SDK.

Set MODEL_DIR to the downloaded checkpoint directory and RVN_API_KEY to your own API key. The examples bind only to loopback; run one backend at a time on a single GPU.

vLLM

vllm serve "$MODEL_DIR" \
  --host 127.0.0.1 --port 18081 --api-key "$RVN_API_KEY" \
  --served-model-name rvn --dtype float16 \
  --max-model-len 4096 --gpu-memory-utilization 0.80 \
  --max-num-seqs 1 --max-num-batched-tokens 512 \
  --enforce-eager --seed 0 \
  --enable-auto-tool-choice --tool-call-parser qwen3_coder \
  --mamba-ssm-cache-dtype float32

SGLang

SGLANG_MAMBA_CONV_DTYPE=float16 python -m sglang.launch_server \
  --model-path "$MODEL_DIR" \
  --host 127.0.0.1 --port 18082 --api-key "$RVN_API_KEY" \
  --served-model-name rvn --dtype float16 \
  --context-length 4096 --mem-fraction-static 0.80 \
  --max-running-requests 1 --disable-cuda-graph --random-seed 0 \
  --tool-call-parser qwen3_coder --mamba-ssm-dtype float32

Do not omit SGLANG_MAMBA_CONV_DTYPE=float16 for the tested SGLang version. Its default convolution state is BF16 even with FP16 model loading, which caused a BF16/FP16 Triton mismatch during warmup. The supported environment setting fixes the cache dtype without modifying model weights.

Tool requests were tested with tool_choice: "auto" and chat_template_kwargs: {"enable_thinking": false}.

Actual verification results

Each backend passed all four short request checks:

  • English: Cat.
  • Turkish translation: Günaydın.
  • JSON: {"animal":"cat"}.
  • One parsed get_weather tool call with {"city":"Ankara"} and finish_reason: "tool_calls". This checks tool-call generation, not execution of a real weather service.

The three completion fixtures used identical recorded input token IDs across backends. The tool fixture exercised /v1/chat/completions. Raw requests/responses are included under evidence/vllm-05/ and evidence/sglang-05/.

Additional checks:

  • Real Transformers 5.12.1 load: 851 state-dict keys, zero missing, unexpected or mismatched keys, and zero loading errors.
  • Exporter and runtime-check unit tests: 17 passed.
  • Every saved tensor was reopened; dtype, shape, finiteness and exported-value hashes were verified.
  • Inverse→forward conversion: 803/851 tensors bit-exact. The remaining 48 A_log tensors involve exp/log reconstruction, with maximum absolute forward-round-trip difference 1.52587890625e-05, within the exporter's numerical gate.
  • All 20 original package files, including all 13 weight shards, passed a full SHA-256 recheck after serving tests.

These are smoke tests, not broad quality, throughput, long-context, refusal, or GGUF-versus-safetensors logit-equivalence benchmarks. They do not establish token-identical output for arbitrary prompts across backends.

The unchanged upstream tokenizer emitted a Transformers regex warning during harness loading. No automatic tokenizer rewrite was applied; the short fixtures passed, but a broad tokenizer regression suite was not run.

Provenance and reconstruction

Source model: main RVN GGUF repository.

The exporter reverses tensor-name mapping, linear-attention value-head grouping, convolution singleton-axis removal, converter-shifted RMS-norm offsets, and -exp(A_log). The latter is an arithmetic reconstruction, not a promise of recovering original BF16 bits.

The main repository documents RVN's original derivation and credits Tim Rohrbaugh's ARA variant. This format-conversion release does not claim a new run of those procedures or newly measured versions of its historical KL/refusal numbers. Thanks to the Qwen team, the llama.cpp contributors, the inference-backend maintainers, and the RVN community.

Files and reproducibility

  • model-*.safetensors, model.safetensors.index.json: all main text weights and their index.
  • config.json, tokenizer files, chat_template.jinja, generation_config.json: runtime assets.
  • export-receipt.json: all per-tensor transformation, source/export/saved hashes and round-trip errors.
  • artifact-manifest.json, SHA256SUMS: independently inventoried original model package.
  • reproduce/: exact exporter, tests, pinned schemas, upstream config/index, and runtime validation code. See reproduce/README.md.
  • evidence/: raw API fixtures and load, test and post-runtime integrity results.
  • publication-manifest.json: file identities for this public payload; excludes itself.

License: Apache-2.0. See LICENSE.