NeoHorse-1-4B-Max-Heretic-Uncensored-TURBO
NeoHorse-1-4B-Max-Heretic-Uncensored-TURBO is an uncensored, heretic-processed variant of TokenRhythm's NeoHorse-1-4B — itself a routing-guided agentic post-training fine-tune of Qwen3.5-4B — with refusal behavior stripped out using the Heretic tool via multi-shot censorship removal, driving refusal rates down to a minimal level while aiming to preserve NeoHorse's original agentic tool-use, coding, and instruction-following gains over the base Qwen3.5-4B. The repository ships the full-precision safetensors checkpoint at root alongside a standard GGUF quantization sweep (BF16, Q6_K, Q5_K_M/S, Q4_K_M/S, Q3_K_L/M) for llama.cpp deployment, and is marked experimental, with the model card noting it may generate artifacts and warranting output review before use. It is released under the Apache License 2.0, inherited through its Qwen3.5-4B and NeoHorse-1-4B lineage.
Repository layout
+-- prithivMLmods/NeoHorse-1-4B-Max-Heretic-Uncensored-TURBO (main)
+-- / (Root: safetensors)
+-- GGUF/ (.gguf files: BF16, Q3_K_L, Q3_K_M, Q4_K_M, Q4_K_S, Q5_K_M, Q5_K_S, Q6_K)
GGUF Model Files
| File Name | Quant Type | File Size | File Link | Description |
|---|---|---|---|---|
| NeoHorse-1-4B-Max-Heretic-Uncensored-TURBO.BF16.gguf | BF16 | 8.42 GB | Link | Full BF16 weights. Highest quality, largest file size. |
| NeoHorse-1-4B-Max-Heretic-Uncensored-TURBO.Q3_K_L.gguf | Q3_K_L | 2.42 GB | Link | Lower quality but usable, good for low RAM availability. |
| NeoHorse-1-4B-Max-Heretic-Uncensored-TURBO.Q3_K_M.gguf | Q3_K_M | 2.26 GB | Link | Low quality. |
| NeoHorse-1-4B-Max-Heretic-Uncensored-TURBO.Q4_K_M.gguf | Q4_K_M | 2.71 GB | Link | Good quality, default size for most use cases, recommended. |
| NeoHorse-1-4B-Max-Heretic-Uncensored-TURBO.Q4_K_S.gguf | Q4_K_S | 2.56 GB | Link | Slightly lower quality with more space savings, recommended. |
| NeoHorse-1-4B-Max-Heretic-Uncensored-TURBO.Q5_K_M.gguf | Q5_K_M | 3.07 GB | Link | High quality, recommended. |
| NeoHorse-1-4B-Max-Heretic-Uncensored-TURBO.Q5_K_S.gguf | Q5_K_S | 2.99 GB | Link | High quality, recommended. |
| NeoHorse-1-4B-Max-Heretic-Uncensored-TURBO.Q6_K.gguf | Q6_K | 3.46 GB | Link | Very high quality, near perfect, recommended. |
llama.cpp
LLM inference in C/C++ — https://github.com/ggml-org/llama.cpp
Usage
Transformers
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
model_id = "prithivMLmods/NeoHorse-1-4B-Max-Heretic-Uncensored-TURBO"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.bfloat16,
device_map="auto",
)
messages = [{"role": "user", "content": "Hello, who are you?"}]
text = tokenizer.apply_chat_template(
messages, tokenize=False, add_generation_prompt=True
)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=512, temperature=1.0, top_p=0.95, top_k=20)
print(tokenizer.decode(out[0][inputs.input_ids.shape[1]:], skip_special_tokens=True))
The root of the repository holds the full-precision safetensors checkpoint. The GGUF/ subdirectory holds the standard llama.cpp quantization sweep, from BF16 down to Q3_K_L.
llama.cpp (GGUF)
1. Download a quant
Pick one GGUF file based on your available RAM/VRAM. Reasonable defaults: Q4_K_M (balanced), Q5_K_M (higher quality), Q6_K/BF16 (near-lossless, more memory).
pip install -U "huggingface_hub[cli]"
hf download prithivMLmods/NeoHorse-1-4B-Max-Heretic-Uncensored-TURBO \
--include "GGUF/NeoHorse-1-4B-Max-Heretic-Uncensored-TURBO.Q4_K_M.gguf" \
--local-dir ./NeoHorse-1-4B-TURBO
2. Build llama.cpp (skip if already installed)
git clone https://github.com/ggml-org/llama.cpp
cd llama.cpp
cmake -B build -DGGML_CUDA=ON # drop -DGGML_CUDA=ON for CPU-only
cmake --build build --config Release -j
3. Serve it (OpenAI-compatible API)
./build/bin/llama-server \
-m ./NeoHorse-1-4B-TURBO/GGUF/NeoHorse-1-4B-Max-Heretic-Uncensored-TURBO.Q4_K_M.gguf \
-c 262144 \
--host 0.0.0.0 \
--port 8080 \
-ngl 999
-c 262144sets the full native context window; lower it (e.g.-c 32768) if you hit memory limits.-ngl 999offloads all layers to GPU; drop it or set a smaller number for CPU/partial offload.
Query it:
curl http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "neohorse-1-4b-turbo",
"messages": [{"role": "user", "content": "Hello, who are you?"}],
"temperature": 1.0,
"top_p": 0.95,
"top_k": 20,
"max_tokens": 512
}'
4. Or run it directly in the terminal (no server)
./build/bin/llama-cli \
-m ./NeoHorse-1-4B-TURBO/GGUF/NeoHorse-1-4B-Max-Heretic-Uncensored-TURBO.Q4_K_M.gguf \
-c 262144 \
-ngl 999 \
-cnv \
--temp 1.0 --top-p 0.95 --top-k 20
-cnv starts an interactive chat session using the model's built-in chat template.
License and Attribution
This model is released under the Apache License 2.0, inherited from the base and fine-tuned models.
- Base Model: Qwen/Qwen3.5-4B
- Agentic Fine-Tune: TokenRhythm/NeoHorse-1-4B
- Decensoring Tool: Heretic