prithivMLmods/NeoHorse-1-4B-Max-Heretic-Uncensored-TURBO

🤗 Hugging Face sourcetext-generationapache-2.04.2B params8.4 GBGGUF✓ 13 checksumsupdated today
Submit in one command

Run it next to your model folder. It makes the torrent, checks your files against Hugging Face, and submits it. You just start seeding and paste your key from your account. It only reads your files and never changes them. Read the script first if you like.

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo prithivMLmods/NeoHorse-1-4B-Max-Heretic-Uncensored-TURBO ./model-folder
Needs a seeder →

NeoHorse-1-4B-Max-Heretic-Uncensored-TURBO

NeoHorse-1-4B-Max-Heretic-Uncensored-TURBO is an uncensored, heretic-processed variant of TokenRhythm's NeoHorse-1-4B — itself a routing-guided agentic post-training fine-tune of Qwen3.5-4B — with refusal behavior stripped out using the Heretic tool via multi-shot censorship removal, driving refusal rates down to a minimal level while aiming to preserve NeoHorse's original agentic tool-use, coding, and instruction-following gains over the base Qwen3.5-4B. The repository ships the full-precision safetensors checkpoint at root alongside a standard GGUF quantization sweep (BF16, Q6_K, Q5_K_M/S, Q4_K_M/S, Q3_K_L/M) for llama.cpp deployment, and is marked experimental, with the model card noting it may generate artifacts and warranting output review before use. It is released under the Apache License 2.0, inherited through its Qwen3.5-4B and NeoHorse-1-4B lineage.

Repository layout

+-- prithivMLmods/NeoHorse-1-4B-Max-Heretic-Uncensored-TURBO (main)
    +-- / (Root: safetensors)
    +-- GGUF/ (.gguf files: BF16, Q3_K_L, Q3_K_M, Q4_K_M, Q4_K_S, Q5_K_M, Q5_K_S, Q6_K)

GGUF Model Files

File Name Quant Type File Size File Link Description
NeoHorse-1-4B-Max-Heretic-Uncensored-TURBO.BF16.gguf BF16 8.42 GB Link Full BF16 weights. Highest quality, largest file size.
NeoHorse-1-4B-Max-Heretic-Uncensored-TURBO.Q3_K_L.gguf Q3_K_L 2.42 GB Link Lower quality but usable, good for low RAM availability.
NeoHorse-1-4B-Max-Heretic-Uncensored-TURBO.Q3_K_M.gguf Q3_K_M 2.26 GB Link Low quality.
NeoHorse-1-4B-Max-Heretic-Uncensored-TURBO.Q4_K_M.gguf Q4_K_M 2.71 GB Link Good quality, default size for most use cases, recommended.
NeoHorse-1-4B-Max-Heretic-Uncensored-TURBO.Q4_K_S.gguf Q4_K_S 2.56 GB Link Slightly lower quality with more space savings, recommended.
NeoHorse-1-4B-Max-Heretic-Uncensored-TURBO.Q5_K_M.gguf Q5_K_M 3.07 GB Link High quality, recommended.
NeoHorse-1-4B-Max-Heretic-Uncensored-TURBO.Q5_K_S.gguf Q5_K_S 2.99 GB Link High quality, recommended.
NeoHorse-1-4B-Max-Heretic-Uncensored-TURBO.Q6_K.gguf Q6_K 3.46 GB Link Very high quality, near perfect, recommended.

llama.cpp

LLM inference in C/C++ — https://github.com/ggml-org/llama.cpp

Usage

Transformers

from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

model_id = "prithivMLmods/NeoHorse-1-4B-Max-Heretic-Uncensored-TURBO"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.bfloat16,
    device_map="auto",
)

messages = [{"role": "user", "content": "Hello, who are you?"}]
text = tokenizer.apply_chat_template(
    messages, tokenize=False, add_generation_prompt=True
)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=512, temperature=1.0, top_p=0.95, top_k=20)
print(tokenizer.decode(out[0][inputs.input_ids.shape[1]:], skip_special_tokens=True))

The root of the repository holds the full-precision safetensors checkpoint. The GGUF/ subdirectory holds the standard llama.cpp quantization sweep, from BF16 down to Q3_K_L.

llama.cpp (GGUF)

1. Download a quant

Pick one GGUF file based on your available RAM/VRAM. Reasonable defaults: Q4_K_M (balanced), Q5_K_M (higher quality), Q6_K/BF16 (near-lossless, more memory).

pip install -U "huggingface_hub[cli]"

hf download prithivMLmods/NeoHorse-1-4B-Max-Heretic-Uncensored-TURBO \
  --include "GGUF/NeoHorse-1-4B-Max-Heretic-Uncensored-TURBO.Q4_K_M.gguf" \
  --local-dir ./NeoHorse-1-4B-TURBO

2. Build llama.cpp (skip if already installed)

git clone https://github.com/ggml-org/llama.cpp
cd llama.cpp
cmake -B build -DGGML_CUDA=ON   # drop -DGGML_CUDA=ON for CPU-only
cmake --build build --config Release -j

3. Serve it (OpenAI-compatible API)

./build/bin/llama-server \
  -m ./NeoHorse-1-4B-TURBO/GGUF/NeoHorse-1-4B-Max-Heretic-Uncensored-TURBO.Q4_K_M.gguf \
  -c 262144 \
  --host 0.0.0.0 \
  --port 8080 \
  -ngl 999
  • -c 262144 sets the full native context window; lower it (e.g. -c 32768) if you hit memory limits.
  • -ngl 999 offloads all layers to GPU; drop it or set a smaller number for CPU/partial offload.

Query it:

curl http://localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "neohorse-1-4b-turbo",
    "messages": [{"role": "user", "content": "Hello, who are you?"}],
    "temperature": 1.0,
    "top_p": 0.95,
    "top_k": 20,
    "max_tokens": 512
  }'

4. Or run it directly in the terminal (no server)

./build/bin/llama-cli \
  -m ./NeoHorse-1-4B-TURBO/GGUF/NeoHorse-1-4B-Max-Heretic-Uncensored-TURBO.Q4_K_M.gguf \
  -c 262144 \
  -ngl 999 \
  -cnv \
  --temp 1.0 --top-p 0.95 --top-k 20

-cnv starts an interactive chat session using the model's built-in chat template.

License and Attribution

This model is released under the Apache License 2.0, inherited from the base and fine-tuned models.