prithivMLmods/NeoHorse-1-4B-Max-Heretic-Uncensored-TURBO

🤗 Hugging Face 来源text-generationapache-2.04.2B 参数8.4 GBGGUF✓ 13 个校验和今天更新
一条命令提交

在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo prithivMLmods/NeoHorse-1-4B-Max-Heretic-Uncensored-TURBO ./model-folder
需要做种者 →

NeoHorse-1-4B-Max-Heretic-Uncensored-TURBO

NeoHorse-1-4B-Max-Heretic-Uncensored-TURBO is an uncensored, heretic-processed variant of TokenRhythm's NeoHorse-1-4B — itself a routing-guided agentic post-training fine-tune of Qwen3.5-4B — with refusal behavior stripped out using the Heretic tool via multi-shot censorship removal, driving refusal rates down to a minimal level while aiming to preserve NeoHorse's original agentic tool-use, coding, and instruction-following gains over the base Qwen3.5-4B. The repository ships the full-precision safetensors checkpoint at root alongside a standard GGUF quantization sweep (BF16, Q6_K, Q5_K_M/S, Q4_K_M/S, Q3_K_L/M) for llama.cpp deployment, and is marked experimental, with the model card noting it may generate artifacts and warranting output review before use. It is released under the Apache License 2.0, inherited through its Qwen3.5-4B and NeoHorse-1-4B lineage.

Repository layout

+-- prithivMLmods/NeoHorse-1-4B-Max-Heretic-Uncensored-TURBO (main)
    +-- / (Root: safetensors)
    +-- GGUF/ (.gguf files: BF16, Q3_K_L, Q3_K_M, Q4_K_M, Q4_K_S, Q5_K_M, Q5_K_S, Q6_K)

GGUF Model Files

File Name Quant Type File Size File Link Description
NeoHorse-1-4B-Max-Heretic-Uncensored-TURBO.BF16.gguf BF16 8.42 GB Link Full BF16 weights. Highest quality, largest file size.
NeoHorse-1-4B-Max-Heretic-Uncensored-TURBO.Q3_K_L.gguf Q3_K_L 2.42 GB Link Lower quality but usable, good for low RAM availability.
NeoHorse-1-4B-Max-Heretic-Uncensored-TURBO.Q3_K_M.gguf Q3_K_M 2.26 GB Link Low quality.
NeoHorse-1-4B-Max-Heretic-Uncensored-TURBO.Q4_K_M.gguf Q4_K_M 2.71 GB Link Good quality, default size for most use cases, recommended.
NeoHorse-1-4B-Max-Heretic-Uncensored-TURBO.Q4_K_S.gguf Q4_K_S 2.56 GB Link Slightly lower quality with more space savings, recommended.
NeoHorse-1-4B-Max-Heretic-Uncensored-TURBO.Q5_K_M.gguf Q5_K_M 3.07 GB Link High quality, recommended.
NeoHorse-1-4B-Max-Heretic-Uncensored-TURBO.Q5_K_S.gguf Q5_K_S 2.99 GB Link High quality, recommended.
NeoHorse-1-4B-Max-Heretic-Uncensored-TURBO.Q6_K.gguf Q6_K 3.46 GB Link Very high quality, near perfect, recommended.

llama.cpp

LLM inference in C/C++ — https://github.com/ggml-org/llama.cpp

Usage

Transformers

from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

model_id = "prithivMLmods/NeoHorse-1-4B-Max-Heretic-Uncensored-TURBO"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.bfloat16,
    device_map="auto",
)

messages = [{"role": "user", "content": "Hello, who are you?"}]
text = tokenizer.apply_chat_template(
    messages, tokenize=False, add_generation_prompt=True
)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=512, temperature=1.0, top_p=0.95, top_k=20)
print(tokenizer.decode(out[0][inputs.input_ids.shape[1]:], skip_special_tokens=True))

The root of the repository holds the full-precision safetensors checkpoint. The GGUF/ subdirectory holds the standard llama.cpp quantization sweep, from BF16 down to Q3_K_L.

llama.cpp (GGUF)

1. Download a quant

Pick one GGUF file based on your available RAM/VRAM. Reasonable defaults: Q4_K_M (balanced), Q5_K_M (higher quality), Q6_K/BF16 (near-lossless, more memory).

pip install -U "huggingface_hub[cli]"

hf download prithivMLmods/NeoHorse-1-4B-Max-Heretic-Uncensored-TURBO \
  --include "GGUF/NeoHorse-1-4B-Max-Heretic-Uncensored-TURBO.Q4_K_M.gguf" \
  --local-dir ./NeoHorse-1-4B-TURBO

2. Build llama.cpp (skip if already installed)

git clone https://github.com/ggml-org/llama.cpp
cd llama.cpp
cmake -B build -DGGML_CUDA=ON   # drop -DGGML_CUDA=ON for CPU-only
cmake --build build --config Release -j

3. Serve it (OpenAI-compatible API)

./build/bin/llama-server \
  -m ./NeoHorse-1-4B-TURBO/GGUF/NeoHorse-1-4B-Max-Heretic-Uncensored-TURBO.Q4_K_M.gguf \
  -c 262144 \
  --host 0.0.0.0 \
  --port 8080 \
  -ngl 999
  • -c 262144 sets the full native context window; lower it (e.g. -c 32768) if you hit memory limits.
  • -ngl 999 offloads all layers to GPU; drop it or set a smaller number for CPU/partial offload.

Query it:

curl http://localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "neohorse-1-4b-turbo",
    "messages": [{"role": "user", "content": "Hello, who are you?"}],
    "temperature": 1.0,
    "top_p": 0.95,
    "top_k": 20,
    "max_tokens": 512
  }'

4. Or run it directly in the terminal (no server)

./build/bin/llama-cli \
  -m ./NeoHorse-1-4B-TURBO/GGUF/NeoHorse-1-4B-Max-Heretic-Uncensored-TURBO.Q4_K_M.gguf \
  -c 262144 \
  -ngl 999 \
  -cnv \
  --temp 1.0 --top-p 0.95 --top-k 20

-cnv starts an interactive chat session using the model's built-in chat template.

License and Attribution

This model is released under the Apache License 2.0, inherited from the base and fine-tuned models.