AMAImedia/Qwen3.5-9B-NeoHorse1-Heretic-NOESIS-BF16-GGUF

🤗 Hugging Face sourcetext-generationapache-2.09B params18 GBGGUF✓ 20 checksumsupdated today
Submit in one command

Run it next to your model folder. It makes the torrent, checks your files against Hugging Face, and submits it. You just start seeding and paste your key from your account. It only reads your files and never changes them. Read the script first if you like.

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo AMAImedia/Qwen3.5-9B-NeoHorse1-Heretic-NOESIS-BF16-GGUF ./model-folder
Needs a seeder →

⚡ Each donation funds the next large quant.

I host free GGUF or MoE quants as independent research.
Local hardware: Mechrevo Kuangshi GM7AG0M — RTX 3060 Laptop 6GB GDDR6, 64GB DDR5, i7-12700H (14C/20T, 4.7GHz), Windows 11, Samsung 990 Pro.
Good for imatrix and 0.6–35B-class work in RAM. 9B+ and searches need rented H200/Blackwell, typically $100 per quant.

🎉 Boosty🦄  |  ☕ Buy Me a Coffee🦄  |  ⭐ DonationAlerts🦄

💚 Thanks to Hugging Face for extra storage.🦄


NOESIS / AMAImedia

Released as part of the NOESIS Professional Multilingual Dubbing Automation Platform.

AMAImedia


This is a decensored version of a model, made using Heretic v1.4.0

Abliteration parameters

Parameter Value
direction_index 16.33
attn.o_proj.max_weight 1.48
attn.o_proj.max_weight_position 19.08
attn.o_proj.min_weight 1.46
attn.o_proj.min_weight_distance 16.49
mlp.down_proj.max_weight 1.44
mlp.down_proj.max_weight_position 18.83
mlp.down_proj.min_weight 1.43
mlp.down_proj.min_weight_distance 13.39

Performance

Metric This model Original model (a model)
KL divergence 0.0181 0 (by definition)
Refusals 18/100 97/100

NeoHorse-1-9B

Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness.

Technical Report

NeoHorse-1-9B is a 9B causal language model and an initial prototype on the path toward recursive self-improvement (RSI). It is post-trained from Qwen3.5-9B for text-based agent harnesses, tool use, coding, and instruction following.

Derived from Qwen/Qwen3.5-9B and fine-tuned by TokenRhythm. This release contains language-model weights only and is repackaged for text-only inference. Vision weights are not included. Repackaging changes configuration and tensor key names, without changing the fine-tuned tensor values.

Highlights

  • Path toward RSI: the routing harness assigns tasks to a heterogeneous model pool, records tool interactions and outcomes, estimates capability demand, and uses capability-level feedback to shape the next training mixture. Updated models can return to the harness, closing a prototype evaluation–selection–update loop; extending this loop across successive iterations is the next step toward RSI.
  • Agentic post-training framework: the associated research explores routing-guided curriculum SFT and routing-guided on-policy distillation to turn execution trajectories into training signal while preserving execution and harness context around each response.
  • Data quality: exact and near-duplicate removal, evaluation decontamination, structural validation, six-dimensional semantic evaluation, and subscene-level Scene/Goal/Outcome labeling.
  • Broad gains: 69.04 macro average across ten benchmarks versus 65.60 for Qwen3.5-9B (+3.44).

Model Details

Property Value
Model family NeoHorse Agent-Native Causal Language Model
Parameters Approximately 9B
Base model Qwen3.5-9B
Post-training Routing-guided agentic post-training
Interface Text input and text output
Context length 262,144 natively and extensible up to 1,010,000 tokens.
Weight format / precision Safetensors / BF16

Evaluation

The 9B track compares NeoHorse-1-9B with five representative open-weight baselines: Granite-4.2-8B, Qwen3.5-9B, Ornith-1.5-9B, Gemma-4-12B-it, and Muse-Glimmer-30B. Results cover ten benchmarks and are grouped by capability. Higher is better; Δ is NeoHorse-1-9B minus Qwen3.5-9B. Bold and underline mark the best and second-best results in each benchmark row, respectively; ties share the same formatting.

Benchmark Granite-4.2-8B Qwen3.5-9B Ornith-1.5-9B Gemma-4-12B-it Muse-Glimmer-30B NeoHorse-1-9B Δ vs Qwen3.5-9B
🤖 Agentic
QwenClawBench37.0144.0447.2743.5346.1148.73+4.69
WorkBuddy Bench35.0739.6029.2929.6545.8540.15+0.55
PinchBench56.9374.5568.2258.8971.3582.25+7.70
VitaBench23.0031.2526.7536.5048.5042.25+11.00
BFCL v452.0664.8865.0362.0653.7467.43+2.55
tau2-Bench62.2888.0483.6859.3776.6490.82+2.78
💻 Coding
HumanEval96.3492.6893.90100.0098.1798.17+5.49
LiveCodeBench v672.0065.1447.4373.1465.7165.14+0.00
📚 Instruction Following
IFBench78.0066.3340.0077.6778.6766.33+0.00
IFEval92.9889.4671.3594.2793.9089.09-0.37
📊 Overall
Ten-benchmark average60.5765.6057.2963.5167.8669.04+3.44

Reported protocol: SGLang v0.5.17 · temperature=1.0 · top_p=0.95 · top_k=20 · min_p=0.0 · presence_penalty=1.5 · repetition_penalty=1.0 · thinking mode enabled with enable_thinking=true and force_nonempty_content=true. QwenClawBench, WorkBuddy Bench, and tau2-Bench use three runs; PinchBench and VitaBench use one run; the remaining benchmarks follow their official protocols. VitaBench uses the DeepSeek-V4-Flash simulator and judge.

Deployment

The examples below are for self-hosted deployment from a downloaded local checkpoint.

Local checkpoint path

The examples below assume the checkpoint has already been downloaded to local disk. Set MODEL_PATH to the directory containing config.json, tokenizer files, and model weights.

MODEL_PATH="/path/to/NeoHorse-1-9B"

The OpenAI-compatible requests below use the server's --served-model-name (for example, neohorse-1-9b), not the filesystem path.

SGLang

The technical report uses SGLang v0.5.17.

pip install "sglang==0.5.17"
MODEL_PATH="/path/to/NeoHorse-1-9B"
python3 -m sglang.launch_server \
  --model-path "$MODEL_PATH" \
  --served-model-name neohorse-1-9b \
  --host 0.0.0.0 \
  --port 30000 \
  --context-length 262144 \
  --reasoning-parser qwen3 \
  --tool-call-parser qwen3_coder

Send an OpenAI-compatible request after the server starts:

curl http://localhost:30000/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"neohorse-1-9b","messages":[{"role":"user","content":"Write a Python function that returns the first n Fibonacci numbers."}],"max_tokens":512}'

vLLM

pip install -U vllm
MODEL_PATH="/path/to/NeoHorse-1-9B"
vllm serve "$MODEL_PATH" \
  --served-model-name neohorse-1-9b \
  --host 0.0.0.0 \
  --port 8000 \
  --max-model-len 262144 \
  --reasoning-parser qwen3 \
  --enable-auto-tool-choice \
  --tool-call-parser qwen3_coder

The server exposes an OpenAI-compatible /v1/chat/completions endpoint. Send a request after the server starts:

curl http://localhost:8000/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"neohorse-1-9b","messages":[{"role":"user","content":"Write a Python function that returns the first n Fibonacci numbers."}],"max_tokens":512}'

The example uses the configured 262,144-token context limit. Actual capacity depends on GPU memory and serving settings; reduce the context limit if needed. These launch examples have not yet been validated on GPU for this repackaged release.

License

NeoHorse-1-9B is released under the Apache License 2.0.

The upstream model is Qwen/Qwen3.5-9B. Its original copyright notice, Copyright 2026 Alibaba Cloud, is retained in the license file. TokenRhythm has modified the model through fine-tuning and repackaging for text-only inference. Modification notices are included in this model card and the released configuration, weight index, and Safetensors metadata.

Citation

@misc{neohorse2026,
  title        = {NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness},
  author       = {NeoHorse Team},
  year         = {2026},
  howpublished = {arXiv preprint}
}

For questions or issue reports, use the NeoHorse project repository.