gbuzhf/Ornith-1.5-35B-A3B-Huihui-Sangreal-MTP-ICE-GGUF

🤗 Hugging Face 来源text-generationmit激活 3B184 GBGGUF✓ 6 个校验和今天更新
一条命令提交

在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo gbuzhf/Ornith-1.5-35B-A3B-Huihui-Sangreal-MTP-ICE-GGUF ./model-folder
需要做种者 →

Ornith-1.5-35B-A3B — Huihui × Sangreal ICE

Five ICE tiers of huihui-ai's abliterated Ornith-1.5-35B-A3B, quantized against Sangreal — a purpose-built 77-bucket calibration corpus — and shipping peculiar-ragdoll's Qwen-Sharp chat template. The abliterated checkpoint already carries Ornith-1.5's trained MTPv2 head, so every tier drafts for speculative decoding out of the box.

This release introduces a new ICE rung: 15G-ICE, the first tier below the 17 GB bound the method was originally scoped to, built for 16 GB cards. The BF16 master all five were cut from is published here.

Best got better — ICE 1.6 (2026-09-25)

Rebuilt tiers vs the files they replace — mean KLD to BF16, paired on the same chunks, one session.

tier code KLD (2K) code @16K paired code t
15G 0.088158 → 0.083796 (−4.9%) −4.4% −6.02 clear win - replaced
19G 0.036887 → 0.034761 (−5.8%) −6.7% −3.07 clear win - replaced
23G 0.018988 → 0.018473 (−2.7%) −4.5% −1.49 replaced · lower KLD at 2K and 16K; code top-1 −0.06 pp
25G 0.015308 → 0.014603 (−4.6%) −5.2% −2.61 clear win - replaced

Changed: output.weight Q8_0 → Q6_K, MTP head at the floor, freed bytes → routed experts, 36,025-chunk Sangreal imatrix, template v22.5.1.

⚠️ Uncensored. This is an abliterated checkpoint — refusal behaviour has been suppressed in the weights. Sandbox it at the OS level and control its network and code-execution access; with no refusal backstop, a prompt injection from a hostile page or third-party code has nothing to stop it.

Measurements

One binary, one reference, one session, 64 chunks at n_ctx 2048. Every file below — including our previous CyberTiel ladder and the Official CyberTiel's UD ladder — was re-measured in this same session; cross-session KLD drifts ~0.8% on identical inputs, which is the same order as the effect being measured.

2026-09-25: a second session (same GPU class, binary, reference, corpora) re-measured every shipped Sangreal tier and every UD row: all 24 values reproduced exactly. Replaced rows show then now. Raw logs: measurements/2026-09-25/.

Table 1 — Code (code.test.raw)

PPL(base) 2.194208 · BF16 = 100

file size mean KLD 99% KLD 99.9% KLD PPL ratio same top-1 active bpw file bpw overall
≈ 25–27 GB
UD-Q5_K_XL 26.98 GB 0.011993 0.1702 0.8170 1.0009 97.25 % 7.833 6.080 98.3
(New) Sangreal 25G-ICE 24.85 GB
24.82 GB
0.015308
0.014603
0.2274
0.2129
1.0218
0.9268
1.0017
1.0014
97.00 %
97.03 %
7.686
7.385
5.599
5.593
98.0
98.1
CyberTiel 25G-ICE 24.85 GB 0.015497 0.2207 1.0820 1.0007 97.06 % 7.686 5.599 98.0
≈ 22.5–23 GB
(New) Sangreal 23G-ICE 22.84 GB
22.81 GB
0.018988
0.018473
0.2831
0.2767
1.2650
1.2717
1.0023
1.0026
96.73 %
96.67 %
7.523
7.215
5.145
5.139
97.7
CyberTiel 23G-ICE 22.84 GB 0.019269 0.2822 1.3248 1.0018 96.66 % 7.523 5.145 97.7
UD-Q4_K_XL 22.75 GB 0.019742 0.2909 1.3855 1.0041 96.61 % 7.474 5.126 97.6
UD-Q4_K_M 22.52 GB 0.020440 0.2880 1.3215 1.0040 96.59 % 7.130 5.075 97.6
≈ 21 GB
Sangreal 21G-ICE 20.85 GB 0.023415 0.3529 1.4214 1.0042 96.35 % 7.357 4.698 97.3
UD-Q4_K_S 21.28 GB 0.023484 0.3466 1.6264 1.0028 96.31 % 7.025 4.795 97.3
CyberTiel 21G-ICE 20.85 GB 0.024669 0.3787 1.6756 1.0044 96.28 % 7.357 4.698 97.2
≈ 18 GB
(New) Sangreal 19G-ICE 18.82 GB
18.76 GB
0.036887
0.034761
0.5614
0.5241
2.4828
2.1468
1.0159
1.0145
95.33 %
95.35 %
7.192
6.871
4.241
4.228
96.1
96.3
CyberTiel 19G-ICE 18.82 GB 0.036910 0.5798 2.3285 1.0154 95.24 % 7.192 4.241 96.1
UD-IQ4_XS 18.12 GB 0.044161 0.6689 2.8208 1.0191 94.80 % 6.757 4.083 95.5
≈ 13.5–15 GB
(New) Sangreal 15G-ICE 14.83 GB
14.81 GB
0.088158
0.083796
1.4526
1.3755
4.7767
4.6790
1.0467
1.0427
92.91 %
92.93 %
6.853
6.536
3.341
3.336
92.2
92.5
UD-IQ3_XXS 13.60 GB 0.105845 1.7903 5.2413 1.0594 91.96 % 5.489 3.064 90.9

Table 2 — Code at long context (code.test.raw, n_ctx 16384 × 8)

Scored on positions 8,193–16,384 of each chunk (65,536 tokens, as in the 2K tables). Compare rows within this table only. (New) rows: shipped file then, ICE 1.6 file now.

PPL(base) 2.121504 · BF16 = 100

file size mean KLD 99% KLD 99.9% KLD PPL ratio same top-1 active bpw file bpw overall
≈ 25–27 GB
UD-Q5_K_XL 26.98 GB 0.037584 0.5148 7.1463 1.0016 96.86 % 7.833 6.080 96.5
(New) Sangreal 25G-ICE 24.85 GB
24.82 GB
0.041354
0.039198
0.5975
0.5428
7.2349
6.8581
0.9993
0.9992
96.60 %
96.68 %
7.686
7.385
5.599
5.593
96.2
96.4
≈ 22.5–23 GB
(New) Sangreal 23G-ICE 22.84 GB
22.81 GB
0.045097
0.043053
0.7004
0.6588
7.2783
7.0492
0.9996
0.9995
96.39 %
96.41 %
7.523
7.215
5.145
5.139
95.9
96.0
UD-Q4_K_XL 22.75 GB 0.045938 0.7000 7.5137 0.9971 96.29 % 7.474 5.126 95.8
UD-Q4_K_M 22.52 GB 0.046128 0.6952 7.2629 0.9989 96.21 % 7.130 5.075 95.8
≈ 21 GB
UD-Q4_K_S 21.28 GB 0.051208 0.8028 8.1942 0.9942 95.97 % 7.025 4.795 95.4
Sangreal 21G-ICE 20.85 GB 0.051563 0.8315 7.3234 0.9949 96.01 % 7.357 4.698 95.4
≈ 18 GB
(New) Sangreal 19G-ICE 18.82 GB
18.76 GB
0.063182
0.058960
1.0950
0.9983
8.1007
6.9570
1.0189
1.0160
95.24 %
95.28 %
7.192
6.871
4.241
4.228
94.4
94.7
UD-IQ4_XS 18.12 GB 0.077803 1.4135 8.5959 1.0284 94.66 % 6.757 4.083 93.3
UD-Q3_K_XL 17.23 GB 0.084033 1.5944 8.1866 1.0209 94.21 % — 3.882 92.8
≈ 13.5–15 GB
(New) Sangreal 15G-ICE 14.83 GB
14.81 GB
0.112369
0.107390
2.1470
2.1159
9.0985
8.3963
1.0393
1.0353
93.09 %
93.31 %
6.853
6.536
3.341
3.336
90.9
91.2
UD-IQ3_XXS 13.60 GB 0.137273 2.7651 9.3111 1.0729 92.11 % 5.489 3.064 89.2

Table 3 — English text (WikiText-2)

PPL(base) 7.574505 · BF16 = 100

file size mean KLD 99% KLD 99.9% KLD PPL ratio same top-1 active bpw file bpw overall
≈ 25–27 GB
UD-Q5_K_XL 26.98 GB 0.024330 0.2425 1.0036 0.9929 93.88 % 7.833 6.080 96.5
(New) Sangreal 25G-ICE 24.85 GB
24.82 GB
0.029038
0.027264
0.2823
0.2677
1.2220
1.1085
0.9857
0.9847
93.27 %
93.21 %
7.686
7.385
5.599
5.593
96.0
96.1
CyberTiel 25G-ICE 24.85 GB 0.027866 0.2793 1.0757 0.9867 93.36 % 7.686 5.599 96.1
≈ 22.5–23 GB
(New) Sangreal 23G-ICE 22.84 GB
22.81 GB
0.033585
0.031606
0.3299
0.3093
1.3758
1.1956
0.9870
0.9907
92.69 %
92.78 %
7.523
7.215
5.145
5.139
95.5
95.7
CyberTiel 23G-ICE 22.84 GB 0.032992 0.3290 1.2449 0.9899 92.75 % 7.523 5.145 95.6
UD-Q4_K_M 22.52 GB 0.035692 0.3576 1.3199 0.9766 92.42 % 7.130 5.075 95.3
UD-Q4_K_XL 22.75 GB 0.035837 0.3753 1.3951 0.9761 92.56 % 7.474 5.126 95.3
≈ 21 GB
CyberTiel 21G-ICE 20.85 GB 0.039658 0.4136 1.4041 0.9848 92.02 % 7.357 4.698 94.9
Sangreal 21G-ICE 20.85 GB 0.039968 0.4156 1.5834 0.9819 92.03 % 7.357 4.698 94.9
UD-Q4_K_S 21.28 GB 0.040169 0.4107 1.5040 0.9808 91.97 % 7.025 4.795 94.9
≈ 18 GB
(New) Sangreal 19G-ICE 18.82 GB
18.76 GB
0.060435
0.059599
0.6098
0.5982
2.1340
2.0323
0.9948
0.9907
90.15 %
90.17 %
7.192
6.871
4.241
4.228
93.1
CyberTiel 19G-ICE 18.82 GB 0.060573 0.6220 2.2733 0.9951 90.15 % 7.192 4.241 93.0
UD-IQ4_XS 18.12 GB 0.070787 0.7346 2.6002 1.0441 89.51 % 6.757 4.083 92.2
≈ 13.5–15 GB
(New) Sangreal 15G-ICE 14.83 GB
14.81 GB
0.126386
0.120439
1.3329
1.2781
3.9267
3.7501
1.0027
0.9987
85.89 %
86.18 %
6.853
6.536
3.341
3.336
87.9
88.3
UD-IQ3_XXS 13.60 GB 0.151623 1.6338 4.3641 1.0920 84.85 % 5.489 3.064 86.2

overall = 0.70/(1 + meanKLD) + 0.30 * sameTop1, ×100. BF16 = 100. Same composite the CyberTiel card uses, so the two are directly comparable: 70% on how close the whole output distribution stays, 30% on agreement about the argmax.

Read the tail columns as shape, not order. 99.9% KLD is roughly the 33rd-worst token of 32,768 — an extreme order statistic with large sampling variance, so it inverts between adjacent files without meaning anything. Rank on mean KLD.

Two bpw columns. Only 8 of 256 experts fire per token, so a bit in ffn_*_exps is worth ~3% of a bit in attention, the shared expert or the output head. active bpw weights by that; file bpw is just size ÷ parameters. It is why a 22.81 GB file computes at ~7.2 bpw.

Don't compare the tables to each other. Code is more predictable text, so every file scores about half the divergence on it. Compare rows within a table.

Every row is measured against this lineage's own BF16 master — the huihui abliterated checkpoint — in one session, one binary, one reference. Quantizations cut from a different trunk are deliberately absent: scoring them here would measure the distance between trunks and call it quantization damage.

What Sangreal is

The calibration corpus these tiers were quantized against. Built from primary sources, not assembled from eaddario's set, bartowski's calibration_datav5, or any other ready-made calibration file. 77 buckets, 9,533 documents, 47,443,549 tokens, sha256 85a6b823a6762bf6….

The full render is 92,663 chunks. 15G, 19G, 23G and 25G use a 36,025-chunk pass over it — 63x the stock Ornith imatrix (573) and 45x bartowski's calibration_datav5 (~800). 21G keeps the earlier 17,225-chunk pass.

block buckets share documents synthetic
code + cyber + spec 36 45.10% 5,858 0.1%
reasoning 5 13.16% 553 100.0%
agentic 7 11.95% 1,018 65.2%
science + domain 11 11.24% 1,183 0.3%
language 11 9.66% 596 0.0%
mathematics 7 8.89% 325 0.0%
total 77 100% 9,533 21.7%

Split: technical 70.2 · science+domain 20.1 · language 9.7.

Synthetic is 21.7% by bytes and sits in exactly three buckets — teacher:k2-horizon, teacher:deepseek-v4-pro, teacher:hermes. Everything outside reasoning and agentic is real text.

Language — 11 languages, 5 non-Latin scripts

bucket share docs bucket share docs
arabic 0.92% 48 russian 0.87% 58
chinese 0.89% 31 spanish 0.87% 72
hindi 0.88% 44 portuguese 0.87% 70
japanese 0.88% 33 turkish 0.87% 56
french 0.88% 60 german 0.86% 73
english 0.87% 51

The spread across the whole block is 0.06 points, and that is deliberate: language's job here is activation, not ranking. Each bucket sits just above the saturation floor at 44.9% consumption, so these buckets select for the first time — the legacy v1 spec had consumed 96% of turkish and 92% of english.

Mathematics — absent from the pool until this cycle

bucket share docs
math (general) 3.16% 170
analysis 1.24% 28
algebra 1.16% 46
geometry 1.12% 30
proof 1.08% 28
topology 1.05% 20
number-theory 0.06% 3

Sourced from the LaTeX algebraic geometry, mathlib4, UniMath, math-comp, open-web-math, AutoMathText and proof-pile.

The imatrix is published here — imatrix/sangreal-36025.imatrix.gguf, the exact file the ICE 1.6 tiers were quantized against. Pass it to llama-quantize --imatrix to rebuild them from the BF16 master, or to cut your own.

What's inside

MTPv2 head, native. huihui's abliterated checkpoint already carries Ornith-1.5's trained multi-token-prediction head; its mtp.* norms are bit-identical to ornith-ai's on all seven tensors. Every tier ships blk.40 with nextn.* intact, so llama.cpp can use it as a draft model for speculative decoding. ICE 1.6 tiers carry it at Q2_K experts / Q4_K matrices; acceptance is unchanged (15G: 94.1% at --spec-draft-p-min 0.75).

Chat template: peculiar-ragdoll's Qwen-Sharp v22.5.0, embedded in the ICE 1.5 tiers. It is a genuinely good template. The ICE 1.6 tiers embed v22.5.1 (even more sharpened): an agent-mode directive when tools are present (off via chat_template_kwargs: {"decisive": false}) and fixed error coaching; plain chat renders as v22.5.0.

ICE = Isolation of Compounding Error: allocate bits by how far a quantization error travels, not by how large the activations are. Routers stay F32, SSM decay gates stay F32, the KV-cached projections stay F16, the always-on dense path stays Q8_0 — together 0.14% of the model, kept exact for ~152 MB — and the entire budget is spent on the routed expert bank, which is 93% of the parameters but only 8-of-256 active per token. ICE 1.6 moves the output head to Q6_K: its error stays flat from 2K to 16K instead of compounding, which is why active bpw falls while KLD improves. Method and the cases where it does not win: gbuzhf/ICE-quantization.

Files

  • measurements/ — raw llama-perplexity output for every 2K row (2026-09-19), and measurements/2026-09-25/ for the ICE 1.6 session: all three columns, the re-measured shipped and UD files, the two control builds, and the MTP acceptance run.
  • recipes/ — the exact llama-quantize tensor map for each tier (443 rules): cfg_<T>-ICE16.txt for 15G, 19G, 23G and 25G; cfg_21G-ICE_ICEbase.txt for 21G.
  • imatrix/ — sangreal-36025.imatrix.gguf (ICE 1.6 tiers: 15G, 19G, 23G and 25G; 36,025 chunks × 512). 1020 tensors. With this plus recipes/, every ICE 1.6 tier here is reproducible byte-for-byte from the published BF16 master.

Credits

KLD measures fidelity to this repo's master and nothing else — not reasoning, tool use or speed. It is also within-lineage: every ladder is measured against its own master, so these values are not comparable to another repo's. Compare across lineages with PPL.