Ornith-1.5-35B-A3B — Huihui × Sangreal ICE
Five ICE tiers of huihui-ai's abliterated Ornith-1.5-35B-A3B, quantized against Sangreal — a purpose-built 77-bucket calibration corpus — and shipping peculiar-ragdoll's Qwen-Sharp chat template. The abliterated checkpoint already carries Ornith-1.5's trained MTPv2 head, so every tier drafts for speculative decoding out of the box.
This release introduces a new ICE rung: 15G-ICE, the first tier below the
17 GB bound the method was originally scoped to, built for 16 GB cards. The BF16
master all five were cut from is published
here.
⚠️ Uncensored. This is an abliterated checkpoint — refusal behaviour has been suppressed in the weights. Sandbox it at the OS level and control its network and code-execution access; with no refusal backstop, a prompt injection from a hostile page or third-party code has nothing to stop it.
Tiers
| file | size | fits (RAM+VRAM) | where it sits |
|---|---|---|---|
…-MTP-15G-ICE.gguf |
14.79 GB | 16 GB | the new bottom rung — fits a 16 GB card |
…-MTP-19G-ICE.gguf |
18.82 GB | 24 GB | most context headroom on a 24 GB card |
…-MTP-21G-ICE.gguf |
20.85 GB | 24 GB | UD-Q4_K_S class |
…-MTP-23G-ICE.gguf |
22.84 GB | 24–32 GB | start here — UD-Q4_K_XL class |
…-MTP-25G-ICE.gguf |
24.85 GB | 32 GB | closest to BF16 in this repo |
Measurements
One binary, one reference, one session, 64 chunks at n_ctx 2048. Every file below — including our previous CyberTiel ladder and the Official CyberTiel's UD ladder — was re-measured in this same session; cross-session KLD drifts ~0.8% on identical inputs, which is the same order as the effect being measured.
Table 1 — English text (WikiText-2)
PPL(base) 7.574505 · BF16 = 100
| file | size | mean KLD | 99% KLD | 99.9% KLD | PPL ratio | same top-1 | active bpw | file bpw | overall |
|---|---|---|---|---|---|---|---|---|---|
| ≈ 25–27 GB | |||||||||
| UD-Q5_K_XL | 26.98 GB | 0.024330 | 0.2425 | 1.0036 | 0.9929 | 93.88 % | 7.833 | 6.080 | 96.5 |
| CyberTiel 25G-ICE | 24.85 GB | 0.027866 | 0.2793 | 1.0757 | 0.9867 | 93.36 % | 7.686 | 5.599 | 96.1 |
| Sangreal 25G-ICE | 24.85 GB | 0.029038 | 0.2823 | 1.2220 | 0.9857 | 93.27 % | 7.686 | 5.599 | 96.0 |
| ≈ 22.5–23 GB | |||||||||
| CyberTiel 23G-ICE | 22.84 GB | 0.032992 | 0.3290 | 1.2449 | 0.9899 | 92.75 % | 7.523 | 5.145 | 95.6 |
| Sangreal 23G-ICE | 22.84 GB | 0.033585 | 0.3299 | 1.3758 | 0.9870 | 92.69 % | 7.523 | 5.145 | 95.5 |
| UD-Q4_K_M | 22.52 GB | 0.035692 | 0.3576 | 1.3199 | 0.9766 | 92.42 % | 7.130 | 5.075 | 95.3 |
| UD-Q4_K_XL | 22.75 GB | 0.035837 | 0.3753 | 1.3951 | 0.9761 | 92.56 % | 7.474 | 5.126 | 95.3 |
| ≈ 21 GB | |||||||||
| CyberTiel 21G-ICE | 20.85 GB | 0.039658 | 0.4136 | 1.4041 | 0.9848 | 92.02 % | 7.357 | 4.698 | 94.9 |
| Sangreal 21G-ICE | 20.85 GB | 0.039968 | 0.4156 | 1.5834 | 0.9819 | 92.03 % | 7.357 | 4.698 | 94.9 |
| UD-Q4_K_S | 21.28 GB | 0.040169 | 0.4107 | 1.5040 | 0.9808 | 91.97 % | 7.025 | 4.795 | 94.9 |
| ≈ 18 GB | |||||||||
| Sangreal 19G-ICE | 18.82 GB | 0.060435 | 0.6098 | 2.1340 | 0.9948 | 90.15 % | 7.192 | 4.241 | 93.1 |
| CyberTiel 19G-ICE | 18.82 GB | 0.060573 | 0.6220 | 2.2733 | 0.9951 | 90.15 % | 7.192 | 4.241 | 93.0 |
| UD-IQ4_XS | 18.12 GB | 0.070787 | 0.7346 | 2.6002 | 1.0441 | 89.51 % | 6.757 | 4.083 | 92.2 |
| ≈ 13.5–15 GB | |||||||||
| Sangreal 15G-ICE | 14.83 GB | 0.126386 | 1.3329 | 3.9267 | 1.0027 | 85.89 % | 6.853 | 3.341 | 87.9 |
| UD-IQ3_XXS | 13.60 GB | 0.151623 | 1.6338 | 4.3641 | 1.0920 | 84.85 % | 5.489 | 3.064 | 86.2 |
Table 2 — Code (code.test.raw)
PPL(base) 2.194208 · BF16 = 100
| file | size | mean KLD | 99% KLD | 99.9% KLD | PPL ratio | same top-1 | active bpw | file bpw | overall |
|---|---|---|---|---|---|---|---|---|---|
| ≈ 25–27 GB | |||||||||
| UD-Q5_K_XL | 26.98 GB | 0.011993 | 0.1702 | 0.8170 | 1.0009 | 97.25 % | 7.833 | 6.080 | 98.3 |
| Sangreal 25G-ICE | 24.85 GB | 0.015308 | 0.2274 | 1.0218 | 1.0017 | 97.00 % | 7.686 | 5.599 | 98.0 |
| CyberTiel 25G-ICE | 24.85 GB | 0.015497 | 0.2207 | 1.0820 | 1.0007 | 97.06 % | 7.686 | 5.599 | 98.0 |
| ≈ 22.5–23 GB | |||||||||
| Sangreal 23G-ICE | 22.84 GB | 0.018988 | 0.2831 | 1.2650 | 1.0023 | 96.73 % | 7.523 | 5.145 | 97.7 |
| CyberTiel 23G-ICE | 22.84 GB | 0.019269 | 0.2822 | 1.3248 | 1.0018 | 96.66 % | 7.523 | 5.145 | 97.7 |
| UD-Q4_K_XL | 22.75 GB | 0.019742 | 0.2909 | 1.3855 | 1.0041 | 96.61 % | 7.474 | 5.126 | 97.6 |
| UD-Q4_K_M | 22.52 GB | 0.020440 | 0.2880 | 1.3215 | 1.0040 | 96.59 % | 7.130 | 5.075 | 97.6 |
| ≈ 21 GB | |||||||||
| Sangreal 21G-ICE | 20.85 GB | 0.023415 | 0.3529 | 1.4214 | 1.0042 | 96.35 % | 7.357 | 4.698 | 97.3 |
| UD-Q4_K_S | 21.28 GB | 0.023484 | 0.3466 | 1.6264 | 1.0028 | 96.31 % | 7.025 | 4.795 | 97.3 |
| CyberTiel 21G-ICE | 20.85 GB | 0.024669 | 0.3787 | 1.6756 | 1.0044 | 96.28 % | 7.357 | 4.698 | 97.2 |
| ≈ 18 GB | |||||||||
| Sangreal 19G-ICE | 18.82 GB | 0.036887 | 0.5614 | 2.4828 | 1.0159 | 95.33 % | 7.192 | 4.241 | 96.1 |
| CyberTiel 19G-ICE | 18.82 GB | 0.036910 | 0.5798 | 2.3285 | 1.0154 | 95.24 % | 7.192 | 4.241 | 96.1 |
| UD-IQ4_XS | 18.12 GB | 0.044161 | 0.6689 | 2.8208 | 1.0191 | 94.80 % | 6.757 | 4.083 | 95.5 |
| ≈ 13.5–15 GB | |||||||||
| Sangreal 15G-ICE | 14.83 GB | 0.088158 | 1.4526 | 4.7767 | 1.0467 | 92.91 % | 6.853 | 3.341 | 92.2 |
| UD-IQ3_XXS | 13.60 GB | 0.105845 | 1.7903 | 5.2413 | 1.0594 | 91.96 % | 5.489 | 3.064 | 90.9 |
overall = 0.70/(1 + meanKLD) + 0.30 * sameTop1, ×100. BF16 = 100. Same
composite the CyberTiel card uses, so the two are directly comparable: 70% on how
close the whole output distribution stays, 30% on agreement about the argmax.
Read the tail columns as shape, not order.
99.9% KLDis roughly the 33rd-worst token of 32,768 — an extreme order statistic with large sampling variance, so it inverts between adjacent files without meaning anything. Rank onmean KLD.Two bpw columns. Only 8 of 256 experts fire per token, so a bit in
ffn_*_expsis worth ~3% of a bit in attention, the shared expert or the output head.active bpwweights by that;file bpwis just size ÷ parameters. It is why a 22.84 GB file computes at ~7.5 bpw.
Don't compare the two tables to each other. Code is more predictable text, so every file scores about half the divergence on it. Compare rows within a table.
Every row is measured against this lineage's own BF16 master — the huihui abliterated checkpoint — in one session, one binary, one reference. Quantizations cut from a different trunk are deliberately absent: scoring them here would measure the distance between trunks and call it quantization damage.
Sangreal vs the previous imatrix
The previous CyberTiel-calibrated ladder sits in both tables at identical recipes on an identical trunk, so the gap between them is purely the calibration.
Tested paired on the common chunks: Sangreal is ahead on code at every tier (−5.1% KLD at 21G, combined p = 0.019) and behind on WikiText-2 at the richer tiers (+4.2% at 25G, combined p = 0.041). The prose penalty grows as the tier gets richer — at the cheap end the bits are the constraint, at the rich end the prior is.
That is what a corpus that is 45% code, cyber and spec should do, and it is the trade
you are picking up here. Reproduce with paired_test.py against measurements/.
What Sangreal is
The calibration corpus these tiers were quantized against. Built from primary
sources, not assembled from eaddario's set, bartowski's calibration_datav5, or any
other ready-made calibration file. 77 buckets, 9,533 documents, 47,443,549 tokens,
sha256 85a6b823a6762bf6….
The full render is 92,663 chunks. These quantizations used a 17,225-chunk
pass over it — 30x the stock Ornith imatrix (573) and 21x bartowski's
calibration_datav5 (~800).
| block | buckets | share | documents | synthetic |
|---|---|---|---|---|
| code + cyber + spec | 36 | 45.10% | 5,858 | 0.1% |
| reasoning | 5 | 13.16% | 553 | 100.0% |
| agentic | 7 | 11.95% | 1,018 | 65.2% |
| science + domain | 11 | 11.24% | 1,183 | 0.3% |
| language | 11 | 9.66% | 596 | 0.0% |
| mathematics | 7 | 8.89% | 325 | 0.0% |
| total | 77 | 100% | 9,533 | 21.7% |
Split: technical 70.2 · science+domain 20.1 · language 9.7.
Synthetic is 21.7% by bytes and sits in exactly three buckets — teacher:k2-horizon,
teacher:deepseek-v4-pro, teacher:hermes. Everything outside reasoning and
agentic is real text.
Language — 11 languages, 5 non-Latin scripts
| bucket | share | docs | bucket | share | docs | |
|---|---|---|---|---|---|---|
arabic |
0.92% | 48 | russian |
0.87% | 58 | |
chinese |
0.89% | 31 | spanish |
0.87% | 72 | |
hindi |
0.88% | 44 | portuguese |
0.87% | 70 | |
japanese |
0.88% | 33 | turkish |
0.87% | 56 | |
french |
0.88% | 60 | german |
0.86% | 73 | |
english |
0.87% | 51 |
The spread across the whole block is 0.06 points, and that is deliberate:
language's job here is activation, not ranking. Each bucket sits just above the
saturation floor at 44.9% consumption, so these buckets select for the first time —
the legacy v1 spec had consumed 96% of turkish and 92% of english.
Mathematics — absent from the pool until this cycle
| bucket | share | docs |
|---|---|---|
math (general) |
3.16% | 170 |
analysis |
1.24% | 28 |
algebra |
1.16% | 46 |
geometry |
1.12% | 30 |
proof |
1.08% | 28 |
topology |
1.05% | 20 |
number-theory |
0.06% | 3 |
Sourced from the LaTeX algebraic geometry, mathlib4, UniMath, math-comp, open-web-math, AutoMathText and proof-pile.
The imatrix is published here — imatrix/sangreal-17225.imatrix.gguf, the exact
file these five tiers were quantized against. Pass it to llama-quantize --imatrix
to rebuild any tier from the BF16 master, or to cut your own.
The corpus outlives this model. An imatrix is tied to one architecture — it is a table of per-channel importance for one specific tensor list, and it transfers to nothing else. The corpus that produced it has no such limit. Sangreal is organic, model-agnostic text, so the same render calibrates an imatrix for any model, any architecture, any size.
That is the durable result here. These five tiers are the first thing built with it, not the reason it exists. The full corpus will be published.
What's inside
MTPv2 head, native. huihui's abliterated checkpoint already carries
Ornith-1.5's trained multi-token-prediction head; its mtp.* norms are bit-identical
to ornith-ai's on all seven tensors. Every tier ships blk.40 with nextn.* intact,
so llama.cpp can use it as a draft model for speculative decoding.
Chat template: peculiar-ragdoll's Qwen-Sharp v22.5.0, embedded in every tier.
It is a genuinely good template — thinking-effort control, tool-call formatting, and
a terseness layer that can be switched off per request via
chat_template_kwargs: {"terse": false}. Credited below.
ICE = Isolation of Compounding Error: allocate bits by how far a quantization error travels, not by how large the activations are. Routers stay F32, SSM decay gates stay F32, the KV-cached projections stay F16, the always-on dense path stays Q8_0 — together 0.14% of the model, kept exact for ~152 MB — and the entire budget is spent on the routed expert bank, which is 93% of the parameters but only 8-of-256 active per token. Method and the cases where it does not win: gbuzhf/ICE-quantization.
Files
measurements/— rawllama-perplexityoutput for every row of both tables.recipes/— the exactllama-quantizetensor map for each tier (443 rules).imatrix/—sangreal-17225.imatrix.gguf, the importance matrix used for all five tiers. 1020 tensors, 17,225 chunks x 512, merged from 12 disjoint partial passes. With this plusrecipes/, every tier here is reproducible byte-for-byte from the published BF16 master.
Credits
| base model | ornith-ai/Ornith-1.5-35B-A3B |
| abliterated checkpoint | huihui-ai/Huihui-Ornith-1.5-35B-A3B-abliterated |
| chat template | peculiar-ragdoll/Qwen-Sharp-Chat-Templates |
| Official CyberTiel UD ladder | peculiar-ragdoll/Cyber-Tiel-Coder-35B-A3B-GGUF-MTP |
| our previous ICE ladder | gbuzhf/…-CyberTiel-Calibrated-MTPv2-ICE-GGUF |
| ICE method | gbuzhf/ICE-quantization |
KLD measures fidelity to this repo's master and nothing else — not reasoning, tool use or speed. It is also within-lineage: every ladder is measured against its own master, so these values are not comparable to another repo's. Compare across lineages with PPL.