⚠️ NEEDS POST TRAINING. A research preview, published as the record of a
compression method and its measurements rather than as a finished assistant. Further healing and instruction tuning runs are planned over time, and these checkpoints will improve as those land.
Qwen3.8-Whittle-v40 (16.8B) best measured cut
🔬 RESEARCH PREVIEW. Smaller than every other cut in the family and
better than all of them. Chosen by measurement, not by assumption. No training.
40 layers of Qwen3.8-27B, kept in the exact [gdn, gdn, gdn, attention] pattern
llama.cpp requires: 30 gated-DeltaNet layers, 10 full-attention stations, a true
3:1 ratio. Which layers survive was decided by a per-layer identity map (cosine
between each layer's input and output residual stream over 80 probes): within
every span between surviving attention stations, the three most load-bearing
recurrent layers are kept and the rest are dropped.
Why this beats the earlier cuts
Measured with an 80-probe, 8-domain knowledge atlas, scored as the mean fraction of the intact model's next-token probability retained:
| shape | params | layers | retention |
|---|---|---|---|
| v40 (this) | 16.8B | 40 | 0.648 |
| 44L cut | 19.2B | 44 | 0.408 |
| same size, attention crowded early | 16.8B | 40 | 0.379 |
| same rule, 36 layers | 15.1B | 36 | 0.441 |
Two lessons are in that table. Choosing layers by measurement instead of removing contiguous bands is worth more than 2 billion parameters. And where the attention stations sit matters more than how many there are: three variants at identical size and identical selection rule score 0.648, 0.544 and 0.379 purely by station placement.
Against the intact 27B, this cut is at or above the original on agentic control (0.0922 vs 0.0885), Python (0.0030 vs 0.0023) and science (0.0035 vs 0.0034), holds 90% on math, and keeps 58% of multilingual recall where the 44L cut kept 3%. Its end-of-turn probability at a finished answer, the quantity that governs looping, stays at 0.46 and 0.44 versus the original's 0.53 and 0.33.
Honest state
No repair training has been done. It is a research preview: expect dulled
confidence on hard recall and the family's known rough edges. C++ recall is the
weakest domain, being distributed across layers no cut preserves. Anti-loop
sampling is still recommended:
--dry-multiplier 0.8 --dry-base 1.75 --dry-allowed-length 4 --repeat-penalty 1.15
Support this work
Independent research on consumer hardware. Every donation becomes A100 hours, and every A100 hour ends up as a public model or a public measurement. ☕ ko-fi.com/davida81328
Acknowledgements
Base model by the Qwen team (Apache 2.0). Cut and measured by David Aylward with Claude (Fable 5, Anthropic) as co-author.