logic65/Qwen3.8-Whittle-tri-14.7B

🤗 Hugging Face 来源text-generationapache-2.0激活 14.7B63 GBGGUF✓ 14 个校验和今天更新
一条命令提交

在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo logic65/Qwen3.8-Whittle-tri-14.7B ./model-folder
需要做种者 →

⚠️ NEEDS POST TRAINING. A research preview, published as the record of a

compression method and its measurements rather than as a finished assistant. Further healing and instruction tuning runs are planned over time, and these checkpoints will improve as those land.

Qwen3.8-Whittle-tri (14.7B)

A 27B compressed to 14.7B by training, not cutting. It writes code, answers factual questions, never loops, and fits on two consumer 8GB cards.

Download q38_tri_repaired_q8.gguf (15.7GB) and run it. That is the whole setup.

llama-server -m q38_tri_repaired_q8.gguf -ngl 24 -c 4096 --jinja

Standard [gdn, gdn, gdn, attention] blocks at exactly 3:1, so it loads in any llama.cpp build with Qwen3.5 support. No fork, no patches, no special flags. On 16GB of VRAM keep about 24 of the 32 layers on GPU; with 24GB it fits whole.

What it is good at

Factual recall is close to intact: capitals, elements, dates, who wrote what. Code is solid across Python, C, JavaScript, SQL and HTML, and markdown fencing is reliable (8/8 on our fencing probe). It is loop-free: zero repetition failures in six generations at both greedy and temperature 0.7, which is not true of the larger models in this family.

What it is bad at

Arithmetic and multi-step reasoning. It scores 3.6% on GSM8K. It will often state the right rule and then compute the wrong answer. Do not use it for maths, data analysis, or anything where a number has to be right.

Non-English is unreliable outside French and Spanish. Long agent sessions were not the target of this model's training.

Honest positioning

39-prompt battery: 31/39. That ties a variant 2.1B larger and beats the 19.2B cut, but two models in the family score higher (36/39 at 16.8B, 35/39 at 20.8B). Pick this one if you want the smallest thing that still works, or if looping has bitten you before. Pick the bigger ones if you want maximum quality.

This is a research preview. It needs real post-training before it reasons, and that is a funding problem rather than a technical one.

How it was made, briefly

The 64 layers of Qwen3.8-27B were distilled into 32: adjacent recurrent layers were merged in pairs, and eight full-attention stations were absorbed into their neighbours rather than deleted. Nothing was dropped. Full methodology, every script, the adapters and the measurements are in the developer card: Whittle dev

Collections: models you can run and research and dev

Support this work

Independent research on consumer hardware. Every donation becomes GPU hours, and every GPU hour ends up as a public model or a public measurement. ☕ ko-fi.com/davida81328

Base model by the Qwen team (Apache 2.0). Built and measured by David Aylward with Claude (Fable 5, Anthropic) as co-author.