logic65/Qwen3.8-Whittle-tri-kd-lora

🤗 Hugging Face 来源apache-2.0488 MBsafetensors✓ 2 个校验和今天更新
一条命令提交

在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo logic65/Qwen3.8-Whittle-tri-kd-lora ./model-folder
需要做种者 →

Whittle 14.7B logit-distillation adapter (research preview)

LoRA r32 trained to match the output distribution of the Qwen3.8-27B parent, not just the reference answers. Targets are the parent's top-128 log-probs per token (recorded once over the heal corpus), and the loss is KL against those plus a small cross-entropy anchor, taken on assistant turns only.

Why distribution rather than labels: the base model already predicts this corpus almost perfectly (cross-entropy ~4e-05) while its distribution sits ~2.5 nats from the parent's. Hard labels had nothing left to teach; the distribution did.

3.90M tokens, ~1.2 epochs, 80 minutes on one A100. The teacher pass ran the parent at nf4, so the deep tail of the targets is approximate.

Research preview, needs further work. Published as a record of method and measurements. See the Whittle collection for the models themselves.