Joakimpalm-Zen/Xyntetik-Kvist-14B-research-checkpoints

🤗 Hugging Face 来源apache-2.0激活 14B58 GBsafetensors✓ 10 个校验和今天更新
需要做种者 →

Xyntetik-Kvist-14B research checkpoints

Research checkpoints, not a release. To run the model, use Xyntetik-Kvist-14B or its GGUF files. These two checkpoints are for anyone who wants to continue or redo its training.

  • What they are: two BF16 safetensors checkpoints of the 14.4 B-parameter Kvist student, in the released model's layout. checkpoint-06000/ is the end of general distillation and the start of every envelope attempt. a11-checkpoint-01320/ is envelope attempt 11, which passed the full preregistered gate and was not released.
  • Weights only. No optimizer state is published. A run resumed from here starts AdamW's moments at zero, where the recorded runs carried them from segment to segment, so it will not replay the recorded loss curve step for step.
  • What resuming takes: the frozen BF16 parent as the teacher, about 97 GB of host memory for the trainer state, a CUDA GPU (the recorded runs used a 24 GB slice), the rebuilt training corpus, and a few path edits in scripts written for one machine. Details below.
  • Not claimed: neither checkpoint is a supported serving target. No GGUF is published for either, and the release checks (serving on Runner, the quantised rows) were run on the released checkpoint only.
  • Parent model: meta-models/Muse-Glimmer-30B at a4e59da52a7bc87ae7251dd5545c0dd437c44b68, cut by width to a dense 14B and distilled back from it
  • Collection: Xyntetik Kvist: distilled agents
  • Evidence: training record (dataset: the trainers, the envelope policy files, preregistrations, gate records, corpus manifests)

The checkpoints

folder step what it is held-out retention against the parent closed loop (E3)
checkpoint-06000/ 6,000 general distillation, 98.3 M tokens, 162 hours K3: KLD 0.786, margin-qualified top-1 83.9% 0 of 60 (the gate's control)
a11-checkpoint-01320/ 1,320 envelope attempt 11, resumed from attempt 10 E4: KLD 0.7585 57 of 60
released Kvist-14B 1,440 envelope attempt 12, resumed from a11-checkpoint-01320/ E4: KLD 0.762, margin-qualified top-1 84.0% 57 of 60

All retention numbers are on the study's 45,056 held-out positions; the parent solves all 60 closed-loop tasks.

checkpoint-06000/ is the width-cut student after 6,000 steps of distillation from the frozen BF16 parent on general text (no tool-calling documents). It passed K3 (margin-qualified top-1 of at least 80%). As the gate's control it wrote 0 of 200 first turns in Runner's muse wire format and solved none of the 60 closed-loop tasks. Start here to try a different agentic phase, a different corpus, or more general distillation.

a11-checkpoint-01320/ passed every arm of the preregistered envelope gate: E1 198 of 200 first turns well formed, E2 99 of 99 tool calls valid, E3 57 of 60 closed-loop tasks, E4 KLD 0.7585, E5 3 of 200 turns looped. Under the rule registered on 2026-09-27 (hold attempt 11, run attempt 12, ship the better full pass), attempt 12 ranked first and was released. Start here to rerun attempt 12's segment or to branch off before it.

The released model's folder has the same layout (config, tokenizer and four safetensors shards with their index), so the same commands take it; its optimizer state is not published either.

Resuming

What it needs.

  • The parent, meta-models/Muse-Glimmer-30B at revision a4e59da52a7bc87ae7251dd5545c0dd437c44b68, BF16 safetensors. It is the teacher at every step; its text decoder is streamed from local disk a layer at a time, not held in GPU memory.
  • Host memory: the trainer keeps fp32 master weights and bf16 AdamW moments for the 12.1 B trainable parameters (embeddings and head frozen) in pinned host memory, about 8 bytes a parameter or 97 GB, before the checkpoint itself is loaded.
  • A CUDA GPU. The recorded runs used a 24 GB slice; general distillation ran at about 97 seconds a step there.
  • Software: PyTorch with CUDA, numpy, safetensors, pyarrow, the gguf package, and a transformers build with Muse-Glimmer support (these checkpoints were written with 5.15.0.dev0).
  • The code and data from the training record: the trainers in code/, the seven study modules they import in code/deps/muse-glimmer-sparsity/scripts/, and the envelope policy files, ready to use, in data/envelope/.

Path edits. The scripts were written for one machine and hard-code its layout. Call the folder that holds the record's code/deps/muse-glimmer-sparsity/ the study folder, and give it a data/ subfolder.

  1. STUDY at the top of code/train_student.py, code/score_student.py and code/build_student_corpus.py: the study folder.
  2. SNAPSHOT and ROOT at the top of scripts/mgcommon.py: the local parent snapshot and the study folder.
  3. The output paths in scripts/build_corpus.py and scripts/run9_build_corpus.py: the study folder's data/.
  4. --tokenizer on train_student_policy.py: the parent snapshot.

The corpus. The envelope phase interleaves general distillation on corpus-v3-muse. The corpus scripts read the datasets as parquet from the local Hugging Face cache, at these revisions:

hf download HuggingFaceTB/smoltalk --repo-type dataset --revision 5feaf2fd3ffca7c237fc38d1861bc30365d48ffa
hf download allenai/tulu-3-sft-mixture --repo-type dataset --revision b14afda60f1bbebe55d5d2fa1e4df5042f97f8be
hf download Salesforce/wikitext --repo-type dataset --include "wikitext-103-raw-v1/*"

build_student_corpus.py excludes the study's earlier corpora and filters against their held-out split, so build those first with build_corpus.py and run9_build_corpus.py (their recorded manifests are the record's evidence/corpus/study-corpus_manifest.json and study-corpus_run9_manifest.json), then:

python build_student_corpus.py --output corpus-v3-muse --render muse \
    --cap-tokens 3000000 --tulu-tokens 6000000 --wikitext-tokens 5000000

A faithful rebuild matches the recorded tokens_sha256 (34a37a0ba8086c1144e437a87a9f02becffe1a2f69ba2afae24ebafb61d52fab, in the record's evidence/corpus/corpus-v3-muse.manifest.json); check it before training.

Download.

hf download Joakimpalm-Zen/Xyntetik-Kvist-14B-research-checkpoints \
    --include "a11-checkpoint-01320/*" --local-dir kvist-checkpoints
hf download Joakimpalm-Zen/Xyntetik-Kvist-14B-training-record --repo-type dataset \
    --include "code/*" "data/envelope/*" --local-dir kvist-record

Run. --student and --resume both take the downloaded folder: the trainer loads config and weights from --resume and copies the config, tokenizer, template, pruning map and license into each checkpoint it saves from --student.

# Attempt 12's segment again, from attempt 11 (the recorded command with local paths).
python train_student_policy.py --student kvist-checkpoints/a11-checkpoint-01320 \
    --resume kvist-checkpoints/a11-checkpoint-01320 --tokenizer <parent snapshot> \
    --corpus corpus-v3-muse --policy kvist-record/data/envelope/examples-muse-a12-mixed.jsonl \
    --output train-w14b-envelope-a12 \
    --steps 1500 --windows 4 --tokens 4096 --policy-windows 8 --policy-tokens 1536 \
    --lr 3e-05 --warmup 20 --seed 20260922 \
    --start-step 1320 --stop-at 1440 --early-stop off --general-per-policy 2 \
    --kind-weights '{"turn:answer": 2, "flight": 2}' --terminator-weight 10

# A new envelope phase from the end of general distillation (attempt 2's settings).
python train_student_policy.py --student kvist-checkpoints/checkpoint-06000 \
    --resume kvist-checkpoints/checkpoint-06000 --tokenizer <parent snapshot> \
    --corpus corpus-v3-muse --policy kvist-record/data/envelope/examples-muse-all.jsonl \
    --output my-envelope \
    --steps 1500 --windows 4 --tokens 4096 --policy-windows 8 --policy-tokens 1536 \
    --lr 3e-05 --warmup 20 --seed 20260922 --early-stop on

--start-step keeps the recorded learning-rate schedule and step numbering and draws a fresh sample stream from --seed plus the start step, so with the same seed, start step, policy file and corpus a resumed run draws the recorded run's samples. Because the moments restart at zero, the first steps of a resumed segment still move the weights differently from the recorded ones.

Measure. code/score_student.py gives the retention row against the parent on the held-out split (the E4 arm); the closed-loop and conformance arms need the checkpoint converted to GGUF and served on Xyntetik Runner, as the Kvist-14B card's Reproduce section shows. Every attempt's command, the gate rules and the disclosures are on the Kvist-14B card.

Files and checksums

Both folders hold the same tokenizer, template, config, pruning map and license; only the weights and STEP differ.

checkpoint-06000/

file bytes sha256
config.json 4,524 bc9f6cfbf7316804894655e240de02d85d3a4c82d63bd785aa442329afd5ef4b
generation_config.json 202 1fa51889b1f8d3659802dedaa27e005b81e5c58483f13ecf13f2d97306bc6e35
chat_template.jinja 9,992 cfc67e5f349f37690dfd31ed1f18bc4442a9dd32fe39a648f993cb4eb3cae678
tokenizer.json 28,129,897 c9dbee66967b58f31a7c27f723c3760da3526ccd0427578e8905b0abb0031c4d
tokenizer_config.json 79,936 781e6c74f571642c71202167b67d9255b28cc439bdda1582ff31346182f5a9c5
model.safetensors.index.json 55,358 d2ca8996136570dbad016f17c49f0825c9b9b21fd31737e54ae5e720d01d5906
PRUNE.json 3,519,819 0356aae40e33a54d4560aae68d7d70d01ea61d873d53db8bc11918664bb4b7ac
LICENSE 11,358 cfc7749b96f63bd31c3c42b5c471bf756814053e847c10f3eb003417bc523d30
STEP 4 8d284220975da66c6cba4a32aaedfa0b897187d86ed9f7a2419c9bd10a103505
model-00000.safetensors 8,613,298,928 21a11296ee7ebd0583edbd259f89752675dd194d4e5b115a36838996ebac1fd9
model-00001.safetensors 8,624,085,248 26ec9a1c9af0150f4569d74f33a40408e17239c4aeb9656749fb46563b868314
model-00002.safetensors 8,618,234,144 c4e166a8dc3f16ad2392bbaef7fc5788e16f87626d9f072da4b442b197e3bccd
model-00003.safetensors 3,032,028,272 369e180fde8289bdeb746f21dbfab2f543c6468d531fc7e0e24cdceaddb949b3

a11-checkpoint-01320/

file bytes sha256
config.json 4,524 bc9f6cfbf7316804894655e240de02d85d3a4c82d63bd785aa442329afd5ef4b
generation_config.json 202 1fa51889b1f8d3659802dedaa27e005b81e5c58483f13ecf13f2d97306bc6e35
chat_template.jinja 9,992 cfc67e5f349f37690dfd31ed1f18bc4442a9dd32fe39a648f993cb4eb3cae678
tokenizer.json 28,129,897 c9dbee66967b58f31a7c27f723c3760da3526ccd0427578e8905b0abb0031c4d
tokenizer_config.json 79,936 781e6c74f571642c71202167b67d9255b28cc439bdda1582ff31346182f5a9c5
model.safetensors.index.json 55,358 d2ca8996136570dbad016f17c49f0825c9b9b21fd31737e54ae5e720d01d5906
PRUNE.json 3,519,819 0356aae40e33a54d4560aae68d7d70d01ea61d873d53db8bc11918664bb4b7ac
LICENSE 11,358 cfc7749b96f63bd31c3c42b5c471bf756814053e847c10f3eb003417bc523d30
STEP 4 5f0fbfc126a9876e177921aeb69a9705447e5b4f29fc96631ab4ac6d31c62717
model-00000.safetensors 8,613,298,928 56949e91d5c1e26461fe400021cb82664a20af816c16a4a5163d14336313ab0a
model-00001.safetensors 8,624,085,248 65fa18dc9a68fab9819b0a7a573d2b1ee1a395fe33c3fcd7755213eeed7305ec
model-00002.safetensors 8,618,234,144 13fbc60619d23abb7f5b9e8076361f11e1301239716e5f932a53a9754b3f1f18
model-00003.safetensors 3,032,028,272 4e23e5b7464d6672b150f8165cb110e39ae3dc0ba1ee7de1062a613b1669c2fa

License

Apache-2.0, as the parent. Muse Glimmer's usage policy applies to these weights as to the parent's.