Xyntetik-Kvist-14B research checkpoints
Research checkpoints, not a release. To run the model, use Xyntetik-Kvist-14B or its GGUF files. These two checkpoints are for anyone who wants to continue or redo its training.
- What they are: two BF16 safetensors checkpoints of the 14.4 B-parameter Kvist student, in the released model's layout.
checkpoint-06000/is the end of general distillation and the start of every envelope attempt.a11-checkpoint-01320/is envelope attempt 11, which passed the full preregistered gate and was not released.- Weights only. No optimizer state is published. A run resumed from here starts AdamW's moments at zero, where the recorded runs carried them from segment to segment, so it will not replay the recorded loss curve step for step.
- What resuming takes: the frozen BF16 parent as the teacher, about 97 GB of host memory for the trainer state, a CUDA GPU (the recorded runs used a 24 GB slice), the rebuilt training corpus, and a few path edits in scripts written for one machine. Details below.
- Not claimed: neither checkpoint is a supported serving target. No GGUF is published for either, and the release checks (serving on Runner, the quantised rows) were run on the released checkpoint only.
- Parent model: meta-models/Muse-Glimmer-30B at
a4e59da52a7bc87ae7251dd5545c0dd437c44b68, cut by width to a dense 14B and distilled back from it- Collection: Xyntetik Kvist: distilled agents
- Evidence: training record (dataset: the trainers, the envelope policy files, preregistrations, gate records, corpus manifests)
The checkpoints
| folder | step | what it is | held-out retention against the parent | closed loop (E3) |
|---|---|---|---|---|
checkpoint-06000/ |
6,000 | general distillation, 98.3 M tokens, 162 hours | K3: KLD 0.786, margin-qualified top-1 83.9% | 0 of 60 (the gate's control) |
a11-checkpoint-01320/ |
1,320 | envelope attempt 11, resumed from attempt 10 | E4: KLD 0.7585 | 57 of 60 |
| released Kvist-14B | 1,440 | envelope attempt 12, resumed from a11-checkpoint-01320/ |
E4: KLD 0.762, margin-qualified top-1 84.0% | 57 of 60 |
All retention numbers are on the study's 45,056 held-out positions; the parent solves all 60 closed-loop tasks.
checkpoint-06000/ is the width-cut student after 6,000 steps of distillation from the frozen BF16 parent on general text (no tool-calling documents). It passed K3 (margin-qualified top-1 of at least 80%). As the gate's control it wrote 0 of 200 first turns in Runner's muse wire format and solved none of the 60 closed-loop tasks. Start here to try a different agentic phase, a different corpus, or more general distillation.
a11-checkpoint-01320/ passed every arm of the preregistered envelope gate: E1 198 of 200 first turns well formed, E2 99 of 99 tool calls valid, E3 57 of 60 closed-loop tasks, E4 KLD 0.7585, E5 3 of 200 turns looped. Under the rule registered on 2026-09-27 (hold attempt 11, run attempt 12, ship the better full pass), attempt 12 ranked first and was released. Start here to rerun attempt 12's segment or to branch off before it.
The released model's folder has the same layout (config, tokenizer and four safetensors shards with their index), so the same commands take it; its optimizer state is not published either.
Resuming
What it needs.
- The parent, meta-models/Muse-Glimmer-30B at revision
a4e59da52a7bc87ae7251dd5545c0dd437c44b68, BF16 safetensors. It is the teacher at every step; its text decoder is streamed from local disk a layer at a time, not held in GPU memory. - Host memory: the trainer keeps fp32 master weights and bf16 AdamW moments for the 12.1 B trainable parameters (embeddings and head frozen) in pinned host memory, about 8 bytes a parameter or 97 GB, before the checkpoint itself is loaded.
- A CUDA GPU. The recorded runs used a 24 GB slice; general distillation ran at about 97 seconds a step there.
- Software: PyTorch with CUDA,
numpy,safetensors,pyarrow, theggufpackage, and atransformersbuild with Muse-Glimmer support (these checkpoints were written with 5.15.0.dev0). - The code and data from the training record: the trainers in
code/, the seven study modules they import incode/deps/muse-glimmer-sparsity/scripts/, and the envelope policy files, ready to use, indata/envelope/.
Path edits. The scripts were written for one machine and hard-code its layout. Call the folder that holds the record's code/deps/muse-glimmer-sparsity/ the study folder, and give it a data/ subfolder.
STUDYat the top ofcode/train_student.py,code/score_student.pyandcode/build_student_corpus.py: the study folder.SNAPSHOTandROOTat the top ofscripts/mgcommon.py: the local parent snapshot and the study folder.- The output paths in
scripts/build_corpus.pyandscripts/run9_build_corpus.py: the study folder'sdata/. --tokenizerontrain_student_policy.py: the parent snapshot.
The corpus. The envelope phase interleaves general distillation on corpus-v3-muse. The corpus scripts read the datasets as parquet from the local Hugging Face cache, at these revisions:
hf download HuggingFaceTB/smoltalk --repo-type dataset --revision 5feaf2fd3ffca7c237fc38d1861bc30365d48ffa
hf download allenai/tulu-3-sft-mixture --repo-type dataset --revision b14afda60f1bbebe55d5d2fa1e4df5042f97f8be
hf download Salesforce/wikitext --repo-type dataset --include "wikitext-103-raw-v1/*"
build_student_corpus.py excludes the study's earlier corpora and filters against their held-out split, so build those first with build_corpus.py and run9_build_corpus.py (their recorded manifests are the record's evidence/corpus/study-corpus_manifest.json and study-corpus_run9_manifest.json), then:
python build_student_corpus.py --output corpus-v3-muse --render muse \
--cap-tokens 3000000 --tulu-tokens 6000000 --wikitext-tokens 5000000
A faithful rebuild matches the recorded tokens_sha256 (34a37a0ba8086c1144e437a87a9f02becffe1a2f69ba2afae24ebafb61d52fab, in the record's evidence/corpus/corpus-v3-muse.manifest.json); check it before training.
Download.
hf download Joakimpalm-Zen/Xyntetik-Kvist-14B-research-checkpoints \
--include "a11-checkpoint-01320/*" --local-dir kvist-checkpoints
hf download Joakimpalm-Zen/Xyntetik-Kvist-14B-training-record --repo-type dataset \
--include "code/*" "data/envelope/*" --local-dir kvist-record
Run. --student and --resume both take the downloaded folder: the trainer loads config and weights from --resume and copies the config, tokenizer, template, pruning map and license into each checkpoint it saves from --student.
# Attempt 12's segment again, from attempt 11 (the recorded command with local paths).
python train_student_policy.py --student kvist-checkpoints/a11-checkpoint-01320 \
--resume kvist-checkpoints/a11-checkpoint-01320 --tokenizer <parent snapshot> \
--corpus corpus-v3-muse --policy kvist-record/data/envelope/examples-muse-a12-mixed.jsonl \
--output train-w14b-envelope-a12 \
--steps 1500 --windows 4 --tokens 4096 --policy-windows 8 --policy-tokens 1536 \
--lr 3e-05 --warmup 20 --seed 20260922 \
--start-step 1320 --stop-at 1440 --early-stop off --general-per-policy 2 \
--kind-weights '{"turn:answer": 2, "flight": 2}' --terminator-weight 10
# A new envelope phase from the end of general distillation (attempt 2's settings).
python train_student_policy.py --student kvist-checkpoints/checkpoint-06000 \
--resume kvist-checkpoints/checkpoint-06000 --tokenizer <parent snapshot> \
--corpus corpus-v3-muse --policy kvist-record/data/envelope/examples-muse-all.jsonl \
--output my-envelope \
--steps 1500 --windows 4 --tokens 4096 --policy-windows 8 --policy-tokens 1536 \
--lr 3e-05 --warmup 20 --seed 20260922 --early-stop on
--start-step keeps the recorded learning-rate schedule and step numbering and draws a fresh sample stream from --seed plus the start step, so with the same seed, start step, policy file and corpus a resumed run draws the recorded run's samples. Because the moments restart at zero, the first steps of a resumed segment still move the weights differently from the recorded ones.
Measure. code/score_student.py gives the retention row against the parent on the held-out split (the E4 arm); the closed-loop and conformance arms need the checkpoint converted to GGUF and served on Xyntetik Runner, as the Kvist-14B card's Reproduce section shows. Every attempt's command, the gate rules and the disclosures are on the Kvist-14B card.
Files and checksums
Both folders hold the same tokenizer, template, config, pruning map and license; only the weights and STEP differ.
checkpoint-06000/
| file | bytes | sha256 |
|---|---|---|
config.json |
4,524 | bc9f6cfbf7316804894655e240de02d85d3a4c82d63bd785aa442329afd5ef4b |
generation_config.json |
202 | 1fa51889b1f8d3659802dedaa27e005b81e5c58483f13ecf13f2d97306bc6e35 |
chat_template.jinja |
9,992 | cfc67e5f349f37690dfd31ed1f18bc4442a9dd32fe39a648f993cb4eb3cae678 |
tokenizer.json |
28,129,897 | c9dbee66967b58f31a7c27f723c3760da3526ccd0427578e8905b0abb0031c4d |
tokenizer_config.json |
79,936 | 781e6c74f571642c71202167b67d9255b28cc439bdda1582ff31346182f5a9c5 |
model.safetensors.index.json |
55,358 | d2ca8996136570dbad016f17c49f0825c9b9b21fd31737e54ae5e720d01d5906 |
PRUNE.json |
3,519,819 | 0356aae40e33a54d4560aae68d7d70d01ea61d873d53db8bc11918664bb4b7ac |
LICENSE |
11,358 | cfc7749b96f63bd31c3c42b5c471bf756814053e847c10f3eb003417bc523d30 |
STEP |
4 | 8d284220975da66c6cba4a32aaedfa0b897187d86ed9f7a2419c9bd10a103505 |
model-00000.safetensors |
8,613,298,928 | 21a11296ee7ebd0583edbd259f89752675dd194d4e5b115a36838996ebac1fd9 |
model-00001.safetensors |
8,624,085,248 | 26ec9a1c9af0150f4569d74f33a40408e17239c4aeb9656749fb46563b868314 |
model-00002.safetensors |
8,618,234,144 | c4e166a8dc3f16ad2392bbaef7fc5788e16f87626d9f072da4b442b197e3bccd |
model-00003.safetensors |
3,032,028,272 | 369e180fde8289bdeb746f21dbfab2f543c6468d531fc7e0e24cdceaddb949b3 |
a11-checkpoint-01320/
| file | bytes | sha256 |
|---|---|---|
config.json |
4,524 | bc9f6cfbf7316804894655e240de02d85d3a4c82d63bd785aa442329afd5ef4b |
generation_config.json |
202 | 1fa51889b1f8d3659802dedaa27e005b81e5c58483f13ecf13f2d97306bc6e35 |
chat_template.jinja |
9,992 | cfc67e5f349f37690dfd31ed1f18bc4442a9dd32fe39a648f993cb4eb3cae678 |
tokenizer.json |
28,129,897 | c9dbee66967b58f31a7c27f723c3760da3526ccd0427578e8905b0abb0031c4d |
tokenizer_config.json |
79,936 | 781e6c74f571642c71202167b67d9255b28cc439bdda1582ff31346182f5a9c5 |
model.safetensors.index.json |
55,358 | d2ca8996136570dbad016f17c49f0825c9b9b21fd31737e54ae5e720d01d5906 |
PRUNE.json |
3,519,819 | 0356aae40e33a54d4560aae68d7d70d01ea61d873d53db8bc11918664bb4b7ac |
LICENSE |
11,358 | cfc7749b96f63bd31c3c42b5c471bf756814053e847c10f3eb003417bc523d30 |
STEP |
4 | 5f0fbfc126a9876e177921aeb69a9705447e5b4f29fc96631ab4ac6d31c62717 |
model-00000.safetensors |
8,613,298,928 | 56949e91d5c1e26461fe400021cb82664a20af816c16a4a5163d14336313ab0a |
model-00001.safetensors |
8,624,085,248 | 65fa18dc9a68fab9819b0a7a573d2b1ee1a395fe33c3fcd7755213eeed7305ec |
model-00002.safetensors |
8,618,234,144 | 13fbc60619d23abb7f5b9e8076361f11e1301239716e5f932a53a9754b3f1f18 |
model-00003.safetensors |
3,032,028,272 | 4e23e5b7464d6672b150f8165cb110e39ae3dc0ba1ee7de1062a613b1669c2fa |
License
Apache-2.0, as the parent. Muse Glimmer's usage policy applies to these weights as to the parent's.