Whittle-Qwen-3.8-35B-A3B
☕ Support this work
Whittle is built by one person on a grocery budget and rented GPU hours. This model is the first in the line whose memory carries knowledge; the next steps — more memory rows, fact-dense training data, the on-policy distillation — are compute we cannot currently pay for. If this research is useful to you, or you want to see it finished: ko-fi.com/davida81328. Every hour of GPU time goes straight into the next checkpoint, and every checkpoint, table and log lands in these repos.
A 35.1 B-parameter, ~3 B-active mixture-of-experts in the Qwen3.8-Flash-Next (qwen4_exp) format whose 10 B-parameter hashed n-gram memory is
load-bearing — the first Whittle where zeroing the memory measurably hurts the model. Distilled from Qwen3.8-27B (thinking on) at the logit level on
25,000 verified teacher traces (published with their top-20 logprobs), built on
Whittle-Next-27B-A3B v4.4, with memory contents transferred from
Qwen3.8-Flash-Next's own table. Runs on stock llama.cpp, no patches.
GGUFs (Q8_0 → Q3_K_M) are on logic65/Whittle-Qwen-3.8-35B-A3B-GGUF. This repo holds the
full weights (model-*.safetensors, 35.1 B incl. the memory), the config/tokenizer, and every probe in eval/.
Which weights are which (28 Sep 2026). The root is Phase-2 step 6000 (p2-s6000): lw5 plus 6,000 steps of offline logit distillation from Qwen3.8-27B
on the teacher's own verified reasoning traces (see Phase 2 below). Its predecessors are preserved unchanged: agentfix2 under bf16-agentfix2/, lw5 under
bf16-lw5/, lw2 under bf16-lw2/ and the first release tbl1 under bf16-tbl1/. The Phase-2 run continues; its checkpoints and memory tables are on the
dev branch under p2/.
Research preview, distillation in progress. This model uses the Qwen3.8-Flash-Next architecture (
qwen4_exp: hyper-connections, gated DeltaNet + attention, the 10 B n-gram memory), and its memory rows come from Qwen3.8-Flash-Next's own table. Phase 2 is the broad distillation: 25,178 complete, graded Qwen3.8-27B answers across maths, code, chat, strict JSON and executed agent episodes, each with the teacher's top-20 distribution at every token. Step 6000 has trained on 3,276 of the 22,950 rows that fit its 4k window (14 %); the 1,470 long-context rows are not used yet, and the on-policy phase has not started. It is not a complete distillation and should not be compared with Qwen3.8 on equal terms. Read Measured and Caveats before relying on it.
What is in the box
| this model | |
|---|---|
| total parameters | 35.1 B = 25.1 B body + 10.0 B n-gram memory |
| active per token | ~3 B (8 of 180 routed experts + shared expert; the memory is a lookup) |
| memory | 8 hash heads × 4,880,000 rows × 256, bigram + trigram, injected before layer 2; rows = exact Qwen3.8-Flash-Next row-sets per visited bucket, zero elsewhere |
| body | 40 layers, hidden 2048, GDN + full attention, hyper-connection residual streams (count 4) |
| context | 262k positions (as the parent); tested to 75k |
| teacher | Qwen3.8-27B, thinking on, complete traces + per-position top-128 distributions |
Phase 2: offline logit distillation (the root, 28 Sep 2026)
Teacher data. Qwen3.8-27B (FP8, 98.1 % top-1 agreement with bf16) answered 19,500 prompts with thinking on at its highest effort; 18,651 (96 %) passed completeness and correctness checks: the answer ended on its own, the thinking closed, maths/GSM/JSON graded exact, long-context answers exact, agent episodes executed in a sandbox and kept only when hidden unit tests passed. That gives 25,178 training rows (agent episodes split into turns), 78 M tokens, and the teacher's top-20 raw log-probabilities at each of the 25.4 M answer tokens. The set is public: logic65/whittle-distill-qwen3.8-27b. Prompts are stored without the teacher's high-effort instruction, so the student learns careful reasoning as its default (context distillation).
Training. From lw5, one 96 GB Blackwell, ~7.1 s/step, 4,096-token windows: forward KL to the teacher's renormalised top-20 on answer tokens plus reply-only cross-entropy on 60 % of steps, code-corpus rows for the memory on the rest, the dependence loss with its margin raised to 7.5 nats, LoRA r8 on all 40 layers plus shared experts and routers, the memory rows trained jointly (lr 0.02), PoSE positions to 131k. Step 6000 = 12 hours.
It thinks longer than its predecessors (≈1,030 reasoning tokens per MATH-60 problem against ≈600): it was trained on the teacher's high-effort traces. Four of the ten
problems that hit the 6,144-token cap were solved when given room, so budget max_tokens accordingly (8k+ for maths).
Structured output follows the instruction now. Asked for a bare JSON object, lw5 wrapped 65 of 72 replies in a markdown code fence (and on the long-context gate earlier checkpoints wrote an unquoted string value in about one reply in three). p2-s6000 returned a bare, valid object on every item, including the two schemas it never trained on (a shipment record and a nested support ticket); the one miss read a quantity wrong.
The memory in Phase 2. The trainer's own in-run probe (memory off minus on, 16 windows per held set: noisy, and not the routine of the memory table below) shows the
memory's contribution on unseen code rising from +20.4 at step 200 to +32.4 at step 6000 and on general text from +7.5 to +10.2, while on chat rows it fell from +5.3
to +0.5 and on science from +3.2 to +0.6. The dependence term trains only on corpus rows, where the gap (+26 to +42) sits far above its margin, so nothing holds the table on
chat-like text while the body learns the teacher's replies. Every eval line is in the training log on the dev branch.
Measured
All numbers are ours, on the same machine with the same scripts; replies and logs are in eval/ (Phase 2 in eval/p2_s6000/).
Served as the Q8_0 GGUF on stock llama.cpp (3× RTX 3060, memory in RAM), serving sampler, thinking on, one sample per item:
| probe | lw2 | lw5 | agentfix2 | p2-s6000 (root) |
|---|---|---|---|---|
| stop/loop battery, 24 prompts | 24/24 clean | 23/24 | 24/24 clean | 24/24 clean, 0 loops, 0 self-turn leaks |
| GSM8K test, 50 unseen problems | 41/50 | 41/50 | 40/50 | 48/50 |
| MATH-60 (MATH train L2–4, in no training set), 6,144-token cap | 44/60 | 44/60 | 41/60 | 44/60 (L2 19, L3 13, L4 12; 10 at the cap) |
| MATH-60, the capped problems re-run with 16,384 tokens | — | — | — | 48/60 |
| strict-JSON hold-out: 72 new items, 6 schemas (2 never trained), bare JSON demanded | — | 7/72 bare (65 fenced), 60/72 correct content | — | 72/72 valid JSON, 71/72 exact |
| tool use: stops after a successful Write, ~2k / ~2.5k / ~11k-token contexts (24 each) | — | 14 / 13 / 7 | 21 / 20 / 20 | 24 / 24 / 24 |
| tool use: the extraction step before that Write, Claude Code-sized context | — | 22/24 | 22/24 | 24/24 |
| long-context gate, 5 seeds × 6 levels, strict JSON | 19.6/36 (16–25) | 13.8/36 (0–26) | — | — |
| same gate, reading only: score when the answer parsed | 4.90 / 6 | 4.93 / 6 | — | — |
lw5, agentfix2 and p2-s6000 were run back to back on one harness (27–28 Sep); the lw2 column and the long-context rows come from lw2's and lw5's release runs (lw5 scored 24/24 and 42/50 there). MATH-60 on earlier checkpoints: tbl1 46/60, v4.4 48/60, v4.3 43/60 — a handful of problems on a 60-problem probe, to be read as such.
Read that long-context row carefully, because we nearly published a wrong number. The gate asks six questions about a real pull request with the diff plus growing amounts of the repository as context, and demands a bare JSON object. It scores all-or-nothing per level: one unquoted value and six correct answers score zero. A single run of it is close to a coin flip — the same checkpoint scored 30/36 and 17/36 on consecutive runs of identical prompts. Across five seeds, reading is indistinguishable between lw2 and lw5 (4.90 vs 4.93 of 6 when the output parses); what differs is how often the output is valid JSON. Earlier versions of this card quoted a single lucky run; these are means with their ranges. If you need structured output from this model, constrain it with a JSON schema at serve time rather than trusting it to punctuate.
The production Qwen3.6-35B-A3B reviewer we run scores 3/3/2/4/3/3 on the single-sample version of the same gate. Replies for every probe are in eval/.
The memory carries knowledge. "Memory gain" = held-out cross-entropy with the memory zeroed minus with it on, on rows never trained on; positive means the body needs the table. Every previous Whittle sat at ±0.003 (the experts had learned around the memory). Relying on the table is the design: the 256→180 expert carve removed knowledge, and the memory is its replacement — so the table must always be served whole.
| held-out rows | tbl1 | lw2 | lw5 | agentfix2 |
|---|---|---|---|---|
| code (unseen files) | +2.12 | +2.84 | +10.79 | +4.93 |
| general text | +0.22 | +0.67 | +3.90 | +1.69 |
| chat rows | +0.017 | +0.08 | +2.99 | +1.02 |
| science cards | +0.03 | +0.41 | +1.44 | +0.18 |
p2-s6000 has not been through this routine yet; it runs on the final Phase-2 checkpoint (see The memory in Phase 2 above for the in-run trend). agentfix2 gave back about half of lw5's gain in 300 steps because the dependence term was already satisfied on every corpus row (gap about +6 nats against a margin of 1.0); Phase 2 raised the margin to 7.5.
The jump at lw5 came from one constant. The training term that rewards routing knowledge through the memory compares cross-entropy with the memory off against on, and only applies when the difference falls below a margin. That margin was 0.3 nats while the actual difference ran 1.9 to 5.4, so the term had been satisfied on essentially every row and had stopped doing anything after the first few hundred steps. Raising it to 1.0 keeps it acting on the weakest tenth of rows for the whole run.
The memory's output is ~0.28× the residual norm at layer 2; the difference from earlier versions is that the layers above now use it. Cross-entropy with the memory did not degrade as the gain grew: on held code it is 1.2511 at lw5 against 1.257 at lw2, so the body is leaning on the table rather than being hollowed out.
Teacher parity (logit-lens agreement of the student's readout with the teacher's, 400 unseen 512-token rows, 3,840 positions per pair; measured on the final Phase-2 checkpoint at the end of the run):
| pair | v4.4 | tbl1 | lw2 | lw5 |
|---|---|---|---|---|
| student L39 ← teacher L63, top-1 agree / teacher-mass | 87.4 % / 97.2 % | 87.5 % / 97.2 % | 87.4 % / 97.2 % | 87.9 % / 97.3 % |
| student L35 ← teacher L59 | 21.2 % / 49.5 % | 40.3 % / 68.7 % | 44.5 % / 71.2 % | 43.0 % / 71.0 % |
Held-out CE with the memory on: tbl1 chat 1.2098 / general 2.1141; lw2 1.2131 / 2.1304; lw5 1.2055 / 2.1299; agentfix2 1.2027 / 2.1271; p2-s6000 1.2262 / 2.1143 (the Phase-2 run's own eval, which scored lw5 1.2041 / 2.1272 at its start; v4.4: 1.2332 / 2.1218). Chat-row CE rising under distillation is the expected direction: the teacher's own cross-entropy on those rows is higher than the student's.
Agent use: the post-success rewrite loop (fixed in agentfix2)
An outside Claude Code evaluation (25 Sep) found the lw5 root repeating identical Write calls until the turn limit. We reproduced it with synthetic tool-use probes (24 items each, graded on the tool-call arguments, serving sampler, Q8_0) and traced it: with thinking on, after a tool result reports success, the model's next reasoning block restarts the task from the user's original request and does the work again. It only happens when the client does not send the earlier reasoning back (the chat template then shows those turns with an empty thinking block), and it grows with the amount of tool output in between.
The two tool use rows of the table above are these probes. With the earlier reasoning passed back in the history, lw5 already stopped correctly 24/24; p2-s6000 does so without it.
agentfix2 is lw5 plus 300 training steps on 212 rejection-sampled episodes of lw5's own tool use — stop after a success, run the next step, write after a read, retry after
an error; only replies our grader accepted — trained on the reply tokens only, 20 % of steps, the rest code corpus. Clients that keep reasoning_content in the history
avoid the loop on either checkpoint: we measured llama.cpp's Anthropic endpoint (what Claude Code talks to), the Vercel AI SDK and LiteLLM all carrying it through.
Claude Code now runs. Claude Code sends its environment block as a system message after the user turn (and another after every tool result). The previous chat template
raised System message must be at the beginning, so every Claude Code request failed; the template now renders a system message wherever it arrives. Conversations with
one leading system message render byte-for-byte as before. Claude Code completed a write-then-read task end to end against llama-server's /v1/messages with this template.
How it was made
- Memory transfer. Qwen3.8-Flash-Next's 320 M-row table (16 heads × 20 M × 160, fp8) was read shard by shard; for every bigram/trigram in a 172 M-token counting corpus (code, the teacher traces, chat), the exact Qwen row-set (4 kept heads × 160) was written into our 5× larger hash geometry. Bigram buckets end up holding single n-grams on average (mean row norm 1.05× a raw Qwen row-set); trigram buckets still average ~15 colliding n-grams. Unvisited buckets are exact zeros. Hash contract unchanged from Whittle-Next (rows-per-head ×5).
- Table-first warm-up (30 min): only the memory's projections and rows learned.
- Joint training with a dependence loss (3.3 h on one 96 GB Blackwell). On a quarter of the corpus rows the trainer runs the same row a
second time with the memory zeroed and penalises
relu(0.3 − (CE_off − CE_on)): the model is punished whenever it is just as good without the memory. It cannot start from nothing — while the memory carries no information the two forwards have identical gradients — but once anything predictive flows it rewards routing knowledge through the table. The gap went +0.001 → +1.69 nats over the run, with the margin met on 97 % of memory rows at the end. (At this margin the term goes quiet once the gap clears 0.3; see step 5.) Alongside: forward-KL distillation on 1,840 complete Qwen3.8-27B thinking traces (top-128 per position) and a gentle layer-wise steer (student layers 35/39 toward teacher layers 59/63, weight 0.15 on a quarter of steps). → tbl1. - lw2: 3,624 more steps, the layer-wise steer at weight 0.2 on half the steps, and PoSE (2k-token rows placed at random position offsets up to 131k, so the rotary positions the model meets at 60–75k context are trained, not extrapolated).
- lw5: 1,643 more steps with three changes. The training mix moved to three quarters teacher reasoning traces. The layer-wise steer dropped to weight 0.1 on a quarter of steps and was pointed at a new teacher-layer cache built from maths reasoning rather than general text, because the old cache was the reason the steer cost arithmetic. And the dependence margin went from 0.3 to 1.0, which is what made the memory load-bearing in earnest.
- agentfix2: 300 more steps from lw5 for the tool-use loop. lw5 was sampled on 235 synthetic agent contexts (Claude Code-style tools and working directories) and only the replies a grader accepted were kept; the 212 that fit whole in a 2,048-token window (so the task and the working directory are always in view) were trained with cross-entropy on the reply tokens only, on 20 % of steps. A first attempt that also trained on the prompts — random paths, hashes and IDs — taught the model to discount the memory within 200 steps and was stopped.
- Phase 2 (the root, step 6000): offline logit distillation from lw5 on 25k verified Qwen3.8-27B traces; see Phase 2 above.
Run it
llama-server -m Whittle-Qwen-3.8-35B-A3B-Q8_0.gguf -ngl 99 -c 16384 --jinja -fa on -ot per_layer_token_embd=CPU
-ot per_layer_token_embd=CPUkeeps the 10 B memory (~10.5 GB at Q8) in system RAM: it is read one row per token per head, so the GPU footprint is that of a 27 B-class model and generation speed is that of a 3 B model.- Sampler:
temperature 0.7, top_p 0.8, top_k 20, repeat_penalty 1.05— sample, do not decode greedily. - Thinking:
"chat_template_kwargs": {"enable_thinking": true}; give it 4k+ tokens for code.--reasoning-format deepseekseparates the block. - Architecture
qwen4exp; if your build reports an unknown architecture, update llama.cpp. - Tool use / agents: keep each assistant turn's
reasoning_contentin the history for the rest of the tool episode. With it, even lw5 did not loop in our probes. Claude Code works against llama-server's Anthropic endpoint (/v1/messages) with the chat template in this repo and in the GGUFs.
Caveats, measured
- Not a complete distillation: 14 % of the usable teacher rows seen, long-context rows unused, no on-policy phase yet (see the preview note at the top).
- In Phase 2 the memory's contribution on chat and science text shrank while code and general text grew (in-run probe above).
- A report of Claude Code writing to a corrupted working-directory path (
/written as-) did not reproduce in our probes, on Q8_0 or on the reporter's exact Q5_K_M file; if you see it, please open a discussion with the transcript. - The memory has learned code first: its gain there is nearly three times its gain on general text and chat, and seven times its gain on science.
- Maths sits at the level of the 27B-A3B line on our probe (41–46/60 at the 6,144-token cap across every checkpoint here; p2-s6000 44/60, 48/60 given 16k tokens).
- Science is the one held set that got worse as the memory got stronger: cross-entropy with the memory on went from 0.756 at lw2 to 0.796 at lw5, the price of a maths-heavy mix.
- Because the body now depends on the memory, the table must be served whole — this release ships every row (no norm pruning). A GGUF that drops or re-hashes the table will behave like v4.4 minus its knowledge.
- Structured output: p2-s6000 returned bare, valid JSON on 72/72 hold-out items (earlier checkpoints fenced it or left a string unquoted). For production
pipelines, constraining the decode with
response_format+json_schemais still the safe choice. - Thinking length is unbounded by default and p2-s6000 thinks longer than earlier checkpoints: a whimsical prompt can spend 2k+ tokens deliberating. Give
max_tokensheadroom or cap the reasoning with your server's reasoning-budget option. - Teacher voice: long graded replies pull toward the 27B's planning register. Greedy decoding loops on this family; use the sampler above.
Provenance
David Aylward (logic65) & Claude (Anthropic). Parent: logic65/Whittle-Next-27B-A3B (its card carries the full v1–v4.4 lineage back to Qwen3.6-35B-A3B). Teacher: Qwen/Qwen3.8-27B. Memory contents: Qwen/Qwen3.8-Flash-Next. All Apache-2.0. No benchmark test sets were trained on; the maths probe problems come from the MATH train split and were excluded from every training corpus.