BoldingBuilds/Qwen3.8-27B-Reasoning-Termination-Fix-GGUF

🤗 Hugging Face sourcetext-generationapache-2.027B activated17 GBGGUF✓ 1 checksumupdated today
Submit in one command

Run it next to your model folder. It makes the torrent, checks your files against Hugging Face, and submits it. You just start seeding and paste your key from your account. It only reads your files and never changes them. Read the script first if you like.

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo BoldingBuilds/Qwen3.8-27B-Reasoning-Termination-Fix-GGUF ./model-folder
Needs a seeder →

Qwen3.8 27B Reasoning Termination Fix (Q4_K_M GGUF)

Fewer empty answers from Qwen3.8 27B within a 4,096-token output budget. This GGUF changes only the output-layer row for </think>. On a 200-problem MATH-500 subset, empty final responses fell from 31 to 0 and correct answers rose from 150 to 169. Accuracy was similar to stock with llama.cpp's --reasoning-budget flag, while the weight edit needs no runtime budget control. Occasional stray </think> tags remain.

September 29 update: Added a reasoning effort comparison run under the release conditions (stock and this model, at the template default and at medium, two seeds each), and the earlier development comparison it supersedes. Clarified untested use cases. The released weights and the main release results are unchanged.

200 MATH-500 problems, 4,096 output tokens correct empty final response
this model 169 0
stock 150 31
stock + --reasoning-budget 2048 167 0
stock + --reasoning-budget 3072 166 0

Paired over the same problems, this model solved 21 that stock missed and missed 2 that stock solved (exact McNemar p = 0.0001). All 31 of stock's empty responses were the model still thinking when the 4,096 tokens ran out. Every p-value on this card is exploratory and unadjusted, including for the fact that this version was chosen from several candidates using these same benchmarks (see below).

Compared with the flag. Against --reasoning-budget 2048 it scored 169 vs 167 (9 vs 7 on the problems where they differ, p = 0.80), and against 3072 it scored 169 vs 166 (9 vs 6, p = 0.61). Neither difference is statistically significant. The point of this build is that it needs no flag, so it should behave the same in apps and APIs that do not expose one. Beyond llama.cpp, only a small LM Studio check was run (see Compatibility); other apps were not tested.

With a larger budget the flag did better. At 16,384 output tokens, on the first 100 of the same problems:

100 MATH-500 problems, 16,384 output tokens correct empty final response stray </think>
stock + --reasoning-budget 2048 88 0 5
this model 84 0 4
stock 83 5 0

The flag solved 4 problems this model missed and this model solved none the flag missed. That is not statistically significant (p = 0.13), but it points one way, so if you run llama.cpp yourself with a large budget, the flag is the better tool. Combining them did not help either: this model plus the 2,048 flag scored 163 of 200 at 4,096 tokens, against 169 for this model alone (11 vs 5, p = 0.21).

The medium reasoning effort setting, compared with this model

Added September 29, 2026. Qwen3.8's chat template takes a reasoning_effort argument (low, medium, xhigh; high is treated as xhigh). If nothing is sent, the template uses xhigh, and every other row on this card was measured that way. Setting medium changes the instructions the template puts in front of the model's thinking, so it is a different lever from the --reasoning-budget token cutoff compared above. Several people have pointed out that medium alone removes most empty answers, so here it is measured under the same conditions as the rest of the card: same 200 MATH-500 problems, same 4,096-token budget, same machine and llama.cpp build, 8 parallel requests, file sampler defaults. Because sampling at temperature 1.0 moves a few problems from run to run, every cell was run twice, with request seed 0 and seed 1.

200 MATH-500 problems, 4,096 output tokens correct, seed 0 correct, seed 1 empty, seed 0 empty, seed 1 stray </think>, seed 0 stray </think>, seed 1
stock, template default 150 148 31 32 0 0
stock, reasoning_effort: medium 156 155 14 12 0 0
this model, template default 169 168 0 0 2 6
this model, reasoning_effort: medium 158 159 7 1 6 1

The seed 0 stock and this-model default rows are the same runs as at the top of this card. Every empty answer in the table is the model still thinking when the budget ran out; none ended the turn with no answer.

What the pairs say, on the same problems (exact McNemar, seed 0 then seed 1):

  • Medium helps stock, but not by much. 156 vs 150 (16 vs 10, p = 0.33) and 155 vs 148 (17 vs 10, p = 0.25). It cuts the empty answers by more than half and the rest of the gain is within run-to-run noise.
  • This model at the default beat stock at medium. 169 vs 156 (20 vs 7, p = 0.019) and 168 vs 155 (19 vs 6, p = 0.015).
  • Do not set medium on top of this model. 158 vs 169 (9 vs 20, p = 0.06) and 159 vs 168 (10 vs 19, p = 0.14): about ten points lower both times, and a few empty answers came back. The </think> row was fitted on the model's thinking at the default effort; medium changes that thinking, and the fitted row no longer lands in the right place.
  • At medium, this model and stock were not distinguishable: 158 vs 156 (13 vs 11, p = 0.84) and 159 vs 155 (7 vs 3, p = 0.34).

For scale, rerunning the same configuration with a different seed flipped 3 to 13 problems in each direction (p between 0.73 and 1.0 for all four same-arm pairs), so single-run differences of that size mean nothing.

Which should you use? If your client lets you set reasoning_effort, medium is a real improvement on stock and costs nothing. If it does not, or you do not want to depend on it, this file gives a larger improvement with nothing set. Use one or the other. Stray </think> tags remain the cost of this file: 2 and 6 of 200 here (1% to 3%), never seen in the stock runs.

Per-problem results for all eight runs are in evidence/per_item/math500_200_at_4096_reasoning_effort.csv and the counts in evidence/summary_counts.csv (benchmark math500_200_at_4096_reasoning_effort). The medium runs sent --chat-template-kwargs '{"reasoning_effort":"medium"}' to llama-server and nothing else changed.

Earlier development comparison: medium reasoning effort

Added September 29, 2026. I tested medium effort during development but omitted this comparison from the initial release write-up. It is kept here as the record of that run; the section above now has the same comparison under the release conditions, with the released file and two seeds, and is the one to cite. Medium effort changes the chat-template instructions; it is different from the --reasoning-budget token cutoff compared above.

These are earlier Qwen3.8 27B development runs, using stock Q4_K_M and an earlier edit (cb2), on one H100 with 16 parallel requests and a 4,096-token output limit. They use the same 200 math prompts as the release tests, but the released edit is round 6 and the main release math runs used an RTX 3090 with 8 parallel requests. Treat these as a separate comparison, not additional rows measured under the final release conditions.

Earlier development run Math correct /200 Empty math answers Stray </think> in math
Stock, default effort 147 32 0
Stock, medium effort 156 14 0
Earlier edit, cb2 167 1 4
Same development campaign HumanEval passed /164 Empty coding answers
Stock, default effort 147 11
Stock, medium effort 156 0
Earlier edit, cb2 154 0

Medium helped, but did not eliminate math blanks in this run. All 14 medium-effort math blanks hit the output limit. On HumanEval, medium eliminated blanks and slightly outscored the earlier edit. These results do not establish that the edit is universally better than medium effort.

I recovered the original launcher and raw records, checked the shared prompt IDs and text, and recomputed math scores with the original scorer. HumanEval counts were recounted from saved execution outcomes; I did not re-execute the generated code for this update. Settings, provenance limits and per-item results are in the development evidence. No medium-effort comparison is reported here for Flash-Next. The final release results above are unchanged.

Files

file size
Qwen3.8-27B-Reasoning-Termination-Fix-Q4_K_M.gguf 16.81 GB

SHA-256: 2da7bb4517e5bd19e61ff2d8b894832d148b720afb1f06b98e5ab6f19f6c844b

Language model only. The vision encoder is not included and vision was not tested.

Usage

llama-server -m Qwen3.8-27B-Reasoning-Termination-Fix-Q4_K_M.gguf -ngl 99 -fa on -c 32768 --jinja

This command sets no sampler options, so llama.cpp uses the defaults stored in the file (temperature 1.0, top-k 20, top-p 0.95) plus its own min-p default of 0.05. Those are the sampler values the main llama.cpp release results were measured with; I read them back from a running eval server. -c is the context window, not the output budget. The budget is the max_tokens your client sends; the results above use 4,096 unless marked 16,384.

Run this file at the template's default reasoning effort. Do not also set reasoning_effort: medium; measured together they scored about ten points lower on MATH than this file alone (see the reasoning effort section).

Compatibility

  • llama.cpp, tested. The main release results except the LM Studio check below were produced with a PrismML llama.cpp fork (commit 5d80cff, 2026-09-17, with local debug patches that only activate through environment variables the evals did not set; the patch is evidence/prism-5d80cff-local.patch). This exact file was also loaded on mainline llama.cpp (ggml-org master, commit 5262471, 2026-09-28) and answered all 30 problems of a 30-problem MATH check (24 correct, one stray tag).

  • LM Studio, small check. LM Studio 0.4.16 on Windows, engine llama.cpp-win-x86_64-nvidia-cuda12-avx2 2.47.0, one RTX 5080 (16 GB) with 60% of the model on the GPU (--gpu 0.6), context 20,480, 4 parallel requests, through its OpenAI-compatible server. Each request sent the same prompt as the MATH eval with max_tokens 4,096 and seed: 0, and no sampler settings, so LM Studio applied its own defaults (I did not check which values it used). The set was the first 12 of the 31 MATH problems where stock gave an empty reply at 4,096 tokens on llama.cpp, so this checks that the fix carries over; it does not measure accuracy.

    12 stock-failure MATH problems, LM Studio answered correct ran out of tokens while thinking
    this model 12 4 0
    stock (same file with the original row) 1 1 11

    Two of this model's 12 answers ran into the 4,096-token limit after the answer had started. No stray tags in either run. On llama.cpp this model also answered all 12 of these problems and got 4 right. Per-problem results: evidence/per_item/math500_12_lmstudio.csv.

  • Ollama and other apps: untested. Ollama's own Qwen3.8 library entry needs Ollama 0.32.12 or newer. On the older Ollama I had (0.20.2), importing this GGUF directly gave a bare prompt template with no thinking setup, so I did not run it there.

What changed from stock

edited 1 row of output.weight (the </think> token, id 248069, Q6_K): 4,161 of its 4,200 bytes changed, 0 bytes changed anywhere else in the file (full byte comparison)
unchanged everything else is byte-identical to the stock file: llama-quantize Q4_K_M, no imatrix, from the BF16 GGUF in unsloth/Qwen3.8-27B-GGUF (revision 4ca72078). Stock file SHA-256 b0e67404dddcb4253804c7a7c77839f16f291fdbd35d8ea70c9ea705508610c5

How the fix works. Only one output-layer row was fitted; no other weights were trained. The change is a fitted delta to the </think> row, intended to improve when the model ends its reasoning. Whether the model closes its reasoning depends on the score the output layer gives </think> at each step. I fit a change to that row, on the model's own final hidden states, so the score rises once the model is ready to answer or is going in circles, and stays low mid-thought and inside the answer. It was fit in rounds. Each round ran the previous version on the fitting prompts and added the places where it went wrong: closing the thinking too late, writing a stray </think> into its answer, or ending its turn while still inside the thinking, with no reply at all.

Why it helps. On 120 hard practice problems (MATH Level 5 training problems, not in any test here), stock left 47 empty at 4,096 tokens, all of them still thinking at the cutoff. In 4 of those 47, the thinking already contained the correct answer in a \boxed{}. In 27 of the 47, the correct answer string appeared somewhere in the thinking. String presence alone does not show that the model had completed a correct solution.

Fitting data vs test data. The fix was fit on 469 prompts that are not in any test here: 150 MATH training problems, 80 MBPP training problems, 40 Alpaca instructions and 199 harmful requests (AdvBench, JailbreakBench). Before fitting, any prompt that matched or closely resembled (word overlap of 50% or more) a prompt from any test on this card, including all of MATH-500, was removed. One defect: in the 150 MATH fitting prompts, the backslash in \boxed{} was lost to a shell escaping bug, so the model saw a control character there. The test prompts were not affected.

The tests did influence which version was released. Fitting never used a test prompt, but I built six rounds and compared them on these same MATH-500, HumanEval and refusal sets before picking this one. The numbers here are therefore for a selected candidate, not a fully held-out evaluation. The round history is in evidence/README.md. One earlier round scored higher on MATH-500 (171) but sent empty replies to 15 of 250 harmful prompts, which is why it was not chosen.

Other results

HumanEval (164 problems), pass@1, 4,096 output tokens, same settings:

build passed empty final response
this model 157 0
stock (same Q4_K_M file) 151 8
stock + --reasoning-budget 3072 157 0

This model vs stock: 7 vs 1 on the problems where they differ (p = 0.07). That leans toward this model but is not statistically significant. All 8 of stock's empty responses ran out of tokens while thinking. Against the flag: 2 vs 2, no difference. An earlier version of this card compared against a different stock quant (UD-Q4_K_XL, 147 passed, 13 empty); that run is kept in the evidence files, labeled, but the row above is the like-for-like comparison.

Refusals: aggregate rates stayed close to stock. Same judge and prompt for every row:

stock this model
StrongREJECT (first 150), judged a refusal 145 (96.7%) 147 (98.0%)
SimpleSafetyTests (100), judged a refusal 82 83
XSTest safe prompts (first 100), full compliance 95 98
XSTest safe prompts, partial refusal 1 2
XSTest safe prompts, full refusal 1 0
XSTest safe prompts, empty (ran out of tokens) 3 0
empty reply that ended the turn, 250 harmful prompts 1 2
empty reply that ran out of tokens, 250 harmful prompts 1 0

These are aggregate rates. Individual prompts moved in both directions; per-prompt labels are in evidence/per_item/refusals_judged.csv. Earlier rounds of this fix turned some refusals into silence: the model ended its turn inside its thinking and sent an empty reply (15 of 250 in one round, 24 in another). Later rounds were fit against exactly that, and this build sends 2 such empty replies, against 1 for stock.

Stray </think> tags. This model wrote a stray </think> into its final answer in 2 of 200 MATH answers at 4,096 tokens and 4 of 100 at 16,384 tokens, and in none of the HumanEval or refusal answers. No stray tags were observed in any tested stock run without the flag. Stock with the 2,048 flag also produced them (5 of 100 at 16,384 tokens).

How this was tested

These settings describe the main release evaluation. The earlier medium-effort comparison has its own settings and provenance linked above.

  • Server: llama-server with -ngl 99 --flash-attn on --jinja, several requests in parallel (MATH: 8 slots of 5,120 tokens at the 4,096 budget, 4 slots at 16,384; other sets: 4 slots of 6,144). Each request sends the messages, max_tokens and seed: 0, nothing else. MATH and the refusal sets send the prompt as a single user message (MATH adds "Put your final answer in \boxed{}."); HumanEval adds a short system prompt asking for code only. Exact prompts are in the scripts in evidence/scripts/.
  • Thinking on, at the chat template's default reasoning effort (xhigh); no reasoning_effort or template arguments sent, except the rows marked medium in the reasoning effort section, which sent --chat-template-kwargs '{"reasoning_effort":"medium"}'.
  • Sampler: the file's defaults as above. Sampling at temperature 1.0 with parallel slots is not bit-reproducible run to run, so expect small differences on a rerun.
  • Empty final response means the reply after the reasoning was empty. It has two causes, counted separately in evidence/summary_counts.csv: the output budget ran out during reasoning (finish_reason: length), or the model ended its turn with no answer (finish_reason: stop).
  • MATH-500: the first 200 problems after a seeded shuffle (seed 7) of the 500, and the first 100 of those at 16,384 tokens. Scored by exact match of the last \boxed{} answer after normalisation. The scorer is strict, so compare rows to each other, not to published MATH-500 numbers.
  • HumanEval: all 164 problems, pass@1, one sample each; the code block is extracted and run against each problem's own tests.
  • Refusals: StrongREJECT (first 150 prompts in file order), SimpleSafetyTests (all 100) and XSTest safe prompts (first 100 in file order). Graded by an LLM judge that is not the model under test: a Q8_0 file from OBLITERATUS/Qwen3.8-27B-OBLITERATED (SHA-256 4cfb568f..., full hash in evidence/hashes.txt), thinking off, temperature 0, JSON-schema output. Harmful sets use the StrongREJECT rubric; XSTest uses its authors' three-way labels. An uncensored judge was used because a safety-tuned judge can refuse to read harmful replies.
  • Every row on this card was run on the same machine (RTX 3090s) and the same llama.cpp build, except the 30-problem mainline load check, the LM Studio check and the earlier development comparison (H100).
  • Statistics: exact two-sided McNemar test on paired per-problem results. The p-values are exploratory: they are not adjusted for multiple comparisons, or for choosing this version from several candidate rounds using these same benchmarks.

Evidence

evidence/ holds the per-problem results for every row on this card (IDs, scores and flags only, no model text); summary_counts.csv with the main release counts on the card; hashes.txt with SHA-256 of the model files, judge, datasets and scripts; the eval, judge, fit and bake scripts as run; the exact server settings; dataset selection rules; judge labels by prompt index; the fitting prompt manifest (source and SHA-256 of each prompt); the bake log with the full-file byte comparison; and the round-by-round selection history.

Not included: the text of the fitting prompts and of the refusal test prompts, since many of them are harmful requests, and the model's own generations. All of these prompts come from public datasets named in the evidence README, so the test sets can be rebuilt with the stated selection rules, but rerunning the fit needs regenerated captures and will not give byte-identical results.

Limitations

  • Most results are at a 4,096-token output budget. With a larger budget stock runs out less often and the gap shrinks (see the 16,384-token result).
  • The fix makes the model stop thinking sooner. Overall scores went up, but not on every problem: it missed 2 MATH problems at 4,096 tokens, 2 at 16,384 and 1 HumanEval problem that stock solved, and its safety labels differ from stock on individual prompts. Tasks that need very long reasoning may lose more.
  • Occasional stray </think> tags in MATH answers: 1% to 3% at 4,096 tokens (2/200 and 6/200 over two seeds), and 4% at 16,384 tokens (4/100).
  • Tested in English, text only, single turn. Agent, tool use and multi-turn use were not tested.
  • Speculative decoding with a draft model was not tested for the released file; its effect on draft acceptance and speedup is unknown.
  • One quant (Q4_K_M). The same row edit could be applied to other quants; none were tested.
  • One seed per configuration, except the reasoning effort section, which was run with two seeds. Between seeds, 3 to 13 problems flipped each way for the same configuration.

Related work

Most fixes for runaway thinking work at run time: budget forcing and llama.cpp's --reasoning-budget; early exit from a probe on the hidden states (Zhang et al. 2025, LYNX); or stopping rules on the </think> score and what follows it (ThinkBrake, Entropy After </think>). Others edit weights inside the network to change how long the model reasons (ThinkEdit, Steer2Edit). Sheng et al. 2025 show that the model plans its reasoning length through the </think> score, and report that simply scaling that score at run time made decoding unstable (their Appendix C.5).

I am not aware of earlier published work that fits the stop decision into the </think> output row itself, as a change that depends on the model's hidden state and ships as a plain GGUF with nothing extra at run time. An earlier version of this row edit is part of my Bonsai 2 V2 release. This is the first release of it as a standalone fix on an unmodified stock model.

Credits

Base model: Qwen/Qwen3.8-27B. Test sets: MATH (MATH-500), HumanEval, StrongREJECT, SimpleSafetyTests, XSTest. Fitting prompts: MATH (train), MBPP (train), Alpaca, AdvBench, JailbreakBench.