sahilchachra/ThinkingCap-Qwen3.6-27B-AWQ

🤗 Hugging Face sourceimage-text-to-textapache-2.027.4B params92 GBsafetensors✓ 2 checksumsupdated today
Submit in one command

Run it next to your model folder. It makes the torrent, checks your files against Hugging Face, and submits it. You just start seeding and paste your key from your account. It only reads your files and never changes them. Read the script first if you like.

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo sahilchachra/ThinkingCap-Qwen3.6-27B-AWQ ./model-folder
Needs a seeder →

ThinkingCap-Qwen3.6-27B-AWQ

AWQ (W4A16) quantization of bottlecapai/ThinkingCap-Qwen3.6-27B — a token-efficient reasoning finetune of Qwen/Qwen3.6-27B by BottleCap AI (qwen3_5: dense hybrid GatedDeltaNet linear-attention + full-attention over 64 text layers in a 3:1 pattern, plus a vision tower and an MTP head; multimodal image-text-to-text). It matches Qwen3.6-27B answer quality while emitting ~50% fewer thinking tokens on average.

Variant: AWQ W4A16 — 4-bit symmetric integer weights, group size 128, with activation-aware scaling. Activations stay BF16. Quantized by: sahilchachra Tooling: llm-compressor (AWQModifier + QuantizationModifier) -> compressed-tensors pack-quantized

This is a quantized derivative. Weights, behavior, and license follow the base model — see the original card for full details, benchmarks, and citation.

What is quantized

Quantized to 4-bit:

  • full-attention self_attn.{q,k,v,o}_proj
  • mlp.{gate,up,down}_proj (all text layers)

Kept in BF16: GatedDeltaNet linear_attn (mamba) layers, vision tower (model.visual.*, 27 blocks), MTP head, token embeddings, lm_head, all norms (incl. q_norm / k_norm).

Note on what's 4-bit vs BF16 (quality/speed tradeoff)

This is a hybrid architecture: 48 of the 64 layers are GatedDeltaNet linear-attention ("Mamba") layers. Only the full-attention self_attn.{q,k,v,o} projections and the per-layer mlp.{gate,up,down} projections are quantized to 4-bit; the Mamba linear_attn layers are deliberately kept in BF16 (along with the vision tower, MTP head, embeddings, lm_head and norms), because they are quantization-sensitive and keeping them full-precision preserves the reasoning quality and the concise <think> behavior.

As a result, roughly two-thirds of the weight bytes read per token stay BF16 (~18 GB BF16 vs ~9 GB of 4-bit weights), so this variant's memory footprint is close to an 8-bit build and it is tuned for quality rather than peak throughput — on Blackwell it can run a little slower than a full W8A8/FP8 build of the base model.

The Mamba layers can also be quantized (to shrink the model further and speed up memory-bound decoding), but accuracy may take a hit — this build intentionally trades that extra speed for output quality.

Calibration

AWQ: 128 sequences x 512 tokens of GSM8K (openai/gsm8k, config main, train split) — the model's headline in-domain reasoning dataset — rendered through the model's own chat template with the <think>…</think> reasoning format, so calibration matches the model's real (thinking) inference distribution. NVFP4 is data-free (no calibration).

Prompt template & sampling

This is a reasoning ("thinking") model. Use the Qwen3.6 chat template — ChatML (<|im_start|>role … <|im_end|>) with a <think>…</think> reasoning trace, thinking enabled by default. Apply it via tokenizer.apply_chat_template(messages, add_generation_prompt=True) (or the processor for image inputs); do not hand-format prompts. See the base Qwen/Qwen3.6-27B for full usage details.

Recommended sampling: temperature=1.0, top_p=0.95, top_k=20, min_p=0.0 with thinking on (the base model's recommended sampling, per the original card).

Usage (vLLM)

from vllm import LLM, SamplingParams

# This is a multimodal checkpoint: the vision tower is kept in BF16
# (only the text / MoE weights are 4-bit). vLLM builds the full model.
llm = LLM(
    model="sahilchachra/ThinkingCap-Qwen3.6-27B-AWQ",
    trust_remote_code=True,
)
out = llm.chat(
    [{"role": "user", "content": "Hello!"}],
    SamplingParams(temperature=0.6, top_p=0.95, max_tokens=512),
)
print(out[0].outputs[0].text)

Serving via the CLI, pass the flag directly:

vllm serve sahilchachra/ThinkingCap-Qwen3.6-27B-AWQ \
    --trust-remote-code \
    --max-model-len 262144 --reasoning-parser qwen3