detrax/Qwen3-4B-Thinking-2507-Qwen3.8-Max-Distillation-Detrax

🤗 Hugging Face 来源text-generationapache-2.0激活 4B10 GBGGUF✓ 4 个校验和今天更新
需要做种者 →

🚀 Qwen3-4B-Thinking-2507-Qwen3.8-Max-Distillation-Detrax - GGUF

This repository contains GGUF format quantizations of the distilled model, built upon Qwen/Qwen3-4B-Thinking-2507. The distilled model delivers stronger coding, reasoning, and constraint-following capabilities in a compact 4B-parameter footprint.

75.7% of benchmark questions improved, with gains up to +93% on difficult algorithmic tasks and +33.2% on constraint-heavy coding.

Benchmark Base Distilled Improvement
Overall Relevancy 0.7389 0.7447 +0.78%
Overall Similarity 0.7317 0.7398 +1.10%
Math Reasoning 0.714 0.742 +3.9%
Code Generation 0.765 0.789 +3.1%
Logical Inference 0.793 0.813 +2.5%

The model excels at multi-step reasoning, code generation, mathematical problem-solving, structured logical inference, and strict instruction-following — making it ideal for developers who want stronger technical behavior from a compact model.


📦 Model Files

Filename Quantization Size Description
Qwen3-4B-Thinking-2507-Qwen3.8-Max-Distillation-Detrax.gguf Full (F16) ~7.5 GB Highest quality / full precision
Qwen3-4B-Thinking-2507-Qwen3.8-Max-Distillation-Detrax-q8_0.gguf Q8_0 ~3.9 GB Best performance — excellent quality with strong efficiency
Qwen3-4B-Thinking-2507-Qwen3.8-Max-Distillation-Detrax-q6_K.gguf Q6_K ~3.1 GB High quality 6-bit quantization
Qwen3-4B-Thinking-2507-Qwen3.8-Max-Distillation-Detrax-q4_K_M.gguf Q4_K_M ~2.3 GB 4-bit medium — good balance of size and quality

Recommendation: Start with Qwen3-4B-Thinking-2507-Qwen3.8-Max-Distillation-Detrax-q8_0.gguf for the best performance while still remaining practical on most hardware.


🚀 Usage with llama.cpp

Option 1: Direct from Hugging Face

llama-cli \
  --model hf:<your-username>/<your-repo-name>/Qwen3-4B-Thinking-2507-Qwen3.8-Max-Distillation-Detrax-q8_0.gguf \
  --prompt "Write a Python function to compute the Fibonacci sequence" \
  --threads 8 \
  -no-cnv

Option 2: Local file

./llama-cli \
  --model Qwen3-4B-Thinking-2507-Qwen3.8-Max-Distillation-Detrax-q8_0.gguf \
  --prompt "Explain the concept of recursion in simple terms" \
  --threads 8

Option 3: With GPU offloading

./llama-cli \
  --model Qwen3-4B-Thinking-2507-Qwen3.8-Max-Distillation-Detrax-q8_0.gguf \
  --n-gpu-layers 40 \
  --cache-type-k q8_0 \
  --cache-type-v q8_0 \
  --threads 16 \
  --prompt "Solve: If x + y = 10 and 2x - y = 5, find x and y" \
  -no-cnv

Option 4: Chat mode with system prompt

./llama-cli \
  --model Qwen3-4B-Thinking-2507-Qwen3.8-Max-Distillation-Detrax-q8_0.gguf \
  --system "You are a helpful coding assistant." \
  --prompt "How do I reverse a linked list in Python?" \
  -no-cnv

🐍 Python (llama-cpp-python)

from llama_cpp import Llama

llm = Llama(
    model_path="Qwen3-4B-Thinking-2507-Qwen3.8-Max-Distillation-Detrax-q8_0.gguf",
    n_ctx=8192,
    n_threads=8,
    n_gpu_layers=40,  # Set to 0 for CPU-only
    verbose=False
)

output = llm.create_chat_completion(
    messages=[
        {"role": "system", "content": "You are a helpful coding assistant."},
        {"role": "user", "content": "Write a quicksort implementation in Python."}
    ],
    max_tokens=1024,
    temperature=0.7
)

print(output["choices"][0]["message"]["content"])

🔧 Conversion Details

These GGUF files were generated using llama.cpp via the GGUF-my-repo space.
The files are compatible with:

  • llama.cpp (all recent versions)
  • LM Studio
  • Jan
  • koboldcpp
  • Any GGUF-compatible runtime

📊 Benchmark Details

  • 148 questions across 8 domains
  • Same prompts evaluated on both base and distilled models
  • Base model: Qwen/Qwen3-4B-Thinking-2507
  • Metrics: Relevancy + Similarity

Biggest Wins:

Task Base Distilled Improvement
First Repeating Element (Algorithmic) 0.3568 0.6887 +93.0%
URL Analyzer (Constraint-Heavy Coding) 0.6329 0.8433 +33.2%

The strongest improvements are concentrated in technical reasoning, coding, mathematics, and constraint-heavy tasks.


💡 When to Use This Model

This distilled model is particularly effective for:

  • 🧠 Multi-step reasoning tasks
  • 💻 Code generation and debugging
  • 📐 Mathematical problem-solving
  • 🎯 Strict instruction and constraint following
  • 🔍 Structured logical inference

📝 Notes

  • The model uses the Qwen chat template (<|im_start|>system, <|im_start|>user, <|im_start|>assistant)
  • Recommended context length: 8192 tokens
  • The -no-cnv flag disables auto-conversation mode for single-turn prompts

📄 License

This model is released under the Apache-2.0 license.


🙏 Acknowledgments


📎 Citation

If you use this model in your research or production, please cite:

@misc{qwen3-4b-thinking-2507-qwen3.8-max-distillation-detrax-gguf,
  author = {Your Name},
  title = {Qwen3-4B-Thinking-2507-Qwen3.8-Max-Distillation-Detrax - GGUF Quantizations},
  year = {2026},
  publisher = {Hugging Face},
  url = {https://huggingface.co/<your-username>/<your-repo-name>}
}

Bottom Line: If your workload is coding, reasoning, mathematics, or structured technical tasks, this distilled model is worth trying.