Qwen3-4B PokerBench GRPO (GGUF)
GGUF quantized versions of YiPz/qwen3-4b-pokerbench-grpo.
Model Description
Qwen3-4B fine-tuned for poker using:
- SFT on high quality reasoning traces
- GRPO reinforcement learning with LLM-as-judge rewards
Available Files
| File | Size | Description |
|---|---|---|
qwen3-4b-pokerbench-grpo-f16.gguf |
~8 GB | Full precision (F16) - highest quality |
qwen3-4b-pokerbench-grpo-q8_0.gguf |
~4.5 GB | 8-bit quantization - high quality |
qwen3-4b-pokerbench-grpo-q4_k_m.gguf |
~2.5 GB | 4-bit quantization - recommended |
Usage with Ollama
# Download Q4_K_M (recommended)
huggingface-cli download YiPz/qwen3-4b-pokerbench-grpo-gguf \
qwen3-4b-pokerbench-grpo-q4_k_m.gguf --local-dir ./
# Create Modelfile
cat > Modelfile << 'EOF'
FROM ./qwen3-4b-pokerbench-grpo-q4_k_m.gguf
PARAMETER temperature 0.6
PARAMETER top_p 0.95
PARAMETER num_ctx 3072
PARAMETER stop "<|im_end|>"
PARAMETER stop "<|endoftext|>"
SYSTEM "You are an expert poker coach. Analyze situations with step-by-step reasoning in <think></think> tags and provide your action in <action></action> tags."
TEMPLATE """{{- if .System }}<|im_start|>system
{{ .System }}<|im_end|>
{{ end }}{{- range .Messages }}<|im_start|>{{ .Role }}
{{ .Content }}<|im_end|>
{{ end }}<|im_start|>assistant
"""
EOF
# Create and run
ollama create pokerbench-grpo -f Modelfile
ollama run pokerbench-grpo "I have AKo on the button with 100bb. UTG raises 3bb. What should I do?"
Output Format
<think>
1. Position: On the button with excellent position...
2. Hand strength: AKo is a premium hand...
3. Stack depth: 100bb effective allows for deep play...
4. Recommendation: 3-bet for value...
</think>
<action>raise 9</action>
Related Models
- Full weights: YiPz/qwen3-4b-pokerbench-grpo