YiPz/qwen3-4b-pokerbench-grpo-gguf

🤗 Hugging Face 来源text-generationapache-2.0激活 4B15 GBGGUF✓ 3 个校验和今天更新
一条命令提交

在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo YiPz/qwen3-4b-pokerbench-grpo-gguf ./model-folder
需要做种者 →

Qwen3-4B PokerBench GRPO (GGUF)

GGUF quantized versions of YiPz/qwen3-4b-pokerbench-grpo.

Model Description

Qwen3-4B fine-tuned for poker using:

  1. SFT on high quality reasoning traces
  2. GRPO reinforcement learning with LLM-as-judge rewards

Available Files

File Size Description
qwen3-4b-pokerbench-grpo-f16.gguf ~8 GB Full precision (F16) - highest quality
qwen3-4b-pokerbench-grpo-q8_0.gguf ~4.5 GB 8-bit quantization - high quality
qwen3-4b-pokerbench-grpo-q4_k_m.gguf ~2.5 GB 4-bit quantization - recommended

Usage with Ollama

# Download Q4_K_M (recommended)
huggingface-cli download YiPz/qwen3-4b-pokerbench-grpo-gguf \
    qwen3-4b-pokerbench-grpo-q4_k_m.gguf --local-dir ./

# Create Modelfile
cat > Modelfile << 'EOF'
FROM ./qwen3-4b-pokerbench-grpo-q4_k_m.gguf

PARAMETER temperature 0.6
PARAMETER top_p 0.95
PARAMETER num_ctx 3072
PARAMETER stop "<|im_end|>"
PARAMETER stop "<|endoftext|>"

SYSTEM "You are an expert poker coach. Analyze situations with step-by-step reasoning in <think></think> tags and provide your action in <action></action> tags."

TEMPLATE """{{- if .System }}<|im_start|>system
{{ .System }}<|im_end|>
{{ end }}{{- range .Messages }}<|im_start|>{{ .Role }}
{{ .Content }}<|im_end|>
{{ end }}<|im_start|>assistant
"""
EOF

# Create and run
ollama create pokerbench-grpo -f Modelfile
ollama run pokerbench-grpo "I have AKo on the button with 100bb. UTG raises 3bb. What should I do?"

Output Format

<think>
1. Position: On the button with excellent position...
2. Hand strength: AKo is a premium hand...
3. Stack depth: 100bb effective allows for deep play...
4. Recommendation: 3-bet for value...
</think>

<action>raise 9</action>

Related Models