HYU-NLP-EVAL/qwen3-4b-rar-medicine-static-r0-matched-seed11-step-033

🤗 Hugging Face 来源text-generationapache-2.04B 参数8.0 GBsafetensors✓ 6 个校验和今天更新
一条命令提交

在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo HYU-NLP-EVAL/qwen3-4b-rar-medicine-static-r0-matched-seed11-step-033 ./model-folder
需要做种者 →

Static-R0 Matched GRPO on RaR-Medicine — step 33

This is the policy after 33 global optimizer updates of the matched static-rubric GRPO run (planned total: 48). It is intentionally separate from the OnlineRubrics/dynamic-rubric checkpoints.

Experiment identity

  • Method: static_r0_matched
  • Reward source: rar_static_r0_only
  • Domain: Medicine
  • Training data: RaR-Medicine, 1,500 prompts
  • Seed: 11
  • Policy: Qwen/Qwen3-4B-Instruct-2507
  • Base revision: cdbee75f17c01a7cc42f958dc650907174af0554
  • Thinking: disabled
  • GRPO global prompt batch: 96
  • Rollouts per prompt: 16
  • Learning rate: 5e-06

The root files are a BF16 Transformers export for inference. The original_checkpoint/ directory contains the exact original veRL/FSDP policy parameter checkpoint and its tokenizer/configuration files. Optimizer, trainer, and data-loader state are intentionally not published; the complete resume checkpoint remains on Daisy.

from transformers import AutoModelForCausalLM, AutoTokenizer

repo_id = "HYU-NLP-EVAL/qwen3-4b-rar-medicine-static-r0-matched-seed11-step-033"
tokenizer = AutoTokenizer.from_pretrained(repo_id)
model = AutoModelForCausalLM.from_pretrained(
    repo_id, torch_dtype="bfloat16", device_map="auto"
)

This is an intermediate research checkpoint, not a clinical model. No medical capability or safety claim is made.

Original actor parameter SHA256: 5a78fc291520c5f4651a9e4dbac6753940e3ab6bde9d7a4b29008546b0dd7f9f