HYU-NLP-EVAL/qwen3-1.7b-rar-medicine-static-r0-step-016

🤗 Hugging Face 来源text-generationapache-2.01.7B 参数3.4 GBsafetensors✓ 2 个校验和今天更新
一条命令提交

在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo HYU-NLP-EVAL/qwen3-1.7b-rar-medicine-static-r0-step-016 ./model-folder
需要做种者 →

Qwen3-1.7B Static-R0 Policy Checkpoint

This public model repository contains one policy checkpoint from the static-rubric discriminability-horizon experiment. The domain and optimizer step are encoded in the repository name.

Experiment

  • Base policy: Qwen/Qwen3-1.7B
  • Base revision: 70d244cc86ccca08cf5af4e1e306ecf908b1ad5e
  • RL algorithm: GRPO
  • Training reward: frozen prompt-specific initial rubric (R0)
  • Domain: RaR Medicine or RaR Science, as named in the repository
  • Training seed: 11
  • Export format: Hugging Face Transformers, BF16 safetensors
  • Checkpoint contents: model weights, configuration, tokenizer, and chat template
  • Excluded: optimizer, scheduler, trainer state, rollouts, rubrics, and evaluation data

Medicine checkpoint repositories cover the ten planned audit points through step 48. Science training was only saved through step 3 at the time of this upload, so the public Science repositories contain steps 0 and 3.

Loading

from transformers import AutoModelForCausalLM, AutoTokenizer

repo_id = "HYU-NLP-EVAL/<repository-name>"
tokenizer = AutoTokenizer.from_pretrained(repo_id)
model = AutoModelForCausalLM.from_pretrained(
    repo_id,
    dtype="bfloat16",
    device_map="auto",
)

Intended use

These checkpoints are research artifacts for studying reward saturation and static-rubric staleness during policy optimization. Medicine checkpoints are not medical devices and must not be used as a substitute for professional medical advice.