HYU-NLP-EVAL/qwen3-4b-healthbench-static-r0-step-030

🤗 Hugging Face 来源text-generationapache-2.04B 参数8.0 GBsafetensors✓ 4 个校验和今天更新
一条命令提交

在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo HYU-NLP-EVAL/qwen3-4b-healthbench-static-r0-step-030 ./model-folder
需要做种者 →

Qwen3-4B HealthBench Static-Rubric R0 — Step 30

This checkpoint is pi_30, after 30 static-rubric RL optimizer updates.

This model belongs to the static-rubric R0 experiment, not a dynamic-rubric training run. Training rewards use a frozen, prompt-specific rubric bank: each training prompt is scored with its own R0(x) throughout optimization.

Provenance

  • Run ID: pilot-static-r0-100step-20260821
  • Optimizer step: 30
  • Base model: Qwen/Qwen3-4B-Instruct-2507
  • Base revision: cdbee75f17c01a7cc42f958dc650907174af0554
  • Reward source: static_r0_only
  • Policy training split: 256 HealthBench prompts
  • Parameterization: full-model RL
  • Export dtype: BF16
  • Original format: VERL FSDP v1, world size 1, FP32 state dict

Loading

from transformers import AutoModelForCausalLM, AutoTokenizer

repo_id = "HYU-NLP-EVAL/qwen3-4b-healthbench-static-r0-step-030"
tokenizer = AutoTokenizer.from_pretrained(repo_id)
model = AutoModelForCausalLM.from_pretrained(repo_id, torch_dtype="bfloat16")

Intended use and limitations

This is a research checkpoint for studying proxy-rubric staleness during policy optimization. It is not a medical device and must not be used as a substitute for professional medical advice. Static-rubric reward improvement does not by itself establish improvement against independent HealthBench ground truth.