HYU-NLP-EVAL/qwen3-4b-rar-medicine-onlinerubrics-seed11-step-027

🤗 Hugging Face 来源text-generationapache-2.04B 参数8.0 GBsafetensors✓ 4 个校验和今天更新
一条命令提交

在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo HYU-NLP-EVAL/qwen3-4b-rar-medicine-onlinerubrics-seed11-step-027 ./model-folder
需要做种者 →

OnlineRubrics RaR-Medicine: step 27, seed 11

Intermediate policy from dynamic OnlineRubrics-Every GRPO training, distinct from static-rubric GRPO. Base model: Qwen/Qwen3-4B-Instruct-2507; thinking disabled. This checkpoint is a historical policy state used by the Phase-1 audit. No downstream medical capability or safety claim is made. Research use only; not validated for clinical decision-making.

Root files are the veRL-exported Hugging Face inference model (BF16). original_checkpoint/ preserves the exact original FSDP parameter checkpoint and tokenizer/configuration files. Optimizer state, training data, responses, rubrics, infrastructure configuration, and credentials are not included. The original is retained because export precision/serialization differs.

Base model revision: cdbee75f17c01a7cc42f958dc650907174af0554 Original actor tree SHA256: 4c83ed5b93a4359a2ca0092ea0094e8b05bcaa5457b4227a85bbff67b1e535f3