nics-efc/VPR-Qwen3-4B-Sokoban

🤗 Hugging Face sourcetext-generationapache-2.04.4B params8.8 GBsafetensors✓ 3 checksumsupdated today
Submit in one command

Run it next to your model folder. It makes the torrent, checks your files against Hugging Face, and submits it. You just start seeding and paste your key from your account. It only reads your files and never changes them. Read the script first if you like.

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo nics-efc/VPR-Qwen3-4B-Sokoban ./model-folder
Needs a seeder →

VPR-Qwen3-4B-Sokoban

This is a Qwen3-4B checkpoint trained with Verifiable Process Rewards (VPR) on Markovian Sokoban interactions. At each visited state, VPR samples four action responses, scores their parsed actions with a task-grounded search-based Sokoban oracle, commits one highest-reward candidate, and optimizes all eligible candidates using locally normalized advantages.

Reported result

Evaluation SR
Sokoban 28.40 ± 2.79

Values are percentages reported under the evaluation protocol in the VPR paper. SR is success rate; CR is completion rate.

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "nics-efc/VPR-Qwen3-4B-Sokoban"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype="auto",
    device_map="auto",
)

inputs = tokenizer("<current Markovian game prompt>", return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=512)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))

Use the environment prompts, parsers, and action conventions in the VPR codebase for reproduction. These task-specific checkpoints are not intended as general-purpose assistants.

Resources

Limitations

Training relies on task-grounded oracle signals and specific Markovian prompts. Performance outside the documented environments and action formats has not been established. Evaluate safety and correctness before open-ended deployment.

Citation

@misc{yuan2026verifiable,
  title         = {Verifiable Process Rewards for Agentic Reasoning},
  author        = {Huining Yuan and Zelai Xu and Huaijie Wang and Xiangmin Yi and Jiaxuan Gao and Xiao-Ping Zhang and Yu Wang and Chao Yu and Yi Wu},
  year          = {2026},
  eprint        = {2605.10325},
  archivePrefix = {arXiv},
  primaryClass  = {cs.AI},
  url           = {https://arxiv.org/abs/2605.10325}
}