nics-efc/VPR-Qwen3-4B-Base-Math-Mixed

🤗 Hugging Face sourcetext-generationapache-2.04.4B params8.8 GBsafetensors✓ 3 checksumsupdated today
Submit in one command

Run it next to your model folder. It makes the torrent, checks your files against Hugging Face, and submits it. You just start seeding and paste your key from your account. It only reads your files and never changes them. Read the script first if you like.

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo nics-efc/VPR-Qwen3-4B-Base-Math-Mixed ./model-folder
Needs a seeder →

VPR-Qwen3-4B-Base-Math-Mixed

This checkpoint is trained from Qwen3-4B-Base with mixed math and VPR game experience across Sokoban, Sudoku, and Minesweeper. VPR supplies action-level process rewards through task-grounded oracles and state-group rollout, while math training follows the mixed-training protocol described in the paper.

Reported results

Evaluation Metric Result
General-reasoning OOD suite Macro average 52.75
ALFWorld SR 16.12 ± 2.87
WebShop Score 42.01 ± 2.14
WebShop SR 1.33 ± 0.76

Values are reported under the VPR paper's evaluation protocol. SR is success rate. The general-reasoning value is the macro average across the reported OOD benchmarks. These OOD results are not a claim of exactly matched total rollout compute across training methods.

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "nics-efc/VPR-Qwen3-4B-Base-Math-Mixed"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype="auto",
    device_map="auto",
)

inputs = tokenizer("Solve the task step by step.", return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=1024)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))

Use the VPR codebase for the exact prompts, environments, and evaluation entry points.

Resources

Limitations

The checkpoint is shaped by the documented math distribution, task-grounded game oracles, prompts, and action formats. Performance and safety outside those settings have not been established.

Citation

@misc{yuan2026verifiable,
  title         = {Verifiable Process Rewards for Agentic Reasoning},
  author        = {Huining Yuan and Zelai Xu and Huaijie Wang and Xiangmin Yi and Jiaxuan Gao and Xiao-Ping Zhang and Yu Wang and Chao Yu and Yi Wu},
  year          = {2026},
  eprint        = {2605.10325},
  archivePrefix = {arXiv},
  primaryClass  = {cs.AI},
  url           = {https://arxiv.org/abs/2605.10325}
}