laion/a3-rl-laion_nemotron-gym-agent-calendar-80-8B

🤗 Hugging Face sourcereinforcement-learningapache-2.08.2B params16 GBsafetensors✓ 9 checksumsupdated today
Submit in one command

Run it next to your model folder. It makes the torrent, checks your files against Hugging Face, and submits it. You just start seeding and paste your key from your account. It only reads your files and never changes them. Read the script first if you like.

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo laion/a3-rl-laion_nemotron-gym-agent-calendar-80-8B ./model-folder
Needs a seeder →

a3-rl-laion_nemotron-gym-agent-calendar (step 80, 8B)

RL-trained (SkyRL) agent checkpoint, selected by 5-period EMA (alpha=1/3) of reward/avg_raw_reward across all 80 steps of the training chain.

The launch RL config is included as rl_config.yaml. Parsed training metrics, plots, and raw logs are under training_logs/.

Training Traces

Training-time Daytona/Harbor rollouts for this run are uploaded as a companion dataset: open-athena/a3-rl-laion_nemotron-gym-agent-calendar

The dataset contains the last episode of each trial (per make_and_upload_trace_dataset --episodes last) — the same rollouts the policy was trained on after rollback / truncation.