qwen3-4b-agent-trajectory-lora
This repository provides a LoRA adapter fine-tuned from Qwen/Qwen2.5-7B-Instruct using LoRA + Unsloth.
This repository contains LoRA adapter weights only. The base model must be loaded separately.
Training Objective
Trained to improve multi-turn agent task performance on ALFWorld (household tasks) and DBBench (database operations). Loss is applied to all assistant turns, enabling the model to learn environment observation, action selection, tool use, and error recovery.
Training Configuration
| Item | Value |
|---|---|
| Base model | Qwen/Qwen2.5-7B-Instruct |
| Method | LoRA (full precision base, bf16) |
| Datasets | ALFWorld v5 + DBBench v4 |
| Max sequence length | 8192 |
| Epochs | 2 |
| Learning rate | 2e-06 |
| LoRA r / alpha | 64 / 128 |
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
import torch
base = "Qwen/Qwen2.5-7B-Instruct"
adapter = "legoskier/Qwen2.5-7B-agent-trajectory-lora_2"
tokenizer = AutoTokenizer.from_pretrained(base)
model = AutoModelForCausalLM.from_pretrained(
base, torch_dtype=torch.bfloat16, device_map="auto"
)
model = PeftModel.from_pretrained(model, adapter)
Sources & Terms
Training datasets:
u-10bei/sft_alfworld_trajectory_dataset_v5u-10bei/dbbench_sft_dataset_react_v4
This repository does not redistribute the datasets. Users must comply with the respective dataset licenses and the base model's terms of use.