Gensyn/Qwen2.5-1.5B-Instruct

🤗 Hugging Face 来源text-generationapache-2.01.5B 参数3.1 GBsafetensors✓ 1 个校验和今天更新
一条命令提交

在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo Gensyn/Qwen2.5-1.5B-Instruct ./model-folder
需要做种者 →

Qwen2.5-1.5B-Instruct

Introduction

This model is intended for use in the Gensyn RL Swarm, to finetune locally using peer-to-peer reinforcement learning post-training.

Once finetuned, the model can be used as normal in any workflow, for details on how to do this please refer to the original model documentation.

For more details on the original model, please refer to the original repository here.

This repo contains an unmodified version of the instruction-tuned 1.5B Qwen2.5 model, which has the following features:

  • Type: Causal Language Models
  • Training Stage: Pretraining & Post-training
  • Architecture: transformers with RoPE, SwiGLU, RMSNorm, Attention QKV bias and tied word embeddings
  • Number of Parameters: 1.54B
  • Number of Paramaters (Non-Embedding): 1.31B
  • Number of Layers: 28
  • Number of Attention Heads (GQA): 12 for Q and 2 for KV
  • Context Length: Full 32,768 tokens and generation 8192 tokens

Requirements

This model is intended for use in the Gensyn RL Swarm system, for details on model requirements when using outside of a swarm, refer to the original Qwen repo here.

Quickstart

To deploy this model into a swarm and/or participate in the Gensyn Testnet, follow the instructions in the RL Swarm repository, read about the testnet, read the RL Swarm overview, and/or read the RL Swarm technical report.