quik-models/upbeat-waterfall-104

🤗 Hugging Face 来源text-generationmit638 MBother✓ 3 个校验和今天更新
已有模型文件?提交模型种子

如果你有完整的模型文件并有权分享,请把示例文件夹路径替换为你的文件路径,再运行这条命令。它会校验文件、制作种子,并将磁力链接和校验和提交给 Pirate Face。请让种子客户端持续做种,方便其他人从节点下载。Pirate Face 不接收模型文件。你可以从账户页面获取社区密钥。也可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo quik-models/upbeat-waterfall-104 ./model-folder
需要做种者 →

AutoResearch-tinystories-depth8

AutoResearch-tinystories-depth8 is a 201.3M parameter decoder-only Transformer trained from scratch on TinyStories (karpathy/tinystories-gpt4-clean).

This model is part of the AutoResearch project, which focuses on training, evaluating, and releasing efficient language models with reproducible research workflows.


Overview

This is a 8-layer decoder-only Transformer trained on the TinyStories (karpathy/tinystories-gpt4-clean) dataset for 0.2 hours of wall-clock training time. The model achieves a validation bits-per-byte (val_bpb) of 1.259341 (perplexity: 2.3939) on the held-out validation set.


References

Papers

  • NanoGPT / NanoChat architecture patterns

Datasets

  • Training: TinyStories (karpathy/tinystories-gpt4-clean)
  • Tokenizer: climbmix-400b-shuffle

Related Projects

WANDB Run


Highlights

  • Trained from scratch
  • 201.3M parameters
  • Trained on 15.7M tokens (30 steps)
  • 8-layer decoder-only Transformer with sliding window attention
  • RoPE positional encoding, RMSNorm, ReLU² activation
  • MuonAdamW optimizer (Muon for matrices, AdamW for embeddings)
  • Hugging Face Transformers compatible

Model Architecture

Property Value
Architecture Decoder-only Transformer
Parameters 201,327,632 (201.3M)
Layers 8
Hidden Size 1024
Attention Heads 8
KV Heads 8
Head Dimension 128
Feed Forward Size 4096
Context Length 2048
Vocabulary Size 16,384
Positional Encoding RoPE
Activation ReLU²
Normalization RMSNorm
Window Pattern SSSL
Weight Tying No

Training

This model was trained from scratch for 0.2 hours (629s) of wall-clock training time.

Training Configuration

Setting Value
Optimizer MuonAdamW (Muon + AdamW)
Precision torch.bfloat16
Learning Rate 0.04 (matrix) / 0.6 (embedding)
Weight Decay 0.2
Batch Size 4 × 2048 = 8,192 tokens/step
Gradient Accumulation 64 steps
Total Batch Size 524,288 tokens
Context Length 2048
Vocabulary 16,384 tokens (BPE)
LR Scheduler Linear warmdown (50%)
Activation Checkpointing Enabled

Hardware

  • GPU: NVIDIA GeForce RTX 4060 Ti
  • VRAM: 16.0 GB
  • Peak VRAM Used: 4.4 GB
  • MFU: 15.66%
  • Framework: PyTorch 2.9.1+cu128

Dataset

  • Name: TinyStories (karpathy/tinystories-gpt4-clean)
  • Language: English

Preprocessing

Data is packed into fixed-length sequences of 2048 tokens using the nanochat-compatible BPE tokenizer (16,384 vocabulary, 9 special tokens). No additional filtering or deduplication is applied beyond what is in the source dataset.


Intended Use

This model is intended for:

  • Educational purposes and research
  • Text generation experiments
  • Studying small language model training dynamics

Not recommended for:

  • Production use or safety-critical applications
  • Tasks requiring factual accuracy

Evaluation

Results

Metric Score
Validation BPB 1.259341
Perplexity 2.3939
Peak VRAM 4.4 GB
MFU 15.66%

Example Generations

Example 1

Prompt

Once upon a time,

Generation

Once upon a time, there was a small cat named Tim was very sad. "It was a big, the sand and played with it. He liked to play with her friends were playing together.
One day, a big snow. Tim loved the bird was very happy.

Example 2

Prompt

A lonely dragon

Generation

A lonely dragon. And from her mommy to take it. A big smile. She asked him.<|bos|>Once upon a time, there was a big, there was a big, there was a little girl named Max and down. Lucy liked to dance

Example 3

Prompt

The opposite of boy is

Generation

The opposite of boy is not know what to get his friends. It was a big, she was happy and Ben are glad you make something with the woods. They had to say goodbye and they can be kind. He did not care of the best friends, that

Example 4

Prompt

The opposite of queen is

Generation

The opposite of queen is friendly elephant, but it too. It was sad. He thought it." Tim saw the park.<|bos|>Once upon a time, there was a little girl named Mommy. Tim was very surprised. They had a tree, but Tim was very

Example 5

Prompt

My name is

Generation

My name is a little boy named Tim. Lily. The purple ball.<|bos|>One day, a big stick. They were sad and pretty flower was very happy. Lily and softly, "Yes, it. Tim was so happy and loved to do

Usage

import torch
import pickle
import json
from train import GPT, GPTConfig, Tokenizer

# Load config
with open('config.json', 'r') as f:
    config_dict = json.load(f)
config = GPTConfig(**{k: v for k, v in config_dict.items() if k in GPTConfig.__dataclass_fields__})

# Load model
model = GPT(config)
state_dict = torch.load('model.pt', map_location='cpu')['state_dict']
model.load_state_dict(state_dict)
model.eval()

# Load tokenizer
with open('tokenizer.pkl', 'rb') as f:
    tokenizer = pickle.load(f)

# Generate
prompt = 'Once upon a time, '
input_ids = tokenizer.encode(prompt)
x = torch.tensor([input_ids], dtype=torch.long)
with torch.no_grad():
    for _ in range(50):
        logits = model(x)
        probs = torch.softmax(logits[:, -1, :] / 0.8, dim=-1)
        next_token = torch.multinomial(probs, num_samples=1)
        input_ids.append(next_token.item())
        x = torch.tensor([input_ids], dtype=torch.long)
print(tokenizer.decode(input_ids))

Repository Structure

model.pt                  # Model weights
config.json               # Model architecture config
dataset.txt               # Dataset name used for training
token_bytes.pt            # Token byte mappings
tokenizer.pkl             # Trained BPE tokenizer
tokenizer_config.json     # Tokenizer configuration
training_metrics.json     # Training metrics
README.md                 # This file

Limitations

  • Small model size limits language understanding and coherence
  • Trained on a single dataset (TinyStories) — limited domain
  • Fixed time budget training — not fully trained to convergence
  • No RLHF or safety alignment

Ethical Considerations

  • This is a research artifact, not a production model
  • The training data consists of synthetic stories (GPT-4 generated)
  • No harmful content filtering was applied
  • Intended for research and educational use only

Citation

@misc{autoresearch_tinystories_depth8,
  title={AutoResearch-tinystories-depth8},
  author={Dustin Loring},
  year={2026},
  howpublished={\url{https://huggingface.co/quik-models/upbeat-waterfall-104}}
}}

Version History

Version Date Notes
v1.0 2026-07-30 Initial release

Acknowledgements

Built with the AutoResearch training framework.

Thanks to:

  • Hugging Face
  • PyTorch
  • The creators of the TinyStories dataset
  • The open-source AI research community

License

This model is released under the MIT License unless otherwise specified.