quik-models/bumbling-wind-67

🤗 On Hugging Facetext-generationmit1.1 GBotherChecksums witnessedupdated today
Magnet

AutoResearch-tinystories-depth8

!AutoResearch Cover

AutoResearch-tinystories-depth8 is a 285.2M parameter decoder-only Transformer trained from scratch on TinyStories (karpathy/tinystories-gpt4-clean).

This model is part of the AutoResearch project, which focuses on training, evaluating, and releasing efficient language models with reproducible research workflows.


Overview

This is a 8-layer decoder-only Transformer trained on the TinyStories (karpathy/tinystories-gpt4-clean) dataset for 0.5 hours of wall-clock training time. The model achieves a validation bits-per-byte (val_bpb) of 0.509277 (perplexity: 1.4233) on the held-out validation set.


References

Papers

  • NanoGPT / NanoChat architecture patterns

Datasets

Related Projects

WANDB Run


Highlights

  • Trained from scratch
  • 285.2M parameters
  • Trained on 39.8M tokens (76 steps)
  • 8-layer decoder-only Transformer with sliding window attention
  • RoPE positional encoding, RMSNorm, ReLU² activation
  • MuonAdamW optimizer (Muon for matrices, AdamW for embeddings)
  • Mixture of Experts (8 routed + 1 shared, top-2 routing)
  • Hugging Face Transformers compatible

Model Architecture

| Property | Value |

|-----------|------:|

| Architecture | Decoder-only Transformer |

| Parameters | 285,245,968 (285.2M) |

| Layers | 8 |

| Hidden Size | 512 |

| Attention Heads | 4 |

| KV Heads | 4 |

| Head Dimension | 128 |

| Feed Forward Size | 2048 (MoE: 8 experts, 1 shared, top-2) |

| Context Length | 2048 |

| Vocabulary Size | 16,384 |

| Positional Encoding | RoPE |

| Activation | ReLU² |

| Normalization | RMSNorm |

| Window Pattern | SSSL |

| Weight Tying | No |


Training

This model was trained from scratch for 0.5 hours (1802s) of wall-clock training time.

Training Configuration

| Setting | Value |

|---------|------:|

| Optimizer | MuonAdamW (Muon + AdamW) |

| Precision | torch.bfloat16 |

| Learning Rate | 0.04 (matrix) / 0.6 (embedding) |

| Weight Decay | 0.2 |

| Batch Size | 4 × 2048 = 8,192 tokens/step |

| Gradient Accumulation | 64 steps |

| Total Batch Size | 524,288 tokens |

| Context Length | 2048 |

| Vocabulary | 16,384 tokens (BPE) |

| LR Scheduler | Linear warmdown (50%) |

| Activation Checkpointing | Enabled |

Hardware

  • GPU: NVIDIA GeForce RTX 4060 Ti
  • VRAM: 16.0 GB
  • Peak VRAM Used: 6.3 GB
  • MFU: 33.08%
  • Framework: PyTorch 2.9.1+cu128

Dataset

  • Name: TinyStories (karpathy/tinystories-gpt4-clean)
  • Language: English

Preprocessing

Data is packed into fixed-length sequences of 2048 tokens using the nanochat-compatible BPE tokenizer (16,384 vocabulary, 9 special tokens). No additional filtering or deduplication is applied beyond what is in the source dataset.


Intended Use

This model is intended for:

  • Educational purposes and research
  • Text generation experiments
  • Studying small language model training dynamics

Not recommended for:

  • Production use or safety-critical applications
  • Tasks requiring factual accuracy

Evaluation

Results

| Metric | Score |

|---------|------:|

| Validation BPB | 0.509277 |

| Perplexity | 1.4233 |

| Peak VRAM | 6.3 GB |

| MFU | 33.08% |


Example Generations

Example 1

Prompt

Once upon a time, 

Generation

Once upon a time, 3 year old boy, Jack, went outside to play. He saw a black and white black dog running around. He pointed and said, "Hi! I'm Jack, this is my dog. He is very black!"
Jack smiled and said,

Example 2

Prompt

A lonely dragon 

Generation

A lonely dragon 3 year old child exploring the new places. She had no one to come along, and no one else was ever seen in the world.<|bos|>Once upon a time, there was a little boy named Tim. Tim had a long rod that he loved

Example 3

Prompt

The opposite of boy is 

Generation

The opposite of boy is 3 years old.
The boy said "I can help you, I can do anything you want. That's very kind of you."
The boy smiled and said "Thank you, boy. I'm so happy to help. Thank you, you're

Example 4

Prompt

The opposite of queen is 

Generation

The opposite of queen is 3 and she is very happy. She hugs queen and kisses her cheek. queen and queen queen were very happy.<|bos|>Once upon a time, there was a little dog named Tim. Tim loved to play in the park. One day, while playing

Example 5

Prompt

My name is 

Generation

My name is 4rayla and I am 3 years 3-year-old. He was very happy with his new friend. 
He was always excited to learn about how to act.<|bos|>Once upon a time, there was a little girl named Lily. She

Usage

import torch
import pickle
import json
from train import GPT, GPTConfig, Tokenizer

# Load config
with open('config.json', 'r') as f:
    config_dict = json.load(f)
config = GPTConfig(**{k: v for k, v in config_dict.items() if k in GPTConfig.__dataclass_fields__})

# Load model
model = GPT(config)
state_dict = torch.load('model.pt', map_location='cpu')['state_dict']
model.load_state_dict(state_dict)
model.eval()

# Load tokenizer
with open('tokenizer.pkl', 'rb') as f:
    tokenizer = pickle.load(f)

# Generate
prompt = 'Once upon a time, '
input_ids = tokenizer.encode(prompt)
x = torch.tensor([input_ids], dtype=torch.long)
with torch.no_grad():
    for _ in range(50):
        logits = model(x)
        probs = torch.softmax(logits[:, -1, :] / 0.8, dim=-1)
        next_token = torch.multinomial(probs, num_samples=1)
        input_ids.append(next_token.item())
        x = torch.tensor([input_ids], dtype=torch.long)
print(tokenizer.decode(input_ids))

Repository Structure

model.pt                  # Model weights
config.json               # Model architecture config
dataset.txt               # Dataset name used for training
token_bytes.pt            # Token byte mappings
tokenizer.pkl             # Trained BPE tokenizer
tokenizer_config.json     # Tokenizer configuration
training_metrics.json     # Training metrics
README.md                 # This file

Limitations

  • Small model size limits language understanding and coherence
  • Trained on a single dataset (TinyStories) — limited domain
  • Fixed time budget training — not fully trained to convergence
  • No RLHF or safety alignment

Ethical Considerations

  • This is a research artifact, not a production model
  • The training data consists of synthetic stories (GPT-4 generated)
  • No harmful content filtering was applied
  • Intended for research and educational use only

Citation

@misc{autoresearch_tinystories_depth8,
  title={AutoResearch-tinystories-depth8},
  author={Dustin Loring},
  year={2026},
  howpublished={\url{https://huggingface.co/quik-models/bumbling-wind-67}}
}}

Version History

| Version | Date | Notes |

|----------|------|------|

| v1.0 | 2026-07-28 | Initial release |


Acknowledgements

Built with the AutoResearch training framework.

Thanks to:

  • Hugging Face
  • PyTorch
  • The creators of the TinyStories dataset
  • The open-source AI research community

License

This model is released under the MIT License unless otherwise specified.