quik-models/iconic-galaxy-102

🤗 On Hugging Facetext-generationmit655 MBotherChecksums witnessedupdated today
Magnet

AutoResearch-tinystories-depth8

!AutoResearch Cover

AutoResearch-tinystories-depth8 is a 184.6M parameter decoder-only Transformer trained from scratch on TinyStories (karpathy/tinystories-gpt4-clean).

This model is part of the AutoResearch project, which focuses on training, evaluating, and releasing efficient language models with reproducible research workflows.


Overview

This is a 8-layer decoder-only Transformer trained on the TinyStories (karpathy/tinystories-gpt4-clean) dataset for 0.2 hours of wall-clock training time. The model achieves a validation bits-per-byte (val_bpb) of 0.917943 (perplexity: 1.8894) on the held-out validation set.


References

Papers

  • NanoGPT / NanoChat architecture patterns

Datasets

  • Training: TinyStories (karpathy/tinystories-gpt4-clean)
  • Tokenizer: climbmix-400b-shuffle

Related Projects

WANDB Run


Highlights

  • Trained from scratch
  • 184.6M parameters
  • Trained on 21.0M tokens (40 steps)
  • 8-layer decoder-only Transformer with sliding window attention
  • RoPE positional encoding, RMSNorm, ReLU² activation
  • MuonAdamW optimizer (Muon for matrices, AdamW for embeddings)
  • Mixture of Experts (4 routed + 1 shared, top-2 routing)
  • Hugging Face Transformers compatible

Model Architecture

| Property | Value |

|-----------|------:|

| Architecture | Decoder-only Transformer |

| Parameters | 184,566,288 (184.6M) |

| Layers | 8 |

| Hidden Size | 512 |

| Attention Heads | 4 |

| KV Heads | 4 |

| Head Dimension | 128 |

| Feed Forward Size | 2048 (MoE: 4 experts, 1 shared, top-2) |

| Context Length | 2048 |

| Vocabulary Size | 16,384 |

| Positional Encoding | RoPE |

| Activation | ReLU² |

| Normalization | RMSNorm |

| Window Pattern | SSSL |

| Weight Tying | No |


Training

This model was trained from scratch for 0.2 hours (619s) of wall-clock training time.

Training Configuration

| Setting | Value |

|---------|------:|

| Optimizer | MuonAdamW (Muon + AdamW) |

| Precision | torch.bfloat16 |

| Learning Rate | 0.04 (matrix) / 0.6 (embedding) |

| Weight Decay | 0.2 |

| Batch Size | 4 × 2048 = 8,192 tokens/step |

| Gradient Accumulation | 64 steps |

| Total Batch Size | 524,288 tokens |

| Context Length | 2048 |

| Vocabulary | 16,384 tokens (BPE) |

| LR Scheduler | Linear warmdown (50%) |

| Activation Checkpointing | Enabled |

Hardware

  • GPU: NVIDIA GeForce RTX 4060 Ti
  • VRAM: 16.0 GB
  • Peak VRAM Used: 4.2 GB
  • MFU: 26.39%
  • Framework: PyTorch 2.9.1+cu128

Dataset

  • Name: TinyStories (karpathy/tinystories-gpt4-clean)
  • Language: English

Preprocessing

Data is packed into fixed-length sequences of 2048 tokens using the nanochat-compatible BPE tokenizer (16,384 vocabulary, 9 special tokens). No additional filtering or deduplication is applied beyond what is in the source dataset.


Intended Use

This model is intended for:

  • Educational purposes and research
  • Text generation experiments
  • Studying small language model training dynamics

Not recommended for:

  • Production use or safety-critical applications
  • Tasks requiring factual accuracy

Evaluation

Results

| Metric | Score |

|---------|------:|

| Validation BPB | 0.917943 |

| Perplexity | 1.8894 |

| Peak VRAM | 4.2 GB |

| MFU | 26.39% |


Example Generations

Example 1

Prompt

Once upon a time,

Generation

Once upon a time, there was a little bird named Tim. Tim was a big bird, so he needed to find a way to help the bird to help, and he went home with his friend.
Tim walked to the park to play and land. They had lots

Example 2

Prompt

A lonely dragon

Generation

A lonely dragon and bravely because it was bright and beautiful. The king promised to be part of the story and that all were never good friends.<|bos|>Once upon a time, there was a big pink bunny named Ollie. Lucy loved

Example 3

Prompt

The opposite of boy is

Generation

The opposite of boy is ignorant. He is scared and he can make others happy.
But, a girl, the boy did not think about the boy. He wanted to make a big boy happy girl. He is very smart and threw the boy.
Sudd

Example 4

Prompt

The opposite of queen is

Generation

The opposite of queen is happy and gentle. The queen is happy. They learned that everyone is different thinks that is the to be patient and has for others.<|bos|>One day, a little boy named Tim saw a happy bird with lots of food. The bird was very imp

Example 5

Prompt

My name is

Generation

My name is Jumpy.<|bos|>Once upon a time, there was a boy named Tim. Tim liked to play with his friends in the park. One day, Tim and his friends were playing outside. He saw a big pile of paper. He wanted to

Usage

import torch
import pickle
import json
from train import GPT, GPTConfig, Tokenizer

# Load config
with open('config.json', 'r') as f:
    config_dict = json.load(f)
config = GPTConfig(**{k: v for k, v in config_dict.items() if k in GPTConfig.__dataclass_fields__})

# Load model
model = GPT(config)
state_dict = torch.load('model.pt', map_location='cpu')['state_dict']
model.load_state_dict(state_dict)
model.eval()

# Load tokenizer
with open('tokenizer.pkl', 'rb') as f:
    tokenizer = pickle.load(f)

# Generate
prompt = 'Once upon a time, '
input_ids = tokenizer.encode(prompt)
x = torch.tensor([input_ids], dtype=torch.long)
with torch.no_grad():
    for _ in range(50):
        logits = model(x)
        probs = torch.softmax(logits[:, -1, :] / 0.8, dim=-1)
        next_token = torch.multinomial(probs, num_samples=1)
        input_ids.append(next_token.item())
        x = torch.tensor([input_ids], dtype=torch.long)
print(tokenizer.decode(input_ids))

Repository Structure

model.pt                  # Model weights
config.json               # Model architecture config
dataset.txt               # Dataset name used for training
token_bytes.pt            # Token byte mappings
tokenizer.pkl             # Trained BPE tokenizer
tokenizer_config.json     # Tokenizer configuration
training_metrics.json     # Training metrics
README.md                 # This file

Limitations

  • Small model size limits language understanding and coherence
  • Trained on a single dataset (TinyStories) — limited domain
  • Fixed time budget training — not fully trained to convergence
  • No RLHF or safety alignment

Ethical Considerations

  • This is a research artifact, not a production model
  • The training data consists of synthetic stories (GPT-4 generated)
  • No harmful content filtering was applied
  • Intended for research and educational use only

Citation

@misc{autoresearch_tinystories_depth8,
  title={AutoResearch-tinystories-depth8},
  author={Dustin Loring},
  year={2026},
  howpublished={\url{https://huggingface.co/quik-models/iconic-galaxy-102}}
}}

Version History

| Version | Date | Notes |

|----------|------|------|

| v1.0 | 2026-07-30 | Initial release |


Acknowledgements

Built with the AutoResearch training framework.

Thanks to:

  • Hugging Face
  • PyTorch
  • The creators of the TinyStories dataset
  • The open-source AI research community

License

This model is released under the MIT License unless otherwise specified.