Limite 1B - Base Soup
Limite 1B - Base Soup is the post-soup checkpoint of the Limite 1B model family.
This model card focuses on loading and running the checkpoint with Hugging Face Transformers. For the latest checkpoint and its full model card, see Limite 1B - Violetto.
Run with Transformers
Limite 1B - Base Soup supports inference through the standard Hugging Face Transformers APIs. The custom architecture code is downloaded from this repository, so loading the model requires trust_remote_code=True.
Validated on an NVIDIA H100 with Python 3.12 · PyTorch 2.11.0 (CUDA 13.0) · Transformers 5.6.2. Other version and hardware combinations have not yet been formally qualified. SDPA is the portable default.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "paradigma-inc/limite-1b-base-soup"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_id,
trust_remote_code=True,
torch_dtype="auto",
device_map="auto",
attn_implementation="sdpa",
)
messages = [{"role": "user", "content": "Solve: If x + 3 = 8, what is x?"}]
text = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True,
)
inputs = tokenizer([text], return_tensors="pt").to(model.device)
with torch.inference_mode():
output_ids = model.generate(**inputs, max_new_tokens=256)
answer = tokenizer.decode(
output_ids[0, inputs.input_ids.shape[1]:],
skip_special_tokens=True,
)
print(answer)
Supported inference paths
| Attention backend | Cache | Status |
|---|---|---|
| SDPA | DynamicCache |
Supported |
| SDPA | mixed global/sliding-window StaticCache |
Supported |
| SDPA | StaticCache + torch.compile |
Supported |
| FlashAttention 2 | DynamicCache |
Supported |
| FlashAttention 2 | StaticCache |
Unsupported; rejected with an explicit error |
FlashAttention 2 with StaticCache is intentionally rejected because that combination does not produce numerically correct logits for Limite's hybrid local/global attention layout. Use SDPA with StaticCache, or FlashAttention 2 with DynamicCache.
FlashAttention 2 was validated through Transformers' kernels-community/flash-attn2 integration with kernels==0.12.3. The separately installed native flash_attn package has not been independently qualified.
Official support currently covers inference. The Transformers implementation is differentiable and exposes the standard causal-language-model loss, but full training, gradient checkpointing, PEFT/LoRA, and distributed-training workflows have not yet been formally validated and are not part of the supported interface.
License
The model weights are released under Apache-2.0.