Sanskrit → English — Custom Transformer (trained from scratch)
A compact encoder–decoder Transformer trained from scratch on 10,000 Sanskrit–English pairs
(NLU Assignment 2), with a joint SentencePiece BPE vocabulary. It is small and fast — meant to be
efficient rather than to match large pretrained models. This is not a 🤗 Transformers
architecture, so it ships with a self-contained modeling.py.
| Parameters | ~9.4M |
| Architecture | 4+4 layer Transformer, d_model 256, 4 heads, tied embeddings |
| Vocabulary | 8,000 (joint SentencePiece BPE) |
| Test BLEU / BERTScore-F1 | 0.089 / 0.346 |
| Inference | ~7 ms/sentence |
Files: pytorch_model.bin (weights), spm.model (tokenizer), config.json (hyperparameters),
modeling.py (model + load/translate helpers).
Install
pip install torch sentencepiece huggingface_hub
Inference
from huggingface_hub import snapshot_download
import sys
d = snapshot_download("krpraveen/sanskrit-en-custom-transformer")
sys.path.insert(0, d)
from modeling import load, translate
model, sp, cfg = load(d) # add device="cuda" on a GPU
print(translate(model, sp, cfg, ["बाल: भवत्सु प्रेमं प्रकटयति ।"]))
# ['Boy displays love in you.']
Use it from an open-source chat UI (Gradio)
import gradio as gr
def respond(message, history):
return translate(model, sp, cfg, [message])[0]
gr.ChatInterface(
respond,
title="Sanskrit → English (custom Transformer)",
description="Type a Sanskrit sentence in Devanagari.",
examples=["बाल: भवत्सु प्रेमं प्रकटयति ।", "अस्तु, इदं सम्यक् दृश्यते ।"],
).launch()
pip install gradio first. To host it, create a Hugging Face Space (SDK: Gradio) with an
app.py (the load + respond code) and a requirements.txt of
torch sentencepiece huggingface_hub gradio.
Notes
Trained only on the provided dataset — no pretrained weights and no external data. Being a
from-scratch model on 10k pairs, quality is modest; a larger model or more data would help.
For higher quality see the fine-tuned IndicTrans2 model
krpraveen/indictrans2-sanskrit-en-finetuned.