krpraveen/indictrans2-sanskrit-en-finetuned

🤗 Hugging Face sourcetranslationmit228M params913 MBsafetensors✓ 3 checksumsupdated today
Submit in one command

Run it next to your model folder. It makes the torrent, checks your files against Hugging Face, and submits it. You just start seeding and paste your key from your account. It only reads your files and never changes them. Read the script first if you like.

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo krpraveen/indictrans2-sanskrit-en-finetuned ./model-folder
Needs a seeder →

IndicTrans2 (200M) fine-tuned — Sanskrit → English

Full fine-tune of ai4bharat/indictrans2-indic-en-dist-200M on 10,000 Sanskrit–English sentence pairs (NLU Assignment 2). Trained only on the provided data — no external parallel corpora.

Direction Sanskrit (san_Deva) → English (eng_Latn)
Parameters 211M
Test BLEU / BERTScore-F1 0.244 / 0.604
Base model IndicTrans2 distilled 200M (MIT)

BLEU is the default NLTK corpus BLEU; BERTScore is F1 with rescale_with_baseline=True.

Install

pip install torch transformers==4.49.0 IndicTransToolkit==1.1.1 sentencepiece sacremoses

Inference

from transformers import AutoModelForSeq2SeqLM, AutoTokenizer
from IndicTransToolkit.processor import IndicProcessor

repo = "krpraveen/indictrans2-sanskrit-en-finetuned"
tok = AutoTokenizer.from_pretrained(repo, trust_remote_code=True)
model = AutoModelForSeq2SeqLM.from_pretrained(repo, trust_remote_code=True).eval()
ip = IndicProcessor(inference=True)

def translate(sentences):
    pre = ip.preprocess_batch(sentences, src_lang="san_Deva", tgt_lang="eng_Latn")
    enc = tok(pre, return_tensors="pt", padding=True, truncation=True, max_length=256)
    out = model.generate(**enc, num_beams=5, max_length=256,
                         length_penalty=1.2, no_repeat_ngram_size=3)
    return ip.postprocess_batch(tok.batch_decode(out, skip_special_tokens=True), lang="eng_Latn")

print(translate(["बाल: भवत्सु प्रेमं प्रकटयति ।"]))
# ['Boy displays affection in you all.']

Runs on CPU or GPU; add .half().cuda() on a GPU for speed.

Use it from an open-source chat UI (Gradio)

The snippet below wraps the model in a Gradio ChatInterface — a free, open-source chat UI you can run locally or host as a Hugging Face Space. Type Sanskrit, get the English translation as the reply.

import gradio as gr

def respond(message, history):
    return translate([message])[0]

gr.ChatInterface(
    respond,
    title="Sanskrit → English translator",
    description="Type a Sanskrit sentence in Devanagari.",
    examples=["बाल: भवत्सु प्रेमं प्रकटयति ।", "अस्तु, इदं सम्यक् दृश्यते ।"],
).launch()

pip install gradio first. To publish it, create a Hugging Face Space (SDK: Gradio) with an app.py containing the inference code above plus this block, and a requirements.txt listing torch transformers==4.49.0 IndicTransToolkit==1.1.1 sentencepiece sacremoses gradio.

Disclosure

Derived from IndicTrans2 (MIT). roberta-large is used only inside bert-score for evaluation, not for translation. No external translation APIs and no extra parallel data were used.