krpraveen/indictrans2-sanskrit-en-finetuned

🤗 Hugging Face 来源translationmit228M 参数913 MBsafetensors✓ 3 个校验和今天更新
一条命令提交

在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo krpraveen/indictrans2-sanskrit-en-finetuned ./model-folder
需要做种者 →

IndicTrans2 (200M) fine-tuned — Sanskrit → English

Full fine-tune of ai4bharat/indictrans2-indic-en-dist-200M on 10,000 Sanskrit–English sentence pairs (NLU Assignment 2). Trained only on the provided data — no external parallel corpora.

Direction Sanskrit (san_Deva) → English (eng_Latn)
Parameters 211M
Test BLEU / BERTScore-F1 0.244 / 0.604
Base model IndicTrans2 distilled 200M (MIT)

BLEU is the default NLTK corpus BLEU; BERTScore is F1 with rescale_with_baseline=True.

Install

pip install torch transformers==4.49.0 IndicTransToolkit==1.1.1 sentencepiece sacremoses

Inference

from transformers import AutoModelForSeq2SeqLM, AutoTokenizer
from IndicTransToolkit.processor import IndicProcessor

repo = "krpraveen/indictrans2-sanskrit-en-finetuned"
tok = AutoTokenizer.from_pretrained(repo, trust_remote_code=True)
model = AutoModelForSeq2SeqLM.from_pretrained(repo, trust_remote_code=True).eval()
ip = IndicProcessor(inference=True)

def translate(sentences):
    pre = ip.preprocess_batch(sentences, src_lang="san_Deva", tgt_lang="eng_Latn")
    enc = tok(pre, return_tensors="pt", padding=True, truncation=True, max_length=256)
    out = model.generate(**enc, num_beams=5, max_length=256,
                         length_penalty=1.2, no_repeat_ngram_size=3)
    return ip.postprocess_batch(tok.batch_decode(out, skip_special_tokens=True), lang="eng_Latn")

print(translate(["बाल: भवत्सु प्रेमं प्रकटयति ।"]))
# ['Boy displays affection in you all.']

Runs on CPU or GPU; add .half().cuda() on a GPU for speed.

Use it from an open-source chat UI (Gradio)

The snippet below wraps the model in a Gradio ChatInterface — a free, open-source chat UI you can run locally or host as a Hugging Face Space. Type Sanskrit, get the English translation as the reply.

import gradio as gr

def respond(message, history):
    return translate([message])[0]

gr.ChatInterface(
    respond,
    title="Sanskrit → English translator",
    description="Type a Sanskrit sentence in Devanagari.",
    examples=["बाल: भवत्सु प्रेमं प्रकटयति ।", "अस्तु, इदं सम्यक् दृश्यते ।"],
).launch()

pip install gradio first. To publish it, create a Hugging Face Space (SDK: Gradio) with an app.py containing the inference code above plus this block, and a requirements.txt listing torch transformers==4.49.0 IndicTransToolkit==1.1.1 sentencepiece sacremoses gradio.

Disclosure

Derived from IndicTrans2 (MIT). roberta-large is used only inside bert-score for evaluation, not for translation. No external translation APIs and no extra parallel data were used.