alakxender/mt5-dhivehi-word-parallel

🤗 Hugging Face 来源mit300M 参数1.2 GBsafetensors✓ 2 个校验和今天更新
一条命令提交

在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo alakxender/mt5-dhivehi-word-parallel ./model-folder
需要做种者 →

MT5 Dhivehi (gatitos__en_dv fine-tuned)

This model is a fine-tuned version of google/mt5-small on the Google smol gatitos__en_dv dataset.

⚠️ This is not a general-purpose translator. This finetune is to test MT5 usage on dhivehi. It is not intended for any other use.

Model Summary

  • Base model: google/mt5-small
  • Task: Translation (English → Dhivehi)
  • Domain: Unknown, This is an experimental finetune, so try words or short phrases only.
  • Dataset: google/smol → gatitos__en_dv
  • Training framework: Hugging Face Transformers
  • Loss target: ~0.01

Training Details

Parameter Value
Epochs 90
Batch size 4
Learning rate 5e-5 (constant)
Final train loss 0.3797
Gradient norm (last) 15.72
Total steps 89,460
Samples/sec ~14.24
FLOPs 2.36e+16
  • Training time: ~6.98 hours (25,117 seconds)
  • Optimizer: AdamW
  • Scheduler: Constant (no decay)
  • Logging: Weights & Biases

Example Usage (Gradio)

from transformers import MT5ForConditionalGeneration, T5Tokenizer

model = MT5ForConditionalGeneration.from_pretrained("alakxender/mt5-dhivehi-word-parallel")
tokenizer = T5Tokenizer.from_pretrained("alakxender/mt5-dhivehi-word-parallel")

text = "translate English to Dhivehi: Hello, how are you?"
inputs = tokenizer(text, return_tensors="pt")
output = model.generate(**inputs, max_length=64)
print(tokenizer.decode(output[0], skip_special_tokens=True))

Intended Use

This model is meant for:

  • Research in low-resource translation
  • Experimentation with Dhivehi-language modeling
  • Exprementation on the tokenizer