sbintuitions/hikari-medium

🤗 Hugging Face 来源automatic-speech-recognitionmit764M 参数1.5 GBsafetensors✓ 12 个校验和今天更新
一条命令提交

在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo sbintuitions/hikari-medium ./model-folder
需要做种者 →

Hikari-medium

Hikari-medium is a streaming speech-to-text translation and transcription model. It performs simultaneous, low-latency translation directly from audio, without waiting for an utterance to finish.

Highlights

  • 🎧 Simultaneous S2TT — emits target-language text while the speaker is still talking.
  • 🌐 Language pairs — EN→JA, EN→RU, EN→DE plus streaming ASR in English.
  • ⚡ Low latency — fully causal Whisper encoder; CUDA-graph captured for fast autoregressive decoding.
  • 🔄 Single model, multiple tasks — task and target language are selectable at runtime.
  • 🖥️ Browser-based demo — WebRTC microphone input through a Gradio client.

Architecture

Hikari is a Whisper-style encoder-decoder with two modifications:

  • The encoder is made fully causal, so it can be unrolled over streaming audio chunks.

Training and the causal-alignment objective are described in the paper.

audio chunks ──▶ Causal Whisper Encoder ──▶ Whisper Decoder ──▶ streaming text

Supported tasks

Task Description
transcribe Simultaneous speech-to-text (English)
translate Simultaneous speech-to-text translation (EN→JA, EN→DE, EN→RU)

Usage

The model is intended to be served with the hikari-server / hikari-client tools from the Hikari repository.

# install
uv venv .venv --python=3.10 && source .venv/bin/activate
uv pip install torch==2.8.0 torchcodec==0.7.0 torchaudio==2.8.0 torchvision==0.23.0 \
  --index-url https://download.pytorch.org/whl/cu126
uv pip install "hikari @ git+https://github.com/sbintuitions/hikari"

# start the server (GPU machine) — the checkpoint is fetched from this repo
hikari-server --port 4440 --checkpoint sbintuitions/hikari-medium --device cuda:0

# start the client (local machine), then open http://localhost:5666
hikari-client --server-port 4440 --app-port 5666 --chunk-ms 80

See the README on GitHub for SSH tunneling, development install, and the full configuration surface.

Requirements

  • Python ≥ 3.10
  • PyTorch ≥ 2.8.0 with CUDA (server)
  • CUDA GPU (tested on A100, H100); the client runs on CPU

Acknowledgements

Built on top of OpenAI's Whisper.

License

Released under the MIT License.

Citation

@misc{koshkin2026streamingtranslationtranscriptionspeechtotext,
      title={Streaming Translation and Transcription Through Speech-to-Text Causal Alignment},
      author={Roman Koshkin and Jeon Haesung and Lianbo Liu and Hao Shi and Mengjie Zhao and Yusuke Fujita and Yui Sudo},
      year={2026},
      eprint={2603.11578},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2603.11578}
}