moonshine-ai/moonshine-streaming-tiny-ja

🤗 Hugging Face sourceautomatic-speech-recognitionmit27M params108 MBsafetensors✓ 1 checksumupdated today
Submit in one command

Run it next to your model folder. It makes the torrent, checks your files against Hugging Face, and submits it. You just start seeding and paste your key from your account. It only reads your files and never changes them. Read the script first if you like.

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo moonshine-ai/moonshine-streaming-tiny-ja ./model-folder
Needs a seeder →

Moonshine Streaming Tiny — Japanese

Japanese streaming speech recognition, 27.0M parameters. Same architecture as moonshine-ai/moonshine-streaming-tiny, trained for Japanese with a 12,288-entry Japanese tokenizer.

Moonshine Streaming pairs a 50 Hz time-domain audio frontend with a sliding-window Transformer encoder, so it transcribes incrementally rather than waiting for an utterance to finish. It is intended for on-device use on edge-class hardware.

Checkpoint identity

This repository is a conversion of a specific training checkpoint, recorded here because the training run that produced it was still in progress when this snapshot was taken and a better one may replace it:

Checkpoint ja12k_tiny_stageC_best.safetensors
Stage C (read-speech mix)
Architecture slinkier_prime_adapted
Tokenizer tokenizer_ja12k.json, vocab 12,288
Snapshot taken 2026-08-23
Tensors / parameters 163 / 27.0M

If you need reproducibility, pin the revision of this repository rather than tracking main.

Usage

pip install --upgrade transformers datasets[audio]
from transformers import MoonshineStreamingForConditionalGeneration, AutoProcessor
import torch

model = MoonshineStreamingForConditionalGeneration.from_pretrained(
    "moonshine-ai/moonshine-streaming-tiny-ja"
).eval()
processor = AutoProcessor.from_pretrained("moonshine-ai/moonshine-streaming-tiny-ja")

inputs = processor(audio, return_tensors="pt", sampling_rate=16000)

# Cap the output length. Like other seq2seq ASR models this one can fall into a
# repetition loop, and short or noisy clips are where it happens.
seq_lens = inputs.attention_mask.sum(dim=-1)
max_new_tokens = int((seq_lens * 6.5 / 16000).max().item()) + 2

generated = model.generate(**inputs, max_new_tokens=max_new_tokens)
print(processor.batch_decode(generated, skip_special_tokens=True)[0])

Pass the attention_mask. The encoder applies its per-layer sliding windows only when it is given one; called without a mask it attends over the whole utterance instead, which is a different model from the one that was trained. The processor returns the mask, so the snippet above is the safe form. The processor also pads audio to a whole number of 80-sample frames, which the frontend requires.

Architecture

Encoder 6 layers, width 320, 8 heads, sliding windows (16, 4) on the first two and last two layers and (16, 0) between
Decoder 6 layers, width 320, 8 heads, RoPE over 32 of each head's 40 dimensions
Frontend 50 Hz features, CMVN, asinh compression, two causal stride-2 convolutions
Adapter learned absolute positional embeddings before the decoder

The lookahead layers give roughly 80 ms of lookahead; the intermediate layers have none.

Training data

Trained on a large-scale automatically labeled Japanese corpus, plus a read-speech mix in the final stage:

  • Podcast crawl, roughly 109,000 hours.
  • YouTube crawl, roughly 50,000 hours.
  • Stage C read-speech mix, including Common Voice Japanese.

The podcast and YouTube transcripts are pseudo-labels: they were produced by running a Whisper-family teacher model over crawled audio, not by human transcription. The model therefore inherits the teacher's error modes, including its handling of proper nouns, numerals and code-switching, and its transcription conventions for a language written without spaces. No human-verified transcript was used for the bulk of training.

Evaluation

Japanese is scored on character error rate with spaces removed (cer_nospace), never WER. Japanese is written without spaces, so tokenization differences alone can read as several hundred percent WER while the characters are correct.

suite_ja is FLEURS Japanese (650 utterances) and ReazonSpeech Japanese (5,263).

Full panels, batch 8

Panel CER
fleurs_ja 11.50
reazonspeech_ja 26.73
macro 19.115

Seeded 400-utterance sample, batch 1

Batch 1 is the honest number for deployment. Batched evaluation zero-pads short clips up to the longest in the batch, and that trailing silence flatters the model by more than a point on spontaneous speech.

Panel CER
fleurs_ja 11.62
reazonspeech_ja 27.77
macro 19.70

This repository against the training checkpoint

These weights were converted from the neo training checkpoint, and the conversion was checked by measurement rather than inspection: same seeded sample, same batch size, same normalizer.

fleurs_ja reazonspeech_ja macro
Training checkpoint 11.62 27.77 19.699
This repository 11.35 28.08 19.712

397/400 and 385/400 transcripts are byte-identical. The residual comes from the frontend's 80-sample frame alignment, which this path pads and the training path does not.

Limitations

  • Machine-labeled training data. See above; the model reproduces its teacher's mistakes as well as its strengths.
  • Repetition loops on short clips. About 1.25% of ReazonSpeech utterances in the batch-1 sample run away, costing 0.34 CER. Cap the output length.
  • Short-utterance sensitivity. Utterances with references under ~15 characters are far harder than the macro number suggests (CER above 60% on that bucket of spontaneous speech) and are where numerically small changes produce large per-utterance swings.
  • Evaluated only on read speech (FLEURS) and spontaneous speech (ReazonSpeech). No evaluation of telephony, children's speech, heavy dialect, or noisy far-field conditions.
  • Snapshot of an in-progress run. See the checkpoint identity above.

Out-of-scope use

Not intended for non-consensual surveillance, speaker identification, or high-stakes decisions.

License

MIT.