ilintar/Dasheng-AudioGen-GGUF

🤗 Hugging Face 来源text-to-audioapache-2.011 GBGGUF✓ 4 个校验和今天更新
一条命令提交

在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo ilintar/Dasheng-AudioGen-GGUF ./model-folder
需要做种者 →

Dasheng-AudioGen — GGUF

GGUF weights for mispeech/Dasheng-AudioGen, a text→audio generation model, packaged for the C++/GGML runtime thinksound.cpp.

The port is verified faithful to the PyTorch reference (stage-by-stage parity, and the generated audio matches the reference character and amplitude). It runs on CPU, ROCm/HIP, and Vulkan.

Pipeline

text → Flan-T5-Large → content adapter (cross-attn + duration predictor)
     → LayerFusionAudioDiT (U-Net flow-matching, 25-step sway sampler, CFG)
     → Vocos decoder (×2 upsampler → ConvNeXt → ISTFT)
     → 16 kHz mono audio

The latent length is the model's predicted duration (from the adapter's global_duration head); the requested duration does not force the output length.

Files

File Size Contents
dasheng-dit.gguf 8.7 GB LayerFusionAudioDiT backbone and the content adapter (loaded via tensor-name filter)
dasheng-decoder.gguf 696 MB Vocos decoder (upsampler + ConvNeXt + ISTFT head)
flan-t5-large-f32.gguf 1.4 GB Flan-T5-Large text encoder (google/flan-t5-large)
flan-t5-tokenizer.gguf 650 KB T5 SentencePiece tokenizer

Put all four in one directory.

Usage

Build thinksound.cpp, then:

# CLI
ts-dasheng-generate --dir /path/to/gguf --caption "a dog barking" -o out.wav

# HTTP server (POST /v1/dasheng/generate)
ts-server --dir /path/to/gguf

Feed the bare caption (e.g. a dog barking) — do not prepend a <|caption|> tag; it is not a special token in this tokenizer and degrades the output.

License

Apache-2.0, inherited from the upstream mispeech/Dasheng-AudioGen. These are format-converted weights; all model credit goes to the original authors.