aufklarer/WeSpeaker-ResNet34-LM-CoreML

🤗 Hugging Face 来源audio-classificationmit13 MBother✓ 3 个校验和今天更新
一条命令提交

在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo aufklarer/WeSpeaker-ResNet34-LM-CoreML ./model-folder
需要做种者 →

WeSpeaker-ResNet34-LM — CoreML

CoreML conversion of WeSpeaker ResNet34-LM for Apple Neural Engine.

Produces 256-dimensional L2-normalized speaker embeddings from audio.

Model Details

Detail Value
Architecture ResNet34 with statistics pooling
Parameters ~6.6M
Input 80-bin log-mel spectrogram (16kHz)
Output 256-dim L2-normalized speaker embedding
BatchNorm Fused into Conv2d at conversion time

Usage

let model = try await WeSpeakerModel.fromPretrained(backend: .coreML)
let embedding = model.embed(audio: samples, sampleRate: 16000)
let similarity = WeSpeakerModel.cosineSimilarity(embeddingA, embeddingB)

Variants

Variant Backend Model ID
MLX GPU aufklarer/WeSpeaker-ResNet34-LM-MLX
CoreML Neural Engine aufklarer/WeSpeaker-ResNet34-LM-CoreML

Links